Country
Microsoft Accelerator startup DefinedCrowd connects machine learning with native speakers
Part of Microsoft Accelerator's batch 3 of startups, DefinedCrowd is filling a niche in the big data and machine learning community, providing near-real-time feeds of rich language data, checked by actual well-informed humans all over the world. The need comes from the Catch-22 that often arrests deep data analysis, in that you have to understand the data to analyze it, but you must analyze it to understand it. The vast landscape of the spoken and written word and its big data counterpart in natural language processing is especially troublesome in this way. "In the artificial intelligence space, to develop virtual assistants like Cortana, or Apple's Siri and things like that, you need large amounts of voice recordings, you need transcriptions of those voices, you need intents and empathy labeling of those voices," said Daniela Braga, co-founder and chief scientist, in an interview with TechCrunch. "The crowd input provides the extra refinement of the data that basically no machine can do."
Can Neuroscience Understand Donkey Kong, Let Alone a Brain?
Even though the duo knew everything about the chip--the state of each transistor and the voltage along every wire--their inferences were trivial at best and seriously misleading at worst. "Most of my friends assumed that we'd pull out some insights about how the processor works," says Jonas. "But what we extracted was so incredibly superficial. We saw that the processor has a clock and it sometimes reads and writes to memory. Awesome, but in the real world, this would be a millions-of-dollars data set." Last week, the duo uploaded their paper, titled "Could a neuroscientist understand a microprocessor?"
The Craziest Predictions At Code Conference This Week
Some of the most powerful people in the world have awfully bold predictions about the future of mankind. At this week's Code Conference, the likes of Elon Musk and Jeff Bezos hit audiences with some pretty crazy thoughts. Billionaires being held accountable for their actions? Less of a prediction and more of a call-out, Gawker CEO Nick Denton had a particular target in mind when he made comments that Silicon Valley billionaires are a thousand times more powerful, and under a fraction of the scrutiny and regulation, than Congressmen. Denton's recent legal strife, which was partly the work of venture capitalist Peter Thiel, might one day be seen as the suit heard round the world.
Beyond CCA: Moment Matching for Multi-View Models
Podosinnikova, Anastasia, Bach, Francis, Lacoste-Julien, Simon
We introduce three novel semi-parametric extensions of probabilistic canonical correlation analysis with identifiability guarantees. We consider moment matching techniques for estimation in these models. For that, by drawing explicit links between the new models and a discrete version of independent component analysis (DICA), we first extend the DICA cumulant tensors to the new discrete version of CCA. By further using a close connection with independent component analysis, we introduce generalized covariance matrices, which can replace the cumulant tensors in the moment matching framework, and, therefore, improve sample complexity and simplify derivations and algorithms significantly. As the tensor power method or orthogonal joint diagonalization are not applicable in the new setting, we use non-orthogonal joint diagonalization techniques for matching the cumu-lants. We demonstrate performance of the proposed models and estimation techniques on experiments with both synthetic and real datasets.
Semidefinite Programs for Exact Recovery of a Hidden Community
Hajek, Bruce, Wu, Yihong, Xu, Jiaming
We study a semidefinite programming (SDP) relaxation of the maximum likelihood estimation for exactly recovering a hidden community of cardinality $K$ from an $n \times n$ symmetric data matrix $A$, where for distinct indices $i,j$, $A_{ij} \sim P$ if $i, j$ are both in the community and $A_{ij} \sim Q$ otherwise, for two known probability distributions $P$ and $Q$. We identify a sufficient condition and a necessary condition for the success of SDP for the general model. For both the Bernoulli case ($P={{\rm Bern}}(p)$ and $Q={{\rm Bern}}(q)$ with $p>q$) and the Gaussian case ($P=\mathcal{N}(\mu,1)$ and $Q=\mathcal{N}(0,1)$ with $\mu>0$), which correspond to the problem of planted dense subgraph recovery and submatrix localization respectively, the general results lead to the following findings: (1) If $K=\omega( n /\log n)$, SDP attains the information-theoretic recovery limits with sharp constants; (2) If $K=\Theta(n/\log n)$, SDP is order-wise optimal, but strictly suboptimal by a constant factor; (3) If $K=o(n/\log n)$ and $K \to \infty$, SDP is order-wise suboptimal. The same critical scaling for $K$ is found to hold, up to constant factors, for the performance of SDP on the stochastic block model of $n$ vertices partitioned into multiple communities of equal size $K$. A key ingredient in the proof of the necessary condition is a construction of a primal feasible solution based on random perturbation of the true cluster matrix.
Degrees of Freedom in Deep Neural Networks
Gao, Tianxiang, Jojic, Vladimir
In this paper, we explore degrees of freedom in deep sigmoidal neural networks. We show that the degrees of freedom in these models is related to the expected optimism, which is the expected difference between test error and training error. We provide an efficient Monte-Carlo method to estimate the degrees of freedom for multi-class classification methods. We show degrees of freedom are lower than the parameter count in a simple XOR network. We extend these results to neural nets trained on synthetic and real data, and investigate impact of network's architecture and different regularization choices. The degrees of freedom in deep networks are dramatically smaller than the number of parameters, in some real datasets several orders of magnitude. Further, we observe that for fixed number of parameters, deeper networks have less degrees of freedom exhibiting a regularization-by-depth.
Robust Ensemble Clustering Using Probability Trajectories
Huang, Dong, Lai, Jian-Huang, Wang, Chang-Dong
Note that V Y L Link set of G w ij W eight between two nodes in G G K -elite neighbor graph (K -ENG) V Node set of G . Note that V Y L Link set of G w ij W eight between two nodes in G p ij (1-step) transition probability fromy i to y j P (1-step) transition probability matrix,P { p ij } N N p T ij T -step transition probability fromy i to y j P T T -step transition probability matrix,P T { p T ij } N N p T i: The i -th row ofP T, p T i: { p T i 1,ยทยทยท,p T i N} PT T i Probability trajectory of a random walker starting fromnode y i with lengthT PTS ij Probability trajectory based similarity betweeny i and y j R (0) Set of the initial regions for PTA,R (0) { R (0) 1,ยทยทยท,R (0) R (0) } S (0) Initial similarity matrix for PTA,S (0) { s (0) ij } R (0) R (0) R ( t) Set of thet -step regions for PTA, R ( t) { R ( t) 1,ยทยทยท,R ( t) R ( t) } S ( t) The t -step similarity matrix for PTA,S ( t) { s ( t) ij } R ( t) R ( t) G Microcluster-cluster bipartite graph (MCBG) N Number of nodes in G V Node set of G L Link set of G w ij W eight between two nodes in G A sparse graph termedK -elite neighbor graph (K -ENG) is then constructed with only a small number of probably reliable links. The ENS strategy is a crucial step in our approach. W e argue that using a small number of probably reliable links may lead to significantly better consensus results than using all graph links regardless of their reliability . The random walk process driven by a new transition probability matrix is performed on theK -ENG to explore the global structure information.
A Graph-Based Semi-Supervised k Nearest-Neighbor Method for Nonlinear Manifold Distributed Data Classification
Tu, Enmei, Zhang, Yaqian, Zhu, Lin, Yang, Jie, Kasabov, Nikola
$k$ Nearest Neighbors ($k$NN) is one of the most widely used supervised learning algorithms to classify Gaussian distributed data, but it does not achieve good results when it is applied to nonlinear manifold distributed data, especially when a very limited amount of labeled samples are available. In this paper, we propose a new graph-based $k$NN algorithm which can effectively handle both Gaussian distributed data and nonlinear manifold distributed data. To achieve this goal, we first propose a constrained Tired Random Walk (TRW) by constructing an $R$-level nearest-neighbor strengthened tree over the graph, and then compute a TRW matrix for similarity measurement purposes. After this, the nearest neighbors are identified according to the TRW matrix and the class label of a query point is determined by the sum of all the TRW weights of its nearest neighbors. To deal with online situations, we also propose a new algorithm to handle sequential samples based a local neighborhood reconstruction. Comparison experiments are conducted on both synthetic data sets and real-world data sets to demonstrate the validity of the proposed new $k$NN algorithm and its improvements to other version of $k$NN algorithms. Given the widespread appearance of manifold structures in real-world problems and the popularity of the traditional $k$NN algorithm, the proposed manifold version $k$NN shows promising potential for classifying manifold-distributed data.
Characterizing Diseases from Unstructured Text: A Vocabulary Driven Word2vec Approach
Ghosh, Saurav, Chakraborty, Prithwish, Cohn, Emily, Brownstein, John S., Ramakrishnan, Naren
Traditional disease surveillance can be augmented with a wide variety of real-time sources such as, news and social media. However, these sources are in general unstructured and, construction of surveillance tools such as taxonomical correlations and trace mapping involves considerable human supervision. In this paper, we motivate a disease vocabulary driven word2vec model (Dis2Vec) to model diseases and constituent attributes as word embeddings from the HealthMap news corpus. We use these word embeddings to automatically create disease taxonomies and evaluate our model against corresponding human annotated taxonomies. We compare our model accuracies against several state-of-the art word2vec methods. Our results demonstrate that Dis2Vec outperforms traditional distributed vector representations in its ability to faithfully capture taxonomical attributes across different class of diseases such as endemic, emerging and rare.