Country
Machine Learning with Micron's Automata Processor - insideBIGDATA
A new survey paper describing Micron's Automata Processor (AP) was recently published. AP has many potential applications in data mining, bioinformatics, natural language processing, etc. Micron has recently stopped developing AP, however other companies such as Natural Intelligence Semiconductor, a spin-off from Micron) and some academic research centers (Center for Automata Processing at the University of Virginia) are leading the development and market-adoption of AP. Problems from a wide variety of application domains can be modeled as "nondeterministic finite automaton" (NFA) and hence, efficient execution of NFAs can improve the performance of several key applications. However, traditional architectures, such as CPU and GPU are not inherently suited for executing NFAs, and hence, special-purpose architectures are required for accelerating them. Micron's automata processor (AP) exploits massively parallel in-memory processing capability of DRAM for executing NFAs and hence, it can provide orders of magnitude performance improvement compared to traditional architectures.
Government Should Address Potential Bias in Artificial Intelligence, Lawmakers Say
Bias in artificial intelligence could critically impact the deployment, adoption and evolution of the technology, Democratic lawmakers said in Washington Wednesday. They also detailed their plans to combat the issue and help America maintain its position as a global leader in AI. "We have a real concern about bias in data--is that data bias in some historical way or in some intentional way?" Rep. Jerry McNerney, D-Calif., told Politico's Steven Overly at an event held by the publication and technology corporation, Intel. "And we want to make sure the data doesn't harm groups of people or sectors of the country." McNerney, who holds a doctorate in mathematics, elaborated on how algorithm developers implicitly assume that whatever they produce will be logical. Often, it isn't until they look back at the results that they realize they've inserted bias through actions like failing to make important considerations during the process or using data that excluded certain groups of people.
Data Science Jobs Report 2019: Python Way Up, TensorFlow Growing Rapidly, R Use Double SAS
In my ongoing quest to track The Popularity of Data Science Software, I've just updated my analysis of the job market. To save you from reading the entire tome, I'm reproducing that section here. One of the best ways to measure the popularity or market share of software for data science is to count the number of job advertisements that highlight knowledge of each as a requirement. Job ads are rich in information and are backed by money, so they are perhaps the best measure of how popular each software is now. Plots of change in job demand give us a good idea of what is likely to become more popular in the future.
15 Open Datasets for Healthcare
Machine Learning is exploding into the world of healthcare. When we talk about the ways ML will revolutionize certain fields, healthcare is always one of the top areas seeing huge strides, thanks to the processing and learning power of machines. There's a good chance you either are or will soon be employed in the healthcare field. A while back, I wrote a list of 25 excellent open datasets for ML and included healthdata.gov and MIMIC Critical Care Database. Here are 15 more excellent datasets specifically for healthcare.
Fortnite World Cup kicks off with $30m at stake
After 10 weeks of open qualifiers attracting more than 40 million competitors, the Fortnite World Cup finals will be held this weekend in New York. Up for grabs for the 100 qualifiers – many of whom are between 12 and 16 years old – is a total prize pot of $30m, the largest ever for an esports event. With more than 250 million players, Fortnite: Battle Royale has become one of the most popular video games in the world since its launch in 2017. The World Cup represents the title's entry into the lucrative world of professional games tournament circuits, where revenues are set to pass $1bn this year, due to exploding sponsorship, advertising and broadcast rights. Epic Games has arranged a three-day festival around the finals, with thousands of fans expected to attend.
Video game streaming: is it worth it?
Streaming video games is an idea with such obvious advantages that like virtual reality, motion controls and 3D screens, it had already hit the market several times before it was technologically possible: witness the untimely demise of OnLive in 2015. The big question facing Microsoft and Google, both of which showed off their entries into the "cloud gaming" market at the E3 video game conference in Los Angeles last month, is whether they've taken the plunge at the right time, or whether they, too, will be chalked up in history as premature entrants. After playing with Microsoft's Project xCloud and Google's Stadia, we can draw some conclusions but others will have to wait. Both services are aiming at different targets, and based on the idealised situations in which they were presented, they each achieve their goals. But not everything is in their hands. No plan survives contact with the enemy, and no streaming service has yet survived contact with the realities of home broadband.
DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks
Wiedemann, Simon, Kirchoffer, Heiner, Matlage, Stefan, Haase, Paul, Marban, Arturo, Marinc, Talmaj, Neumann, David, Nguyen, Tung, Osman, Ahmed, Marpe, Detlev, Schwarz, Heiko, Wiegand, Thomas, Samek, Wojciech
The field of video compression has developed some of the most sophisticated and efficient compression algorithms known in the literature, enabling very high compressibility for little loss of information. Whilst some of these techniques are domain specific, many of their underlying principles are universal in that they can be adapted and applied for compressing different types of data. In this work we present DeepCABAC, a compression algorithm for deep neural networks that is based on one of the state-of-the-art video coding techniques. Concretely, it applies a Context-based Adaptive Binary Arithmetic Coder (CABAC) to the network's parameters, which was originally designed for the H.264/AVC video coding standard and became the state-of-the-art for lossless compression. Moreover, DeepCABAC employs a novel quantization scheme that minimizes the rate-distortion function while simultaneously taking the impact of quantization onto the accuracy of the network into account. Experimental results show that DeepCABAC consistently attains higher compression rates than previously proposed coding techniques for neural network compression. For instance, it is able to compress the VGG16 ImageNet model by x63.6 with no loss of accuracy, thus being able to represent the entire network with merely 8.7MB. The source code for encoding and decoding can be found at https://github.com/fraunhoferhhi/DeepCABAC.
Dilated FCN: Listening Longer to Hear Better
Gong, Shuyu, Wang, Zhewei, Sun, Tao, Zhang, Yuanhang, Smith, Charles D., Xu, Li, Liu, Jundong
Deep neural network solutions have emerged as a new and powerful paradigm for speech enhancement (SE). The capabilities to capture long context and extract multi-scale patterns are crucial to design effective SE networks. Such capabilities, however, are often in conflict with the goal of maintaining compact networks to ensure good system generalization. In this paper, we explore dilation operations and apply them to fully convolutional networks (FCNs) to address this issue. Dilations equip the networks with greatly expanded receptive fields, without increasing the number of parameters. Different strategies to fuse multi-scale dilations, as well as to install the dilation modules are explored in this work. Using Noisy VCTK and AzBio sentences datasets, we demonstrate that the proposed dilation models significantly improve over the baseline FCN and outperform the state-of-the-art SE solutions.
Alternative Blockmodelling
Correa, Oscar, Chan, Jeffrey, Nguyen, Vinh
Many approaches have been proposed to discover clusters within networks. Community finding field encompasses approaches which try to discover clusters where nodes are tightly related within them but loosely related with nodes of other clusters. However, a community network configuration is not the only possible latent structure in a graph. Core-periphery and hierarchical network configurations are valid structures to discover in a relational dataset. On the other hand, a network is not completely explained by only knowing the membership of each node. A high level view of the inter-cluster relationships is needed. Blockmodelling techniques deal with these two issues. Firstly, blockmodelling allows finding any network configuration besides to the well-known community structure. Secondly, blockmodelling is a summary representation of a network which regards not only membership of nodes but also relations between clusters. Finally, a unique summary representation of a network is unlikely. Networks might hide more than one blockmodel. Therefore, our proposed problem aims to discover a secondary blockmodel representation of a network that is of good quality and dissimilar with respect to a given blockmodel. Our methodology is presented through two approaches, (a) inclusion of cannot-link constraints and (b) dissimilarity between image matrices. Both approaches are based on non-negative matrix factorisation NMF which fits the blockmodelling representation. The evaluation of these two approaches regards quality and dissimilarity of the discovered alternative blockmodel as these are the requirements of the problem.
A Matrix--free Likelihood Method for Exploratory Factor Analysis of High-dimensional Gaussian Data
Dai, Fan, Dutta, Somak, Maitra, Ranjan
This paper proposes a novel profile likelihood method for estimating the covariance parameters in exploratory factor analysis of high-dimensional Gaussian datasets with fewer observations than number of variables. An implicitly restarted Lanczos algorithm and a limited-memory quasi-Newton method are implemented to develop a matrix-free framework for likelihood maximization. Simulation results show that our method is substantially faster than the expectation-maximization solution without sacrificing accuracy. Our method is applied to fit factor models on data from suicide attempters, suicide ideators and a control group.