Genre
Deep Recurrent Neural Network for Protein Function Prediction from Sequence
As high throughput biological sequencing becomes faster and cheaper, the need to extract useful information from sequencing becomes ever more paramount, often limited by low throughput experimental characterizations. For proteins, accurate prediction of their functions directly from their primary amino acid sequences has been a long standing challenge. Here, machine learning using artificial recurrent neural networks (RNN) was applied towards classification of protein function directly from primary sequence without sequence alignment, heuristic scoring or feature engineering. The RNN models containing long short term memory (LSTM) units trained on public, annotated datasets from UniProt achieved high performance for in class prediction of four important protein functions tested, particularly compared to other machine learning algorithms using sequence derived protein features. RNN models were used also for out of class predictions of phylogenetically distinct protein families with similar functions, including proteins of the CRISPR associated nuclease, ferritin like iron storage and cytochrome P450 families. Applying the trained RNN models on the partially unannotated UniRef100 database predicted not only candidates validated by existing annotations but also currently unannotated sequences. Some RNN predictions for the ferritin like iron sequestering function were experimentally validated, even though their sequences differ significantly from known, characterized proteins and from each other and cannot be easily predicted using popular bioinformatics methods. As sequencing and experimental characterization data increases rapidly, the machine learning approach based on RNN could be useful for discovery and prediction of homologues for a wide range of protein functions. Introduction As the cost of DNA sequencing is decreasing drastically over the last decade, the volume of biological sequences particularly for new proteins is also increasing rapidly. Discovering the functions of these new proteins not only could allow one to better understand their roles in their native contexts, but also utilize them in synthetic biology to assembled new biological circuits and pathways for useful applications such as production of valuable compounds or treating disease. However, the experimental characterization of proteins' properties such as structure and function can be slow and resource demanding using techniques such as x ray crystallography, cryo TEM, or functional assays, significantly outpaced by sequencing.
Multiclass MinMax Rank Aggregation
Rankings, a special form of ordinal data, have received significant attention in the machine learning community as they arise in a number of important application domains, such as recommender systems, social voting and product placement platforms. Of particular importance are rankings of the form of linear orders (permutations) and partial rankings (weak orders), which are frequently obtained through conversion from ratings. One of the main processing tasks for rankings is rank aggregation, which often involves evaluating the median of a set of permutations or partial rankings under a suitably chosen distance function [2], [4], [7], [9], [11], [12], [16]. The median rank aggregation problem under the Kendall τ distance was introduced by Kemeny [11], and was proved to be NPhard by Bartholdi et al. [4]. A number of approximation algorithms for the problem have been described in [2], mostly pertaining to permutations; a corresponding PTAS (polynomial time approximation scheme) was proposed in [12]. In the context of partial ranking aggregation, known solutions include the results of [1], [10]. Median aggregation under other distance functions has received less attention, one notable exception being the Spearman rank aggregation problem [7], which is known to provide a constant approximation for Kendall τ aggregation using a polynomial time algorithm based on weighted bipartite matching [9]. We propose to investigate a broad new family of rank aggregation problems in which the median is replaced by a minmax type of function and where the rankings are grouped in classes.
The Opacity of Backbones
Hemaspaandra, Lane A., Narváez, David E.
A backbone of a boolean formula $F$ is a collection $S$ of its variables for which there is a unique partial assignment $a_S$ such that $F[a_S]$ is satisfiable [MZK+99,WGS03]. This paper studies the nontransparency of backbones. We show that, under the widely believed assumption that integer factoring is hard, there exist sets of boolean formulas that have obvious, nontrivial backbones yet finding the values, $a_S$, of those backbones is intractable. We also show that, under the same assumption, there exist sets of boolean formulas that obviously have large backbones yet producing such a backbone $S$ is intractable. Further, we show that if integer factoring is not merely worst-case hard but is frequently hard, as is widely believed, then the frequency of hardness in our two results is not too much less than that frequency.
The value AI brings to marketing
Artificial intelligence will come to the forefront this year in marketing departments across the world. But do marketers really understand what AI is in marketing and how to best implement it in their marketing strategies? A Demandbase and Wakefield Research AI survey asked just these kinds of questions of marketers and found some very interesting results. To understand the results of this survey at a deeper level, I spoke with Aman Naimat, SVP of Technology at Demandbase. Naimat has a very extensive background in data science.
Panasonic introduces service robots to Tokyo Airport
If you're passing through Tokyo Narita Airport, there's a good chance your pre-flight dinner could be cleared up by a robot. Panasonic has started testing its HOSPI service robots at the airport to help combat'labour shortages' in Japan. The firm hopes if the Dalek-style robots are a success in the airport, they could be drafted in to help deal with the influx of tourists for the 2020 Olympics. Panasonic has begun testing service robots at the international airport to help combat'labour shortages' in Japan The robots were originally developed to be used in healthcare, delivering drugs around hospitals. Pre-installed mapping information allows HOSPI to move autonomously.
Apple joins Amazon, Google and Facebook in AI research group
Apple published its first paper on AI last month and now the company is set to join five others in a newly-formed research group. The Partnership on AI announced today that Apple would become its sixth founding member, adding to a lineup that already touts Amazon, Facebook, Google, IBM and Microsoft. The group was first formed last September as a means of supporting research, establishing ethical guidelines and promoting both transparency and privacy when it comes to AI studies. In today's announcement, the Partnership on AI explained that Apple has already been working with the group before it was made official last fall, but now the company is a full member alongside those other tech titans. Part of today's news was also that the group selected its board of trustees that will oversee the initiative.
Virtualitics Launches as First Platform to Merge Artificial Intelligence, Big Data and Virtual/Augmented Reality
WIRE)--Virtualitics, LLC, a data analytics in virtual reality (VR) and augmented reality (AR) startup, today announced the launch of its new tool that combines powerful data visualization in VR/AR and artificial intelligence to provide insights and discover actionable knowledge hidden in big and complex data. The company also announced they closed a $3 million investment seed round from angel investors. Virtualitics combines VR/AR with machine learning and natural language in a data exploration, collaborative environment suitable for both data scientists and non-expert users. The technology is the only one of its kind that can provide a simultaneous rendering of up to 10 dimensions, revealing multidimensional relationships present in the data, which may not be discoverable in any other way. The company is based on over a decade of research at Caltech and NASA's Jet Propulsion Laboratory (JPL) and was founded by: professor George Djorgovski, the founding director of Caltech's Center for Data-Driven Discovery; Michael Amori, former Deutsche Bank managing director, Harvard MBA and Caltech alumnus; Dr. Ciro Donalek, computational scientist at Caltech; and Dr. Scott Davidoff, manager of the Human Interfaces Group at NASA's JPL.
The Deep Learning Market Map: 60 Startups Working Across E-Commerce, Cybersecurity, Sales, And More
Increased investor interest in AI startups – from around 10 deals in Q1'11 to over 120 in Q2'16 – can be attributed to recent advances in machine learning algorithms, particularly "deep learning" technology, a souped up version of AI. Just this week, Google integrated deep learning into its Google Translate tool; Baidu announced the launch of DeepBench, an "open source benchmarking tool for evaluating deep learning performance across different hardware platforms"; and NVIDIA introduced Xavier, a deep learning-based supercomputer for driverless cars. In the private market, Google put deep learning in the spotlight back in 2014 when it acquired 4 startups focused on this AI tech in quick succession: DeepMind, Vision Factory, Dark Blue Labs, and DNNresearch. Apple, which joined the race in 2015, most recently acquired Turi, which has developed a deep learning toolkit, among other AI-based solutions. Not to be outdone, Intel has acquired around 5 AI startups since January 2015, including deep learning startup Nervana Systems and, more recently, Movidius.
The Proliferation of AI and Automation in Banking
As investigations into Wells Fargo's systemic banking fraud continue, new findings show that retail branches were given notice of internal audits a day or three in advance of them taking place. This allowed for managers and employees to scrub, purge and modify documents and data, which, by most firsthand accounts, probably looked something like this. With new oversights and practices well underway at Wells Fargo, many banks are having a broader discussion of just what guaranteed transparency and compliance looks like beyond just throwing people and money at the situation. Many institutions are asking themselves, especially in the wake of the Wells Fargo mess, "What does the future of controls, compliance and processes look like?" The answer to this question, like so many others in a rapidly changing, operationally strained industry, is increasingly rooted in the adoption of technologies like automation and artificial intelligence (AI).
Artificial intelligence uncovers new insight into biophysics of cancer
Their machine-learning platform predicted a trio of reagents that was able to generate a never-before-seen cancer-like phenotype in tadpoles. The research, reported in Scientific Reports on January 27, shows how artificial intelligence (AI) can help human researchers in fields such as oncology and regenerative medicine control complex biological systems to reach new and previously unachievable outcomes. The researchers had previously shown that pigment cells (melanocytes) in developing frogs could be converted to a cancer-like, metastatic form by disrupting their normal bioelectric and serotonergic signaling and had used AI to reverse-engineer a model that explained this complex process. However, during these extensive experiments, the biologists observed something remarkable: All the melanocytes in a single frog larva either converted to the cancer-like form or remained completely normal. Conversion of only some of the pigment cells in a single tadpole was never seen; how, the researchers asked, could such an all-or-none coordination of cells across the tadpole body be explained and controlled?