Goto

Collaborating Authors

 Scientific Discovery


Top 10 Capabilities for Exploring Complex Relationships in Data for Scientific Discovery

@machinelearnbot

With all of the discussion about Big Data these days, there is frequest reference to the 3 V's that represent the top big data challenges: Volume, Velocity, and Variety. These 3 V's generally refer to the size of the dataset (Volume), the rate at which data is flowing into (or out of) your systems (Velocity), and the complexity (dimensionality) of the data (Variety). Most practitioners agree that big data volume is indeed huge, but that is not necessarily big data's biggest challenge, at least not in terms of data storage capacities, which are growing rapidly also and keeping pace with data volume. The velocity of big data is also a very big challenge, though primarily for applications and use cases that specifically demand near-real-time analysis and response to dynamic data streams. However, unlike volume and velocity, most will agree that the variety (complexity) of the data is truly big data's biggest mega-challenge at all scales and in most applications.


Big Data Discovery Is The Next Big Trend In Analytics ZDNet

#artificialintelligence

According to Gartner, "Big Data Discovery" is the next big trend in analytics. Each of these areas has seen explosive growth, but there are clear upsides and downsides to each. For example, Data Discovery excels in ease of use, but allows only limited depth of exploration, while Data Science provides powerful analysis but is slow, complex, and difficult to implement. Since the disadvantages of the three technologies map to nicely to the advantages of the others, they are now starting to blend, and Gartner believes Big Data Discovery will be a distinct new market category by 2017. The emerging Big Data Discovery tools will be simpler to use than data science products and accessible to a wider ranger of users, with more powerful manipulation of a wider variety of data sources. According to Gartner Analyst Joao Tapadinhas, these tools will be used by new "Citizen Data Scientists" who marry the skills of traditional business analysts with some of the expertise of expert statisticians.


Artificial Intelligence to Win the Nobel Prize and Beyond: Creating the Engine for Scientific Discovery

AI Magazine

This article proposes a new grand challenge for AI reasearch: to develop AI system to make major scientific discoveries in biomedical sciences that worth Nobel Prize. There are a series of human cognitive limitations that prevents us from making accerlated scientific discoveries, particularity in biomedical sciences. As a result, scientific discoveries are left behind at the level of cottage industry. AI systems can transform scientific discoveries into highly efficient practice, thereby enable us to expand our knowledge in unprecedented way.


Artificial Intelligence to Win the Nobel Prize and Beyond: Creating the Engine for Scientific Discovery

AI Magazine

This article proposes a new grand challenge for AI reasearch: to develop AI system to make major scientific discoveries in biomedical sciences that worth Nobel Prize. There are a series of human cognitive limitations that prevents us from making accerlated scientific discoveries, particularity in biomedical sciences. As a result, scientific discoveries are left behind at the level of cottage industry. AI systems can transform scientific discoveries into highly efficient practice, thereby enable us to expand our knowledge in unprecedented way. Such system may out-compute all possible hypotheses and may redefine the nature of scientific intuition, hence scientific discovery process.


The Silent Rockstar of BigData: Machine Learning

#artificialintelligence

Sure, world is crying out loud that big-data's biggest problem will be resources. Demand has skyrocketed and everyone in the world is going into tailspin in meeting that demands. Companies are going frantic and overspending to hire data scientists to secure themselves from any upcoming shortfall. This is nothing but a sign that world needs our robot algorithm friends to pacify some demand and increase credibility to new paradigms. Who could forget Steve Balmer's famous quote comparing Big Data as a Machine Learning problem.


Key-Object โ€“ A New Paradigm in Search?

@machinelearnbot

Summary: The premise of this new Key Object architecture is that search is broken, at least as it applies to complex merchandise like computers, printers, and cameras. An innovative and workable solution is described. The question remains, is the pain sufficient to justify a switch? As we are all fond of saying, innovation follows pain points. Are we missing something in our uber-critical search capabilities that needs to be resolved?


Accelerating Science: A Computing Research Agenda

arXiv.org Artificial Intelligence

The emergence of "big data" offers unprecedented opportunities for not only accelerating scientific advances but also enabling new modes of discovery. Scientific progress in many disciplines is increasingly enabled by our ability to examine natural phenomena through the computational lens, i.e., using algorithmic or information processing abstractions of the underlying processes; and our ability to acquire, share, integrate and analyze disparate types of data. However, there is a huge gap between our ability to acquire, store, and process data and our ability to make effective use of the data to advance discovery. Despite successful automation of routine aspects of data management and analytics, most elements of the scientific process currently require considerable human expertise and effort. Accelerating science to keep pace with the rate of data acquisition and data processing calls for the development of algorithmic or information processing abstractions, coupled with formal methods and tools for modeling and simulation of natural processes as well as major innovations in cognitive tools for scientists, i.e., computational tools that leverage and extend the reach of human intellect, and partner with humans on a broad range of tasks in scientific discovery (e.g., identifying, prioritizing formulating questions, designing, prioritizing and executing experiments designed to answer a chosen question, drawing inferences and evaluating the results, and formulating new questions, in a closed-loop fashion). This calls for concerted research agenda aimed at: Development, analysis, integration, sharing, and simulation of algorithmic or information processing abstractions of natural processes, coupled with formal methods and tools for their analyses and simulation; Innovations in cognitive tools that augment and extend human intellect and partner with humans in all aspects of science.


Zaloni's new Mica release makes data discovery, curation and self-service data preparation more collaborative and intuit

#artificialintelligence

Zaloni, the data lake company, released today a new version of its Mica self-service data preparation platform at Strata Hadoop World. Mica provides users with an on-ramp for self-service data discovery, curation, and preparation of data in the data lake. With Mica, business users have the tools they need for rapidly discovering data sets, interacting with them and uncovering needed business insights. According to a January 2016 report, entitled Overcoming Obstacles That Prevent the Deployment and Use of a Modern BI and Analytics Platform, Gartner predicts that by "2017, most business users and analysts in organizations will have access to self-service tools to prepare data for analysis." "Data preparation can no longer exclusively be an IT function," said Ben Sharma, Zaloni's co-founder and CEO.


Is data science a new paradigm, or recycled material?

@machinelearnbot

Data science is the result of a new paradigm taking place in IT. The question was raised recently, and here I explain how and why data science is part of this new paradigm, and not recycled material. Many data science techniques are very different, if not the opposite of old techniques that were designed to be implemented on abacus, rather than computers. These new tools are often model-free. Indeed, old techniques such as logistic regression and classification trees don't even belong to data science, more stable techniques are used in data science.


Bayesian hypothesis testing for one bit compressed sensing with sensing matrix perturbation

arXiv.org Machine Learning

This letter proposes a low-computational Bayesian algorithm for noisy sparse recovery in the context of one bit compressed sensing with sensing matrix perturbation. The proposed algorithm which is called BHT-MLE comprises a sparse support detector and an amplitude estimator. The support detector utilizes Bayesian hypothesis test, while the amplitude estimator uses an ML estimator which is obtained by solving a convex optimization problem. Simulation results show that BHT-MLE algorithm offers more reconstruction accuracy than that of an ML estimator (MLE) at a low computational cost.