Scientific Discovery
Thomas Kuhn Threw an Ashtray at Me - Issue 63: Horizons
Errol Morris feels that Thomas Kuhn saved him from a career he was not suited for--by having him thrown out of Princeton. In 1972, Kuhn was a professor of philosophy and the history of science at Princeton, and author of The Structure of Scientific Revolutions, which gave the world the term "paradigm shift." As Morris tells the story in his recent book, The Ashtray, Kuhn was antagonized by Morris' suggestions that Kuhn was a megalomaniac and The Structure of Scientific Revolutions was an assault on truth and progress. To say the least, Morris, then 24, was already the iconoclast who would go on to make some of the most original documentary films of our time. After launching the career he was suited for with The Gates of Heaven in 1978, a droll affair about pet cemeteries, Morris earned international acclaim with The Thin Blue Line, which led to the reversal of a murder conviction of a prisoner who had been on death row. In 2004, Morris won an Academy Award for The Fog of War, a dissection of former Secretary of Defense Robert McNamara, a major architect of the Vietnam War. His 2017 film, Wormwood, a miniseries on Netflix, centers on the mystery surrounding a scientist who in 1975 worked on a biological warfare program for the Army, and suspiciously fell to his death from a hotel room. The Ashtray--Morris explains the title in our interview below--is as arresting and idiosyncratic as Morris' films.
AI for code encourages collaborative, open scientific discovery
We have seen significant recent progress in pattern analysis and machine intelligence applied to images, audio and video signals, and natural language text, but not as much applied to another artifact produced by people: computer program source code. In a paper to be presented at the FEED Workshop at KDD 2018, we showcase a system that makes progress towards the semantic analysis of code. By doing so, we provide the foundation for machines to truly reason about program code and learn from it. The work, also recently demonstrated at IJCAI 2018, is conceived and led by IBM Science for Social Good fellow Evan Patterson and focuses specifically on data science software. Data science programs are a special kind of computer code, often fairly short, but full of semantically rich content that specifies a sequence of data transformation, analysis, modeling, and interpretation operations.
Data Discovery Evolving Into Information Relationship Mapping Leveraging Machine Learning
What once started as early analysis of singular data sources has now evolved into far more robust ways of analyzing information and the relationships between different fields and information sources. Data discovery is another area where machine learning (ML) is beginning to make inroads. Twenty years ago, data discovery was a term used to define the early analytics needed to better understand data. For instance, Evoke Software was a company that analyzed large volumes of customer data. It both used metadata to understand field content to find trends and exceptions, and also looked at raw data and used algorithms to identify field boundaries in older or less documented data sources.
A look at the leading data discovery software and vendors
Turning data into business insight is the ultimate goal. It's not about gathering as much data as possible, it's about applying tools and making discoveries that help a business succeed. The data discovery software market includes a range of software and cloud-based services that can help organizations gain value from their constantly growing information resources. These products fall within the broad BI category, and at their most basic, they search for patterns within data and data sets. Many of these tools use visual presentation mechanisms, such as maps and models, to highlight patterns or specific items of relevance.
Request-and-Reverify: Hierarchical Hypothesis Testing for Concept Drift Detection with Expensive Labels
Yu, Shujian, Wang, Xiaoyang, Principe, Jose C.
One important assumption underlying common classification models is the stationarity of the data. However, in real-world streaming applications, the data concept indicated by the joint distribution of feature and label is not stationary but drifting over time. Concept drift detection aims to detect such drifts and adapt the model so as to mitigate any deterioration in the model's predictive performance. Unfortunately, most existing concept drift detection methods rely on a strong and over-optimistic condition that the true labels are available immediately for all already classified instances. In this paper, a novel Hierarchical Hypothesis Testing framework with Request-and-Reverify strategy is developed to detect concept drifts by requesting labels only when necessary. Two methods, namely Hierarchical Hypothesis Testing with Classification Uncertainty (HHT-CU) and Hierarchical Hypothesis Testing with Attribute-wise "Goodness-of-fit" (HHT-AG), are proposed respectively under the novel framework. In experiments with benchmark datasets, our methods demonstrate overwhelming advantages over state-of-the-art unsupervised drift detectors. More importantly, our methods even outperform DDM (the widely used supervised drift detector) when we use significantly fewer labels.
Faster data discovery and access - Forrester Names IBM a Leader in Machine Learning Data Catalogs - Watson
The promise of AI is that it will deliver digital transformation and improve productivity and efficiency across businesses. For many of our customers, IBM Watson has already helped deliver on this promise โ by enriching customer interactions, accelerating research and discovery, empowering employees, and mitigating risk. The next step for businesses is to make AI ubiquitous by operationalizing their workflows across the full AI lifecycle. IBM is committed to delivering these fundamental, end-to-end AI capabilities and giving enterprises everything they need. For example, consider the critical step of understanding and preparing data for productive and speedy use in analytical tools, machine learning and deep learning.
Astroinformatics - Wikipedia
Astroinformatics is primarily focused on developing the tools, methods, and applications of computational science, data science, and statistics for research and education in data-oriented astronomy.[1] Early efforts in this direction included data discovery, metadata standards development, data modeling, astronomical data dictionary development, data access, information retrieval,[3] data integration, and data mining[4] in the astronomical Virtual Observatory initiatives.[5][6][7] Further development of the field, along with astronomy community endorsement, was presented to the National Research Council (United States) in 2009 in the Astroinformatics "State of the Profession" Position Paper for the 2010 Astronomy and Astrophysics Decadal Survey.[8] That position paper provided the basis for the subsequent more detailed exposition of the field in the Informatics Journal paper Astroinformatics: Data-Oriented Astronomy Research and Education.[1] Astroinformatics as a distinct field of research was inspired by work in the fields of Bioinformatics and Geoinformatics, and through the eScience work[9] of Jim Gray (computer scientist) at Microsoft Research, whose legacy was remembered and continued through the Jim Gray eScience Awards.[10]
Errol Morris Refutes It Thus
The 18th-century Irish philosopher Bishop George Berkeley concluded that, since all we know of the universe is what our senses convey to us, things in the world exist only to the extent that we perceive them. They have no material reality, but are phenomena in and of our minds, or the mind of God. Samuel Johnson famously countered this philosophy by kicking a large stone and saying, "I refute it thus!" Two hundred years later, while American campuses roiled with protests against the Vietnam War, the philosopher, historian, and physicist Thomas Kuhn met with a grad student at Princeton's legendary Institute for Advanced Study to discuss the student's paper. The professor and student disagreed on some fundamental ideas, and the conversation grew heated.
Questioning Truth, Reality, and the Role of Scientific Progress
It's an interesting time to be making a case for philosophy in science. On the one hand, some scientists working on ideas such as string theory or the multiverse--ideas that reach far beyond our current means to test them--are forced to make a philosophical defense of research that can't rely on traditional hypothesis testing. On the other hand, some physicists, such as Richard Feynman and Stephen Hawking, were notoriously dismissive of the value of the philosophy of science. Original story reprinted with permission from Quanta Magazine, an editorially independent publication of the Simons Foundation whose mission is to enhance public understanding of science by covering research developments and trends in mathematics and the physical and life sciences. That value is asserted with gentle but firm assurance by Michela Massimi, the recent recipient of the Wilkins-Bernal-Medawar Medal, an award given annually by the UK's Royal Society. Massimi's prize speech, delivered earlier this week, defended both science and the philosophy of science from accusations of irrelevance.
How AI and NLP can broaden data discovery, accessibility and maintain governance. - ODBMS.org
The challenge of controlling and protecting data is a big one, but the bigger question is how to make people more productive with corporate information while maintaining standards of compliance and governance for broad access and use in the age of digital business. Applying AI and Natural Language Processing within the various stages of data analytics is a key way to democratize data and build in safeguards for broad use. Below is a Q&A with Ayush Parashar, a Co-Founder and Vice President of Engineering with Unifi Software. Often the quest for security can eclipse data usability. How can applying AI be used to both discover data and ensure information is not being seen or used by those that shouldn't have access to certain kinds of data?