Goto

Collaborating Authors

 smile



Representation of Molecules via Algebraic Data Types : Advancing Beyond SMILES & SELFIES

arXiv.org Artificial Intelligence

We introduce a novel molecular representation through Algebraic Data Types (ADTs) - composite data structures formed through the combination of simpler types that obey algebraic laws. By explicitly considering how the datatype of a representation constrains the operations which may be performed, we ensure meaningful inference can be performed over generative models (programs with sample} and score operations). This stands in contrast to string-based representations where string-type operations may only indirectly correspond to chemical and physical molecular properties, and at worst produce nonsensical output. The ADT presented implements the Dietz representation for molecular constitution via multigraphs and bonding systems, and uses atomic coordinate data to represent 3D information and stereochemical features. This creates a general digital molecular representation which surpasses the limitations of the string-based representations and the 2D-graph based models on which they are based. In addition, we present novel support for quantum information through representation of shells, subshells, and orbitals, greatly expanding the representational scope beyond current approaches, for instance in Molecular Orbital theory. The framework's capabilities are demonstrated through key applications: Bayesian probabilistic programming is demonstrated through integration with LazyPPL, a lazy probabilistic programming library; molecules are made instances of a group under rotation, necessary for geometric learning techniques which exploit the invariance of molecular properties under different representations; and the framework's flexibility is demonstrated through an extension to model chemical reactions. After critiquing previous representations, we provide an open-source solution in Haskell - a type-safe, purely functional programming language.


'Doctor Who' is about to ruin emoji for you forever

Mashable

Doctor Who, like the best horror films, has always been at its strongest when it performs this one weird trick. The long-running BBC sci-fi show thrives on taking a single, simple, harmless, ordinary thing you see in everyday life (stone statues, showroom dummies), then tweaking it into a creature so terrifying (the Weeping Angels, the Autons) you will never in your life see that object the same way again. SEE ALSO: 'Doctor Who' fans will be thrilled and delighted by the latest Oxford Dictionary addition This Saturday, in the new Season 10 episode called "Smile," Doctor Who is about to perform that same trick with that most simple, harmless, ordinary facet of 21st century life: emoji. If you watch it -- and you should, because it's among the best episodes of new Who -- the smiley face at the end of your texts will take on sinister new meaning. SEE ALSO: About time: 'Doctor Who' to feature first openly gay TARDIS resident If you're an old-school Whovian like me, you probably rolled your eyes at that point in the Season 10 trailer when new companion Bill Potts (Pearl Mackie) encounters a robot with a flatscreen face composed of a smiley and a thumbs-up.


SMILe: Shuffled Multiple-Instance Learning

AAAI Conferences

Resampling techniques such as bagging are often used in supervised learning to produce more accurate classifiers. In this work, we show that multiple-instance learning admits a different form of resampling, which we call "shuffling." In shuffling, we resample instances in such a way that the resulting bags are likely to be correctly labeled. We show that resampling results in both a reduction of bag label noise and a propagation of additional informative constraints to a multiple-instance classifier. We empirically evaluate shuffling in the context of multiple-instance classification and multiple-instance active learning and show that the approach leads to significant improvements in accuracy.


SMILE: An Informality Classification Tool for Helping to Assess Quality and Credibility in Web 2.0 Texts

AAAI Conferences

The data made available by Web 2.0 applications such as social networks, on-line chats or blogs have give access to multiples sources of information. Due to this dramatic increase in available information, the perception of quality and credibility plays an important role in social media, thus making necessary to discard low quality and uninteresting content. Moreover, the informal features of Web 2.0 texts such as emoticons, typos, slang or loss of formatting impact negatively on user perception regarding content quality and credibility. For this reason, this paper proposes the SMILE system, a novel unsupervised real-time tool for assessing user-generated content quality and credibility using informality levels. As a test case, we focus on Yahoo! Answers, a relevant Web 2.0 application by its amount of users, content and textual diversity. The results of our study show that informality analysis can be used as criteria to help assess the credibility and quality of Web 2.0 information sources.