Country
Representation of Protein-Sequence Information by Amino Acid Subalphabets
Andersen, Claus A. F., Brunak, Soren
Within computational biology, algorithms are constructed with the aim of extracting knowledge from biological data, in particular, data generated by the large genome projects, where gene and protein sequences are produced in high volume. In this article, we explore new ways of representing protein-sequence information, using machine learning strategies, where the primary goal is the discovery of novel powerful representations for use in AI techniques. In the case of proteins and the 20 different amino acids they typically contain, it is also a secondary goal to discover how the current selection of amino acids -- which now are common in proteins -- might have emerged from simpler selections, or alphabets, in use earlier during the evolution of living organisms.
Calendar of Events
NASA Ames Research Center Polish Academy of Sciences URL: www.taai.org.tw/announce/ (PRICAI 2004). (ICKEDS 2004). This book looks at some of the results of the synergy among AI, cognitive science, and education. Examples include virtual students whose misconceptions force students to reflect on their own knowledge, intelligent tutoring systems, and speech recognition technology that helps students learn to read.
AI in the News
"Over the last decade, This eclectic keepsake provides a sampling to design vision and navigation some jobs have vanished, others are fading of what can be found (with links to the full systems based on the honeybee.... But lots of new ones have appeared. Please key lies in understanding how insects With the help of Toronto Star keep in mind that (1) the mere mention of perceive their world…." "From the Luddites the articles were initially available and thinking.... * Bio-informatician: Not Such collection--updated, hyperlinked, and molecular biology and computer science. The robot scientist developed (www.dailynewstribune.com). TIME online edition from the point of view of the human researcher, has participated in various fund-raising (www.time.com). "President Bush will announce it does so as effectively as a person.... events -- including selling Krispy Kreme later this week a plan to resume One question is, if their robot does doughnuts on Moody Street the day after missions to the Moon and send humans make an important discovery, will it be Thanksgiving and raffling off a new DVD to Mars within 20 years, with international eligible to win a Nobel prize?" player donated by Watch City Appliance. Such a plan is likely ... Students created two robots from to be tremendously expensive, and some January 15: A New Robot Makes a Leap scratch that were not manipulated by remote argue that manned space missions are unnecessary in Brainpower. Philadelphia control, but, rather, programmed to with the level of sophistication Inquirer (www.philly.com). "A new robot is compete in the Botball competition.
AI and Bioinformatics
Glasgow, Janice, Jurisica, Igor, Rost, Burkhard
Undoubtedly, bioinformatics is Michael Waddell, David Page, and Jude a truly interdisciplinary field: Although some Shavlik ("Using Machine Learning to Design researchers continuously affect wet labs in life and Interpret Gene-Expression Microarrays") science through collaborations or provision of introduces some background information and tools, others are rooted in the theory departments provides a comprehensive description of how of exact sciences (physics, chemistry, or techniques from machine learning can be used engineering) or computer sciences. This wide to help understand this high-dimensional and variety creates many different perspectives and prolific gene-expression data.
Annotating Protein Function through Lexical Analysis
We now know the full genomes of more than 60 organisms. The experimental characterization of the newly sequenced proteins is deemed to lack behind this explosion of naked sequences (sequencefunction gap). The rate at which expert annotators add the experimental information into more or less controlled vocabularies of databases snails along at an even slower pace. Most methods that annotate protein function exploit sequence similarity by transferring experimental information for homologues. A crucial development aiding such transfer is large-scale, work- and management-intensive projects aimed at developing a comprehensive ontology for gene-protein function, such as the Gene Ontology project. In parallel, fully automatic or semiautomatic methods have successfully begun to mine the existing data through lexical analysis. Some of these tools target parsing controlled vocabulary from databases; others venture at mining free texts from MEDLINE abstracts or full scientific papers. Automated text analysis has become a rapidly expanding discipline in bioinformatics. A few of these tools have already been embedded in research projects.
The Semantic Web and Language Technology, Its Potential and Practicalities: EUROLAN-2003
Cristea, Dan, Ide, Nancy, Tufis, Dan
Later in the school, the focus turned to ontologies, which is where the true power of the semantic web lies. EUROLAN lecturers treated its potential in terms of what the topic of ontology development it might--and might not--bring to us in the future. This year's and how great its impact will really start somewhere, somehow, even if school was organized by the Faculty be. Although it is not yet clear what emerges is a variety of ontological of Computer Science at the A. I. Cuza whether the current vision of the semantic stores from which to choose. University of Iasi, the Research Institute web will indeed reach its expectations, The EUROLAN summer school also for Artificial Intelligence at the there are more and more included a workshop on ontologies Romanian Academy in Bucharest, opinions that it represents a major and information extraction, a student and the Department of Computer technological step that will permanently workshop on applied natural Science at Vassar College.
Semantic Integration Workshop at the Second International Semantic Web Conference (ISWC-2003)
Doan, AnHai, Halevy, Alon Y., Noy, Natalya F.
In numerous distributed environments, including today's World Wide Web, enterprise data management systems, large science projects, and the emerging semantic web, applications will inevitably use the information described by multiple ontologies and schemas. We organized the Workshop on Semantic Integration at the Second International Semantic Web Conference to bring together different communities working on the issues of enabling integration among different resources. The workshop generated a lot of interest and attracted more than 70 participants.
Distribution of Mutual Information from Complete and Incomplete Data
Hutter, Marcus, Zaffalon, Marco
Mutual information is widely used, in a descriptive way, to measure the stochastic dependence of categorical random variables. In order to address questions such as the reliability of the descriptive value, one must consider sample-to-population inferential approaches. This paper deals with the posterior distribution of mutual information, as obtained in a Bayesian framework by a second-order Dirichlet prior distribution. The exact analytical expression for the mean, and analytical approximations for the variance, skewness and kurtosis are derived. These approximations have a guaranteed accuracy level of the order O(1/n^3), where n is the sample size. Leading order approximations for the mean and the variance are derived in the case of incomplete samples. The derived analytical expressions allow the distribution of mutual information to be approximated reliably and quickly. In fact, the derived expressions can be computed with the same order of complexity needed for descriptive mutual information. This makes the distribution of mutual information become a concrete alternative to descriptive mutual information in many applications which would benefit from moving to the inductive side. Some of these prospective applications are discussed, and one of them, namely feature selection, is shown to perform significantly better when inductive mutual information is used.
A Personalized System for Conversational Recommendations
Thompson, C. A., Goker, M. H., Langley, P.
Searching for and making decisions about information is becoming increasingly difficult as the amount of information and number of choices increases. Recommendation systems help users find items of interest of a particular type, such as movies or restaurants, but are still somewhat awkward to use. Our solution is to take advantage of the complementary strengths of personalized recommendation systems and dialogue systems, creating personalized aides. We present a system -- the Adaptive Place Advisor -- that treats item selection as an interactive, conversational process, with the program inquiring about item attributes and the user responding. Individual, long-term user preferences are unobtrusively obtained in the course of normal recommendation dialogues and used to direct future conversations with the same user. We present a novel user model that influences both item search and the questions asked during a conversation. We demonstrate the effectiveness of our system in significantly reducing the time and number of interactions required to find a satisfactory item, as compared to a control group of users interacting with a non-adaptive version of the system.
Representation Dependence in Probabilistic Inference
Non-deductive reasoning systems are often representation dependent: representing the same situation in two different ways may cause such a system to return two different answers. Some have viewed this as a significant problem. For example, the principle of maximum entropyhas been subjected to much criticism due to its representation dependence. There has, however, been almost no work investigating representation dependence. In this paper, we formalize this notion and show that it is not a problem specific to maximum entropy. In fact, we show that any representation-independent probabilistic inference procedure that ignores irrelevant information is essentially entailment, in a precise sense. Moreover, we show that representation independence is incompatible with even a weak default assumption of independence. We then show that invariance under a restricted class of representation changes can form a reasonable compromise between representation independence and other desiderata, and provide a construction of a family of inference procedures that provides such restricted representation independence, using relative entropy.