Machine Translation
On the accuracy of self-normalized log-linear models
Andreas, Jacob, Rabinovich, Maxim, Klein, Dan, Jordan, Michael I.
Calculation of the log-normalizer is a major computational obstacle in applications of log-linear models with large output spaces. The problem of fast normalizer computation has therefore attracted significant attention in the theoretical and applied machine learning literature. In this paper, we analyze a recently proposed technique known as "self-normalization", which introduces a regularization term in training to penalize log normalizers for deviating from zero. This makes it possible to use unnormalized model scores as approximate probabilities. Empirical evidence suggests that self-normalization is extremely effective, but a theoretical understanding of why it should work, and how generally it can be applied, is largely lacking. We prove generalization bounds on the estimated variance of normalizers and upper bounds on the loss in accuracy due to self-normalization, describe classes of input distributions that self-normalize easily, and construct explicit examples of high-variance input distributions. Our theoretical results make predictions about the difficulty of fitting self-normalized models to several classes of distributions, and we conclude with empirical validation of these predictions.
Exploring Key Concept Paraphrasing Based on Pivot Language Translation for Question Retrieval
Zhang, Wei-Nan (Harbin Institute of Technology) | Ming, Zhao-Yan (Digipen Institute of Technology) | Zhang, Yu (Harbin Institute of Technology) | Liu, Ting (Harbin Institute of Technology) | Chua, Tat-Seng (National University of Singapore)
Question retrieval in current community-based question answering (CQA) services does not, in general, work well for long and complex queries. One of the main difficulties lies in the word mismatch between queries and candidate questions. Existing solutions try to expand the queries at word level, but they usually fail to consider concept level enrichment. In this paper, we explore a pivot language translation based approach to derive the paraphrases of key concepts. We further propose a unified question retrieval model which integrates the keyconcepts and their paraphrases for the query question. Experimental results demonstrate that the paraphrase enhanced retrieval model significantly outperforms the state-of-the-art models in question retrieval.
Use of Modality and Negation in Semantically-Informed Syntactic MT
Baker, Kathryn, Bloodgood, Michael, Dorr, Bonnie J., Callison-Burch, Chris, Filardo, Nathaniel W., Piatko, Christine, Levin, Lori, Miller, Scott
This paper describes the resource- and system-building efforts of an eight-week Johns Hopkins University Human Language Technology Center of Excellence Summer Camp for Applied Language Exploration (SCALE-2009) on Semantically-Informed Machine Translation (SIMT). We describe a new modality/negation (MN) annotation scheme, the creation of a (publicly available) MN lexicon, and two automated MN taggers that we built using the annotation scheme and lexicon. Our annotation scheme isolates three components of modality and negation: a trigger (a word that conveys modality or negation), a target (an action associated with modality or negation) and a holder (an experiencer of modality). We describe how our MN lexicon was semi-automatically produced and we demonstrate that a structure-based MN tagger results in precision around 86% (depending on genre) for tagging of a standard LDC data set. We apply our MN annotation scheme to statistical machine translation using a syntactic framework that supports the inclusion of semantic annotations. Syntactic tags enriched with semantic annotations are assigned to parse trees in the target-language training texts through a process of tree grafting. While the focus of our work is modality and negation, the tree grafting procedure is general and supports other types of semantic information. We exploit this capability by including named entities, produced by a pre-existing tagger, in addition to the MN elements produced by the taggers described in this paper. The resulting system significantly outperformed a linguistically naive baseline model (Hiero), and reached the highest scores yet reported on the NIST 2009 Urdu-English test set. This finding supports the hypothesis that both syntactic and semantic information can improve translation quality.
SESSION 4B PAPER 2 THE MECHANIZATION OF LITERATURE SEARCHING
I am quite ready to subscribe to the already mentioned slogan that "whatever a human being can do,an appropriate machine can do, too"; but I do this only because.I regard the slogan as utterly trivial. At the moment, I am not talking about what maohines could do in principle but only about what actually existing or blueprinted machines could do, and it Is with regard to these that I utter my definite opinions. If someone wishes to write sciencefiction about information-processing centres of the (undetermined) future, let him do so and I shall discuss it with him over a glass of beer and even offer some startling suggestions of my own. If he is interested in improving the literature search process today, I would strongly advise him to forget about mechanizing abstracting or indexing. May I add that it is with a good amount of sorrow that I have come to this conclusion which is quite counter, to my temperament and my convictions (never published) of a few years ago.
SESSION 2 PAPER 5 TIGRIS AND EUPHRATES - A COMPARISON BETWEEN HUMAN AND MACHINE TRANSLATION
An unsophisticated translation of such a sentence will therefore not be a good translation. Again, contrary to Mr. Richensi opinion, I believe that the problem involved is serious. There is no simple procedure to find out which, and in what way, the words of the English language are context-dependent. And I don't think that the issue can be belittled for tae reason that contextdependent words do not occur in scientific discussions and writings. They might not be too abundant in ordinary scientific papers on matters physical or chemical, but there would surely be plenty of them in discussions of matters linguistic, for instance. This might be one reason why so far hardly anybody has tried to machine translate papers in linguistics. As soon as this is attempted, the seriousness of the problem will become immediately evident.
Stanford Heuristic Programming Project July 1979 Memo HPP-79-21 Computer Science Department Report No. STAN-CS-79-754
Theorem Proving Vision Robotics Information Processing Psychology Learning and Inductive Inference Planning and Related Problem-solving Techniques A. Natural Language Processing Ovnrview The most common way that human beings communicate Is by speaking or writing In one of the "natural" languages, like English, French, or Chinese. Computer programming languages, on the other hand, seem awkward to humans. These "artificial" languages are designed to have a rigid format, or syntax, so that a computer program reading and compiling code written In an artificial language can understand what the programmer means. In addition to being structurally simpler than natural languages, the artificial languages can express easily only those concepts that are important In programming: "Do this then do that," "See it such and such Is true," etc. The things that can be expressed In a language are referred to as the semantics of the language. The research on understanding natural language described in this section of the Handbook is concerned with programs that deal with the full range of meaning of languages like English.
An Autoencoder Approach to Learning Bilingual Word Representations
P, Sarath Chandar A, Lauly, Stanislas, Larochelle, Hugo, Khapra, Mitesh, Ravindran, Balaraman, Raykar, Vikas C., Saha, Amrita
Cross-language learning allows us to use training data from one language to build models for a different language. Many approaches to bilingual learning require that we have word-level alignment of sentences from parallel corpora. In this work we explore the use of autoencoder-based methods for cross-language learning of vectorial word representations that are aligned between two languages, while not relying on word-level alignments. We show that by simply learning to reconstruct the bag-of-words representations of aligned sentences, within and between languages, we can in fact learn high-quality representations and do without word alignments. We empirically investigate the success of our approach on the problem of cross-language text classification, where a classifier trained on a given language (e.g., English) must learn to generalize to a different language (e.g., German). In experiments on 3 language pairs, we show that our approach achieves state-of-the-art performance, outperforming a method exploiting word alignments and a strong machine translation baseline.
Bucking the Trend: Large-Scale Cost-Focused Active Learning for Statistical Machine Translation
Bloodgood, Michael, Callison-Burch, Chris
We explore how to improve machine translation systems by adding more translation data in situations where we already have substantial resources. The main challenge is how to buck the trend of diminishing returns that is commonly encountered. We present an active learning-style data solicitation algorithm to meet this challenge. We test it, gathering annotations via Amazon Mechanical Turk, and find that we get an order of magnitude increase in performance rates of improvement.