AITopics

1810.11497

Genre: Research Report (0.64)

Industry: Media (0.34)

Technology:

Information Technology > Artificial Intelligence > Speech > Speech Recognition (1.00)
Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (1.00)

arXiv.org Artificial IntelligenceOct-25-2018

SyntaxSQLNet: Syntax Tree Networks for Complex and Cross-DomainText-to-SQL Task

Yu, Tao, Yasunaga, Michihiro, Yang, Kai, Zhang, Rui, Wang, Dongxu, Li, Zifan, Radev, Dragomir

Most existing studies in text-to-SQL tasks do not require generating complex SQL queries with multiple clauses or sub-queries, and generalizing to new, unseen databases. In this paper we propose SyntaxSQLNet, a syntax tree network to address the complex and cross-domain text-to-SQL generation task. SyntaxSQLNet employs a SQL specific syntax tree-based decoder with SQL generation path history and table-aware column attention encoders. We evaluate SyntaxSQLNet on the Spider text-to-SQL task, which contains databases with multiple tables and complex SQL queries with multiple SQL clauses and nested queries. We use a database split setting where databases in the test set are unseen during training. Experimental results show that SyntaxSQLNet can handle a significantly greater number of complex SQL examples than prior work, outperforming the previous state-of-the-art model by 7.3% in exact matching accuracy. We also show that SyntaxSQLNet can further improve the performance by an additional 7.5% using a cross-domain augmentation method, resulting in a 14.8% improvement in total. To our knowledge, we are the first to study this complex and cross-domain text-to-SQL task.

artificial intelligence, module, natural language, (17 more...)

1810.05237

Country: North America > United States (1.00)

Genre:

Research Report > New Finding (0.48)
Research Report > Promising Solution (0.34)

Technology: Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (1.00)

Sakhadeo, Archit, Srivastava, Nisheeth

Effective extractive summarization using frequency-filtered entity relationship graphs

arXiv.org Artificial IntelligenceOct-24-2018

Word frequency-based methods for extractive summarization are easy to implement and yield reasonable results across languages. However, they have significant limitations - they ignore the role of context, they offer uneven coverage of topics in a document, and sometimes are disjointed and hard to read. We use a simple premise from linguistic typology - that English sentences are complete descriptors of potential interactions between entities, usually in the order subject-verb-object - to address a subset of these difficulties. We have developed a hybrid model of extractive summarization that combines word-frequency based keyword identification with information from automatically generated entity relationship graphs to select sentences for summaries. Comparative evaluation with word-frequency and topic word-based methods shows that the proposed method is competitive by conventional ROUGE standards, and yields moderately more informative summaries on average, as assessed by a large panel (N 94) of human raters.

artificial intelligence, information, natural language, (20 more...)

1810.10419

Country:

Europe (0.68)
Asia > India (0.28)

Genre: Research Report (0.64)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (0.66)
Information Technology > Artificial Intelligence > Natural Language > Text Processing (0.46)

#artificialintelligenceOct-9-2018, 13:22:00 GMT

Banking and Investment Text Analytics Tool Amenity Analytics

The Investment Arm of a Financial Corporation used Amenity's API to create a dataset on the biggest players in the autonomous car market to support its strategy and business development efforts. Amenity extracted the most critical autonomous car industry news on a daily basis, identifying actionable patterns across the industry over time. The industry analysis model featured custom event types and modified versions of core taxonomies to best identify insights that are meaningful to the autonomous car industry. Amenity was able to acheive a high degree of accuracy using its proprietary NLP API including tokenization, lemmatization, named entity recognition (NER), dependency parsing and semantic role labeling. The result was a comprehensive industry analysis that provided a 360 degree view on the Autonomous Car industry financials, the supply chain, retail, OEM, suppliers, and technological trends.

artificial intelligence, natural language, text analytics tool amenity analytic, (1 more...)

#artificialintelligence

Industry: Automobiles & Trucks > Manufacturer (1.00)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Text Processing (1.00)
Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (0.68)

arXiv.org Artificial IntelligenceOct-8-2018

Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task

Yu, Tao, Zhang, Rui, Yang, Kai, Yasunaga, Michihiro, Wang, Dongxu, Li, Zifan, Ma, James, Li, Irene, Yao, Qingning, Roman, Shanelle, Zhang, Zilin, Radev, Dragomir

We present Spider, a large-scale, complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 college students. It consists of 10,181 questions and 5,693 unique complex SQL queries on 200 databases with multiple tables, covering 138 different domains. We define a new complex and cross-domain semantic parsing and text-to-SQL task where different complex SQL queries and databases appear in train and test sets. In this way, the task requires the model to generalize well to both new SQL queries and new database schemas. Spider is distinct from most of the previous semantic parsing tasks because they all use a single database and the exact same programs in the train set and the test set. We experiment with various state-of-the-art models and the best model achieves only 14.3% exact matching accuracy on a database split setting. This shows that Spider presents a strong challenge for future research. Our dataset and task are publicly available at https://yale-lily.github.io/spider

artificial intelligence, database, natural language, (15 more...)

1809.08887

Country:

Europe (0.93)
North America > United States (0.93)

Genre: Research Report > Promising Solution (0.48)

Industry: Education (1.00)

Technology: Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (1.00)

arXiv.org Artificial IntelligenceOct-1-2018

IncSQL: Training Incremental Text-to-SQL Parsers with Non-Deterministic Oracles

Shi, Tianze, Tatwawadi, Kedar, Chakrabarti, Kaushik, Mao, Yi, Polozov, Oleksandr, Chen, Weizhu

We present a sequence-to-action parsing approach for the natural language to SQL task that incrementally fills the slots of a SQL query with feasible actions from a pre-defined inventory. To account for the fact that typically there are multiple correct SQL queries with the same or very similar semantics, we draw inspiration from syntactic parsing techniques and propose to train our sequence-to-action models with non-deterministic oracles. We evaluate our models on the WikiSQL dataset and achieve an execution accuracy of 83.7% on the test set, a 2.1% absolute improvement over the models trained with traditional static oracles assuming a single correct target SQL query. When further combined with the execution-guided decoding strategy, our model sets a new state-of-the-art performance at an execution accuracy of 87.1%.

artificial intelligence, machine learning, natural language, (19 more...)

1809.05054

Country: North America (0.28)

Genre: Research Report (0.50)

Industry: Health & Medicine (0.46)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (0.74)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.72)
Information Technology > Artificial Intelligence > Natural Language > Machine Translation (0.69)

Kasewa, Sudhanshu, Stenetorp, Pontus, Riedel, Sebastian

Wronging a Right: Generating Better Errors to Improve Grammatical Error Detection

arXiv.org Machine LearningSep-26-2018

Grammatical error correction, like other machine learning tasks, greatly benefits from large quantities of high quality training data, which is typically expensive to produce. While writing a program to automatically generate realistic grammatical errors would be difficult, one could learn the distribution of naturallyoccurring errors and attempt to introduce them into other datasets. Initial work on inducing errors in this way using statistical machine translation has shown promise; we investigate cheaply constructing synthetic samples, given a small corpus of human-annotated data, using an off-the-rack attentive sequence-to-sequence model and a straight-forward post-processing procedure. Our approach yields error-filled artificial data that helps a vanilla bi-directional LSTM to outperform the previous state of the art at grammatical error detection, and a previously introduced model to gain further improvements of over 5% $F_{0.5}$ score. When attempting to determine if a given sentence is synthetic, a human annotator at best achieves 39.39 $F_1$ score, indicating that our model generates mostly human-like instances.

artificial intelligence, machine learning, natural language, (15 more...)

1810.00668

Country:

Europe (1.00)
North America > United States > Texas (0.14)
North America > United States > Maryland (0.14)

Genre: Research Report (0.40)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Machine Translation (1.00)
Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (1.00)

Bacciu, Davide, Bruno, Antonio

Text Summarization as Tree Transduction by Top-Down TreeLSTM

arXiv.org Machine LearningSep-24-2018

Extractive compression is a challenging natural language processing problem. This work contributes by formulating neural extractive compression as a parse tree transduction problem, rather than a sequence transduction task. Motivated by this, we introduce a deep neural model for learning structure-to-substructure tree transductions by extending the standard Long Short-Term Memory, considering the parent-child relationships in the structural recursion. The proposed model can achieve state of the art performance on sentence compression benchmarks, both in terms of accuracy and compression rate.

artificial intelligence, machine learning, natural language, (20 more...)

1809.09096

Country: North America > United States (0.94)

Genre: Research Report (0.64)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.93)
Information Technology > Artificial Intelligence > Machine Learning > Learning Graphical Models > Undirected Networks > Markov Models (0.46)

Korchev, Dmitriy, Jammalamadaka, Aruna, Bhattacharyya, Rajan

Automatic Rule Learning for Autonomous Driving Using Semantic Memory

arXiv.org Machine LearningSep-24-2018

Abstract-- This paper presents a novel approach for automatic rule learning applicable to an autonomous driving system using real driving data. We represent the actions of other agents (provided by sensors) in the scene via temporal sequences called "episodes". The proposed method adaptively creates new rules automatically by extracting and segmenting valuable information about other agents and their interactions. These rules, which take the form of a "spatiotemporal grammar" or "episodic memory" are stored in a "semantic memory" module for later use. During the testing phase, the system segments constantly changing situations, finds the corresponding parse tree for the current state of the self-car and other agents, and applies the rules stored in semantic memory to stop, yield, continue driving, etc. The method also allows for continues online training during agent driving. Unlike traditional deep driving and machine learning methods that require significant amount of training data to achieve desired quality, the proposed method demonstrates good results with just a few training examples.

machine learning, natural language, time step, (19 more...)

1809.07904

Genre: Research Report (0.70)

Industry:

Health & Medicine > Consumer Health (1.00)
Transportation > Ground > Road (0.86)
Education > Educational Setting > Online (0.54)
Education > Educational Technology > Educational Software > Computer Based Training (0.34)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (1.00)
Information Technology > Artificial Intelligence > Machine Learning (1.00)

Patnaikuni, Shrinivasan R Patnaik, Gengaje, Dr. Sachin R

Syntactico-Semantic Reasoning using PCFG, MEBN, and PR-OWL

arXiv.org Artificial IntelligenceSep-20-2018

Probabilistic context free grammars (PCFG) have been the core of the probabilistic reasoning based parsers for several years especially in the context of the NLP. Multi entity bayesian networks (MEBN) a First Order Logic probabilistic reasoning methodology and is widely adopted and used method for uncertainty reasoning. Further upper ontology like Probabilistic Ontology Web Language (PR-OWL) built using MEBN takes care of probabilistic ontologies which model and capture the uncertainties inherent in the domain's semantic information. The paper attempts to establish a link between probabilistic reasoning in PCFG and MEBN by proposing a formal description of PCFG driven by MEBN leading to usage of PR-OWL modeled ontologies in PCFG parsers.

machine learning, mebn, natural language, (18 more...)

1809.07607

Genre: Research Report (0.40)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Grammars & Parsing (1.00)
Information Technology > Artificial Intelligence > Representation & Reasoning > Uncertainty > Bayesian Inference (0.39)
Information Technology > Artificial Intelligence > Machine Learning > Learning Graphical Models > Directed Networks > Bayesian Learning (0.39)