Generating Factoid Questions With Recurrent Neural Networks: The 30M Factoid Question-Answer Corpus
Published 1 January 2016Open access
Iulian Vlad Serban, Alberto García-Durán, Çağlar Gülçehre, Sungjin Ahn, Sarath Chandar, Aaron Courville
Citations39
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The 30M Factoid Question-Answer Corpus is presented, an enormous question answer pair corpus produced by applying a novel neural network architecture on the knowledge base Freebase to transduce facts into natural language questions.
Abstract
Iulian Vlad Serban, Alberto García-Durán, Caglar Gulcehre, Sungjin Ahn, Sarath Chandar, Aaron Courville, Yoshua Bengio. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Keywords
Computer Science
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
50,318 Citations2021Mandi, Jayanta, Canoy, Rocsildes +2 more
A simple numeric simulation of DNA-co-polymerized hydrogel shape change and a genetic algorithm that generates and selects large batches of material designs that compete with one another to evolve and converge on optimal objective-matching designs are constructed.
arXiv (Cornell University)Distributed Representations of Words and Phrases and their Compositionality
18,086 Citations2013Tomáš Mikolov, Ilya Sutskever +3 more
This paper presents a simple method for finding phrases in text, and shows that learning good vector representations for millions of phrases is possible and describes a simple alternative to the hierarchical softmax called negative sampling.
arXiv (Cornell University)Sequence to Sequence Learning with Neural Networks
13,362 Citations2014Ilya Sutskever, Oriol Vinyals +1 more
IEEE Transactions on Neural Networks and Learning SystemsLSTM: A Search Space Odyssey
6,837 Citations2016Klaus Greff, Rupesh K. Srivastava +3 more
This paper presents the first large-scale analysis of eight LSTM variants on three representative tasks: speech recognition, handwriting recognition, and polyphonic music modeling, and observes that the studied hyperparameters are virtually independent and derive guidelines for their efficient adjustment.
Translating embeddings for modeling multi-relational data
5,178 Citations2015Antoine Bordes, Nicolas Usunier +2 more
TransE is proposed, a method which models relationships by interpreting them as translations operating on the low-dimensional embeddings of the entities, which proves to be powerful since extensive experiments show that TransE significantly outperforms state-of-the-art methods in link prediction on two knowledge bases.
Freebase
4,892 Citations2008Kurt Bollacker, Colin Evans +3 more
MQL provides an easy-to-use object-oriented interface to the tuple data in Freebase and is designed to facilitate the creation of collaborative, Web-based data-oriented applications.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments
3,723 Citations2005Satanjeev Banerjee, Alon Lavie
METEOR is described, an automatic metric for machine translation evaluation that is based on a generalized concept of unigram matching between the machineproduced translation and human-produced reference translations and can be easily extended to include more advanced matching strategies.
arXiv (Cornell University)Microsoft COCO Captions: Data Collection and Evaluation Server
1,627 Citations2015Xinlei Chen, Hao Fang +5 more
The Microsoft COCO Caption dataset and evaluation server are described and several popular metrics, including BLEU, METEOR, ROUGE and CIDEr are used to score candidate captions.
Building Large Knowledge-Based Systems: Representation and Inference in the Cyc Project
1,281 Citations1990Douglas B. Lenat, Ramanathan V. Guha
This review has been difficult for me to write, because my thoughts about Cyc have changed a great deal since I first read the book in the spring of 1990, and my thoughts about the project today have changed a great deal.
Learning Structured Embeddings of Knowledge Bases
861 Citations2011Antoine Bordes, Jason Weston +2 more
Addressing the Rare Word Problem in Neural Machine Translation
660 Citations2015Thang Luong, Ilya Sutskever +3 more
This paper proposes and implements an effective technique to address the problem of end-to-end neural machine translation's inability to correctly translate very rare words, and is the first to surpass the best result achieved on a WMT’14 contest task.
arXiv (Cornell University)Large-scale Simple Question Answering with Memory Networks
562 Citations2015Antoine Bordes, Nicolas Usunier +2 more
This paper studies the impact of multitask and transfer learning for simple question answering; a setting for which the reasoning required to answer is quite easy, as long as one can retrieve the correct evidence given a question, which can be difficult in large-scale conditions.
Semantic Parsing via Paraphrasing
535 Citations2014Jonathan Berant, Percy Liang
This paper presents two simple paraphrase models, an association model and a vector space model, and trains them jointly from question-answer pairs, improving state-of-the-art accuracies on two recently released question-answering datasets.
Lecture notes in computer scienceOpen Question Answering with Weakly Supervised Embedding Models
301 Citations2014Antoine Bordes, Jason Weston +1 more
This paper empirically demonstrate that the model can capture meaningful signals from its noisy supervision leading to major improvements over paralex, the only existing method able to be trained on similar weakly labeled data.
Web question answering
289 Citations2002Susan Dumais, Michele Banko +3 more
This paper describes a question answering system that is designed to capitalize on the tremendous amount of data that is now available online, and uses the redundancy available in large corpora as an important resource to simplify the query rewrites and support answer mining from returned snippets.
arXiv (Cornell University)Practical recommendations for gradient-based training of deep architectures
267 Citations2012Yoshua Bengio
A Comparison of Greedy and Optimal Assessment of Natural Language Student Input Using Word-to-Word Similarity Metrics
122 Citations2012Vasile Rus, Mihai Lintean
A novel, optimal semantic similarity approach based on word-to-word similarity metrics to solve the important task of assessing natural language student input in dialogue-based intelligent tutoring systems.
Question Generation from Paragraphs at UPenn: QGSTEC System Description
67 Citations2010Prashanth Mannem, Rashmi Prasad +1 more
The question generation system uses predicate argument structures of sentences along with semantic roles for the question generation task from paragraphs to identify relevant parts of text before forming questions over them.
Dialogue & DiscourseQuestion Generation from Concept Maps
65 Citations2012Andrew M. Olney, Arthur C. Graesser +1 more
The purpose of the study is to generate and evaluate pedagogically-appropriate questions at varying levels of specificity across one or more sentences to understand the pedagogical nature of questions in tutoring.
Edinburgh Research Explorer (University of Edinburgh)Generating Natural Language from Linked Data: Unsupervised template extraction
60 Citations2013Daniel Duma, Ewan Klein
An architecture for generating natural language from Linked Data that automatically learns sentence templates and statistical document planning from parallel RDF datasets and text and significantly outperforms the baseline on two of three measures: non-redundancy and structure and coherence.
Question Generation with Minimal Recursion Semantics
45 Citations2010Xuchen Yao
The performance of proposed method is compared against other syntax and rule based systems, and the result reveals the challenges of current research on question generation and indicates direction for future work.
Computers & Mathematics with ApplicationsQUEST: A model of question answering
41 Citations1992Arthur C. Graesser, Sallie E. Gordon +1 more
How knowledge is represented by QUEST's conceptual graph structures and how the procedural mechanisms operate on the knowledge structures during question atmwering are described.
Dialogue & DiscourseQuestion Generation based on Lexico-Syntactic Patterns Learned from the Web
26 Citations2012Sérgio Curto, Ana Cristina Mendes +1 more
The question generation task as performed by T HE -M ENTOR is detailed and several filters are applied in order to discard low quality items.
