A Dataset of Syntactic-Ngrams over Time from a Very Large Corpus of English Books
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A dataset of syntactic-ngrams (counted dependency-tree fragments) based on a corpus of 3.5 million English books includes temporal information, facilitating new kinds of research into lexical semantics over time.
Abstract
We created a dataset of syntactic-ngrams (counted dependency-tree fragments) based on a corpus of 3.5 million English books. The dataset includes over 10 billion distinct items covering a wide range of syntactic configurations. It also includes temporal information, facilitating new kinds of research into lexical semantics over time. This paper describes the dataset, the syntactic representation, and the kinds of information provided. 1
