What is disputed on the web?
Published 27 April 2010
Rob Ennals, Dan Byler, John Mark Agosta, Barbara Rosario
Citations48
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A method for automatically acquiring a corpus of disputed claims from the web by searching the web for patterns such as "falsely claimed that X" and then using a statistical classifier to select text that appears to be making a disputed claim.
Abstract
We present a method for automatically acquiring of a corpus of disputed claims from the web. We consider a factual claim to be disputed if a page on the web suggests both that the claim is false and also that other people say it is true.
Keywords
Computer ScienceSocial Sciences
Mining and summarizing customer reviews
7,714 Citations2004Minqing Hu, Bing Liu
This research aims to mine and to summarize all the customer reviews of a product, and proposes several novel techniques to perform these tasks.
Advanced textbooks in control and signal processingPattern Classification
5,015 Citations2001Leo H. Chiang, Evan L. Russell +1 more
CERN Document Server (European Organization for Nuclear Research)Natural Language Processing with Python
3,449 Citations2009Steven Bird, Ewan Klein +1 more
A sentimental education
3,343 Citations2004Bo Pang, Lillian Lee
A novel machine-learning method is proposed that applies text-categorization techniques to just the subjective portions of the document, which greatly facilitates incorporation of cross-sentence contextual constraints.
Automatic acquisition of hyponyms from large text corpora
3,283 Citations1992Marti A. Hearst
A set of lexico-syntactic patterns that are easily recognizable, that occur frequently and across text genre boundaries, and that indisputably indicate the lexical relation of interest are identified.
A comparison of event models for naive bayes text classification
3,224 Citations1998Andrew McCallum, Kamal Nigam
It is found that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes--providing on average a 27% reduction in error over the multi -variateBernoulli model at any vocabulary size.
Lingvisticae InvestigationesA survey of named entity recognition and classification
2,488 Citations2007David R. Nadeau, Satoshi Sekine
Observations about languages, named entity types, domains and textual genres studied in the literature, along with other critical aspects of NERC such as features and evaluation methods, are reported.
Pattern Classification
2,474 Citations2001Duda, Richard O, Hart, Peter E +1 more
Lecture notes in computer scienceThe PASCAL Recognising Textual Entailment Challenge
1,626 Citations2006Ido Dagan, Oren Glickman +1 more
Meme-tracking and the dynamics of the news cycle
1,547 Citations2009Jure Leskovec, Lars Bäckström +1 more
This work develops a framework for tracking short, distinctive phrases that travel relatively intact through on-line text; developing scalable algorithms for clustering textual variants of such phrases, and identifies a broad class of memes that exhibit wide spread and rich variation on a daily basis.
A Bayesian Approach to Filtering Junk E-Mail
1,153 Citations1998Mehran Sahami, Susan Dumais +2 more
This work examines methods for the automated construction of filters to eliminate such unwanted messages from a user’s mail stream, and shows the efficacy of such filters in a real world usage scenario, arguing that this technology is mature enough for deployment.
Communications of the ACMOpen information extraction from the web
1,014 Citations2008Oren Etzioni, Michele Banko +2 more
Open IE (OIE), a new extraction paradigm where the system makes a single data-driven pass over its corpus and extracts a large set of relational tuples without requiring any human input, is introduced.
Wikify!
921 Citations2007Rada Mihalcea, Andras Csomai
This paper introduces the use of Wikipedia as a resource for automatic keyword extraction and word sense disambiguation, and shows how this online encyclopedia can be used to achieve state-of-the-art results on both these tasks.
Neural Information Processing SystemsLearning Syntactic Patterns for Automatic Hypernym Discovery
676 Citations2004Rion Snow, Daniel Jurafsky +1 more
This paper presents a new algorithm for automatically learning hypernym (is-a) relations from text, using "dependency path" features extracted from parse trees and introduces a general-purpose formalization and generalization of these patterns.
How do users evaluate the credibility of Web sites?
555 Citations2003B. J. Fogg, Cathy Soohoo +4 more
Comments in the top 18 areas that people noticed when evaluating Web site credibility are shared, and reasons for the prominence of design look are discussed.
Automatically constructing a dictionary for information extraction tasks
453 Citations1993Ellen Riloff
Using AutoSlog, a system that automatically builds a domain-specific dictionary of concepts for extracting information from text, a dictionary for the domain of terrorist event descriptions was constructed in only 5 person-hours and the overall scores were virtually indistinguishable.
Automatic construction of a hypernym-labeled noun hierarchy from text
373 Citations1999Sharon A. Caraballo
This work goes a step further by automatically creating not just clusters of related words, but a hierarchy of nouns and their hypernyms, akin to the hand-built hierarchy in WordNet.
The Tradeoffs Between Open and Traditional Relation Extraction
321 Citations2008Michele Banko, Oren Etzioni
A new model for Open IE called O-CRF is presented and it is shown that it achieves increased precision and nearly double the recall than the model employed by TEXTRUNNER, the previous stateof-the-art Open IE system.
Size matters
274 Citations2008Joshua Blumenstock
A simple metric -- word count -- is proposed for measuring article quality and it is shown that this metric significantly outperforms the more complex methods described in related work.
Learning semantic constraints for the automatic discovery of part-whole relations
236 Citations2003Roxana Gîrju, Adriana Badulescu +1 more
This paper presents a method and its results for learning semantic constraints to detect part-whole relations and the targeted part-Whole relations were detected with an accuracy of 83%.
Finding Contradictions in Text
230 Citations2008Marie-Catherine de Marneffe, Anna N. Rafferty +1 more
It is demonstrated that a system for contradiction needs to make more fine-grained distinctions than the common systems for entailment, and is argued for the centrality of event coreference and therefore incorporate such a component based on topicality.
Assigning trust to Wikipedia content
206 Citations2008B. Thomas Adler, Krishnendu Chatterjee +4 more
A system that computes quantitative values of trust for the text in Wikipedia articles; these trust values provide an indication of text reliability, and it is shown that text labeled as low-trust has a significantly higher probability of being edited in the future than text labeled as high-trust.
Houghton Mifflin Harcourt
162 Citations2011Stuart C. Gilson, Sarah L. Abbott
National Conference on Artificial IntelligenceNegation, contrast and contradiction in text processing
131 Citations2006Sanda M. Harabagiu, Andrew Hickl +1 more
A framework for recognizing contradictions between multiple text sources by relying on three forms of linguistic information: (a) negation; (b) antonymy; and (c) semantic and pragmatic information associated with the discourse relations.
What's in Wikipedia?
124 Citations2009Aniket Kittur, Ed H. +1 more
A mapping technique is introduced that takes advantage of socially-annotated hierarchical categories while dealing with the inconsistencies and noise inherent in the distributed way that they are generated in Wikipedia.
Highlighting disputed claims on the web
110 Citations2010Rob Ennals, Beth Trushkowsky +1 more
The design of Dispute Finder is explained, and the trade-offs between the various design decisions that were explored are explained.
Cohere: Towards Web 2.0 Argumentation
105 Citations2008Simon Buckingham Shum
The current state of the art in Web Argumentation is reviewed, key features of the Web 2.0 orientation are described, and some of the tensions that must be negotiated in bringing these worlds together are identified.
IEEE Intelligent SystemsLinking Documents to Encyclopedic Knowledge
102 Citations2008Andras Csomai, Rada Mihalcea
Embedded in each Wikipedia article is an abundance of links connecting the most important words or phrases in the text to other pages, thereby letting users quickly access additional information.
Entailment, intensionality and text understanding
101 Citations2003Cleo Condoravdi, Dick Crouch +3 more
A contexted clausal representation is described, derived from approaches in formal semantics, that permits an extended range of intensional entailments and contradictions to be tractably detected.
Bootstrapping for text learning tasks
89 Citations1999Ellen Riloff
This paper presents bootstrapping as an alternative approach to learning from large sets of labeled data, using a small amount of seed information and a large collection of easily-obtained unlabeled data.
It's a contradiction---no, it's not
72 Citations2008Alan Ritter, Doug Downey +2 more
It is shown that background knowledge about meronyms, synonyms, functions, and more is essential for success in the CD task, and a domain-independent algorithm is presented that automatically discovers phrases denoting functions with high precision.
Statement map
24 Citations2009Koji Murakami, Eric Nichols +5 more
The need to address issues of information credibility on the internet is discussed, the development of Statement Map generators for Japanese and English are outlined, the technical issues that are being addressed are discussed, and the construction of the resources necessary to meet the project's goals are reported on.
Grasping Major Statements and Their Contradictions Toward Information Credibility Analysis of Web Contents
12 Citations2008Daisuke Kawahara, Sadao Kurohashi +1 more
This paper proposes a method for providing a bird's eye view of major statements on a given topic and their contradictions, and evaluates the effectiveness of the approach.
DOAJ (DOAJ: Directory of Open Access Journals)Dawkins: The God delusion
4 Citations2007Javier Monserrat
He is a man who takes his atheism seriously, so much so that, in contrast, the grand Scottish philosopher of the XVIII century, David Hume, seems moderate, Michael Ruse offers.
