Crowdsourcing systems on the World-Wide Web
Communications of the ACMPublished 22 March 2011
AnHai Doan, Raghu Ramakrishnan, Alon Halevy
Citations1,349
SJR quartileQ1
SJR score1.15
SNIP3.34
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
The practice of crowdsourcing is transforming the Web and giving rise to a new field.
Keywords
Computer Science
Item-based collaborative filtering recommendation algorithms
9,017 Citations2001Badrul Sarwar, George Karypis +2 more
This paper analyzes item-based collaborative ltering techniques and suggests that item- based algorithms provide dramatically better performance than user-based algorithms, while at the same time providing better quality than the best available userbased algorithms.
The VLDB JournalA survey of approaches to automatic schema matching
3,322 Citations2001Erhard Rahm, Philip A. Bernstein
A taxonomy is presented that distinguishes between schema-level and instance-level, element- level and structure- level, and language-based and constraint-based matchers and is intended to be useful when comparing different approaches to schema matching, when developing a new match algorithm, and when implementing a schema matching component.
Labeling images with a computer game
2,222 Citations2004Luis von Ahn, Laura Dabbish
A new interactive system: a game that is fun and can be used to create valuable output that addresses the image-labeling problem and encourages people to do the work by taking advantage of their desire to be entertained.
Mining knowledge-sharing sites for viral marketing
1,635 Citations2002Matthew Richardson, Pedro Domingos
This research optimize the amount of marketing funds spent on each customer, rather than just making a binary decision on whether to market to him, and takes into account the fact that knowledge of the network is partial, and that gathering that knowledge can itself have a cost.
Communications of the ACMDesigning games with a purpose
1,169 Citations2008Luis von Ahn, Laura Dabbish
Data generated as a side effect of game play also solves computational problems and trains AI algorithms.
SciencereCAPTCHA: Human-Based Character Recognition via Web Security Measures
1,126 Citations2008Luis von Ahn, Benjamin Maurer +3 more
This research explored whether human effort can be channeled into a useful purpose: helping to digitize old printed material by asking users to decipher scanned words from books that computerized optical character recognition failed to recognize.
Digital Repository at the University of Maryland (University of Maryland College Park)Computing and applying trust in web-based social networks
839 Citations2005Jennifer Golbeck, James Hendler
It is shown that, in the case where the user's opinion is divergent from the average, the trust-based recommended ratings are more accurate than several other common collaborative filtering techniques.
Knowledge sharing and yahoo answers
730 Citations2008Lada A. Adamic, Jun Zhang +2 more
This paper analyzes YA's forum categories and cluster them according to content characteristics and patterns of interaction among the users, finding that lower entropy correlates with receiving higher answer ratings, but only for categories where factual expertise is primarily sought after.
CrowdDB
630 Citations2011Michael J. Franklin, Donald Kossmann +3 more
The design of CrowdDB is described, a major change is that the traditional closed-world assumption for query processing does not hold for human input, and important avenues for future work in the development of crowdsourced query processing systems are outlined.
TurKit
239 Citations2009Greg Little, Lydia B. Chilton +2 more
Communications of the ACMAdaptive Web sites
201 Citations2000Mike Perkowitz, Oren Etzioni
A Web management assistant is proposed: a system that can process massive amounts of data about site usage and the potential use of automated adaptation to improve Web sites for visitors.
Using the wisdom of the crowds for keyword generation
150 Citations2008Ariel Fuxman, Panayiotis Tsaparas +2 more
This work identifies queries related to a campaign by exploiting the associations between queries and URLs as they are captured by the user's clicks, and proposes algorithms within the Markov Random Field model to solve this problem.
Proceedings of the VLDB EndowmentHuman-assisted graph search
148 Citations2011Aditya Parameswaran, Anish Das Sarma +3 more
This work provides the first formal algorithmic study of the optimization of human computation for graph search by asking an omniscient human questions of the form "Is there a target node that is reachable from the current node?".
Answering Queries using Humans, Algorithms and Databases
112 Citations2011Aditya Parameswaran, Neoklis Polyzotis
The design of the first declarative language involving human-computable functions, standard relational operators, as well as algorithmic computation is described, which can act as a roadmap for new area of data management research where human computation is routinely used in data analytics.
ACM SIGMOD RecordThe YAGO-NAGA approach to knowledge discovery
104 Citations2009Gjergji Kasneci, Maya Ramanath +2 more
The architecture of the YAGO extractor toolkit, its distinctive approach to consistency checking, its provisions for maintenance and further growth, and the query engine for YAGA, coined NAGA are presented.
Matching Schemas in Online Communities: A Web 2.0 Approach
95 Citations2008Robert McCann, Warren Shen +1 more
This work proposes to enlist the multitude of users in the community to help match the schemas of the data sources, in a Web 2.0 fashion, to address the problem of integrating data from multiple sources.
ORCHESTRA : Rapid, collaborative sharing of dynamic data
92 Citations2005Zachary G. Ives, Nitin Khandelwal +2 more
This paper describes the initial design and prototype of the ORCHESTRA system, which focuses on managing disagreement among multiple data representations and instances and represents an important evolution of the concepts of peer-to-peer data sharing, which considers revision, disagreement, authority, and intermittent participation.
Building large knowledge bases by mass collaboration
80 Citations2003Matthew Richardson, Pedro Domingos
This paper uses first-order probabilistic reasoning techniques to combine potentially inconsistent knowledge sources of varying quality, and it uses machine-learning techniques to estimate the quality of knowledge.
Lecture notes in computer scienceCollecting Community-Based Mappings in an Ontology Repository
65 Citations2008Natalya F. Noy, Nicholas Griffith +1 more
A model for representing mappings collected from the user community and the metadata associated with the mapping is developed and used to bring together more than 30,000 mappings from 7 sources.
Lecture notes in computer scienceMangrove: Enticing Ordinary People onto the Semantic Web via Instant Gratification
50 Citations2003Luke K. McDowell, Oren Etzioni +6 more
Mangrove demonstrates a concrete path for enabling and enticing non-technical people to enter the semantic web, by transferring some of the burden of schema design, data cleaning, and data structuring from content authors to the programmers who create semantic services.
IEEE Intelligent SystemsThe CKC Challenge: Exploring Tools for Collaborative Knowledge Construction
49 Citations2008Natalya F. Noy, Abhita Chugh +1 more
A new generation of tools supports the integration of Web 2.0 and semantic Web approaches, and fully fledged ontology editors such as pOWL support the distributed and collaborative development of ontologies.
Efficiently incorporating user feedback into information extraction and integration programs
47 Citations2009Xiaoyong Chai, Ba-Quy Vuong +2 more
This paper proposes a solution for users to directly provide feedback and for IE/II programs to automatically process such feedback, and shows how to automatically propagate F to the rest of P, and to seamlessly combine F with prior user feedback.
Building data integration systems: a mass collaboration approach
33 Citations2003AnHai Doan, Robert McCann
This paper describes a conceptually new solution to this problem: that of mass collaboration: that of mass collaboration to the problem of schema matching in the context of data integration.
Building Community Wikipedias: A Machine-Human Partnership Approach
24 Citations2008Pedro DeRose, Xiaoyong Chai +5 more
This paper considers combining the above two approaches to building community portals, and proposes a new hybrid machine-human approach that enables building "community wikipedias", backed by an underlying structured database that is continuously updated using automatic techniques.
Proceedings of the International AAAI Conference on Web and Social MediaCourseRank: A Closed-Community Social System through the Magnifying Glass
16 Citations2009Georgia Koutrika, Benjamin Bercovitz +3 more
An analysis of 12 months worth of CourseRank data including user contributed information, such as ratings and comments, as well as information extracted from the user logs, provides useful insights with respect to the potential of closed-community social sites.
Mass collaboration and data mining
9 Citations2001Raghu Ramakrishnan
This talk will introduce Mass Collaboration and discuss some important data mining related issues and the promise of increased business intelligence, and also improved user experiences, leading to increased participation and greater quality in the knowledge that is captured.
ACM SIGMOD RecordThe amateur search
3 Citations2008Michael Olson
In the days that followed Jim's disappearance at sea, his family, friends and colleagues came together in a remarkable effort to help find him.
ComputerChanging the SCE focus from processes to people [Letters]
1 Citations2000Dennis Cox
It’s time to address the taboo of never asking or requiring information about how well prepared people are to produce software and apply questions and tests to apply to software capability evaluations.
