What next?
Published 21 May 2012
Surajit Chaudhuri
Citations102
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Six data management research challenges relevant for Big Data and the Cloud are described, some of which are not new, but their importance is amplified by Big data and Cloud Computing.
Abstract
In this short paper, I describe six data management research challenges relevant for Big Data and the Cloud. Although some of these problems are not new, their importance is amplified by Big Data and Cloud Computing.
Keywords
Computer ScienceDecision SciencesBusiness, Management and Accounting
Lecture notes in computer scienceCalibrating Noise to Sensitivity in Private Data Analysis
7,028 Citations2006Cynthia Dwork, Frank McSherry +2 more
Lecture notes in computer scienceDifferential Privacy
5,027 Citations2006Cynthia Dwork
A general impossibility result is given showing that a formalization of Dalenius' goal along the lines of semantic security cannot be achieved, which suggests a new measure, differential privacy, which, intuitively, captures the increased risk to one's privacy incurred by participating in a database.
Pig latin
1,744 Citations2008Christopher Olston, Benjamin Reed +3 more
A new language called Pig Latin is described, designed to fit in a sweet spot between the declarative style of SQL, and the low-level, procedural style of map-reduce, which is an open-source, Apache-incubator project, and available for general use.
Proceedings of the VLDB EndowmentHive
1,539 Citations2009Ashish Thusoo, Joydeep Sen Sarma +7 more
Hadoop is a popular open-source map-reduce implementation which is being used as an alternative to store and process extremely large data sets on commodity hardware.
Communications of the ACMMapReduce
1,182 Citations2009Jay B. Dean, Sanjay Ghemawat
MapReduce advantages over parallel databases include storage-system independence and fine-grain fault tolerance for large jobs.
Privacy integrated queries
1,102 Citations2009Frank McSherry
PINQ's unconditional structural guarantees require no trust placed in the expertise or diligence of the analysts, substantially broadening the scope for design and deployment of privacy-preserving data analysis, especially by non-experts.
Online aggregation
923 Citations1997Joseph M. Hellerstein, Peter J. Haas +1 more
A new online aggregation interface is proposed that permits users to both observe the progress of their aggregation queries and control execution on the fly, and a suite of techniques that extend a database system to meet these requirements are presented.
Robust Disambiguation of Named Entities in Text
862 Citations2011Johannes Hoffart, Mohamed Amir Yosef +2 more
A robust method for collective disambiguation is presented, by harnessing context from knowledge bases and using a new form of coherence graph that significantly outperforms prior methods in terms of accuracy, with robust behavior across a variety of inputs.
Communications of the ACMAn overview of business intelligence technology
745 Citations2011Surajit Chaudhuri, Umeshwar Dayal +1 more
BI technologies are essential to running today's businesses and this technology is going through sea changes, so how do you protect yourself against these changes?
Proceedings of the VLDB EndowmentSCOPE
731 Citations2008Ronnie Chaiken, B. Keith Jenkins +5 more
A new declarative and extensible scripting language, SCOPE (Structured Computations Optimized for Parallel Execution), targeted for this type of massive data analysis, designed for ease of use with no explicit parallelism, while being amenable to efficient parallel execution on large clusters.
Collective annotation of Wikipedia entities in web text
447 Citations2009Sayali Kulkarni, Amit Singh +2 more
This work gives formulations for the trade-off between local spot-to-entity compatibility and measures of global coherence between entities, and investigates practical solutions based on local hill-climbing, rounding integer linear programs, and pre-clustering entities followed by local optimization within clusters.
Communications of the ACMMapReduce and parallel DBMSs
427 Citations2009Michael Stonebraker, Daniel J. Abadi +5 more
MapReduce complements DBMSs since databases are not designed for extract-transform-load tasks, a MapReduce specialty.
ACM SIGMOD RecordRipple joins for online aggregation
367 Citations1999Peter J. Haas, Joseph M. Hellerstein
ACM SIGMOD RecordJoin synopses for approximate query answering
335 Citations1999Swarup Acharya, Phillip B. Gibbons +2 more
ACM SIGMOD RecordOn random sampling over joins
303 Citations1999Surajit Chaudhuri, Rajeev Motwani +1 more
Join synopses for approximate query answering
197 Citations1999Swarup Acharya, Phillip B. Gibbons +2 more
This paper proposes join synopses as an effective solution for this problem and shows how precomputing just one join synopsis for each relation suffices to significantly improve the quality of approximate answers for arbitrary queries with foreign key joins.
Ripple joins for online aggregation
169 Citations1999Peter J. Haas, Joseph M. Hellerstein
It is shown how ripple joins can be implemented in an existing DBMS using iterators, and an overview of the methods used to compute confidence intervals and to adaptively optimize the ripple join “aspect-ratio” parameters are given.
Google fusion tables
165 Citations2010Héctor González, Alon Halevy +6 more
This paper characterizes such users and applications and highlights the resulting principles, such as seamless Web integration, emphasis on ease of use, and incentives for data sharing, that underlie the design of Fusion Tables.
On random sampling over joins
164 Citations1999Surajit Chaudhuri, Rajeev Motwani +1 more
A detailed study of the inefficiency of sampling the output of a query, based on new insights into the interaction between join and sampling, and develops join sampling techniques for the settings where negative results do not apply.
Adaptive self-tuning memory in DB2
113 Citations2006Adam Storm, Christian Garcia-Arellano +3 more
This work believes this is the first known use of cost-benefit analysis and control theory in database memory tuning across heterogeneous memory consumers.
Targeted disambiguation of ad-hoc, homogeneous sets of named entities
52 Citations2012Chi Wang, Kaushik Chakrabarti +2 more
This paper develops novel techniques that require no knowledge about the entities except their names and proposes a graph-based model, called MentionRank, for that purpose, to leverage the homogeneity constraint and disambiguate the candidate mentions collectively across all documents.
IEEE Transactions on Knowledge and Data EngineeringEntity Synonyms for Structured Web Search
50 Citations2011Tao Cheng, Hady W. Lauw +1 more
This paper proposes an offline, data-driven approach that mines query logs for instances where content creators and web users apply a variety of strings to refer to the same webpages, and generates an expanded set of equivalent strings (entity synonyms) for each entity.
Elsevier eBooksSQL Memory Management in Oracle9i
47 Citations2002Benoît Dageville, Mohamed Zaït
This paper presents a new model used in Oracle9i to manage memory for database operators, which is automatic, adaptive and robust, and presents the architecture of the memory manager, the internal algorithms, and a performance study showing its superiority.
Query optimizers
45 Citations2009Surajit Chaudhuri
It is argued that it is worth rethinking this prevalent model of the optimizer to benefit from leveraging rich usage data and from application input to further advance query optimization technology.
