Gene Ontology: tool for the unification of biology
Nature GeneticsPublished 1 May 2000Open access
Michael Ashburner, Catherine A. Ball, Judith A. Blake, David Botstein, H. Butler, J. Michael Cherry
Citations44,650
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The goal of the Gene Ontology Consortium is to produce a dynamic, controlled vocabulary that can be applied to all eukaryotes even as knowledge of gene and protein roles in cells is accumulating and changing.
Abstract
This FAIRsharing record describes: The Gene Ontology (GO) is a structured vocabulary for use by the research community for the annotation of genes, gene products and sequences. The GO defines concepts/classes used to describe gene function and relationships between these concepts.
Keywords
Biochemistry, Genetics and Molecular Biology
Nature GeneticsGene Ontology: tool for the unification of biology
44,650 Citations2000Michael Ashburner, Catherine A. Ball +18 more
The goal of the Gene Ontology Consortium is to produce a dynamic, controlled vocabulary that can be applied to all eukaryotes even as knowledge of gene and protein roles in cells is accumulating and changing.
Proceedings of the National Academy of SciencesCluster analysis and display of genome-wide expression patterns
16,395 Citations1998Michael B. Eisen, Paul T. Spellman +2 more
A system of cluster analysis for genome-wide expression data from DNA microarray hybridization is described that uses standard statistical algorithms to arrange genes according to similarity in pattern of gene expression, finding in the budding yeast Saccharomyces cerevisiae that clustering gene expression data groups together efficiently genes of known similar function.
Nucleic Acids ResearchThe Pfam Protein Families Database
14,211 Citations2002Alex Bateman
In addition to secondary structure, Pfam multiple sequence alignments now contain active site residue mark-up and new search tools, including taxonomy search and domain query, greatly add to the functionality and usability of the Pfam resource.
ScienceThe Genome Sequence of <i>Drosophila melanogaster</i>
6,026 Citations2000Mark D. Adams, S Celniker +193 more
The nucleotide sequence of nearly all of the approximately 120-megabase euchromatic portion of the Drosophila genome is determined using a whole-genome shotgun sequencing strategy supported by extensive clone-based sequence and a high-quality bacterial artificial chromosome physical map.
Molecular Biology of the CellComprehensive Identification of Cell Cycle–regulated Genes of the Yeast<i>Saccharomyces cerevisiae</i>by Microarray Hybridization
4,799 Citations1998Paul T. Spellman, Gavin Sherlock +7 more
A comprehensive catalog of yeast genes whose transcript levels vary periodically within the cell cycle is created, and it is found that the mRNA levels of more than half of these 800 genes respond to one or both of these cyclins.
Nucleic Acids ResearchThe COG database: a tool for genome-scale analysis of protein functions and evolution
4,683 Citations2000Roman L. Tatusov
The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes.
ScienceLife with 6000 Genes
4,262 Citations1996A. Goffeau, B. G. Barrell +14 more
The genome of the yeast Saccharomyces cerevisiae has been completely sequenced through a worldwide collaboration and provides information about the higher order organization of yeast's 16 chromosomes and allows some insight into their evolutionary history.
ScienceGenome Sequence of the Nematode <i>C. elegans</i> : A Platform for Investigating Biology
3,881 Citations1998The C. elegans Sequencing Consortium*
The 97-megabase genomic sequence of the nematode Caenorhabditis elegans reveals over 19,000 genes and the distinctive distribution of some repeats and highly conserved genes provides evidence for a regional organization of the chromosomes.
Nucleic Acids ResearchThe Gene Ontology resource: enriching a GOld mine
3,880 Citations2020Seth Carbon, Eric Douglass +179 more
A historical archive covering the past 15 years of GO data with a consistent format and file structure for both the ontology and annotations is made available to maintain consistency with other ontologies.
Nucleic Acids ResearchThe SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1999
3,238 Citations1999Amos Bairoch, Rolf Apweiler
Current developments of the database include: cross-references to additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT.
Nucleic Acids ResearchThe SWISS-PROT protein sequence database and its supplement TrEMBL in 2000
3,133 Citations2000Amos Bairoch
The Human Proteomics Initiative (HPI), a major project to annotate all known human sequences according to the quality standards of SWISS-PROT, is described.
ScienceComparative Genomics of the Eukaryotes
1,701 Citations2000Gerald M. Rubin, Mark Yandell +54 more
The fly has orthologs to 177 of the 289 human disease genes examined and provides the foundation for rapid analysis of some of the basic processes involved in human disease.
Nucleic Acids ResearchThe Pfam Protein Families Database
1,285 Citations2000Alex Bateman
The latest version (4.3) of Pfam contains 1815 families, which match 63% of proteins in SWISS-PROT 37 and TrEMBL 9.
Nucleic Acids ResearchMIPS: a database for genomes and protein sequences
1,281 Citations2000Hans‐Werner Mewes
This report describes growing databases reflecting the progress of sequencing the Arabidopsis thaliana and Neurospora crassa genomes (MNCDB), the yeast genome database (MYGD) extended by functional analysis data, the database of annotated human EST-clusters (HIB) and the database of the complete cDNA sequences from the DHGP (German Human Genome Project).
Nucleic Acids ResearchThe ENZYME database in 2000
1,140 Citations2000Amos Bairoch
The ENZYME database is a repository of information related to the nomenclature of enzymes that became an indispensable resource for the development of metabolic databases in recent years.
Nucleic Acids ResearchThe EMBL Nucleotide Sequence Database
1,108 Citations2000William Baker
The European Molecular Biology Laboratory (EMBL) Nucleotide Sequence Database (EBI) is maintained at the European Bioinformatics Institute in an international collaboration with the DNA Data Bank of Japan (DDBJ) and GenBank (USA).
Nucleic Acids ResearchSCOP: a Structural Classification of Proteins database
960 Citations1997Tim Hubbard, Alexey G. Murzin +2 more
The Structural Classification of Proteins (SCOP) database provides a detailed and comprehensive description of the relationships of known protein structures that provide the basis of the ASTRAL sequence libraries that can be used as a source of data to calibrate sequence search algorithms and for the generation of statistics on, or selections of, protein structures.
Annual Review of BiochemistryMCM Proteins in DNA Replication
668 Citations1999Bik K. Tye
New evidence suggests that the MCM2-7 proteins may be involved not only in the initiation but also in the elongation of DNA replication, which results in initiation of DNA synthesis once every cell cycle.
Nucleic Acids ResearchSCOP: a Structural Classification of Proteins database
590 Citations2000Loredana Lo Conte
The Structural Classification of Proteins (SCOP) database provides a detailed and comprehensive description of the relationships of all known proteins structures.
Science<i>Arabidopsis thaliana</i> : A Model Plant for Genome Analysis
575 Citations1998David W. Meinke, J. Michael Cherry +3 more
The entire genome of Arabidopsis thaliana is scheduled to be sequenced by the end of the year 2000, and reaching this milestone should enhance the value ofArabidopsis as a model for plant biology and the analysis of complex organisms in general.
CellFunctional homology of mammalian and yeast RAS genes
325 Citations1985Tohru Kataoka, Scott Powers +5 more
The results indicate that the biochemical function of RAS proteins is essential for vegetative haploid yeast and that this function has been conserved in evolution since the progenitors of yeast and mammals diverged.
Nucleic Acids ResearchThe FlyBase Database of the Drosophila Genome Projects and community literature
324 Citations1999
The salient features of the current databases and how to interrogate and navigate the extensive data sets are discussed.
Nucleic Acids ResearchThe Yeast Proteome Database (YPD) and Caenorhabditis elegans Proteome Database (WormPD): comprehensive resources for the organization and comparison of model organism protein information
286 Citations2000Maria C. Costanzo
YPD is extended to create a database containing complete proteome information about the model organism Caenorhabditis elegans (WormPDtrade mark), and YPD and WormPD are designed for use not only by their respective research communities but also by the broader scientific community.
ScienceYeast: an Experimental Organism for Modern Biology
285 Citations1988David Botstein, Gerald R. Fink
The yeasts Saccharomyces cerevisiae and Schizosac charomyces pombe have become popular and successful model systems for understanding eukaryotic biology at the cellular and molecular levels.
Nucleic Acids ResearchThe Protein Information Resource (PIR)
208 Citations2000Winona C. Barker
The Protein Information Resource (PIR) produces the largest, most comprehensive, annotated protein sequence database in the public domain, the PIR-International Protein Sequence Database, in collaboration with the Munich Information Center for Protein Sequences (MIPS) and the Japan International protein Sequence Database (JIPID).
BioinformaticsAutomated genome sequence analysis and annotation.
199 Citations1999Miguel A. Andrade‐Navarro, Nigel P. Brown +9 more
An automatic system for preliminary functional annotation of protein sequences that has been applied to the analysis of sets of sequences from complete genomes, both to refine overall performance and to make new discoveries comparable to those made by human experts is presented.
Nucleic Acids ResearchThe Mouse Genome Database (MGD): expanding genetic and genomic resources for the laboratory mouse
113 Citations2000Judith A. Blake
The Mouse Genome Database is a comprehensive public database of mouse genomic, genetic and phenotypic information that represents standardized mouse nomenclature for genes and alleles, incorporates links to other genomic resources such as sequence data, and includes a variety of additional information about the laboratory mouse.
Nucleic Acids ResearchIntegrating functional genomic information into the Saccharomyces Genome Database
100 Citations2000Catherine A. Ball
The Saccharomyces Genome Database (SGD) stores and organizes information about the nearly 6200 genes in the yeast genome around the 'locus page' and directs users to the detailed information they seek.
BioinformaticsA novel method for automatic functional annotation of proteins.
90 Citations1999Wolfgang Fleischmann, S Möller +2 more
A method of automatic annotation that produces highly reliable functional prediction using the language and the syntax of SWISS-PROT is developed and successfully used for the automatic annotation of a testset of unknown proteins.
Nature GeneticsGenome cross-referencing and XREFdb: Implications for the identification and analysis of genes mutated in human disease
89 Citations1997Douglas E. Bassett, Mark S. Boguski +5 more
A project is described that is systematically identifying novel expressed sequence tag (EST) sequences that are highly related to genes in model organisms and mapping them to positions on the mouse and human maps, facilitating the identification of genes mutated in human disease states via the positional candidate approach.
Mammalian GenomeConservation of the Caenorhabditis elegans timing gene clk-1 from yeast to human: a gene required for ubiquinone biosynthesis with potential implications for aging
78 Citations1999Zoltán Vajó, Tanya Jonassen +5 more
The protein similarities and the conservation of function of the CLK-1/clk-1/, COQ7/COQ7 gene products suggest a potential link between the production of ubiquinone and aging.
Molecular and Cellular BiologyMyb-Related <i>Schizosaccharomyces pombe</i>cdc5p Is Structurally and Functionally Conserved in Eukaryotes
74 Citations1998Ryoma Ohi, Anna Feoktistova +5 more
It is demonstrated that Cef1p is not involved in transcriptional activation of a class of G2/M-regulated genes typified by SWI5, and results suggest that Cdc5 family members participate in a novel pathway to regulate G 2/M progression.
Nucleic Acids ResearchDNA Data Bank of Japan (DDBJ) in collaboration with mass sequencing teams
73 Citations2000Yoshio Tateno
The authors at DDBJ process and publicise the massive amounts of data submitted mainly by Japanese genome projects and sequencing teams, which is so large that it alone exceeds the total amount submitted in the preceding 10 years.
Nucleic Acids ResearchUsing the Saccharomyces Genome Database (SGD) for analysis of protein similarities and structure
68 Citations1999Stephen A. Chervitz, Erich T. Hester +15 more
Comparison of proteins from complete eukaryotic proteomes will be an extremely powerful way to learn more about a particular protein's structure, its function, and its relationships with other proteins.
Nucleic Acids ResearchGXD: a Gene Expression Database for the laboratory mouse: current status and recent enhancements
35 Citations2000Martin Ringwald
The Gene Expression Database (GXD) is a community resource of gene expression information for the laboratory mouse designed as an open-ended system that can integrate different types of expression data.
Molecular and Cellular BiologyBiochemical and Genetic Conservation of Fission Yeast Dsk1 and Human SR Protein-Specific Kinase 1
32 Citations2000Zhaohua Tang, Tiffany Kuo +2 more
The finding that similar SR networks exist in Schizosaccharomyces pombe consisting of RS domain-containing proteins and SR protein-specific kinases is presented to establish the importance of the networks in eucaryotic organisms.
