login

Harnessing Knowledge from Structural Genomics

StructurePublished 1 January 2008Open access
Helen M. Berman
Citations7
SJR quartileQ1
SJR score2.01
SNIP0.98
View PDF

TL;DR

As the central repository for all macromolecular structures, the Protein Data Bank (PDB) started collaborating with the worldwide structural genomics projects from their inception and it is becoming clear that these efforts will make a significant impact on how the authors do structural biology.

Abstract

As the central repository for all macromolecular structures, the Protein Data Bank (PDB) started collaborating with the worldwide structural genomics projects from their inception (Berman et al., 2000Berman H.M. Westbrook J. Feng Z. Gilliland G. Bhat T.N. Weissig H. Shindyalov I.N. Bourne P.E. The Protein Data Bank.Nucleic Acids Res. 2000; 28: 235-242Crossref PubMed Scopus (24374) Google Scholar, Berman et al., 2003Berman H. Henrick K. Nakamura H. Announcing the worldwide Protein Data Bank.Nat. Struct. Biol. 2003; 10: 980Crossref PubMed Scopus (1619) Google Scholar). From the beginning, it was clear that structural genomics, including the U.S.-funded Protein Structure Initiative (PSI), would change the ways in which we think about publishing and data sharing. As time has gone on it is becoming clear that these efforts will make a significant impact on how we do structural biology. By creating an appropriate infrastructure in the form of a Knowledgebase, the fruits of the PSI effort can enable a new kind of biology. Since 1989, it has become the norm to submit coordinates as a condition for publishing articles describing structure determinations (International Union of Crystallography, 1989International Union of CrystallographyPolicy on publication and the deposition of data from crystallographic studies of biological macromolecules.Acta Crystallogr. A. 1989; 45: 658Crossref Google Scholar). For PSI projects, it has been mandatory to deposit and release the coordinate and structure factor data within one month of completing a structure, prior to any journal publication. The impact of this policy raises some interesting questions. Would this mean that PSI research could be no longer published in standard journals? How would journal publication practices change? Two things have emerged so far. First, a PDB entry can itself be thought of as a publication. The PDB now assigns a digital object identifier (DOI) to every structure, and these are beginning to appear as references in published articles. Second, more than 600 papers describing the results of structure determinations have been authored at PSI centers—subsequent to data release—and many more are in the pipeline. Whether or not this will become a trend for non-PSI structures, which are typically released after journal publication, remains to be seen. An important aspect of the charter of the PSI is the suggestion of a new paradigm for the information sharing in support of the advancement of science. In addition to sharing the results of structure determinations, the PSI projects provide the sequence as well as information about the status of each target under investigation. It is very unusual in conventional structural biology for these types of data to be made public in advance of publication of the structure. TargetDB (http://targetdb.pdb.org) tracks status indicators for each step in the structure determination pipeline (Chen et al., 2004Chen L. Oughtred R. Berman H.M. Westbrook J. TargetDB: A target registration database for structural genomics projects.Bioinformatics. 2004; 20: 2860-2862Crossref PubMed Scopus (145) Google Scholar). Along with information such as protocols for protein production, PepcDB (http://pepcdb.pdb.org) provides the reasons why work on a particular target has stopped (Kouranov et al., 2006Kouranov A. Xie L. de la Cruz J. Chen L. Westbrook J. Bourne P.E. Berman H.M. The RCSB PDB information portal for structural genomics.Nucleic Acids Res. 2006; 34: D302-D305Crossref PubMed Scopus (257) Google Scholar). The information in these resources provides methods to facilitate experimental design, not only for the PSI projects but also for the biological community at large. Data sharing that includes the disclosure of sequences, tracking, and protocol details in advance of publication or deposition into the PDB and the early release of coordinate and experimental data is far ahead of current practices in structural biology. This represents a significant leap from where we were 25 years ago, when some investigators worried about making their coordinates available to the rest of the research community! A review of the progress of the PSI since it began in 2001 demonstrates that it has been tremendously successful in achieving the initial goals of selecting, producing, and determining the structures of many novel proteins in a high throughput manner. More than 2700 structures have been determined; most remarkably, about half of these have been determined in the two years since the second phase, PSI-2, began. Of the structures determined, more than 68% are novel, meaning they have less than a 30% sequence identity with those in the PDB. Our understanding of structure space has been transformed in that the conservation of overall polypeptide chain folds is greater than had been anticipated. With the clever targeting of structures for analysis, the coverage of sequence space that can now be modeled is ever-increasing. In June 2007, I was selected to lead the development of the PSI Structural Genomics Knowledgebase (PSI_SGKB). The idea was to make the products of the PSI widely available to the broader community of biologists. Although I was very aware of the success of the initiative with respect to the production of many structures, I needed to investigate the full scope of activities of the PSI centers before accepting this new challenge. In reviewing all of the PSI center progress reports and websites, I discovered a treasure trove. Indeed, the PSI projects have done more than simply determine many novel structures. In order to develop high throughput methods, they have significantly removed bottlenecks in all aspects of the structure determination pipeline, including protein production, crystallization, data collection, structure determination, and refinement. The result is that PSI structures are determined in a short period of time at a cost much lower than average. The partnerships formed with synchrotron staff scientists have provided process improvements that benefit all of the structural biology community who rely on beamlines for research. New validation procedures developed and used by these initiatives have resulted in PSI structures that are of the same or higher quality than many others in the PDB, and certainly not of the low quality that had originally been anticipated. These procedures can and in fact are being used by others in the structural biology community. Methods are also being developed to annotate structures in an attempt to help discover the functions of the approximately half of the structures with unknown function. Finally, work is in progress to create new methods to leverage these structures and model more sequence space. The challenge then for the PSI_SGKB is to organize this vast amount of data and information for use by a broad spectrum of researchers. The goal is to offer a marketplace of ideas that connect protein sequence information to three-dimensional structures and homology models, enhance functional annotations, and provide access to new experimental protocols and materials. By making all of these products accessible to the greater community, the PSI_SGKB will become an increasingly empowering resource for biologists, biochemists, functional genomists, pharmacologists, educators, and physicians. To achieve these goals, the PSI_SGKB is being developed as a portal that will offer a wide variety of services. The products of the PSI centers will be integrated with external resources and more readily searchable. The key components of the PSI_SGKB currently include experimental data tracking, a materials repository, homology modeling, annotation, technology development, metrics, and outreach. TargetDB and PepcDB were originally established as part of the RCSB PDB to track the progress of targets studied by PSI Centers. TargetDB gives the status of each target and PepcDB provides information about the protocols used for protein production and the reasons for stopping work on any target. Data are regularly collected, tracked, and made available via the TargetDB and PepcDB websites. These resources have both query and report functionality, and provide crossreferences to the RCSB PDB, Pfam, Superfamily, TIGR Families, ProDom, iProClass, and Prosite (Berman et al., 2000Berman H.M. Westbrook J. Feng Z. Gilliland G. Bhat T.N. Weissig H. Shindyalov I.N. Bourne P.E. The Protein Data Bank.Nucleic Acids Res. 2000; 28: 235-242Crossref PubMed Scopus (24374) Google Scholar, Corpet et al., 1998Corpet F. Gouzy J. Kahn D. The ProDom database of protein domain families.Nucleic Acids Res. 1998; 26: 323-326Crossref PubMed Scopus (165) Google Scholar, Gough et al., 2001Gough J. Karplus K. Hughey R. Chothia C. Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure.J. Mol. Biol. 2001; 313: 903-919Crossref PubMed Scopus (886) Google Scholar, Haft et al., 2001Haft D.H. Loftus B.J. Richardson D.L. Yang F. Eisen J.A. Paulsen I.T. White O. TIGRFAMs: a protein family resource for the functional identification of proteins.Nucleic Acids Res. 2001; 29: 41-43Crossref PubMed Scopus (252) Google Scholar, Hulo et al., 2006Hulo N. Bairoch A. Bulliard V. Cerutti L. De Castro E. Langendijk-Genevaux P.S. Pagni M. Sigrist C.J. The PROSITE database.Nucleic Acids Res. 2006; 34: D227-D230Crossref PubMed Scopus (623) Google Scholar, Sonnhammer et al., 1998Sonnhammer E.L. Eddy S.R. Birney E. Bateman A. Durbin R. Pfam: multiple sequence alignments and HMM-profiles of protein domains.Nucleic Acids Res. 1998; 26: 320-322Crossref PubMed Scopus (539) Google Scholar, Wu et al., 2001Wu C.H. Xiao C. Hou Z. Huang H. Barker W.C. iProClass: an integrated, comprehensive and annotated protein classification database.Nucleic Acids Res. 2001; 29: 52-54Crossref PubMed Scopus (26) Google Scholar). PepcDB is an indispensable and truly unique resource for biologists who are expressing and purifying proteins for their own experiments. A Materials Repository (http://www.hip.harvard.edu/PSIMR/index.htm) has been established at Harvard University under the leadership of Josh La Baer. A mechanism for storing and distributing clones is in place. When the repository is operational, it will be possible to determine whether clones are available for any particular target. For every structure determined by the PSI Centers, hundreds of models can be made using a variety of established methods. At the Workshop on Biological Macromolecular Structure Models held in 2005, it was proposed that a portal for models be launched (Berman et al., 2006Berman H.M. Burley S.K. Chiu W. Sali A. Adzhubei A. Bourne P.E. Bryant S.H. Dunbrack Jr., R.L. Fidelis K. Frank J. et al.Outcome of a workshop on archiving structural models of biological macromolecules.Structure. 2006; 14: 1211-1217Abstract Full Text Full Text PDF PubMed Scopus (45) Google Scholar). This would allow access to a variety of models predicted by different methods for any target. This portal (http://www.proteinmodelportal.org), in development by Torsten Schwede and his team at the Swiss Institute of Bioinformatics, gives access to prebuilt models from PSI centers and also to models calculated from the contents of UniProt (The UniProt Consortium, 2007The UniProt ConsortiumThe Universal Protein Resource (UniProt).Nucleic Acids Res. 2007; 35: D193-D197Crossref PubMed Scopus (418) Google Scholar). In the future it will be possible to build models on the fly. For each target, many different annotations are possible: structure determination and validation details; sequence information, including possible domain assignments; structure information, including surface characteristics, cavities, potential and actual active sites; fold classification; protein-protein interactions; protein-ligand interactions; structure-function relationships; and many others. Some of these annotations are available through the PSI centers and others made available through a large variety of resources, including the RCSB PDB, Gene Ontology, ProFunc, ProSite, CATH, and SCOP (Berman et al., 2000Berman H.M. Westbrook J. Feng Z. Gilliland G. Bhat T.N. Weissig H. Shindyalov I.N. Bourne P.E. The Protein Data Bank.Nucleic Acids Res. 2000; 28: 235-242Crossref PubMed Scopus (24374) Google Scholar, Conte et al., 2000Conte L. Bart A. Hubbard T. Brenner S. Murzin A. Chothia C. SCOP: a structural classification of proteins database.Nucleic Acids Res. 2000; 28: 257-259Crossref PubMed Scopus (514) Google Scholar, Hulo et al., 2006Hulo N. Bairoch A. Bulliard V. Cerutti L. De Castro E. Langendijk-Genevaux P.S. Pagni M. Sigrist C.J. The PROSITE database.Nucleic Acids Res. 2006; 34: D227-D230Crossref PubMed Scopus (623) Google Scholar, Laskowski et al., 2005Laskowski R.A. Watson J.D. Thornton J.M. ProFunc: a server for predicting protein function from 3D structure.Nucleic Acids Res. 2005; 33: W89-W93Crossref PubMed Scopus (467) Google Scholar, Orengo et al., 1997Orengo C.A. Michie A.D. Jones S. Jones D.T. Swindells M.B. Thornton J.M. CATH–a hierarchic classification of protein domain structures.Structure. 1997; 5: 1093-1108Abstract Full Text Full Text PDF PubMed Google Scholar, The Gene Ontology Consortium, 2000The Gene Ontology ConsortiumGene Ontology: tool for the unification of biology.Nat. Genet. 2000; 25: 25-29Crossref PubMed Scopus (23524) Google Scholar). Many of these annotations will be made available at the PSI_SGKB site. A workshop will be held in March 2008 to determine which additional annotations should also be made available. Of particular interest will be the development of a process that will allow users to add their own annotations, perhaps using Wiki technology. The PSI centers have developed cutting-edge technologies for all stages of the structure determination pipeline. The descriptions of these technologies would be enormously useful to the broader community for use in other research. Paul Adams at Lawrence Berkley National Laboratory leads the effort developing the module that provides information about these technologies (http://cci.lbl.gov/kb-tech). It will also contain descriptions of the software, access to applications, and the software itself, where possible. The ready availability of a variety of metrics will make it possible to fully appreciate the productivity of the PSI projects. The first version of this module (http://targetdb.pdb.org/MilestonesTables.html) was developed in collaboration with the Intercenter Bioinformatics Group. Examples of metrics include the number of structures, the numbers of unique structures, modeling leverage, publications, and the citations to the PSI initiative. A vigorous outreach and education program is being developed to make the products of the PSI efforts well known and understood by a broad community. This initiative will include partnerships with journals and professional societies. The PSI_SGKB Portal has entry points into each of the Module and PSI center websites. Special features of the site include: news about the PSI program, featured PSI structures, and technology highlights. The Functional Sleuth section highlights proteins whose function is not yet known. Users are encouraged to explore the known annotations for these proteins and attempt to determine the function by doing further experiments. Portal queries are currently supported for protein sequence, PDB ID, or keyword. Reports contain: 1) the protocols that have been used for protein production; 2) characteristics of structures that have been determined; 3) models that have been generated or could be predicted; 4) domain classifications using a variety of methods; and 5) functional and biomedical annotations. When a keyword query is performed, the reports contain links to documents at PSI Centers, PSI Modules, and the PSI_SGKB portal. By providing a “one-stop shop” for PSI products and open forums for interchange of ideas, the PSI_SGKB Portal will create a new venue for connecting structural biology with the broader biological community. The PSI_SGKB is available for public testing at http://kb-test.psi-structuralgenomics.org/KB/. We encourage the community to explore this test site and send comments and suggestions to: [email protected] . The funding for the PSI_SGKB and TargetDB/PepcDB is provided by NIGMS. The RCSB PDB is supported by funds from NSF, NIGMS, DOE, NLM, NCI, NCRR, NIBIB, NINDS, and NIDDK.

Keywords

Materials ScienceBiochemistry, Genetics and Molecular Biology