Multiple sequence comparison — a peptide matching approach
Theoretical Computer SciencePublished 1 June 1997
Marie-France Sagot, Alain Viari, Henry Soldano
Citations22
SJR quartileQ2
SJR score0.49
SNIP0.94
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A peptide matching approach to the multiple comparison of a set of protein sequences by looking for all the words that are common to q of these sequences, where q is a parameter.
Abstract
International audience
Keywords
ChemistryComputer ScienceBiochemistry, Genetics and Molecular Biology
The Design and Analysis of Computer Algorithms
9,456 Citations1974Alfred V. Aho, John E. Hopcroft
This text introduces the basic data structures and programming techniques often used in efficient algorithms, and covers use of lists, push-down stacks, queues, trees, and graphs.
Journal of Molecular EvolutionProgressive sequence alignment as a prerequisitetto correct phylogenetic trees
1,904 Citations1987Da-Fei Feng, Russell F. Doolittle
A progressive alignment method that utilizes the Needleman and Wunsch pairwise alignment algorithm iteratively to achieve the multiple alignment of a set of protein sequences and to construct an evolutionary tree depicting their relationship is described.
Proteins Structure Function and BioinformaticsA workbench for multiple alignment construction and analysis
993 Citations1991Gregory D. Schuler, Stephen F. Altschul +1 more
An interactive program, MACAW (Multiple Alignment Construction and Analysis Workbench), that allows the user to construct multiple alignments by locating, analyzing, editing, and combining “blocks” of aligned sequence segments.
SIAM Journal on Applied MathematicsMinimal Mutation Trees of Sequences
591 Citations1975David Sankoff
Journal of Theoretical BiologyThe classification of amino acid conservation
563 Citations1986William R. Taylor
A classification of amino acid type is described which is based on a synthesis of physico-chemical and mutation data in the form of a Venn diagram from which sub-sets are derived that include groups of amino acids likely to be conserved for similar structural reasons.
SIAM Journal on Applied MathematicsThe Multiple Sequence Alignment Problem in Biology
526 Citations1988Humberto Carrillo-Calvet, David J. Lipman
It is proved here that knowledge of the measure of an arbitrarily chosen alignment can be used in combination with information from the pairwise alignments to considerably restrict the size of the region of the lattice in consideration.
Nucleic Acids ResearchPredictive motifs derived from cytosine methyltransferases
493 Citations1989János Pósfai, Ashok S. Bhagwat +2 more
Five highly conserved motifs occur in a mammalian methyltransferase responsible for the formation of 5-methylcytosine within CG dinucleotides, and can be used to discriminate the known 5- methylcytOSine forming methyltransferases from all other methyl transferases of known sequence, and from allother identified proteins in the PIR, GenBank and EMBL databases.
Journal of Molecular BiologyA strategy for the rapid multiple alignment of protein sequences
481 Citations1987Geoffrey J. Barton, Michael J.E. Sternberg
The multiple alignment algorithm yields an assignment of disulphide connectivity in mammalian serotransferrin that is consistent with crystallographic data, whereas pairwise alignments give an alternative assignment.
Nucleic Acids ResearchAutomated assembly of protein blocks for database searching
476 Citations1991Steven Henikoff, Jorja G. Henikoff
BiometricsIntroduction to Computational Biology: Maps, Sequences and Genomes.
453 Citations1998Michael S. Waterman
This chapter discusses mapping with Real Data Cloning and Clone Libraries, and physical maps and clone libraries, and the challenges faced in mapping with real data.
Rapid identification of repeated patterns in strings, trees and arrays
276 Citations1972Richard M. Karp, Raymond E. Miller +1 more
This paper describes a strategy for constructing efficient algorithms for solving two types of matching problems and develops explicit algorithms for these two problems applied to strings and arrays.
Journal of Molecular BiologyEmpirical and Structural Models for Insertions and Deletions in the Divergent Evolution of Proteins
210 Citations1993Steven A. Benner, Mark A. Cohen +1 more
This model provides theoretical support for using indels as part of "parsing algorithms", important in the de novo prediction of the folded structure of proteins from the sequence data.
Bulletin of Mathematical BiologyEfficient methods for multiple sequence alignment with guaranteed error bounds
190 Citations1993Dan Gusfield
Molecular Biology and EvolutionComparative analysis of multiple protein-sequence alignment methods.
162 Citations1994Marcella A. McClure, T K Vasi +1 more
Evaluated methods' ability to correctly identify the ordered series of motifs found among all members of a given protein family was evaluated and global methods generally performed better than local methods in the detection of motif patterns.
Bulletin of Mathematical BiologyGeneral methods of sequence comparison
162 Citations1984Michael S. Waterman
SIAM Journal on Applied MathematicsTrees, Stars, and Multiple Biological Sequence Alignment
141 Citations1989Stephen F. Altschul, David J. Lipman
This paper presents an extension of Carrillo and Lipman's algorithm to the definition of mulness, which requires the cost of a multiple alignment to be a weighted sum of the costs of its projected pairwise alignments.
Journal of Theoretical BiologyGap costs for multiple sequence alignment
138 Citations1989Stephen F. Altschul
This paper argues that, since gap and substitution costs together specify optimal alignments, they should be defined using a common rationale and proposed new definition of gap costs for multiple alignments is proposed and compared with previous ones.
Statistics in MedicineINTRODUCTION TO COMPUTATIONAL BIOLOGY: MAPS, SEQUENCES AND GENOMES.
129 Citations1996Susan R. Wilson
Bulletin of Mathematical BiologyEfficient methods for multiple sequence alignment with guaranteed error bounds
123 Citations1993Dan Gusfield
This paper considers two previously proposed measures, and given two computationaly efficient multiple alignment methods whose deviation from the optimal value is guaranteed to be less than a factor of two, gives a related randomized method which gives, with high probability, multiple alignments with fairly small error bounds.
Bulletin of Mathematical BiologyA survey of multiple sequence comparison methods
123 Citations1992Siu‐Lai Chan, Andrew K. C. Wong +1 more
A survey of the exhaustive and heuristic methods developed for the comparison of multiple macromolecular sequences and the use of entropy is proposed as a simple measure for the evaluation of the optimality of an alignment in the absence of any a priori knowledge about the structures of the sequences being compared.
Computer applications in the biosciencesImproved sensitivity of biological sequence database searches
95 Citations1990Douglas L. Brutlag, Jean-Pierre Dautricourt +2 more
The sensitivity of DNA and protein sequence database searches is increased by allowing similar but non-identical amino acids or nucleotides to match and one can match k-tuples or words instead of matching individual residues in order to speed the search.
Bulletin of Mathematical BiologyA survey of multiple sequence comparison methods
81 Citations1992S.C. Chan, Andrew K. C. Wong +1 more
A survey of the exhaustive and heuristic methods developed for the comparison of multiple macromolecular sequences and the use of entropy is proposed as a simple measure for the evaluation of the optimality of an alignment in the absence of anya priori knowledge about the structures of the sequences being compared.
Journal of Molecular BiologyA method for multiple sequence alignment with gaps
80 Citations1989S. Subbiah, S.C. Harrison
Given many sequences of low pairwise similarity, the proposed multiple sequence method can extract any familial similarity and so produce a sequence alignment consistent with the underlying structural homology.
Nucleic Acids ResearchA flexible multiple sequence alignment program
74 Citations1988Hugo M. Martínez
The 'regions' method for multisequence alignment used in the previously reported program MALIGN has been generalized to include recursive refinement so that unaligned portions between two regions at the current level of resolution can be handled with increased resolution.
Journal of Molecular EvolutionProtein export in prokaryotes and eukaryotes: Indications of a difference in the mechanism of exportation
49 Citations1986Olivier Gascuel, Antoine Danchin
Investigation of possible variations between prokaryotic and eukaryotic signal sequences of exported proteins has revealed unexpected differences, suggesting that the mode of secretion is rather different in the two types of organisms, in spite of the common features of the signal sequences.
Proceedings of the National Academy of SciencesMultiple-alphabet amino acid sequence comparisons of the immunoglobulin kappa-chain constant domain.
40 Citations1985Samuel Karlin, Ghassan Ghandour
The comparison reveals three regions of pronounced similarity across the three species, independent of allotype, that entail a high degree of identity at the DNA level and are distinguished from the rest of the constant domain in codon usage and in the dinucleotide sequence at abutting sites of adjacent codons.
Molecular Biology and EvolutionComparative Analysis of Multiple Protein-Sequence Alignment Methods
38 Citations1994Matthew McClure, T K Vasi +1 more
Biocomputing: informatics and genome projects.
38 Citations1994Douglas W. Smith
A Primer on Rapid Prototyping of Genomic Databases in Prolog and Predictions of Protein Secondary and Tertiary Structure.
Trends in Biochemical SciencesNitrogenases without molybdenum
37 Citations1989Richard N. Pau
For 50 years molybdenum had been considered to have an indispensable catalytic function for nitrogen fixation, but two nitrogenases recently isolated from the bacterium Azotobacter have changed this view.
Pattern Recognition LettersSearching for flexible repeated patterns using a non-transitive similarity relation
35 Citations1995Henry Soldano, Alain Viari +1 more
This work gives some general properties of maximal subsets of related objects and proposed algorithms for identifying them in the particular case of k-length substrings in a string derive from the Karp, Miller and Rosenberg algorithms for the identification of repeated patterns.
Searching for Repeated Words in a Text Allowing for Mismatches and Gaps
30 Citations1995M-F. Sagot, VINCENT ESCALIER +3 more
An algorithm that locates similar words common to a set of strings deened over an alphabet by using a reference object called a model which is a word over to perform a multiple comparison of the strings as opposed to pairwise comparisons.
Lecture notes in computer scienceCombinatorial Pattern Matching
25 Citations1994Gerhard Goos, Juris Hartmanis
IEEE Transactions on Pattern Analysis and Machine IntelligenceAn algorithm for finding a common structure shared by a family of strings
22 Citations1989Anne M. Landraud, J.-F. Avril +1 more
An algorithm is presented for extracting and localizing a common structure in a family of strings with time complexity O(N/sup 2/L/Sup 2/ log/sub 2/ L) where N is the number of strings and L their maximum length.
Bulletin of Mathematical BiologyA multiple sequence comparison method
13 Citations1993Andrew K. C. Wong, Siu‐Lai Chan +1 more
A new method is presented based on a hierarchical sequence synthesis procedure that does not require any a priori knowledge of the molecular structure of the sequences or the phylogenetic relations among the sequences and produces superior results when compared with some existing methods.
Lecture notes in computer scienceFast identification of approximately matching substrings
5 Citations1994Archie L. Cobbs
An efficient algorithm for finding all maximal matches between S and T, an arbitrary relation on a finite alphabet Σ, with main application is identifying homologous regions of protein sequences.
BiocomputingComparative Sequence Analysis: Finding Genes
5 Citations1994Steven Henikoff
This chapter discusses the various approaches for detecting sequence similarities, with emphasis on their applications to genome sequence analysis, and describes the use of blocks for representing the most highly conserved regions between the gaps.
FINDING FLEXIBLE PATTERNS IN A TEXT - AN APPLICATION TO 3D MOLECULAR MATCHING
5 Citations2005Marie-France Sagot, Alain Viari +2 more
The main motivation for the new algorithm presented in this paper is that of finding the patterns common to a set of protein structures, and it shows how the whole process of searching for all common words, in sequences or structures, rests on an operation of set intersection.
