An introduction to hidden Markov models for biological sequences
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This chapter discusses the hidden Markov models (HMM) for biological sequences, a statistical model which is very well suited for many tasks in molecular biology, although they have been mostly developed for speech recognition.
Abstract
This chapter discusses the hidden Markov models (HMM) for biological sequences. A hidden Markov model (HMM) is a statistical model, which is very well suited for many tasks in molecular biology, although they have been mostly developed for speech recognition. The most popular use of the HMM in molecular biology is as a probabilistic profile of a protein family, which is called a profile HMM. From a family of proteins (or DNA), a profile HMM can be made for searching a database for other members of the family. These profile HMMs resemble the profile and weight matrix methods, and the main contribution is that the profile HMM treats gaps in a systematic way. The HMM is particularly well suited for problems with a simple grammatical structure, such as gene finding. In gene finding several signals must be recognized and combined into a prediction of exons and introns and the prediction must conform to various rules to make it a reasonable gene prediction. An HMM can combine recognition of the signals and it can be made such that the predictions always follow the rules of a gene.
