Tagging sentence boundaries
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper describes an extension of the traditional POS tagging by combining it with the document-centered approach to proper name identification and abbreviation handling that made the resulting system robust to domain and topic shifts.
Abstract
In this paper we tackle sentence boundary disam- biguation through a part-of-speech (POS) tagging framework. We describe necessary changes in text tokenization and the implementation of a POS tagger and provide results of an evaluation of this system on two corpora. We also describe an extension of the traditional POS tagging by combining it with the document-centered approach to proper name identification and abbreviation handling. This made the resulting system robust to domain and topic shifts.
