Filtering noisy parallel corpora of web pages
Published 13 November 2002
Jian‐Yun Nie, Jian Cai
Citations22
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
In our previous study, we successfully built an automatic mining system for parallel texts from the Web - PTMiner that is able to determine a large number of parallel Web pages for different language pairs. However, there are a number of non-parallel text pairs in this corpus. This paper proposes a filtering approach to clean up the corpus. Our experiments show that once the corpus is cleaned, both the translation accuracy of the resulting translation models and the effectiveness of cross-language information retrieval (CLIR) using these models are improved significantly.
Keywords
Computer Science
The mathematics of statistical machine translation: parameter estimation
4,125 Citations1993Peter F. Brown, Vincent J. Della Pietra +2 more
It is reasonable to argue that word-by-word alignments are inherent in any sufficiently large bilingual corpus, given a set of pairs of sentences that are translations of one another.
A program for aligning sentences in bilingual corpora
956 Citations1993William A. Gale, Kenneth Church
This paper will describe a method and a program for aligning sentences based on a simple statistical model of character lengths, which uses the fact that longer sentences in one language tend to be translated into longer sentence in the other language, and that shorter sentences tend to been translated into shorter sentences.
Cross-language information retrieval based on parallel texts and automatic mining of parallel texts from the Web
310 Citations1999Jian‐Yun Nie, Michel Simard +2 more
It is shown that using a probabilistic model, it is able to obtain performances close to those using an MT system, and the possibility of automatically gather parallel texts from the Web in an attempt to construct a reasonable training corpus is investigated.
Using cognates to align sentences in bilingual corpora
293 Citations1993Michel Simard, George Foster +1 more
It is discussed how cognates provide for a cheap and reasonably reliable source of linguistic knowledge, and how better and more efficient results may be obtained by combining the two criteria length and "cogneteness".
A program for aligning sentences in bilingual corpora
258 Citations1991William A. Gale, Kenneth Church
A method for aligning sentences in parallel texts, such as the Canadian Hansards, based on a simple statistical model of character lengths is described, developed and tested on a small trilingual sample of Swiss economic reports.
Aligning sentences in bilingual corpora using lexical information
230 Citations1993Stanley F. Chen
A fast algorithm for aligning sentences with their translations in a bilingual corpus that constructs a simple statistical word-to-word translation model on the fly during alignment and finds the alignment that maximizes the probability of generating the corpus with this translation model.
Aligning a parallel English-Chinese corpus statistically with lexical criteria
185 Citations1994Dekai Wu
This report concerns three related topics: progress on the HKUST English-Chinese Parallel Bilingual Corpus; experiments addressing the applicability of Gale & Church's length-based statistical method to the task of alignment involving a non-Indo-European language; and an improved statistical method that also incorporates domain-specific lexical cues.
Automatic construction of parallel English-Chinese corpus for cross-language information retrieval
60 Citations2000Chen Jiang, Jian‐Yun Nie
A parallel text mining system that finds parallel texts automatically on the Web and the generated Chinese- English parallel corpus is used to train a probabilistic translation model which translates queries for Chinese-English cross-language information retrieval (CLIR).
TREC-5 English and Chinese Retrieval Experiments using PIRCS.
31 Citations1996K. L. Kwok, Laszlo Grunfeld +1 more
Two English automatic ad-hoc runs have been submitted: pircsAAS uses a short and pircSAAL employs long topics, and the new avtf*ildf term weighting was used for short queries.
