login

Paraphrase Identification by Text Canonicalization

Published 1 December 2005
Yitao Zhang, Jon Patrick
Citations62

TL;DR

An approach to sentence-level paraphrase identification by text canon-icalization by text canon-icalization is proposed, which suggests that the MSR Para-phrase Corpus might not be a rich source for learning paraphrasing patterns.

Abstract

This paper proposes an approach to sentencelevel paraphrase identification by text canonicalization. The source sentence pairs are first converted into surface text that approximates canonical forms. A decision tree learning module which employs simple lexical matching features then takes the output canonicalized texts as its input for a supervised learning process. Experiments on the Microsoft Research (MSR) Paraphrase Corpus give comparable performance to other systems that are equipped with more sophisticated lexical semantic and syntactic matching components, with a Confidence-weighted Score of 0.791. An ancillary experiment using the occurrence of nominalizations suggests that the MSR Paraphrase Corpus might not be a rich source for learning paraphrasing patterns. 1

Keywords

Computer Science