A Cluster-Based Plagiarism Detection Method - Lab Report for PAN at CLEF 2010.
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A cluster-based plagiarism detection method is described, which has been used in the learning management system of SCUT to detect plagiarism in the network engineering related courses and was used to detect external plagiarisms in the PAN-10 competition.
Abstract
Abstract. In this paper we describe a cluster-based plagiarism detection method, which we have used in the learning management system of SCUT to detect plagiarism in the network engineering related courses. And we also used this method to detect external plagiarism in the PAN-10 competition. The method is divided into three steps: the first step, called pre-selecting, is to narrow the scope of detection using the successive same fingerprint; the second step, called locating, is to find and merge all fragments between two documents using cluster method; the third step, called post-processing, is to deal with some merging errors. Our method ran 19 hours in the PAN-10 competition, and the result ranked the second place, which met our expectation.
