XRCE Participation to the Book Structure Task
Lecture notes in computer sciencePublished 1 January 2009
Hervé Déjean, Jean-Luc Meunier
Citations8
SJR quartileQ2
SJR score0.35
SNIP0.55
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The XRCE participation to the Structure Extraction task of the INEX Book track 2009 is presented and the four methods used for detecting the book structure in the book body are explained.
Abstract
We present here XRCE participation to the Structure Extraction task of the INEX Book track. After briefly explaining the method used for detecting table of contents and their corresponding entries in the book body, we will mainly discuss the evaluation and the main issues we faced, and eventually we will propose improvements for our method as well as for the evaluation framework/method.
Keywords
Computer Science
International Journal on Document Analysis and Recognition (IJDAR)Recognition of table of contents for electronic library consulting
38 Citations2001Abdel Belaı̈d
A labelling approach for the automatic recognition of tables of contents (ToC) is described, used for the electronic consulting of scientific papers in a digital library system named Calliope, and operates by text labelling without using any a priori model.
International Journal on Document Analysis and Recognition (IJDAR)Detection and analysis of table of contents based on content association
32 Citations2005Xiaofan Lin, Yan Q. Xiong
A new method to detect and analyze TOCs based on content association that fully leverages the text information throughout the whole multi-page document and can be directly applied to a wide range of documents without the need to build or learn the models for individual documents.
Structuring documents according to their table of contents
30 Citations2005Hervé Déjean, Jean-Luc Meunier
This paper presents a method for structuring a document according to the information present in its Table of Contents using a series of generic properties characterizing any ToC, while its hierarchization is achieved using clustering techniques.
International Journal on Document Analysis and Recognition (IJDAR)On tables of contents and how to recognize them
26 Citations2009Hervé Déjean, Jean-Luc Meunier
This method is based on a two-step approach that leverages functional and formal (layout-based) kinds of knowledge and is improved in a second step by automatically learning the form of the table of contents.
Analysis of book documents' table of content based on clustering
14 Citations2009Liangcai Gao, Zhi Tang +3 more
It is observed that book documents are multi-page documents with intrinsic local format consistency, and an automatic TOC analysis method is introduced through clustering that generates TOC entries and extracts their hierarchical structure under the guidance of the model.
Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE<title>Versatile page numbering analysis</title>
6 Citations2007Hervé Déjean, Jean-Luc Meunier
This work proposes here a novel method, based on the notion of sequence, which goes beyond any previous described work, and reports on an extensive evaluation of its performance.
