TyPTex: generic features for text profiler
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work is implementing profiling tools and developing an associated methodology within the ELRA benchmark named Contribution to the construction of corpora of contemporary French that yields constraints for corpus profiling architectures.
Abstract
Very large corpora are increasingly exploited to improve Natural Language Processing (NLP) Systems. This however implies that the lexical, morpho-syntactic and syntactic homogeneity of the data used are mastered. This control in turn requires the development of tools aimed at text calibration or profiling. We are implementing such profiling tools and developing an associated methodology within the ELRA benchmark named Contribution to the construction of corpora of contemporary French. The first results of this approach -- applied to a sample of the main sections of Le Monde newspaper -- yields constraints for corpus profiling architectures. Rsum Le recours croissant aux "trs grands corpus" pour amliorer les systmes de Traitement Automatique des Langues (TAL) suppose de matriser l'homognit lexicale, morpho-syntaxique et syntaxique des donnes utilises. Cela implique en amont le dveloppement d'outils de calibrage de textes. Nous mettons en place de tels outils et la mthodologie associe ...
