login

Two Regimes in the Frequency of Words and the Origins of Complex Lexicons: Zipf’s Law Revisited<sup>∗</sup>

Journal of Quantitative LinguisticsPublished 1 December 2001Open access
Ramon Ferrer‐i‐Cancho, Ricard V. Solé
Citations227
SJR quartileQ1
SJR score0.60
SNIP1.42
View PDF

TL;DR

It is made evident that word frequency as a function of the rank follows two different exponents, ˜(-)1 for the first regime and ™(-)2 for the second.

Abstract

Zipf's law states that the frequency of a word is a power function of its rank. The exponent of the power is usually accepted to be close to (-)1. Great deviations between the predicted and real number of different words of a text, disagreements between the predicted and real exponent of the probability density function and statistics on a big corpus, make evident that word frequency as a function of the rank follows two different exponents, ~(-)1 for the first regime and ~(-)2 for the second. The implications of the change in exponents for the metrics of texts and for the origins of complex lexicons are analyzed.

Keywords

Social SciencesBiochemistry, Genetics and Molecular Biology