Linguistic Anthropology

The study of language has been part of anthropology since the discipline started in the 1ate 1870s. This site is a place for linguistic anthropologists to post their work and discuss important events and trends in the field.

Thursday, April 09, 2009

Eggcorns and Fuzzy Spots

The term eggcorn was coined in 2003 by linguist-bloggers Geoffrey Pullum and Mark Liberman, and has spawned something of a cottage industry of eggcorn-hunters on the Web. An eggcorn is defined as "an idiosyncratic substitution of a word or phrase for a word or words that sound similar or identical in the speaker's dialect [which] introduces a meaning that is different from the original, but plausible in the same context." The eponymous eggcorn, for example, is used by some speakers in place of acorn. Though not standard, the substitution is easy to understand: acorns are seeds (corns) that are sort of shaped like eggs. On web sites such as The Eggcorn Database and Eggcorn Forum, Wikipedia, and various blogs, scores of amateur lexicographers (including yours truly) track and discuss the production and use of these lexical innovations.

Recently Eggcorn Forum contributor Kem Luther, noting uses of both brute and blunt in place of brunt, wrote
I'm beginning to think that people have large fuzzy spots in their brains where the semantic contents of the words "brute," "blunt," "brunt," and "butt" are stored. Speakers do not have a good handle on the meaning of these terms.
I think he is quite right. Furthermore, Kem's observation about "large fuzzy spots" has lead me to create a half-baked theory about the role of word frequency in eggcorn formation.

If, as some linguists suggest (e.g. Bybee 1995, Clark 1987, Tomasello 2005 inter alia), language acquisition and language change are sensitive to frequency, then it is unsurprising that lower frequency words should be less clearly differentiated even in the minds of adult speakers. In other words, structures that are heard more frequently will be learned faster and "better" in the sense that the hearer can use them in the standard way that other speakers do; words that are heard less frequently will be learned less well. (Such claims can be somewhat controversial for morphology and syntax, but I think much less so for words.)

Add to low frequency a competition with phonologically similar words, and certain forms seem doomed to reside in "large fuzzy spots" of the mental lexicon. Low-frequency near-homonyms are not heard often enough to be remembered clearly, and they shade into one another in the hearer's memory.

None of the words brute, blunt, brunt or butt appear on Kilgarriff's lemmatized BNC word frequency list, which means that none of them occur more than 800 times in the hundred-million word British National Corpus.* It is therefore unsurprising that this complex is confounded for many speakers.

In comparison, the complex of and, ant, and amp are not confounded despite their phonological similarity, since they are the 4th, 5539th, and 5960th most frequent words in the BNC, respectively. Conversely, I would not expect moan to be part of such a complex even though it is only the 6309th most frequent word, since it has few near homophones. (Of course, having written that, I fully expect someone to find an eggcorn of moan within a few days. Comments are open.)

Bybee, Joan. 1995. Regular morphology and the lexicon. Language and Cognitive Processes 10, 425-455.

Clark, Eve. 1987. The principle of contrast: a constraint on language acquisition. In B. MacWhinney (ed) Mechanisms of Language Acquisition, 1-33. Mahwah, NJ: Lawrence Erlbaum Associates.

Tomasello, Michael. 2005. Constructing a Language. Cambridge, MA: Harvard University Press.

* Even if the word butt is relatively common in speech, it probably occurs most frequently in the anatomical sense. A Google search for "butt of" (as in "butt of the joke" etc.) returns more than a million and a half raw hits, but "my butt" returns nearly four and a half million. On the other hand, the conjunction but is the 23rd most frequent word on the Kilgarriff list; I suspect that it is frequent enough to avoid conflation with its homophones. Speaking of homophones, by the way, Meriam Webster's 10th Collegiate Dictionary lists six separate head words for butt and five for but.

Labels: , , ,

Tuesday, February 12, 2008

Punctuated Language Evolution

In the past few years, a number of biologists have published studies suggesting various facts about the evolution of languages. This is not terribly surprising, I suppose, since historical linguistics and evolutionary biology share a long lineage. Darwin's (1859) Origin of Species was probably influenced by studies such as Franz Bopp's (1816) Über das Conjugationssystem der Sanskritsprache (On the Conjugation System of Sanskrit) and others in historical linguistics.1

Some of these studies have been noted with approval; others have been roundly criticized.

A brief report by Atkinson, Meade, Venditti, Greenhill & Pagel published in Science falls into the former category.2 The authors looked for evidence of punctuated evolution in three language families: Bantu, Austronesian and Indo-European. As in evolutionary biology, some historical linguists theorize that the development of languages remains relatively stable for a time, punctuated by rapid changes, such as the development of new dialects or languages.

Atkinson et al. compared the apparent rate of lexical change within language families by looking for cognate terms for basic vocabulary items. Using Swadesh lists for each language family, they calculated the number of cognates in various languages within the family.

Atkinson et al. assume that if there is no punctuated evolution, the percentage of cognates should show no relation to the number of branches in the linguistic family tree. On the other hand, if those branches (the development of new languages) cause a burst of new vocabulary, there should be greater lexical diversity and fewer cognates along portions of the family tree with more branching.

The authors find "significantly more lexical change along paths in which more new languages have emerged," which they take to be evidence of punctuated evolution. They conclude,
Our results, representing thousands of years of language evolution, identify a general tendency for newly formed sister languages to diverge in their fundamental vocabulary initially at a rapid pace, followed by longer periods of slower and gradual divergence. Punctuational bursts in phonology, morphology, and syntax, or at later times of language contact, may also occur.

There is (as is always the case in science) room to disagree with Atkinson et al. Some historical linguists, for example, tend to doubt any such broad-brush theories of language change. Atkinson et al. do, however, seem more careful and presumably therefore more reliable than some widely reported work.

Atkinson et al. mention two possible explanations for such rapid change: a founder effect, which might occur when a few speakers of a language move to a new area (see, e.g. Mufwene 1996 for a use of the founder effect in sociolinguistics); or the development of distinction for social reason. For more on this second point, see most of the field of sociolinguistics. The paper's own references to Labov 1994 and Chambers 1995 are not bad places to start.

Atkinson, Quentin D., Andrew Meade, Chris Venditti, Simon J. Greenhill & Mark Pagel. 2008. Languages evolve in punctuational bursts. Science 319(5863), 588.

Chambers, J.K. 1995. Sociolinguistic Theory: Language Variation and its Social Significance. Cambridge, MA: Blackwell.

Labov, William. 1994. Principles of Linguistic Change: Internal Factors. Oxford: Blackwell.

Mufwene, Salikoko. 1996. The founder principle in creole genesis. Diachronica 13(1), 83-134.


1. Scholars disagree about how directly Darwin may have been influenced by philology and historical linguistics, but the notion of "descent with modification" emerged in both biology and philology at around the same time. As a linguist, I'm inclined to note that Bopp's work appeared a generation before Darwin's. A biologist, on the other hand, might suggest that Bopp and his contemporaries were in turn influenced by Carl Linneaus.

2. Pagel, Atkinson & Meade are similarly responsible for the letter titled "Frequency of word-use predicts rates of lexical evolution throughout Indo-European history," which I noted last year.

Labels: ,

Tuesday, October 23, 2007

Quantifying lexical change

Two recent papers in the journal Nature deal with rates of language change. (Thanks are due to Dave for expressing interest in the topic.)

Coverage of these studies in the news media has generally been pretty good. For example, the Independent says, Common words 'less likely to change' and Telegraph.co.uk reports, Scientists chart how words are changing.

As the popular press points out, both recent studies show that frequently used words are relatively resistant to change, compared to words that are used less frequently. Thus, for instance, while the large number of irregular verbs in Old English have gradually become regular (so that healp/halp/holp became help/helped/helped), the most frequent verbs remain irregular (begin/began/begun).

What might not be clear from these news reports is that linguists have noted this link between frequency and regularity for many years. As early as 1935, George Zipf suggested that there is an inverse relationship between frequency and complexity.

What the new studies accomplish is a far more sophisticated analysis of the regularity of language change that earlier scholars noted or theorized.

Lieberman et al. identify 177 irregular verbs in Old English (the language of Beowulf, spoken about 1000-1500 years ago), and compare them with Middle English (the language of Chaucer's Canterbury Tales, spoken about 500-1000 years ago) and Modern English. They divide the verbs into six classes, based on how frequently each word appears in a corpus of Modern English, and note that more frequent verbs are more likely to have become regular over the past millennium or so. Moreover, this tendency follows a fairly regular pattern, so that the rate of change is proportional to the square root of a verb's frequency. In other words, since words such as beat, freeze and melt occur about 100 times more frequently than words such as blend, fret and milk, the latter are about 10 times as likely to be regular. (All six were irregular in Old English.) The authors go so far as to predict the next irregular verb to regularize: wed, the least frequent irregular verb in their sample, may soon be completely regular.

Even more interesting is work by Pagel et al. on cognates in four Indo-European languages. Most languages spoken in Europe, as well as many spoken in central and southern Asia, are part of a large family of related languages called Indo-European. Cognates, similar-sounding words with the same meaning such as Spanish dos, Russian dva, Greek dio and English two, are one source of evidence that all of these languages descend from a common ancestor.

Pagel and his colleagues calculate the frequency of 200 words in English, Spanish, Russian and Greek. The frequency of these words is similar in each language - an interesting finding on its own. But wait, it gets better. The "rate of lexical replacement" (roughly, the likelihood of a new word replacing a cognate) varies inversely with a words frequency. Thus, frequently used words like two are more likely to be cognate in the related languages, while less frequent words tend to differ from language to language. The rate of lexical replacement is strikingly similar for each of the four languages analyzed.

The authors suggest two possibilities for this relationship. It may be that words that occur very frequently are learned relatively easily. Fewer 'errors' therefore lead to less variation, and less likelihood that a new word will replace the old one. Alternately, words that are used more frequently may be more useful, and users may be more conservative in producing them so as not to be misunderstood.

These papers may prove inspiring to cognitive scientists, evolutionary anthropologists, sociolinguists, linguistic anthropologists and others interested in how languages (or by extension, other elements of culture or group practice) change over time.


Lieberman, E., J. Michel, J. Jackson, T. Tang & M.A. Nowak. 2007. Quantifying the evolutionary dynamics of language. Nature 449, 713-716.

Pagel, M., Q.D. Atkinson & A. Meade. 2007. Frequency of word-use predicts rates of lexical evolution throughout Indo-European history. Nature 449, 717-720.

Zipf, G. 1935. The Psycho-Biology of Language. Boston: Houghton Mifflin.

Labels: