Where's le Boeuf?
I was watching a video from a teacher of English whom I generally respect who teaches foreign students. He was regurgitating the old canard about how native English words form only a minority of the vocabulary of English. Indeed, Wikipedia lists the following percentages of word origins in a pie chart:
Well. There is no doubt that there are many, many (MANY) words from all kinds of sources in English. Modern English has the largest wordstock of any language in the history of the world. But those are just the words that are available to us. Nobody uses all of those words in their personal expression, nor does every English reader recognize all of them. The question then is, does the way we talk and write actually bear out the idea that words of English origin form a minority of the vocabulary we actually use?
I’ve done this kind of analysis before, but it bears repeating. Looking for a handy sample of ordinary English, I took down my copy of Aesop’s Fables and turned to this familiar tale.
Before we go on to further analysis, we have to answer: What counts as a word? I’m counting egg, eggs as one word; likewise, getting, got. On the other hand, I’m counting gold, golden and day, daily as being two different words in each case. But it’s not just different forms to consider. Do we count each headword (lemma)? Or do we count repeated words in the sample (total occurrences)? I will count total words both ways so as to give fullest accounting.
The results:
Now, wait a minute, you say. That’s not fair! You chose a children’s book! I would dispute that Aesop’s Fables is a children’s book, but it is written in very simple, direct language. We call that being effective in the writing trade. But, okay, we’ll take another sample of similar size from a non-technical source.
I went to my bookshelves and pulled down God in the Dock, by C.S. Lewis. He has a short meditation therein, which goes as follows.
What counts as a word? Besides counting headwords and total occurrences separately, I’m treating beauty, beautiful as two words. Contractions are counted as single words (OE had contractions, too, witness nolde for ne wolde). Hyphenated words count as two, but compound words count as one. I’m giving woodcuts to OE, despite “cut” being of ON origin; however, I’m crediting turn to Latin, even though it was a very early borrowing and thoroughly nativized as OE tyrnen. Just trying to be fair: no special pleading for native words.
The results:
But, but, but . . . Okay, I’ll give the denigrators of Old English one more chance.
I went to my shelves and took down a little book called The Curmudgeon’s Guide to Getting Ahead, by Charles Murray – a social scientist and purveyor of educated gobbledygook. Surely, we ought to see what English looks like today from him, ja? From the Introduction:
What counts as a word? Well, different grammatical forms of the same word are counted together. This handicaps Old Norse (they, their, them). (N.B. That is here not the singular of those, so these are two different words.) Meanwhile, graduate(s) is counted twice, once as an adjective and once as a noun. (I’m really accepting all handicaps in order to refute accusations of favoritism here.)
And the results?
It’s a blowout.
I don’t care what people tell you about how English is really mostly something else. It’s not. Not only is our grammar solidly Germanic, but words of native English wordstock still form the base of the vocabulary we actually use.
Q.E.F. (quod erat faciendum, which was to be demonstrated)
Or, if you prefer it in native English words: That’s how it is, folks.
Latin 29%Other linguists are more generous; they estimate that as much as one-third of English words are derived from Old English. But is this really true?
French or Anglo-Norman 29%
Other 16%
Germanic only 26% (of which native English words would be rather less than that)
Well. There is no doubt that there are many, many (MANY) words from all kinds of sources in English. Modern English has the largest wordstock of any language in the history of the world. But those are just the words that are available to us. Nobody uses all of those words in their personal expression, nor does every English reader recognize all of them. The question then is, does the way we talk and write actually bear out the idea that words of English origin form a minority of the vocabulary we actually use?
I’ve done this kind of analysis before, but it bears repeating. Looking for a handy sample of ordinary English, I took down my copy of Aesop’s Fables and turned to this familiar tale.
The Goose that laid the Golden Eggs
A Man and his Wife had the good fortune to possess a Goose which laid a Golden Egg every day. Lucky though they were, they soon began to think they were not getting rich fast enough, and, imagining the bird must be made of gold inside, they decided to kill it in order to secure the whole store of precious metal at once. But when they cut it open they found it was just like any other goose. Thus, they neither got rich all at once, as they had hoped, nor enjoyed any longer the daily addition to their wealth.
Much wants more and loses all.
Before we go on to further analysis, we have to answer: What counts as a word? I’m counting egg, eggs as one word; likewise, getting, got. On the other hand, I’m counting gold, golden and day, daily as being two different words in each case. But it’s not just different forms to consider. Do we count each headword (lemma)? Or do we count repeated words in the sample (total occurrences)? I will count total words both ways so as to give fullest accounting.
The results:
Latin, Old French or Old French < Latin = 13 of 76 headwords = only 17.1%It gets worse for Classical sources when you count occurrences rather than headwords
Meanwhile, Old Norse = 4 of 76 headwords = 5.3%
Old English or Middle English = 59 of 76 headwords = 77.6%
Total headwords of Germanic origin = 82.9%, which thus skunks those of Classical origin pretty severely
L, OF, or OF < L = 13 of 112 occurrences = 11.6%
ON 12 of 112 occurrences = 10.7%
OE, ME 87 of 112 occurences = 77.7%
Total Germanic origin (OE + ON) 88.4%
Now, wait a minute, you say. That’s not fair! You chose a children’s book! I would dispute that Aesop’s Fables is a children’s book, but it is written in very simple, direct language. We call that being effective in the writing trade. But, okay, we’ll take another sample of similar size from a non-technical source.
I went to my bookshelves and pulled down God in the Dock, by C.S. Lewis. He has a short meditation therein, which goes as follows.
’Yes,’ my friend said. ‘I don’t see why there shouldn’t be books in Heaven. But you will find that your library in Heaven contains only some of the books you had on earth.’ Which?’ I asked. ‘The ones you gave away or lent.’ ‘I hope the lent ones won’t still have all the borrowers’ dirty thumb marks’ said I. ‘Oh yes they will,’ said he. ‘But just as the wounds of the martyrs will have turned into beauties, so you will find that the thumb-marks have turned into beautiful illuminated capitals or exquisite marginal woodcuts.’Okay, so here we have a professor of English literature just being himself, using words as he would use them. You can see a lot more words here of Classical or Romance origin, okay? So let’s see how the analysis turns out.
What counts as a word? Besides counting headwords and total occurrences separately, I’m treating beauty, beautiful as two words. Contractions are counted as single words (OE had contractions, too, witness nolde for ne wolde). Hyphenated words count as two, but compound words count as one. I’m giving woodcuts to OE, despite “cut” being of ON origin; however, I’m crediting turn to Latin, even though it was a very early borrowing and thoroughly nativized as OE tyrnen. Just trying to be fair: no special pleading for native words.
The results:
L, OF, or OF < L = 11 of 60 headwords = 18.3% (thought it’d be higher, didn’t you?)Even in someone with Classics and Philosophy backgrounds from Oxford, the native English word count is extraordinarily high. How ‘bout that?
ON 4 of 60 = headwords = 6.7%
OE, ME = 45 of 60 headwords = 75% (!)
Total headwords of Germanic origin = 81.7%
L, OF, or OF < L = 12 of 94 occurrences = 12.8%
ON = 10 of 94 occurrences = 10.6%
OE, ME = 72 of 94 occurrences = 74.5%
Total Germanic origin = 87.2%
But, but, but . . . Okay, I’ll give the denigrators of Old English one more chance.
I went to my shelves and took down a little book called The Curmudgeon’s Guide to Getting Ahead, by Charles Murray – a social scientist and purveyor of educated gobbledygook. Surely, we ought to see what English looks like today from him, ja? From the Introduction:
The transition from college to adult life is treacherous. It is easy for new graduates to go directly to graduate studies that lock them into careers they will come to regret. Those who go directly to work are often in their first real jobs, not knowing how an office environment operates or how their supervisors are evaluating them. They often are emerging from universities that have ignored what used to be a central theme of university education: thinking about what it means to live a good life.
What counts as a word? Well, different grammatical forms of the same word are counted together. This handicaps Old Norse (they, their, them). (N.B. That is here not the singular of those, so these are two different words.) Meanwhile, graduate(s) is counted twice, once as an adjective and once as a noun. (I’m really accepting all handicaps in order to refute accusations of favoritism here.)
And the results?
Latin = 16 of 58 headwords = 27.6%English wins again, handily, and it gets worse:
French = 7 of 58 headwords = 12.1%
Which means the total Classical origin = 39.7%
Meanwhile, ON = 1 of 58 headwords = 1.7%
OE, ME 34 of 58 headwords = 58.6%
So, total Germanic origin = 60.3%
Latin = 19 of 80 occurrences = 23.8%
French = 7 of 80 occurrences = 8.8%
Total Classical origin = 32.5%
ON = 6 of 80 occurrences = 7.5%
OE, ME = 48 of 80 occurrences = 60%
Total Germanic origin = 67.5%
It’s a blowout.
I don’t care what people tell you about how English is really mostly something else. It’s not. Not only is our grammar solidly Germanic, but words of native English wordstock still form the base of the vocabulary we actually use.
Q.E.F. (quod erat faciendum, which was to be demonstrated)
Or, if you prefer it in native English words: That’s how it is, folks.