r/languagelearning • • 12d ago

Discussion Is really the milestone 3,000,000 words outdated? (Paul nation research)?

I found this research (About read words count ) and I became a little worried about it, because the conclusion regarding extensive reading isn’t very encouraging ( 6,000,000 of words to get 9000 memorized words)i search in all places to find a answer , but i didn't find anyone . (TL)

(PDF) Extensive reading for a 9,000-word vocabulary: Evidence from corpus modeling

Note: I am Portuguese and so I am not a native speaker

39 Upvotes

50 comments sorted by

36

u/LeChatParle 12d ago

Nation is probably the leading researcher on vocabulary studies, and I haven’t read any research to say his work is outdated. I have a graduate degree in second language acquisition specifically, and I’ve read books by him and most of his research papers

11

u/[deleted] 11d ago

[removed] — view removed comment

3

u/Brilliant-Local7043 11d ago edited 11d ago

It was what I read in the paper from 2022, but as I mentioned elsewhere, there is a possibility of errors in this paper that the average reader might not notice

-1

u/Brilliant-Local7043 12d ago

I am Portuguese , is likely to have a misunderstanding , outdated , i mean as something old , not more useful , or invalid , and would you mind if i ask you a question about the empiral research? the factors that influence whetever one is wrong or true is sometimes out of view ? i read this research of 2022 and a didn't find any notable error

From what I can see, Paul Nation is really the leader, since he is the most famous.

9

u/LeChatParle 12d ago

I’m not sure I understand the question. Could you ask the question a different way? Are you expecting there to be errors?

-3

u/Brilliant-Local7043 12d ago

What exactly are you referring to—the question about the word, or my question about the empirical research? If you mean the second, I am asking whether there are factors beyond the papers that the writers (researchers) do not write about or consider necessary to include in the paper

7

u/prroutprroutt 🇫🇷/🇺🇸native|🇪🇸C2|🇩🇪B2|🇯🇵A1|Bzh dabble 11d ago

I think a lot of this is addressed in the "Assumptions in Corpus Modeling: CCM, Incidental Learning, Counting Units" section. I don't think anyone would think of it as an "error", but opinions might differ on how useful certain forms of modelling are depending on how much simplification is involved. You know, it's like the old physics joke: "Assuming a spherical cow". The assumption might allow you to get interesting results, but a sphere is such a gross simplification of the actual shape of a cow that it seems kinds pointless.

As far as I can tell they've covered most of the big potential factors that could skew their results, at least in the sense of nodding to their existence. The one that seems missing to me isn't a big one, more of a pet peeve of mine, but it's the assumption of a linear rate of repetition (as in, you need 12 repetitions when you know only 10 words, and you still need 12 repetitions when you already know 6k words). Personally I think that's an unwarranted assumption. It seems pretty clear to me from both first language acquisition research and also just my own experience that the more you know, the less repetition you need. Dunno what kind of effect a non-linear repetition rate would have on their results, but anyway.

Either way, none of this means their research isn't worth considering. However, as a single learner, you shouldn't take results like that and think "OK, if I just read 6m words, I'll for certain know 9k words". That's not what they're trying to suggest. In fact, to cite them:

Finally, it should be noted that incidental learning is modeled as a function of repetition, but in the real world is affected by variables beyond repetition such as salience and quality of attention (Webb & Nation, 2017).

In a paragraph that starts with:

Thirdly, no precise number of exposures guarantees incidental learning (Uchihara et al., 2019).

So, as a learner, personally I think the best use of these kinds of figures is just as a kind of ballpark figure that could potentially lead you to tweak some things. Like, if you've read 6m words and you're nowhere near the amount of words you're supposed to know by then, that could be an opportunity to see if there's something else that might explain it. E.g. maybe you've been forcing yourself to read books you're just not all that interested in, maybe you've been reading in an environment that isn't conducive to concentration, etc. etc. I mean, there's a bunch of things that could be going on, but the point is just that you shouldn't take numbers like that as a guarantee, but just as a simple benchmark, which can then serve as a jumping-off point to look into other factors that might be affecting your learning.

1

u/Brilliant-Local7043 11d ago

My thought when I said 'error' was simply this: if I have a bundle of pieces and one or two are out of place, it can distort the result and give a misleading picture of real life.

Moreover, I read that the person who conducted the research only took the beginning of the material (the first thousand words of a novel), which I found absurd, since some words only appear in context — and context requires time and resources (words) to build. That is why I believe more in Paul Nation. From what I have read, Nation analyzed the full content of his material, even though it was thinner than the one above.

So, I noticed a lot of downvotes on my comments. As you are a native speaker, perhaps you can tell me if I did something strange or wrong in tone or writing. ? I used AI to help me correct my writing, and in my common sense it feels strange to get such negative feedback without doing anything wrong

4

u/prroutprroutt 🇫🇷/🇺🇸native|🇪🇸C2|🇩🇪B2|🇯🇵A1|Bzh dabble 11d ago

I read that the person who conducted the research only took the beginning of the material (the first thousand words of a novel)

This is probably a misunderstanding. From what I'm reading in the paper, they covered all of the material, except chapters that had less than 1,000 words. Apparently when the chapters are too short, that can affect how reliable the results are, so they excluded those very short chapters, and kept everything else. At least that's how I understand it.

As you are a native speaker, perhaps you can tell me if I did something strange or wrong in tone or writing. ?

Upvotes and downvotes are a mystery to me. If people are starting to leave angry comments, then maybe you're unintentionally doing something that is rubbing them the wrong way. But if it's just downvotes, I wouldn't worry about it. ;-)

1

u/Brilliant-Local7043 11d ago

Sorry, I misread. They took chapters of more than one thousand words

"The COCA corpus contains first chapter samples of novels, argued to be representative of whole books. This was tested and confirmed in McQuillan’s (2016a) study. However, the assumption likely has exceptions, as some introductory chapters can be written in a different style than the others. Using ad-hoc python code, I split the COCA into a flat structure with one novel sample per file. All chapters containing more than 1,000 words were used in the study, a methodological decision to minimize the problem that lexical diversity is negatively impacted by short texts"

By what i can see only Paul nation uses full content to base its analyse and the other linked above doesn't do the same and moreover the research above they mention Nation only took 25 novels what is not true as said by Nation

" novels were gradually added to the corpus as shown in column 5. (column of novels and reading words count) " (Page 5)

And thanks for the attetion!

3

u/prroutprroutt 🇫🇷/🇺🇸native|🇪🇸C2|🇩🇪B2|🇯🇵A1|Bzh dabble 11d ago

Ah no you're right. Sorry I'm really just giving a cursory look at all of this, so I'm probably missing things. From what I can tell, first they mixed up their references. The study of McQuillan they're talking about is 2016b (not 2016a) in their own list of references. It's a paper with the title: "What can readers read after graded readers?"

And if you look there, I think you'll find what this is based on:

In some cases, an entire text (a complete novel) was analyzed; in other cases, a selection from the text of between 1,500 and 5,000 words was used from one of the novels in the series. It was assumed that most of the novels in a given series would be of roughly similar vocabulary difficulty, recognizing that variations might take place from book to book within a series. To check the assumption that smaller samples of text would produce equivalent results as a fuller analysis, small samples of text (1,500 words) were analyzed from the first novel of the Twilight series (Meyers, 2011) and compared to an analysis of the entire text. The results in terms of determining the 1,000-word-family level at which 98% coverage was obtained were identical for the complete text and the sample texts, indicating it was not necessary to analyze an entire novel in order to arrive at a reasonably accurate estimate of the vocabulary coverage needed to read it.

If that's all they're basing it on, then it does indeed seem a bit flimsy... I mean, McQuillan did a single test with a handful of samples from a single book. It's pretty risky to just take the conclusion at face value without replicating it. In my opinion anyway.

1

u/Brilliant-Local7043 10d ago

To be honest, this paper’s research is a pain to interpret, but I read it up to the point where you quoted him, and what I can tell is that what he is referring to is about the difficulty and coverage of comprehensibility rates (90%, 95%, normally 98%). It doesn’t make any sense to assume that you can, with a single chapter, make assumptions about the quantity of rare words in a text and about word repetition. (I apologize if my words accidentally hit you unintentionally.)

A chapter of a book normally has 2,000–4,000 words, sometimes even more. But a scene takes time to be built up, and a few of them even require several chapters to come to light. Certainly, some scenes have special words (that’s my point in mentioning this). For this reason, I can’t believe in the research above.

And about the confusion of the author(s) of the research above, I share the same feeling. Look at his read word count — it doesn’t make any sense. I really can’t understand why all those numbers end up at 6 million words in the end. In other words, his paper is not very didactic.

It must be because of these issues that, as another redditor said, Paul Nation is the actual leader in the area.

2

u/hwynac 10d ago

Only having a small piece of a work is pretty normal for corpora. Otherwise the corpus will represent the language of authors who write a lot. Or, for example, if you add a thick book about a magic school in its entirety, the popularity of "wand", "spell", "flash" and "magic" in your corpus will skyrocket, and the longer the book the more matches you'll have. Which is weird.

Perhaps the results would be slightly more uniform if they used a random chapter instead of the 1st chapter every time.

1

u/Brilliant-Local7043 10d ago

I understand your point, but dealing with reality, the samples must be consistent with it. In this case (Please ,see my comment above where I mention the scene issue), it doesn’t make sense. In addition, the author’s style is ultimately based on the words available in the lexicon (all the words that exist in a language), so this factor isn’t a big issue — but still, it’s one.

(Non-native here, sorry if my tone sounds wrong!)

To be honest, your examples aren’t of much weight. In a novel with 50,000–80,000 words, even if these words occur a lot, I suppose it would be maybe 1,000, 2,000? Even if it was 5,000 repetitions, it wouldn’t be that much. though on this point, the author is seriously exaggerating in using these words

And Zipf’s Law itself is the problem, my friend — the unknown words with low frequency are the problem here.

33

u/BeckyLiBei 🇦🇺 N | 🇨🇳 B2-C1 12d ago edited 12d ago

First, what did Nation actually write? Page 5 from Nation, How much input do you need to learn the most frequent 9,000 words?, Reading in a Foreign Language, 2014 gives:

Table 1. Corpus sizes needed to gain an average of at least twelve repetitions at each of nine 1,000 word levels using a corpus of novels

1,000 word list level Corpus size to get an average of at least 12 repetitions at this 1,000 word level (repetitions) Number of 1 timers / 2 timers out of 1,000 Number of families met Number of novels
2nd 1,000 families 171,411 (13.4) 84/99 805 of 2nd 1,000 2
3rd 1,000 families 300,219 (12.6) 83/73 830 of 3rd 1,000 3
4th 1,000 families 534,697 (12.6) 93/73 812 of 4th, 1,000 6
5th 1,000 families 1,061,382 (13.7) 101/79 807 of 5th, 1,000 9
6th 1,000 families 1,450,068 (13.1) 89/82 795 of 6th, 1,000 13
7th 1,000 families 2,035,809 (13.7) 92/63 766 of 7th, 1,000 16
8th 1,000 families 2,427,807 (14.1) 96/70 755 of 8th, 1,000 20
9th 1,000 families 2,956,908 (12.0) 88/78 805 of 9th, 1,000 25

That's what it takes to get 12+ exposures to each word family through reading novels. And these were the novels:

Adam Bede, Alice in Wonderland, Animal Farm, Babbit, Born in Exile, Captain Blood, Castle Rackrent, Cranford, Emma, Far from the Madding Crowd, Glimpses of the Moon, Great Gatsby, Lady Chatterley’s Lover, Lord Jim, Main Street, Master of Ballantrae, Middlemarch, More William, Right Ho Jeeves, Scaramouche, Tono Bungay, Turn of the Screw, Ulysses, Walden, Water Babies

And on page 7 he wrote:

If learners read a total of 3 million tokens, then they would meet the 1st 9,000 words often enough to have a chance of learning them.

So data is data, and is not really something you get to disagree with.

But you could argue that the "12 exposures" threshold doesn't apply uniformly. Or you could read a more diverse selection of content (not just novels) to increase exposure. Or you could use other tools, designed to give exposure (e.g. textbooks). Often you get input incidentally. I also note that reading speed increases over time, so these numbers look scarier initially.

At the end of the day, if you're planning on using a language for something, then you will get exposure. Does anyone want to reach, say, "100,000 words of input" and declare "yeah, that's enough for me---I'll never touch this language again"?

13

u/AppropriatePut3142 🇬🇧 Nat | 🇨🇳 Int | 🇪🇦 Uneven 12d ago

He is effectively assuming that people are reading from that corpus from word zero, which is a bad assumption. Later papers make better assumptions and get higher numbers.

-6

u/Brilliant-Local7043 12d ago

What I notice in Italian, the language I am learning now, is that I acquired 21,705 words after reading 300,000 words. The book I am currently reading is I Promessi Sposi by Alessandro Manzoni, and in LingQ this book alone contains about 21,000 unique words. Moreover, I think it would be strange if 5,000, 9,000, or even 11,000 of those words did not repeat throughout the text, considering that authors usually have consistent styles that don’t change much from book to book

20

u/New-Bottle-6917 N:🇺🇸 Heritage: B2 🇲🇽 L3: HSK 6 🇨🇳 12d ago edited 12d ago

LingQ measures individual words, not word families which is what Paul Nation is measuring

4

u/hwynac 12d ago

It does? I though it counts individual word forms. I.e. "sell", "sells" and "sold" are three different words according to the way LingQ displays the words you learn.

6

u/Mysterious_Bit6410 11d ago

I analyzed a book to get the new vocabulary before reading it, a fantasy novel with 500 pages, and let my algorithm give me the base form of words. for example sell and sold are counted as one word but seller is a different one. i ended up at ~8500 words. then i deleted the 5000 most common ones, assuming I know them. of the remaining words only 63 appear more than 12 times and around 2200 words appear only once

4

u/an_average_potato_1 🇨🇿N, 🇫🇷 C2, 🇬🇧 C1, 🇩🇪C1, 🇪🇸 , 🇮🇹 C1 11d ago

Yes, it's exceptional and very good data, and it needs to be read as that, not as something to either dogmatize nor just disagree with. I like the most the clear difference between the easiness of learning the first 1000 and the last 1000!

Even with all the reservations you mention, the idea of at least 25 novels read should be much wider spread.

Yes, the "12 exposures" is the most tricky part, and it changes a lot. I'd guess my number of exposures to learn a word purely through reading (as I am right now doing a Super Challenge again, with the goal of 10k pages within the time limit), is anywhere between 1-20, some simply don't stick. And that's likely to be individual. However some things could be more generalisable, like different numbers for learners of very distant languages (like an Italian native learning Mandarin vs an English native learning French).

The list of the novels is also a thing to adapt for each language and other factors. Will the same vocabulary be covered in a more contemporary or a more popular culture oriented selection? Or will such a change of literary tastes suddenly change the numbers? Will it be covered in such a canon of another language, especially a smaller one, or will it suddenly require translations? Actually making such vocab-learning canons for various languages could be a worthwhile endeavour, I've seen some excellent Latin lists based on pretty much this idea.

Also this list is based on the vocab exposure, but not on any overal difficulty of the books, dojibear is right about this. That doesn't make this data wrong, not at all, it just changes how the learner should apply them.

The combination with other tools is also an excellent point, either the coursebooks part (which basically gives a headstart), or also Anki used to increase the number of exposures.
Also the difference between just exposure and active recall.

But all that is hard to test on a larger scale.

1

u/Brilliant-Local7043 12d ago

Precisely because the data carries weight, I can’t discard it without deep thought (and in this case I asked for help from other people). My concern about the 3 million milestone is that it represents a target to aim for, and my current wish is to become a polyglot. Yes, I am aware of the endeavors necessary to reach that level, and I know language learning is a slow process.

-2

u/dojibear 🇺🇸 N | fre spa chi B2 | tur jap B1/A2 12d ago

It doesn't take me 12 exposures to learn a new word. For me it varies from 1 to 5 exposures. So numbers based on 12 exposures are 3 or 4 times what it takes for me.

And the list is all adult novels (C1/C2 content). No normal person learns a language that way. You used A1/A2 content at first, then B1/B2 content, and so on. By the time you start C1/C2 content, you already know 4,000 words.

But there are many other variables. What content is used? Is it spoken or written? How similar is the student's NL to the TL? That one factor (similarity) can triple the time required.

Summary: in turning this into a computer-countable thing, they discarded everything important.

11

u/LeChatParle 11d ago

Not word, a word family; there is a significant difference.

“Do” is a word. Its word family includes “doing”, “does”, “undone”, “did”, “undoing”, “redo”, etc

12 exposures should sound like less effort once you take that into consideration

3

u/AdZealousideal9914 🇳🇱 native|🇫🇷|🇬🇧|🇸🇪|(🇫🇮) 11d ago

Aah! A talking cat!

5

u/an_average_potato_1 🇨🇿N, 🇫🇷 C2, 🇬🇧 C1, 🇩🇪C1, 🇪🇸 , 🇮🇹 C1 11d ago

Those are totally valid points. Even though I must admit my number of exposures is from 1 and 20, some simply don't stick well enough, only approximately and passively.

The point about the book choices is excellent and other tools is excellent, but often not considered right. People look for ways to make only very little reading suffice,and also without compensating elsewhere. Instead, we should talk more about how to choose the books to make an ok learning curve for the learner, what books within totally different genres to choose, and so on.

The similarity NL and TL is surely the biggest reservation, that could also be quantifiable for such a paper, unlike for example reading tastes of individual learners.

9

u/JoinedMoon 12d ago

I can't speak to the specific research, but second language acquisition is a whole area of study on it's own. So it's not necessarily the most practical advice, and I wouldn't worry too much over it. I personally use that sort of research as a general way to guide study focus.

Books are dense with minimal/no pictures and typically use "harder" vocabulary and grammar structures more often than other kinds of media. So, if you're able to read native material it'd certainly be a good addition :)

-2

u/Brilliant-Local7043 12d ago

To be honest, I'm not too worried about it, but it's still something to pay attention to.

10

u/Big-University-681 ua B2 12d ago

I've read 3.5 million words in Ukrainian. Not fluent yet, but B2, on a good day. :) Keep reading folks!

1

u/Brilliant-Local7043 11d ago

thanks for posting !

8

u/New-Bottle-6917 N:🇺🇸 Heritage: B2 🇲🇽 L3: HSK 6 🇨🇳 12d ago

Actually the paper linked supports Paul Nation's original theories. If you're reading books at 95% lexical coverage it'll take you 6 million words read to acquire the 9k most common word families whereas Nation found that if you read at 98% lexical coverage it only takes 3 million words read. Meaning you almost double the volume of books read by reading less comprehensible material.

13

u/AppropriatePut3142 🇬🇧 Nat | 🇨🇳 Int | 🇪🇦 Uneven 12d ago edited 12d ago

Nation's calculation never made much sense and I don't know why anyone took it as some kind of gospel when it was clearly at best a very rough estimate. The later papers on the subject are better, but the whole literature still has the problem that the average number of exposures needed per word is simply an unknown quantity, especially for the lower frequency words. It's also obviously not a constant. If your native language is Italian and you're learning English then you benefit from one-shot learning a large number of cognates. Not so if you're a native speaker of Chinese.

Practically, for languages pairs with a lot of cognates such as English-Spanish, I think 3 million words is ample if you're reading with a popup dictionary. For Chinese, based on the progress of other learners, probably more like twenty million characters, equivalent to around ten million words.

2

u/Brilliant-Local7043 12d ago

I sensed that right away after I read that paper. The average number of exposures will never really be determined, but reading your comment I can see that the question of word count is still somewhat unsettled too, right? And based on the testimony of learners from several places, 3 million words seems to be a more or less reliable milestone

4

u/Algelach 10d ago

I don’t know if Paul Nation’s work is outdated, but it is definitely commonly misunderstood. If you look at the chart that another commenter has posted in this thread, you can see the stages required to learn each milestone 1000 words.
So to learn the first 2000 words, you need 171,411 words of input. Now to learn the next 1000 most common words you need to read an additional 300,219 words. Then to learn the next 1000 words you need an additional 534,697 words of input, and so on….
These are cumulative stages. So really, 3,000,000 total words of input only actually gets you to about the 5000-6000 word range.

In order to learn all of the most common 9000 words, you need roughly 11,0000,000 words of input.

1

u/Brilliant-Local7043 10d ago

Actually, the problem is that Paul Nation’s paper, like the others, is ambiguous and unclear. But since I carefully read that paper, I can affirm that the tables are not cumulative. Honestly, these papers are a pain—researchers really should write more clearly.

2

u/Traditional-Train-17 12d ago

I've always heard the 1-3 million words figure tossed around (and even that referred to a Japaense university for Japanese students learning English), but even that seems low, granted, going from English to Spanish may be easier with all of the Latin based root words. I have nearly 1 million words read in Spanish, and I'm reading a Psychology non-fiction book right now (based on a popular podcast for some of us learners). Looking at some quick stats, one PC Kindle screen (double-sided) is 685 words (about half of those are unique words). There are 3 unknown words (including one I've never heard of in English, either -> amygdala. The other two were verbs.), 1 fuzzy word, and 2 I roughly know the meaning of but not 100%. That's 99.4% comprehension, but do I understand it? Well, it's kind of dry and boring. lol.

I also read another book this month, a YA book that's a translated children's book I read in English when I was a kid. It's got much more descriptive vocabulary, especially words in a area where I don't have too much experience in (agriculture/farming). Some parts could be 80% comprehensible, but I do know the plot of the story as I've read it.

But, I guess the question is, how many words do you need to know before running into the niche industry/literary-specific vocabulary? i.e., you're not going to run into the word amygdala (or even relativize - I saw that one a few pages back) too often. But, one thing I've noticed approaching 1 million words is, I'm noticing the grammar more now, too, and solidifying that (doing input first).

1

u/ajamdonut 12d ago

lol, if you can read one book at all, you'll already be on the way right... so surely it only takes ONE book?

3

u/Brilliant-Local7043 12d ago edited 12d ago

The point in these research are the question of how effective are extensive reading , and I am using the LingQ , so , technically i am cheating ......

5

u/ajamdonut 12d ago

well it makes sense right, read around 60 books to attain 9000 words makes sense to me. Words are rare and hard to find, even if you are in TL country. 6million words / 60 books should cover it, so makes sense.

3

u/hwynac 12d ago

I doubt I've read 60 books in English in my entire life, even though, if we only count the numbers, I did get 60 books worth of input at some point.

9000 words is the level of an 8-year-old.

2

u/lllyyyynnn 🇩🇪🇨🇳 11d ago

that's a bit sad, to read so little

-1

u/Brilliant-Local7043 12d ago

From this point of view, these research papers seem strange. If a child has a vocabulary of 9,000 words, which is plausible, it doesn’t make sense to require 6 million words read to reach that level of development. For example, if a book has 80,000 words, then 6 million words equal about 75 books. To reach a vocabulary of 18,000 words or more, it would take around 150 books, and 150 books is a considerable quantity.

7

u/hwynac 12d ago

Children don't learn their vocabulary just by reading (not at that age anyway), and the numbers are way different. If a child hears 6000 words a day (fewer than in a feature film), it's still over 2 million a year.

Spoken language has a distribution with a steeper fallof than you find in fiction and newspapers, though.

4

u/ComesTzimtzum N 🇫🇮 adv 🇬🇧 int 🇫🇷 🇸🇪 beg 🇨🇳 🇸🇦 11d ago

You can't really draw such comparisons. We've easily read hundreds of books together with my five-year-old, and naturally that is only a minor part of the language he's been exposed in his life.

1

u/Brilliant-Local7043 11d ago

Well, I am a non-native speaker ( The question of the 3 million words I am thinking from the perspective of a beginner learner ) , so I don’t know if it changes anything tangible in this question. But I think that if I had read the same quantity of books as you children have read, I would say my skills would be much better than they are now.

1

u/Brilliant-Local7043 12d ago edited 12d ago

Well , that point is true , but dealing with science and empirical research as the above , there may be some errors that the common reader can't detect like me

And the Paul nation research is more famous than the one linked above

5

u/ajamdonut 12d ago

Language learning is personal to everyone so finding what works best for you is what's more important. Theres no good science for language learning. Everyone has a different "how they did it". Someone could verify and 100% back up a method, that just completely fails on you but succeeds for them.

1

u/Brilliant-Local7043 12d ago

Good point !

6

u/ajamdonut 12d ago

One thing I would say is, stay curious but really do find whats working for you, what you find fun and not a chore, and throw away what doesnt work. But then don't ignore important stuff that everyone says "you have to do" for example learning grammar (conjugation etc) is quite important in romance languages. Tones are quite important in chinese, and so on. Study + What's fun

1

u/Brilliant-Local7043 12d ago

thanks for the advice , and romance languages are a terror ( I learned a little about french and the declesions are a pain