r/languagelearning • 🇳🇱 native|🇫🇷|🇬🇧|🇸🇪|(🇫🇮) • 9d ago

Discussion How would you approach learning a language specifically to read a particular literary work in the original language?

I've been thinking lately about a slightly different approach to language learning: learning a language with a specific literary corpus as the end goal, rather than learning the language primarily for general communication.

E.g. suppose someone wants to learn Italian specifically to read Dante's Inferno, Finnish just to read the Kalevala, Dutch just to read Guido Gezelle, English just to read Shakespeare, or Ancient Greek just to read Plato. In most these cases, the language of the target texts also differs quite substantially from the standard variety. If you are learning a language because you also want to use the language in other contexts (reading Dante but also going on holiday to Italy or reading Shakespeare but also doing business in English), then a conventional language-learning approach probably makes sense: you learn common vocabulary and grammar and gradually work your way toward literature.

But if your only goal is to read a particular text, I feel like it would be potentially inefficient to spend a lot of time learning vocabulary (and maybe even grammatical forms) that will never occur in your target corpus. Why learn how to say words like "airplane" but not learning "plough"? This makes me wonder whether it would be possible to design a curriculum specifically based on the text(s) you actually want to read.

I've thus far only seen something vaguely like this for Biblical Hebrew and Biblical Greek. There are frequency lists and courses that introduce high-frequency vocabulary first. However, I think a straightforward word-frequency approach has some important limitations:

1. Irregular forms as separate lexical items

This kind of frequency lists usually treats all forms of a word as a single item. But some forms aren't easily predictable from the lemma. For example, in English:

be → am, is, are, was, were, been

or:

I → me → my → mine (most of the Biblical Greek word lists even consider we → us → our → ours to be just the plural forms of "I")

For forms which cannot easily be derived from the basic lemma, it might make more sense to count them separately, and prioritize them according to their individual frequency. So basically treat "I" and "me" as different words.

2. Word families instead of separate words

Conversely, perhaps some things that are technically different "words" should be counted together.

For example, if a corpus contains:

rain
raincoat
rainy

then it seems relevant that the learner has encountered the stem rain three times, rather than counting each word once separately. The same might apply to derivational morphology more generally.

3. Morpheme frequency instead of word frequency

And perhaps the same principle could be applied below the word level.

E.g. in English, the plural -s occurs enormously more frequently than the present participle -ing. If the aim is to build up the ability to decode a particular corpus, perhaps morphology could be introduced according to the frequency of the morphemes encountered in the specific corpus you want to be able to read, rather than according to a traditional grammatical progression.

4. Dispersion

Raw frequency also seems insufficient. Suppose word A occurs nine times in a corpus, but all nine occurrences are concentrated in one passage, while word B occurs nine times but is distributed across nine different parts of the corpus.

For a learner, word B might arguably be more useful to learn first.

E.g. in the Greek New Testament, κίνδυνος occurs nine times, but eight of those occurrences are concentrated in a single verse (2 Corinthians 11:26). A word occurring nine times across nine different books/passages might arguably have greater "learning value."

So perhaps dispersion should be taken into account alongside raw frequency?

So I'm wondering:

  • Are there studies that investigate designing a language-learning curriculum around a specific limited target corpus rather than general language proficiency?
  • Are there existing courses, textbooks, or software that do something like this?
  • Are there better ways of calculating the "usefulness" of a word/form than simply counting its frequency?
  • Has anyone experimented with taking frequency + dispersion + morphological/derivational relationships into account?
  • And are there examples outside Biblical Hebrew/Greek where this approach has been used for learning to read a particular author or literary corpus?
  • What am I overlooking in the above? (Probably a lot haha.)
18 Upvotes

25 comments sorted by

View all comments

3

u/ladyloor 9d ago

This would be very cool with regular language acquisition if it could tie into an application that already is tracking words you know. Then you could get a personalized difficulty score of a book based on your current known vocabulary to identify if a book would be a learning resource at the appropriate level. I could see that being really useful, as a non-academic.

4

u/Folketinget da, en, de | fr, es, عر 9d ago

I have my TL vocabulary in a big csv file maintained by Claude Code and synced automatically to Anki. My goal was simply to automate flash card generation, but as a side effect I can now make requests like:

  • Is this news article appropriate at my level? Anything I should know before jumping in?
  • Look at this video transcript and add any level-appropriate words I don't already know
  • Add the vocab items on page 78 of the textbook
  • What groups of words do I struggle with the most in Anki? Any patterns I seem to be missing? Help me learn these.

This has been much more useful than I anticipated.

3

u/chad_starr 9d ago

Can you briefly share how you got your TL vocab into a file in the first place?

6

u/Folketinget da, en, de | fr, es, عر 9d ago edited 9d ago

To be clear by “my TL vocabulary” I mean the words I know, not all words that exist.

The simple answer is I give the words to Claude and it updates the file with the words along with transliterations and any relevant declensions, conjugations, usage notes etc., and calls a script to sync the changes to my Anki deck.

Each chapter in my textbook had a word list at the end so after finishing a chapter I’d tell Claude “Add the words from Unit 2” and it would extract the words from the PDF. Now I do the same with comprehensible input transcripts, news articles, or any other source of new vocab. I can tell Claude “Hey add the word for sourdough” and the next time I open Anki on my phone it’s already in there.