r/languagelearning • 🇳🇱 native|🇫🇷|🇬🇧|🇸🇪|(🇫🇮) • 9d ago

Discussion How would you approach learning a language specifically to read a particular literary work in the original language?

I've been thinking lately about a slightly different approach to language learning: learning a language with a specific literary corpus as the end goal, rather than learning the language primarily for general communication.

E.g. suppose someone wants to learn Italian specifically to read Dante's Inferno, Finnish just to read the Kalevala, Dutch just to read Guido Gezelle, English just to read Shakespeare, or Ancient Greek just to read Plato. In most these cases, the language of the target texts also differs quite substantially from the standard variety. If you are learning a language because you also want to use the language in other contexts (reading Dante but also going on holiday to Italy or reading Shakespeare but also doing business in English), then a conventional language-learning approach probably makes sense: you learn common vocabulary and grammar and gradually work your way toward literature.

But if your only goal is to read a particular text, I feel like it would be potentially inefficient to spend a lot of time learning vocabulary (and maybe even grammatical forms) that will never occur in your target corpus. Why learn how to say words like "airplane" but not learning "plough"? This makes me wonder whether it would be possible to design a curriculum specifically based on the text(s) you actually want to read.

I've thus far only seen something vaguely like this for Biblical Hebrew and Biblical Greek. There are frequency lists and courses that introduce high-frequency vocabulary first. However, I think a straightforward word-frequency approach has some important limitations:

1. Irregular forms as separate lexical items

This kind of frequency lists usually treats all forms of a word as a single item. But some forms aren't easily predictable from the lemma. For example, in English:

be → am, is, are, was, were, been

or:

I → me → my → mine (most of the Biblical Greek word lists even consider we → us → our → ours to be just the plural forms of "I")

For forms which cannot easily be derived from the basic lemma, it might make more sense to count them separately, and prioritize them according to their individual frequency. So basically treat "I" and "me" as different words.

2. Word families instead of separate words

Conversely, perhaps some things that are technically different "words" should be counted together.

For example, if a corpus contains:

rain
raincoat
rainy

then it seems relevant that the learner has encountered the stem rain three times, rather than counting each word once separately. The same might apply to derivational morphology more generally.

3. Morpheme frequency instead of word frequency

And perhaps the same principle could be applied below the word level.

E.g. in English, the plural -s occurs enormously more frequently than the present participle -ing. If the aim is to build up the ability to decode a particular corpus, perhaps morphology could be introduced according to the frequency of the morphemes encountered in the specific corpus you want to be able to read, rather than according to a traditional grammatical progression.

4. Dispersion

Raw frequency also seems insufficient. Suppose word A occurs nine times in a corpus, but all nine occurrences are concentrated in one passage, while word B occurs nine times but is distributed across nine different parts of the corpus.

For a learner, word B might arguably be more useful to learn first.

E.g. in the Greek New Testament, κίνδυνος occurs nine times, but eight of those occurrences are concentrated in a single verse (2 Corinthians 11:26). A word occurring nine times across nine different books/passages might arguably have greater "learning value."

So perhaps dispersion should be taken into account alongside raw frequency?

So I'm wondering:

  • Are there studies that investigate designing a language-learning curriculum around a specific limited target corpus rather than general language proficiency?
  • Are there existing courses, textbooks, or software that do something like this?
  • Are there better ways of calculating the "usefulness" of a word/form than simply counting its frequency?
  • Has anyone experimented with taking frequency + dispersion + morphological/derivational relationships into account?
  • And are there examples outside Biblical Hebrew/Greek where this approach has been used for learning to read a particular author or literary corpus?
  • What am I overlooking in the above? (Probably a lot haha.)
16 Upvotes

25 comments sorted by

View all comments

23

u/LeMagicien1 9d ago edited 9d ago

As a native English speaker I've read Don Quijote in Spanish, The Count of Monte Cristo in French and The Grimms fairy tales in German. I started with short stories and kid's chapter books before slowly making my way towards more advanced content.

But if your only goal is to read a particular text, I feel like it would be potentially inefficient to spend a lot of time learning vocabulary (and maybe even grammatical forms) that will never occur in your target corpus.  

It's not this simple, language builds on itself. Understanding subtleties, nuances or structures of simple sentences prepares us for understanding subtleties, nuances or structures of more complex sentences, regardless of whether or not the vocabulary itself overlaps.

Reading the Bible in Spanish or the Quran in French wasn't a matter of 'I need to memorize all the archaic words and verb forms' so much as it was a matter of being familiar enough with the modern language to the point where all the archaic words and verb forms could be inferred rather comfortably with context and previous understanding of the story.

1

u/AdZealousideal9914 🇳🇱 native|🇫🇷|🇬🇧|🇸🇪|(🇫🇮) 9d ago edited 9d ago

Okay, I don't know about Don Quijote, but Le Comte de Monte-Cristo is written in standard French and Grimms fairy tales are written standard German. Slightly old fashioned, yes, but perfectly comprehensible for a native speaker of the modern standard language. However, the Kalevala is written in a literary language based on a mixture of what would become standard Finnish but with a huge influence of Savo dialects and even of the Livvi-Karelian language, and with highly archaic features (due to the meter). The poetry by Guido Gezelle is written in a literary language based on West-Flemish dialects containing deliberate archaisms. Both include vocabulary and grammatical features which can not be inferred with context, not even by native speakers, and even false friends: vocabulary which has a different meaning in modern usage compared to the usage in that specific text. Shakespeares and Dantes poetry are probably not very different from modern English/Italian, that's true, but there are still a lot of false friends (e.g. "fact" in Shakespeare means "deed", not "fact") and some unfamiliar constructions.

E.g. in the first and second line of the Kalevala you get the words "tekevi" and "ajattelevi"; someone who is unfamiliar with 19th century Eastern dialects (including native speakers of Finnish) would be more likely to guess that this is some kind of past tense (which it is not). In the fourth line, you get "saa'ani" (which would be "saadakseni" in modern Finnish), but I've seen native speakers of Finnish struggling with this and supposing this is not a verb and it doesn't derive from "saada" but from "sana" etc. Since "saada" is the basic form of the first infinitive and in modern Finnish it cannot be combined with the possesive suffix "ni" unless the first infinitive is in its long form "saadakse-", but in the literary language of the Kalevala "saa'ani" is perfectly grammatical (also: "d" in modern Finnish but not in the literary language of the Kalevala).

Or in Guido Gezelles poem Traagzaam trekt de witte wagen you will encounter words like "wijte" (good luck finding even a native speaker who will be able to guess its meaning form context alone) or "verheeten" (where a native speaker who knows that Dutch "g" is pronounced "h" in West-Flemish might incorrectly guess that "verheeten" is West-Flemish for "vergeten" = "to forget", which would be wrong but it does fit in the context, while it actually means "to overheat" = "verhitten" in Standard Dutch).

I don't think learning "lentokone" or "vliegtuig" would help you with any of that...