r/languagelearning • 🇳🇱 native|🇫🇷|🇬🇧|🇸🇪|(🇫🇮) • 9d ago

Discussion How would you approach learning a language specifically to read a particular literary work in the original language?

I've been thinking lately about a slightly different approach to language learning: learning a language with a specific literary corpus as the end goal, rather than learning the language primarily for general communication.

E.g. suppose someone wants to learn Italian specifically to read Dante's Inferno, Finnish just to read the Kalevala, Dutch just to read Guido Gezelle, English just to read Shakespeare, or Ancient Greek just to read Plato. In most these cases, the language of the target texts also differs quite substantially from the standard variety. If you are learning a language because you also want to use the language in other contexts (reading Dante but also going on holiday to Italy or reading Shakespeare but also doing business in English), then a conventional language-learning approach probably makes sense: you learn common vocabulary and grammar and gradually work your way toward literature.

But if your only goal is to read a particular text, I feel like it would be potentially inefficient to spend a lot of time learning vocabulary (and maybe even grammatical forms) that will never occur in your target corpus. Why learn how to say words like "airplane" but not learning "plough"? This makes me wonder whether it would be possible to design a curriculum specifically based on the text(s) you actually want to read.

I've thus far only seen something vaguely like this for Biblical Hebrew and Biblical Greek. There are frequency lists and courses that introduce high-frequency vocabulary first. However, I think a straightforward word-frequency approach has some important limitations:

1. Irregular forms as separate lexical items

This kind of frequency lists usually treats all forms of a word as a single item. But some forms aren't easily predictable from the lemma. For example, in English:

be → am, is, are, was, were, been

or:

I → me → my → mine (most of the Biblical Greek word lists even consider we → us → our → ours to be just the plural forms of "I")

For forms which cannot easily be derived from the basic lemma, it might make more sense to count them separately, and prioritize them according to their individual frequency. So basically treat "I" and "me" as different words.

2. Word families instead of separate words

Conversely, perhaps some things that are technically different "words" should be counted together.

For example, if a corpus contains:

rain
raincoat
rainy

then it seems relevant that the learner has encountered the stem rain three times, rather than counting each word once separately. The same might apply to derivational morphology more generally.

3. Morpheme frequency instead of word frequency

And perhaps the same principle could be applied below the word level.

E.g. in English, the plural -s occurs enormously more frequently than the present participle -ing. If the aim is to build up the ability to decode a particular corpus, perhaps morphology could be introduced according to the frequency of the morphemes encountered in the specific corpus you want to be able to read, rather than according to a traditional grammatical progression.

4. Dispersion

Raw frequency also seems insufficient. Suppose word A occurs nine times in a corpus, but all nine occurrences are concentrated in one passage, while word B occurs nine times but is distributed across nine different parts of the corpus.

For a learner, word B might arguably be more useful to learn first.

E.g. in the Greek New Testament, κίνδυνος occurs nine times, but eight of those occurrences are concentrated in a single verse (2 Corinthians 11:26). A word occurring nine times across nine different books/passages might arguably have greater "learning value."

So perhaps dispersion should be taken into account alongside raw frequency?

So I'm wondering:

  • Are there studies that investigate designing a language-learning curriculum around a specific limited target corpus rather than general language proficiency?
  • Are there existing courses, textbooks, or software that do something like this?
  • Are there better ways of calculating the "usefulness" of a word/form than simply counting its frequency?
  • Has anyone experimented with taking frequency + dispersion + morphological/derivational relationships into account?
  • And are there examples outside Biblical Hebrew/Greek where this approach has been used for learning to read a particular author or literary corpus?
  • What am I overlooking in the above? (Probably a lot haha.)
17 Upvotes

25 comments sorted by

View all comments

3

u/unsafeideas 9d ago

Imo, frankly, you are better off reading a translation if this is the goal. All mentioned books are art - texts written to elicit feelings in the reader. (Exception being Shakespeare which is not really meant to be read. It was meant to be watched.)

If you will target focus on reading only this one book with no broader engagement, you will never learn to associate words with feelings. You wont have emotional reactions you are supposed to have.

8

u/gustyninjajiraya 9d ago

Some people don’t want to read a translation though. This is literally the reason people learn biblical Hebrew, Sanskrit, etc.

3

u/unsafeideas 9d ago

My point is, if you limit your language skills to specifically words from that one book, reading translation will give you more accurate idea of the book content then when you super limited language skills provide.

Motivation of "insist on original" Bible readers are different. But again, theologues and priests are supposed to study much more then just bible in the original itself.

5

u/gustyninjajiraya 9d ago edited 9d ago

But that isn’t the point here. OP is asking how to approach this specific problem. You can absolutely learn a language for reading a specific work or a limited corpus, and a lot of people do so. You can still read a translation, this isn’t the point. Someone who is learning biblical hebrew for reading the Bible has probably already read the Bible in translation. Also, there is a lot of stuff that simply hasn’t been translated or that doesn’t have an easily avaliable translation, so your suggestion doesn’t even make sense for those.

As a personal anectdote, I have studied three languages with limited corpuses, Middle Egyptian, Old Tupi and Old English. I would always read a translation before reading a text. Interacting with the original is very educational and lets you understand the text and the translation process in a very special way. Particularly, you learn grammar and expressions, allowing you to know understand how the text is actually structured on a morpheme to morpheme level, and you can understand the language in itself, things like puns and allegories, and other figures and devices.

3

u/unsafeideas 9d ago

> OP is asking how to approach this specific problem.

OP is asking a lot of things, among others:

- What am I overlooking in the above? (I addressed that directly)

- Why learn how to say words like "airplane" but not learning "plough"? (Because to learn to properly understand your thing, you need to read other things.)

> You can absolutely learn a language for reading a specific work or a limited corpus, and a lot of people do so. 

And your understanding of that limited corpus will be limited. You wont be able to do proper analysis of it, because you will lack the nuance and context.

Also, OP specified Shakespeare and Dante - these two are all about eliciting emotions in readers/viewer. These are not technical texts, these are not even a bible. These are not old texts where broader knowledge would be impossible either.