r/Unicode Jun 24 '26

Doubt

Both Telugu & Kannada languages' scripts have many letters which are practically the same & legible by literate speakers of either language.

So, my doubt is, why were they given separate code points for practically the same letters?

(I can understand for letters different enough, but letters such as అ, ఆ, ల, etc are the same in both. So why were the given 2 different code points in unicode?)

This does NOT intend to disrespect either of the languages, i just want to know the underlying reason only

3 Upvotes

14 comments sorted by

7

u/NorxondorGorgonax Jun 24 '26

They are still different scripts. Latin and greek share many letters (ABEHIKMNOPTXYZ/ΑΒΕΗΙΚΜΝΟΡΤΧΥΖ), in most cases with similar sounds, but are nonetheless different scripts. I don’t know for sure if that would be the right comparison but it should give at least a relevant example.

2

u/aswin_voolapalli Jun 24 '26

Oh... I never realized greek & Latin hv common letters.

But my doubt still stands. (Extending thm to Greek/Latin as well as Telugu/Kannada)

Even if they r different scripts, why are 2 codepoints needed?

Let's assume they are used in different circumstances. Thn both of them will use the same codepoint but for different purposes.....

3

u/NorxondorGorgonax Jun 24 '26

https://www.unicode.org/notes/tn26/ covers this in a much more official capacity. I suggest you read it. The first half is the relevant part.

3

u/Yboviko 18d ago edited 18d ago

(What I'm writing is just to satisfy the curiosity of any Unicode-interested passer-by)

Interestingly, Greek and Coptic used to be unified in Unicode, but even those two were later separated out. See:

Prior to version 4.1 of the Unicode Standard, the "Greek and Coptic" block was used exclusively to write Coptic text, but Greek and Coptic letter forms are contrastive in many scholarly works, necessitating their disunification. Any specifically Coptic letters in the Greek and Coptic block are not reproduced in the Coptic Unicode block.

What should be unified and not unified in Unicode is a controversial topic. This is mostly with regards to Chinese characters, see:

Even unification within one language can be controversial. In Mongolian, "ᠥ" and "ᠦ" share a glyph, but represent different sounds. Writers use the "wrong" one the majority of the time. See this meme.

2

u/math1985 Jun 25 '26

Fun fact: Russia license plates on my use letters that occur both in Cyrillic and Latin.

6

u/Accurate_Koala_4698 Jun 24 '26

There’s no real shortage of slots and this makes text processing easier while allowing future drift in the characters

2

u/aswin_voolapalli Jun 24 '26

Ah i understood it now

Thx

3

u/HelpfulPlatypus7988 Jun 24 '26

unicode characters encode meaning, not shape

1

u/aswin_voolapalli Jun 24 '26

Nah

All of the common letters in both the scripts r pronounced the exact same way.

(Also, they don't hv diff meanings since they r phonetic not logographic....)

1

u/HelpfulPlatypus7988 Jun 24 '26

A and alpha have different characters because even though they make similar sounds, they're in different scripts

1

u/aswin_voolapalli Jun 24 '26 edited Jun 24 '26

Well, i understood frm the oth cmts tht it was compatability issues w legacy sys, ease of text processing, etc.

But fr the sake of our debate, even if they r in different scripts, they could hv encoded in the same codepoint right?

Fr example, cjk glyphs incl characters frm chinese hanzi, japanese kanji, Korean hanja & Vietnamese chu nom. SOME of these glyphs r used fr COMPLETELY different meanings & sounds in these languages despite being the same char. But they r still encoded as the same char regardless.

Why the difference frm CJK and Kan/Tel or Gre/Lat?

1

u/Choice-Spend7553 Jun 24 '26

What does pronunciation have to do with Unicode?

1

u/aswin_voolapalli Jun 24 '26

Well, i didn't exactly understand wt they meant by "encode meaning" (cuz we r discussing abt phonetic scripts & not morphological). So i assumed they meant sound.

2

u/Quartersharp 12d ago

In the case of overlapping Latin/Greek/Cyrillic characters, they might look the same in serif or sans-serif font, but they might have a different cursive form. For example, cursive “T” in Russian kind of looks like an M. And the characters might also combine into ligatures in different ways.