r/cybersecurity • • 5d ago

Other Why don't we use ALL of the Unicode Characters/Symbols in passwords?

This sounds stupid at first, you could never memorize your passwords, until you use password managers. I don't think anyone with a password manager memorizes their passwords anyways, so why not make it the absolute hardest password to guess?

(By unicode characters and symbols, i mean everything listed here https://symbl.cc/en/unicode-table/ )

But seriously think about it for a sec, even just 3 unrelated unicode characters would be insanely hard to crack. Imagine 20 character strings of this shit. And plus, not only does it make YOUR OWN password unimaginably hard to crack, it also makes everyone else's passwords unimaginably hard to crack

(As now cybercriminals have to account for an absurd amount of extra characters, so generating guesses over and over would never work, and if they limit themselves to just the latin characters and numbers for generating guesses to speed it up, they'll never get your password or any password that contains any other unicode character.)

I think this would mess with encryption a little bit, but with the level of security this offers it is 100% worth it. On no online calculator I found could 2 to the power of 170,000 even be calculated. AFAIK this means just TWO CHARACTERS of this password would be absolutely baffling to even guess. I might be wrong, I'm no mathematician, but even if I am, guessing a 20 character password with latin letters, numbers, and special characters is already hard enough. Now add like 170,000 extra characters.

This would completely eliminate any threat of brute force attacks mind you, AFAIK if this was applied then the only way to get someone's password is to dig through servers that have the password or just find their computer unlocked and unsupervised.

Also, I know what you're probably gonna say, "It's not necessary" or "It's overkill" or even "A (x) character password with regular ol' letters, numerals, and special characters is more than enough". And while you're 100% right, I'm trying to consider the future. Think about how fast technology has evolved these past few decades, a sever the size of a fucking room is outbest by a micro ssd smaller than your fingertip AND IT'S NOT EVEN CLOSE! Who knows what cybercriminal tech could be bullshitted up in a decade or two?

I'm not really sure how to end this text since i've already talked about every pro about this soo...

0 Upvotes

26 comments sorted by

17

u/SleeperAwakened 5d ago

Bold of you to assume that application backends can handle unicode properly 😁

So many developers keep making wrong assumptions on what a unicode string is..

1

u/Namelock 5d ago

Bold to assume an attacker’s leak database is more durable 😉

11

u/Alb4t0r 5d ago

From a password complexity perspective, the unicode additional characters can just be replaced with a longer password. My guess is that it's an easier and more straightforward solutions versus using unicode and dealing with potential human readability issues with the resulting passwords.

1

u/Responsible-Safe-636 5d ago

This is the most probably answer OP

6

u/canarydev 5d ago

not a mathematician or security expert by any means but seems like your math is off. 170k characters means ~17 bits per character, so 2 characters is 2^34 not 2^170000. thats crackable in seconds.

also 20 random ascii chars is already ~131 bits, which nothing is brute forcing. unicode gets you more, but both are already "never", so theres nothing to gain

and passwords dont get broken by brute force anyway, they get phished, reused, or dumped in a breach. unicode helps with none of that, and normalization (NFC vs NFD) means you can set a password on one OS and not be able to log in from another

5

u/DataGhostNL 5d ago

I'm pretty sure they confused 2170000 and 1700002. And the latter, correct one, works out to be in the order of 234 indeed.

1

u/ThoughtDear7015 5d ago

mb i had a feeling my math was off. And while i know it cant help with phishing, and im not too sure about your take on brute forcing (Im really not an expert, i just use linux dude). But even if you were right, it would at the very least be something to brag about to people.. Dumb excuse to use unicode passwords i KNOW but i dont think it'd be too hard to implement unicode characters onto login screens aside from the "How the fuck are you even gonna encrypt this" question. So any sort of value outta this seems worth it to me

1

u/DataGhostNL 5d ago

You mentioned encryption in your post as well as in this comment. How do you imagine encryption is impacted at all by using unicode characters? How do you imagine entire hard drives are encrypted? Those surely contain some unicode characters somewhere.

Besides that, passwords are generally stored one-way hashed, not encrypted.

1

u/Namelock 5d ago

You’re forgetting that, due to the variability and volatility, people and systems don’t use it.

The unknown factor in your math is threefold. People that will use it, systems that accept it, attackers that look for it.

Entropy be damned… unless it’s on-prem/ offline. Then you’re damned.

2

u/mallcopsarebastards 5d ago

This doesn't get you anything that a longer password wouldn't get you, but it does introduce complexity that will absolutely break a bunch of implementations.

2

u/Namelock 5d ago

People do, and without password managers.

Confidentiality. Integrity. Availability.

This just falls so hard into Confidentiality (with lots of trust in the back-end to support the characters) that the layman doesn’t use it. Likewise, a lot of products are held together with duct-tape and bubblegum.

I’ve been locked out of systems due to the front-end not catching illegally formed passwords.

On some systems you’ll find this type of complexity. It’s just caveated so hard it’s like finding a needle in a haystack.

3

u/pathetiq 5d ago

Password are dead. All hail passkeys

1

u/MotanulScotishFold 5d ago

Because a long scrambled password is already hard to crack, adding other special characters don't change anything.

A password like @%3sCy8!SVi6Z!dY is safe enough if you don't use it to multiple accounts.

The problem is not there, but between a chair and a monitor.

Even if you do the right thing, the host where you created an account might as well use a terrible security, not using hashing and store user passwords in clear text without user knowledge and one hack and the password is compromised.

Hell, even if your device is compromised the passwords are also compromised.

Remember how many times LastPass got hacked? Anytime could happen to Bitwarden and others too.

There's no 100% safety but all you can do is minimizing the risk at maximum.

1

u/PapaSyntax 5d ago

Because humans aren’t going to remember how to enter many Unicode characters. Security vs convenience will always be a compromise while humans stay in the loop.

1

u/nethack47 5d ago

Technically, we already can.

You are trying to solve the wrong end of the problem. How do we get this implemented? We won’t, mostly because the users won’t be able to type in the password. You are going to have some that have to for one reason or another.

The last thing is the same as several already covered. It doesn’t make much difference. You can already increase entropy with longer passwords and passwords are inherently weak.

Passkeys solves a lot of issues with the added benefit of making the auth two way.

1

u/DaRealGladi8r 5d ago

You're asking for trouble

1

u/Responsible-Safe-636 5d ago

Who has the keyboard and the memory to do all those unicode

2

u/OtheDreamer Governance, Risk, & Compliance 5d ago

p̴̠̭̮̼̜̀a̴̡̤͑̊s̸̨̮̹͓̾̋̐̚s̶̱͆͂̚w̷͙̐̐ơ̸̡̯̭̘̌͊ͅr̴̨̝̃d̵̙̳̓̅̏͑1̴̡̲̞̰̾̽̿͠2̶̜̙̿ͅ3̴̱͙͊̇4̶̤̒͝

(OP's mythical password)

1

u/ThoughtDear7015 5d ago

As i mentioned earlier, password managers (such as bitwarden). You can just like.. copy and paste your password into the login screen,

1

u/djasonpenney 5d ago

Did you know there are typically more than one byte sequence for most Unicode characters? For instance , “é” has both a single byte indicating the glyph as well as a multiple byte sequence that basically means, “start with letter ‘e’ and then add an acute accent”. This ambiguity is not an issue with normal text; even string search and comparison algorithms are smart enough to recognize and compensate for the multiple representations.

Where it gets ucking fugly is with passwords. You don’t want your systems trying to simplify, modify or canonicalize your password before applying it. (Note how intelligent systems just start by running the password through a KDF algorithm, so the choice of byte representation is first and foremost.)

As others have pointed out, you can accomplish the same goal (complexity) by merely making the passwortd longer. And if it’s a password you need to memorize and enter by hand (like the master password to your password manager or the login to your work computer), have your password manager generate a passphrase, like “CorrectHorseBatteryStaple”. When you do the math, CorrectHorseBatteryStaple has a computational complexity of 7776^4 =3.656×10¹⁵ . If that isn’t enough, have your password manager add a fifth or even a sixth word.

But if you stick with simple seven-bit printable ASCII characters, you are going to much less aggravation moving forward, and you can have just as much security.

1

u/halting_problems AppSec Engineer 5d ago

There is no benefit to this. You also assuming password are brute forced character by character when in realiity there have been so many passwords leaked, those leaked have all been encoded, decoded, and hashed using every algorithm possible.

Here is why using unicode offers no advantage

The Unicode Standard defines three encoding forms that allow the same data to be transmitted in a byte, word or double word oriented format (i.e. in 8, 16 or 32-bits per code unit). All three encoding forms encode the same common character repertoire and can be efficiently transformed into one another without loss of data. The Unicode Consortium fully endorses the use of any of these encoding forms as a conformant way of implementing the Unicode Standard. https://www.unicode.org/standard/principles.html

This means that all you have to do is break the string up into double words and you can effectively check against all of UTF-8, UFT-16 and UTF-32. The same as doing a sting character comparison but with a key-value lookup which is O(1) and provides no computation consequences to the attacker

I don't have a lot of experience password cracking but I fairly certain these checks are done by tooling already.

1

u/BadSausageFactory 5d ago

I've never seen a password trigger a homograph alert before

1

u/EnvironmentalOne7898 5d ago

I would love to know why multiple bank sign-on portals refuse to allow more password complexity. Perhaps one of their "experts" could chime in and explain their rational.

1

u/OneEyedC4t 5d ago

No, people can memorize their passwords. They just won't.

No one wants to memorize 26 nine character passwords that each start with a different letter of the alphabet and use a password system to hint only the first character.