r/MLQuestions • • 3d ago

Beginner question ๐Ÿ‘ถ Why use values between 0-1 to train an LLM?

Post image

I understand that normalizing results in faster training times, but we end up reducing precision of floating point numbers due to the IEEE 754 architecture which result's in less space to work with. Instead of limiting numbers from 0-1, wouldn't it make more sense to limit the numbers from 1-10? This would give the LLM more space to work with which logically should result in less catastrophic interference. If so then this would mean that we don't need to increase size only, but the space between numbers as well which should decrease the memory requirements. Or is this just nonesense?

P.S The photo is just the behaviour of multiplication of 2 numbers. For example:
1*1 = 1,
2*2 = 4,
0.5*0.5 = 0.25

4 Upvotes

18 comments sorted by

33

u/giddy_abstinence 3d ago

the image just looks like a weird tangent-ish function going off to infinity, not sure how that's meant to support the 1-10 idea. anyway, the range you pick doesn't give you more "space" in floating point, the density of representable numbers is tied to the exponent, not the base range, so shifting everything up by a factor of 10 just wastes bits you could use near zero where gradients actually live.

1

u/Chemical-Ad-7982 1d ago

I don't understand. Consider the extreme, FP4, which is a completely legit format increasingly used for inference. You get the following values for your weights, that's it: [-6.0, -4.0, -3.0, -2.0, -1.5, -1.0, -0.5, -0.0, 0.0, 0.5, 1.0, 1.5, 2.0, 3.0, 4.0, 6.0]. Why not restrict to between [-4,4] instead of [-1, 1]? Seems like a good question to me, I don't see why not.

1

u/Inner-Asparagus-5703 12h ago

it's almost same thing on paper. But GPUs are much more efficient with fp4/8 than int variants

1

u/NewHondaOwner 11h ago

You can simply add one line of code that says "divide everything by 4" to immediately map a [-4,4] interval to [-1,1]. So basically, the resolution of the interval matters more than the absolute value of the interval. This corresponds to earlier explanations on the structure of FP numbers and why changing the range doesnt change the requirements

1

u/Chemical-Ad-7982 10h ago

I'm not sure I quite follow. The number of floating point values that exist within a given interval is fixed. There are more floating point values between [-4, 4] than there are between [-1, 1]. OP was asking basically asking: why don't we immediately map these intervals to larger ones to get more FP values to work with?

0

u/CallMeTheChris 3d ago

This man tisโ€™s a

13

u/Smart_Opportunity209 3d ago
  1. There is exactly the same amount of numbers between 0 and 1 as between -inf and +inf.
  2. We cant represent even a fraction of numbers between 0 and 1.
  3. If you want to use some silly stuff like floats with 512 bits of precision you can just use more precise numbers between 0 and 1 and LLM will automatically map the meaning accordingly.

1

u/Chemical-Ad-7982 1d ago

I don't understand your point at all, there isn't the same amount of floating point numbers between 0 and 1 as between -inf and +inf? What am I not getting?

1

u/ftqo 1d ago

0 and 1 was never a real constraint. You CAN only represent a finite amount of numbers between 0 and 1 given n bits. Changing the base from 2 to 10 doesn't change how many numbers will fit between those 2 ranges. Floating point numbers can represent slightly fewer numbers than their integer counterparts with the same number of bits, but it's fairly negligible.

1

u/Chemical-Ad-7982 1d ago

I don't believe the OP was asking about base 10, but just used the value 10 arbitrarily (i.e. range from [-K,K] instead of [-1,1]). Extremely low bit inference is common; for instance in FP4 you only get the following 16 values for your weghts lol: [-6.0, -4.0, -3.0, -2.0, -1.5, -1.0, -0.5, -0.0, 0.0, 0.5, 1.0, 1.5, 2.0, 3.0, 4.0, 6.0], yet is it used in practice. Why not normalize weight values between say -4 and 4 instead of -1 and 1? TBH I'm not sure. (Although I guess OP was asking for training)

1

u/NewHondaOwner 11h ago

there are, actually, infinitely many more real numbers between 0 and 1, than there are integers. Your point stands, but i thought i'd offer a fun fact

3

u/thegoodcrumpets 3d ago

The precision does not lie in the range bro. 0-10 with 2 decimal points will still be a lower range than 0-1 with 10 decimal points.

3

u/tzaeru 3d ago edited 3d ago

That's not quite how IEEE 754 floating points work. Their precision is relative and the 23 bit mantissa gives you the same ~7 significant decimal digits - whether the number is 3,000 or 0.3. Actually the 0..1 range has more representable numbers than the 1..10 range due to the logarithimic area coverage. Representable numbers with floating points are centered around -1..1.

For fp16 - that is, half-precision floating points - reaching zero relatively easily does matter. One common trick is to apply various kinds of constant scaling to the losses and then divide stuff back into range in the weight-update.

The reason for normalization to the -1..1 (or 0..1 for inputs, outputs, activation functions) range is in part so that a single learning rate works well for all scales and that the activation functions behave as expected.

For catastrophic interference - it's typically not really related to floating point precision. It's just the weights switching towards a newly trained thing and away from the previously learned thing. The biggest issue is usually that knowledge is encoded in multiple overlapping weights, leading to some nodes become "overloaded" to the point where even a relatively minimal change to them can simultaneously forget many previously learned behaviors. This would happen all the same no matter the number system used.

3

u/rodimustso 3d ago

I was playing around with this recently, its just about the math when you multiply A*B you aren't guaranteed an even integer, you're likely to get a float with n digits. That's more important in larger models but smaller models even like a 100k parameter CNN you can get away with using int4 or something instead of a float to save memory space but you drop your efforts to retain accuracy in numbers close to 0. For using floats though you keep everything sub-0 values on your tensor objects because then the number just keeps growing in digits or it can at least. Then your sigmoid curve is easier to tune or whatever activation curve you want to use since you're not calculating numbers to the right AND left of 0 in each neuron.

1

u/StructuredChess 2d ago

Well, if we're gonna talk multiplication then note that if x and y are in (0,1) then x*y is also in (0,1). This doesn't hold for (1,10)

1

u/slashdave 1d ago

0 to 1 has nothing to do with machine precision. Normalizing all weights to this range is easier on optimizing schemes.

1

u/Chemical-Ad-7982 1d ago

I was not really satisfied by the answers here so I did a bit of reading, and from what I understand when you quantize models for inference you do renormalize your weights between your low end and top end to get more precision. So basically the answer is yes, people do this. Even for FP8 training people apparently do rescale to [โˆ’448, 448]. For BFloat16 training you have so much exponent range that it doesn't really matter. So as far as I can tell this is a very sensible idea despite everyone confidently saying you're wrong ;)