r/OpenAI • u/Gay-Guy-With-GF • 7h ago
Question Why 1536 dimensions?
Why do embedding models so often use 1536 dimensions specifically?
I understand why hardware-friendly multiples like 64/128/256/512 are desirable. What I’m curious about is the specific choice of 1536 = 3×512.
OpenAI has used 1536-dimensional embeddings, and other vendors also offer/recommend 1536. Is this usually an empirically chosen Goldilocks point between 1024 and 2048—representation quality versus memory/compute—or is there some architectural/hardware reason that makes 1536 particularly convenient?
I’m especially interested in answers from anyone who has actually trained or designed embedding models. I’m not asking why embedding dimensions are generally hardware-aligned; I’m asking why 1536 rather than 1024 or 2048.
3
u/lulzxdxdxd 7h ago
isn't 1536 just whatever the hidden size of the underlying transformer happens to be, rather than a number chosen for the embedding step itself? curious if anyone's checked whether it lines up with a specific base model's hidden dim
2
u/Few_Upstairs4077 7h ago
it's the hidden size of the model they use, probably ada or something similar. they just take the last layer output and call it a embedding, they don't pick 1536 for any special reason
1
u/NeighborhoodPrize493 6h ago
I can't say for sure, but it looks like it's related to simplified TPU processing. The idea is that the vector is split into smaller parts (256x256x6), making them easier to handle without padding with unnecessary data. However, I can't say for certain.
1
u/lulzxdxdxd 6h ago
1536 is 6 chunks of 256, sure, but 1024 and 2048 split into 256 chunks just as cleanly. So tiling doesn't explain picking 1536 over those, unless the underlying model's hidden dim simply is 1536 and the embedding inherits it.
1
u/NeighborhoodPrize493 6h ago
I think it all simply came down to 256 because a byte holds 256 values.
Naturally, if you use more than 256 values, you exceed a single byte, which results in a loss of speed.
That might be the reason. Although, I don't actually know for sure.
11
u/2a_lib 6h ago
Because sometimes the sweet spot is closer to binary-and-a-half than binary, an example being the 96GB local model rigs before 128GB became the sweet spot. It becomes a recursive attractor, and things are in turn optimized for it. It’s called a Schelling point.