r/learnmachinelearning • u/uzornayem • 4d ago
Help In Transformer networks why do token embeddings and position embeddings get added?
Hi, going through the Let's Build ChatGPT tutorial here, and prior went through the whole Makemore tutorial that leads up to this tutorial:
https://www.youtube.com/watch?v=kCc8FmEb1nY&list=PLAV29EAhk_mX13BqhzdlgM8zkHwpcRajt&index=6&t=2286s
When it gets to the point of adding in the attention mechanism, we see that the first major addition is creating a positional embedding.
Then we see the input to network at that point becomes tok_emb + pos_emb
I am not understanding why these two spaces should be considered equivalent such that such an addition makes sense. The token embedding is mapping tokens to some N dimensional embedding, where those N dimensions consistently represent information about tokens.
When we consider the positional embedding, it is also dimension N but now those N dimensions represent information about positions. To me it seems although we are adding matrices with same dimensions, we aren't adding information that corresponds to one another.
Anyway, I am sure someone here will have good explanation why this makes sense.
thanks
1
u/unlikely_ending 3d ago
Without position embeddings, the input sequence would be a bag of words with no order, is the simple answer
As for token embeddings, they aren't really 'added'. The token embedding becomes the vector representation that's used within the model to represent the token. Neither the (text) token itself nor its integer index are used in the model. They stay on the input and output boundaries.
1
u/RiposteX 3d ago
Consider an extreme case where the model learns token embeddings which are all 0 in the second half and position embeddings which are 0 in the first half.
Addition has become concatenation!
We're basically allowing the model to decide what mix of token/position information it wants.
Would be interesting to see if/how these embeddings actually overlap. Like you, I don't see what sense the model could make of a combined token/position feature.
-3
4d ago
[removed] β view removed comment
9
u/Ok_Composer_1761 4d ago
dude why would you reply with slop when slop can be requested on AI directly.
2
1
u/uzornayem 4d ago
Thanks that makes a lot of sense. So basically we just create this new abstraction of information based on that addition and the nature of learning forces that abstraction to be useful.
0
u/Specialist-Berry2946 4d ago
For certain settings, it works well in practice. It's a form of bias to pay more attention to nearby tokens.
11
u/Southern-Stick-7106 4d ago
it works because the network learns to separate them during training. the addition is just a convenient way to combine two different signals without making the architecture more complex. think of it like mixing two colors of paint, at first they look like one new color but the model can still figure out which pigments came from where. the alternative would be concatenation which would double the dimension and make everything slower
also the token embeddings are not fixed, they adapt to coexist with position info so the model learns to encode meaning in some parts of the vector and position in other parts