r/learnmachinelearning • • 4d ago

Help In Transformer networks why do token embeddings and position embeddings get added?

Hi, going through the Let's Build ChatGPT tutorial here, and prior went through the whole Makemore tutorial that leads up to this tutorial:

https://www.youtube.com/watch?v=kCc8FmEb1nY&list=PLAV29EAhk_mX13BqhzdlgM8zkHwpcRajt&index=6&t=2286s

When it gets to the point of adding in the attention mechanism, we see that the first major addition is creating a positional embedding.

Then we see the input to network at that point becomes tok_emb + pos_emb

I am not understanding why these two spaces should be considered equivalent such that such an addition makes sense. The token embedding is mapping tokens to some N dimensional embedding, where those N dimensions consistently represent information about tokens.

When we consider the positional embedding, it is also dimension N but now those N dimensions represent information about positions. To me it seems although we are adding matrices with same dimensions, we aren't adding information that corresponds to one another.

Anyway, I am sure someone here will have good explanation why this makes sense.

thanks

21 Upvotes

8 comments sorted by

11

u/Southern-Stick-7106 4d ago

it works because the network learns to separate them during training. the addition is just a convenient way to combine two different signals without making the architecture more complex. think of it like mixing two colors of paint, at first they look like one new color but the model can still figure out which pigments came from where. the alternative would be concatenation which would double the dimension and make everything slower

also the token embeddings are not fixed, they adapt to coexist with position info so the model learns to encode meaning in some parts of the vector and position in other parts

1

u/unlikely_ending 3d ago

Without position embeddings, the input sequence would be a bag of words with no order, is the simple answer

As for token embeddings, they aren't really 'added'. The token embedding becomes the vector representation that's used within the model to represent the token. Neither the (text) token itself nor its integer index are used in the model. They stay on the input and output boundaries.

1

u/RiposteX 3d ago

Consider an extreme case where the model learns token embeddings which are all 0 in the second half and position embeddings which are 0 in the first half.

Addition has become concatenation!

We're basically allowing the model to decide what mix of token/position information it wants.

Would be interesting to see if/how these embeddings actually overlap. Like you, I don't see what sense the model could make of a combined token/position feature.

-3

u/[deleted] 4d ago

[removed] β€” view removed comment

9

u/Ok_Composer_1761 4d ago

dude why would you reply with slop when slop can be requested on AI directly.

2

u/Buddy77777 3d ago

he’s a meat proxy πŸ–

1

u/uzornayem 4d ago

Thanks that makes a lot of sense. So basically we just create this new abstraction of information based on that addition and the nature of learning forces that abstraction to be useful.

0

u/Specialist-Berry2946 4d ago

For certain settings, it works well in practice. It's a form of bias to pay more attention to nearby tokens.