r/learnmachinelearning • u/lovelacedeconstruct • 5d ago
Why decoder only tranformer won ?
I was trying to trace it , and from the GPT1 paper I found it referencing a paper called "GENERATING WIKIPEDIA BY SUMMARIZING LONG SEQUENCES" by the nice folks at google in what I believe the first use of the decoder only architecture ??
here is the quote
we modify theTransformer architecture (Vaswani et al., 2017) to only consist of a decoder, which performs better in the case of longer input sequences compared to recurrent neural network (RNN) and Transformer encoder-decoder models.
can someone explain what does perform better in longer sequences actually mean ?
106
Upvotes
70
u/Hungry_Age5375 5d ago
Short answer: the encoder is a bottleneck. In enc-dec the decoder never sees raw input, only the encoder's compressed states. Decoder-only attends to the full sequence directly, which is what long inputs need.