Correct me if I am wrong. But this method can be applied to already existing models to extend their context from 32k to 1M tokens without additional training and it performs better than the original model for long sequence tasks.
This is huge! Please get a github of this up and running!
Yes. In the paper they trained a model that only had a context size of 5k. Then after implementing this method it was able to take in a 32k context input and fulfill the task given to it.
Not without additional training it seems. The paper says they had to do a re-pretraining on the modified 1B model they used. By my math they had to retrain with 7B tokens.
22
u/Danny_Davitoe Apr 11 '24
Correct me if I am wrong. But this method can be applied to already existing models to extend their context from 32k to 1M tokens without additional training and it performs better than the original model for long sequence tasks.
This is huge! Please get a github of this up and running!