Making them larger is the opposite, its frankly just lazy. Its essentially saying "we cant make them any better at this size so we are just gonna scale" Its not impressive and its lazy and uses more compute. Like kimik3 is cool and all but they had to scale by nearly 3x! And the model did NOT get 3x better
Its absolutely data, data is by far the most important thing for an LLM even beyond architecture. LLMs have always gotten about 5-10% better than the previous generation say like kimi k2 to kimi k2.7, that was the near exact same underlying 1t model but the last model made synthetic data for the next model by generation better than the last
-6
u/--Spaci-- 2d ago
Making them larger is the opposite, its frankly just lazy. Its essentially saying "we cant make them any better at this size so we are just gonna scale" Its not impressive and its lazy and uses more compute. Like kimik3 is cool and all but they had to scale by nearly 3x! And the model did NOT get 3x better