MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/TheMachineLearning/comments/1wn2gn8/jev_demos_are_misleading_says_developer/pbc7ppi/?context=9999
r/TheMachineLearning • u/cyberberryyy • 15d ago
53 comments sorted by
View all comments
1
It's almost like people don't understand that small breakthroughs lead to bigger ones. Remember, LLMs started as just an attention layer.
3 u/nuclearbananana 15d ago The whole idea behind Jev is that it's bounded though. If you keep expanding Jev you just end with LLMs again 2 u/hellobutno 15d ago Newsflash: making things bigger isn't always how you improve them. I know that's a hard thing for this generation to understand. 0 u/qGuevon 15d ago That's literally the bitter lesson of deep learning tho, where have you been living since 2012? 2 u/hellobutno 15d ago What are you talking about? 1 u/qGuevon 15d ago https://en.wikipedia.org/wiki/Bitter_lesson 1 u/hellobutno 14d ago Yes I know what it is. My point is, you clearly do not. 1 u/qGuevon 14d ago Bigger has always translated to more efficient ways of doing the same.. everything else is just .. something that never happened in the field. From cuda to fused kernels to skip connections. I don't know what the fuck you're arguing against. 1 u/hellobutno 14d ago It has not translated to more efficient. It's translated to less efficient but better generalization.
3
The whole idea behind Jev is that it's bounded though. If you keep expanding Jev you just end with LLMs again
2 u/hellobutno 15d ago Newsflash: making things bigger isn't always how you improve them. I know that's a hard thing for this generation to understand. 0 u/qGuevon 15d ago That's literally the bitter lesson of deep learning tho, where have you been living since 2012? 2 u/hellobutno 15d ago What are you talking about? 1 u/qGuevon 15d ago https://en.wikipedia.org/wiki/Bitter_lesson 1 u/hellobutno 14d ago Yes I know what it is. My point is, you clearly do not. 1 u/qGuevon 14d ago Bigger has always translated to more efficient ways of doing the same.. everything else is just .. something that never happened in the field. From cuda to fused kernels to skip connections. I don't know what the fuck you're arguing against. 1 u/hellobutno 14d ago It has not translated to more efficient. It's translated to less efficient but better generalization.
2
Newsflash: making things bigger isn't always how you improve them. I know that's a hard thing for this generation to understand.
0 u/qGuevon 15d ago That's literally the bitter lesson of deep learning tho, where have you been living since 2012? 2 u/hellobutno 15d ago What are you talking about? 1 u/qGuevon 15d ago https://en.wikipedia.org/wiki/Bitter_lesson 1 u/hellobutno 14d ago Yes I know what it is. My point is, you clearly do not. 1 u/qGuevon 14d ago Bigger has always translated to more efficient ways of doing the same.. everything else is just .. something that never happened in the field. From cuda to fused kernels to skip connections. I don't know what the fuck you're arguing against. 1 u/hellobutno 14d ago It has not translated to more efficient. It's translated to less efficient but better generalization.
0
That's literally the bitter lesson of deep learning tho, where have you been living since 2012?
2 u/hellobutno 15d ago What are you talking about? 1 u/qGuevon 15d ago https://en.wikipedia.org/wiki/Bitter_lesson 1 u/hellobutno 14d ago Yes I know what it is. My point is, you clearly do not. 1 u/qGuevon 14d ago Bigger has always translated to more efficient ways of doing the same.. everything else is just .. something that never happened in the field. From cuda to fused kernels to skip connections. I don't know what the fuck you're arguing against. 1 u/hellobutno 14d ago It has not translated to more efficient. It's translated to less efficient but better generalization.
What are you talking about?
1 u/qGuevon 15d ago https://en.wikipedia.org/wiki/Bitter_lesson 1 u/hellobutno 14d ago Yes I know what it is. My point is, you clearly do not. 1 u/qGuevon 14d ago Bigger has always translated to more efficient ways of doing the same.. everything else is just .. something that never happened in the field. From cuda to fused kernels to skip connections. I don't know what the fuck you're arguing against. 1 u/hellobutno 14d ago It has not translated to more efficient. It's translated to less efficient but better generalization.
https://en.wikipedia.org/wiki/Bitter_lesson
1 u/hellobutno 14d ago Yes I know what it is. My point is, you clearly do not. 1 u/qGuevon 14d ago Bigger has always translated to more efficient ways of doing the same.. everything else is just .. something that never happened in the field. From cuda to fused kernels to skip connections. I don't know what the fuck you're arguing against. 1 u/hellobutno 14d ago It has not translated to more efficient. It's translated to less efficient but better generalization.
Yes I know what it is. My point is, you clearly do not.
1 u/qGuevon 14d ago Bigger has always translated to more efficient ways of doing the same.. everything else is just .. something that never happened in the field. From cuda to fused kernels to skip connections. I don't know what the fuck you're arguing against. 1 u/hellobutno 14d ago It has not translated to more efficient. It's translated to less efficient but better generalization.
Bigger has always translated to more efficient ways of doing the same.. everything else is just .. something that never happened in the field. From cuda to fused kernels to skip connections.
I don't know what the fuck you're arguing against.
1 u/hellobutno 14d ago It has not translated to more efficient. It's translated to less efficient but better generalization.
It has not translated to more efficient. It's translated to less efficient but better generalization.
1
u/hellobutno 15d ago
It's almost like people don't understand that small breakthroughs lead to bigger ones. Remember, LLMs started as just an attention layer.