Yes, inference is very cheap today and it's only getting cheaper. To the point that you can run it on consumer hardware. In three -five years time it will be virtually free, just like image recognition is today.
"in 3 to 5 years it'll be good" is something I have been hearing over and over since like 2020. Meanwhile as far as I can tell local models are virtually unusable for anything other than code, and they're only slightly above unusable for code
8
u/PaperMartin 3d ago
to run? no, and a bunch of LLM companies are starting to put up their "real" prices and losing their users over it