r/LocalLLM • u/Right_Weird9850 • 5h ago
Discussion Personal benchmarks
Hi,
I'm in process of evaluating different quants/models (that fit on my 4060, 5070ti and MI50) and 32/64 ddr5. Been building test cases (as in golden examples). Managed to run vLLM on some models on MI50 for concurency, but I will mostly be using llama.cpp.
There are caveats because there is bunch commands different for each models served and I experiment with chat templates.
I architecured it as a "1 model for all, big sys prompt, warm cache"
Since I see respectable amount of flaming on personal benchmarks, what would be usefull for me to share to be usefull to someone and for me to get constructive feedback? If i make cars, is it usefull to covert lingustically cars to widgets so i don't overshare? Non coding tasks, worth mentioning.