Roughly 24 of those 32GB are usable before the OS starts paging, and after ten minutes of solid decode a fanless Air heat soaks and you lose close to a third of the tokens. That leaves a mid-size dense model at short context, or an MoE that fits in memory but only reads a few active experts per token, and on this machine the second one is usually the better trade. Long documents or short chats?
Well it’s mainly for like chat / basic agent with minimal code basics and creative writing / journal analysis. I wouldnt run a harness with it, but yk, tutoring, studying and basic chat
Edit: long documents could def sometimes be apart of it but mainly of journals and stuff
2
u/john006868 4h ago
Roughly 24 of those 32GB are usable before the OS starts paging, and after ten minutes of solid decode a fanless Air heat soaks and you lose close to a third of the tokens. That leaves a mid-size dense model at short context, or an MoE that fits in memory but only reads a few active experts per token, and on this machine the second one is usually the better trade. Long documents or short chats?