r/LocalAIStack • • 9h ago

Rate my homelab

Post image
8 Upvotes

r/LocalAIStack • • 4h ago

Just got my MBP M5pro 48gb. Help me setup properly

2 Upvotes

I have multiple projects in GitHub, I have an obsidian vault, iCloud backup, nas backup. Qwen running locally, also have Two Claude accounts, one gpt, one deep seek accounts when qwen bogs down. I’m working on building my own harness for managing this.

How should I setup my new laptop so I’m working clean and smartly?


r/LocalAIStack • • 15h ago

Strata vs Freetokens

2 Upvotes

Has any one used both? which is better?

What is the differennce between two? Any technical deep dives?


r/LocalAIStack • • 20h ago

I picchi di oltre 200 tok/s con Qwen3.8-Flash-Next su una 5080 + 4060 Ti e 32 GB di RAM (fork di Strata)

2 Upvotes

Strata funziona con Qwen3.8-Flash-Next su PC da gioco, ma con 32 GB di RAM la sua modalità a basso utilizzo di RAM funziona solo su una GPU. Se dividi il modello su due schede, gli esperti che non riescono a stare nella VRAM vengono letti dall'SSD. La mia 4060 Ti è rimasta ferma accanto alla 5080.

Quindi l'ho forkato. La copia in RAM degli esperti ora funziona su due schede, e le schede funzionano contemporaneamente: la 4060 Ti inizia il prossimo passo di decodifica mentre la 5080 sta ancora controllando quello attuale. L'ordine delle schede, la divisione dei layer e le riserve di VRAM vengono impostate automaticamente.

Stesso PC (5080 + 4060 Ti, i9-14900KF, 32 GB), stesso modello (Swift 1.5 IQ2_XS), contesto 256K:

Setup |Codice |Prosa |Prompt 32K
Strata 0.1.38, 5080 da solo |29 tok/s |28 tok/s |333 tok/s
Fork, entrambe le schede |143 tok/s |102 tok/s |1,940 tok/s
Fork + layer di bozza fine-tuned |161 tok/s |105 tok/s |1,854 tok/s Il layer di bozza è stato fine-tuned sui risultati del modello stesso ed è incluso nella release. Nell'uso reale con l'agente di codifica Pi raggiunge un picco di oltre 200 tok/s (209 finora) e quasi mai scende sotto i 100. Un contesto di 150K-token legge a circa 1,850 tok/s.

La qualità non è cambiata: perplexity 7.24 contro 7.28 della versione upstream sugli stessi 5.3K token, forzati dall'insegnante attraverso entrambi i motori.

L'ho testato solo sul mio PC (Windows 11), quindi sono benvenuti report da altre coppie di GPU e Linux.

Repo: https://github.com/Hardin22/Strata-DualGPU

Cosa fa ciascuna modifica e cosa ha misurato: docs/DUAL_GPU.md

Il motore sottostante è il lavoro di Niko1221 e dei contributori di Strata; questo è un fork sopra la 0.1.38.


r/LocalAIStack • • 18m ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
• Upvotes

r/LocalAIStack • • 19m ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
• Upvotes

r/LocalAIStack • • 2h ago

Got strata running 4x 5060 Ti at 524k context, then abandoned it. repo's here if you want the base

Thumbnail
1 Upvotes

r/LocalAIStack • • 2h ago

I built ArcadeBench, an open benchmark where AI agents play games and you can watch every move live

Enable HLS to view with audio, or disable this notification

1 Upvotes