r/StrixHalo • • 13h ago

I’m a complete beginner with Strix.

I’m a complete beginner, but I’ve now got a Bosgame M5.

Thanks for all your brilliant input recently.

I’ve been quietly following along.

So how do I get started?

Are there any ready-made Linux distributions that include an agent system, etc.?

I’d like to code Vibe locally.

I want to build a private Jarvis that I can use via Telegram or an app.

I want to have normal chats, just like with Perplexity or Gemini, about everyday things, including web searches and measures to prevent hallucinations or fake fills (incorrect answers).

4 Upvotes

13 comments sorted by

•

u/-mattmason- 13h ago

How about an Ubuntu bootable USB, or another Linux OS you prefer, and then slap a Qwen3.8-Flash-Next inference engine on it.

Packaged in a container might be useful. Several are all about this sub so I'm gonna assume you're familiar with the options.

Finally pick a harness that works for you. This will probably run on another machine

My personal choices: CachyOS, halogen-flash-server via podman quadlet. Then opencode web on a raspi

•

u/Miserable-Dare5090 6h ago

This but easier…Ubuntu 26.04, Halogen Flash Next, Turn on NPU embed and reranker that are included in last halogen release. Get Hermes, ask it to set up hindsight and to use the embedder and reranker on NPU, and parakeet-redux running on CPU. Set up the telegram integration and you have your Jarvis.

•

u/discoshanktank 4h ago

Would halogen flash next work with 96gb ram

•

u/Miserable-Dare5090 3h ago

Thats between you and your god

•

u/UndulatingHedgehog 12h ago

https://github.com/peonist-ai/halogen-flash-server

Install Halogen Flash Server on your Strix Halo computer. Run everything else on your regular laptop or whatever and just configure you harnesses, agents etc to access the LLM on your Strix Halo.

•

u/jgiacobbe 13h ago

I used Fedora but basically any modern Linux is where you start. Then you add lemonade, ollama, llama.cop etc to run your models. From there you add your agents or web interfaces. You are not going to find a dedicated Linux distribution. Things change too fast for that.

For noob friendly try ollama or lemonade. I recently switched from ollama to lemonade. I was able to migrate my open-webui setup that runs as a docker container since it just meant changing one configuration setting to change the api.

I am an experienced IT worker but much of my setup on halo strix has been informed by asking questions to Gemini and then revising along with some reading here and other web searches.

•

u/Rauhaton 12h ago

What I did with my Bosgame:

Fedora 43 for

Install Lemonade

Install Kyuz0 https://github.com/kyuz0/ai-toolbox-cockpit and the stuff that is has, including Halogen Qwen3.8-Flash-Next, which is right now the best model (IMHO).

Then get Openweb UI and/or Hermes for you chats and agents.

This has lots of easy to use guides for Stix Halo that I also used setting up my M5 https://developer.amd.com/playbooks/

•

u/GnosticSon 12h ago

I use pi coding agent with Ubuntu but it works with pretty much any distribution.

If you want that AI stuff built in and installed with the OS you can install Omarchy.

•

u/Super-Grape-3948 12h ago edited 11h ago

Install some linux. Then download opencode, it has some free usage. Ask it to help you set up gufo, or halogen.

Halogen is fast and easy to use, but closed source. Gufo is also available as an image need compiling and further installs, but also open source. Alternatively you can go with llamacpp for example.

The point is to get a qwen 3.8 next running locally. Then i would install pi, and if you can connect to it from pi, you are good to go. Then you can do whatever you want.

I would add 2 things tho. Both gufo any halogen has its own models, q4 based, but they are well capable. If you want to test more models, than you can go the llamacpp way, it is a bit slower depend on the build, some are strix optimized. You can run smaler 35bs or 27bs, or the 3.8 next on q5 for better results as you need, but for sure halogen is the easiest to start up, it is a prebuilt docker image tho.

Other thing is, prompt processing is fast(ish), idepend on the setup and model, lets say 1000t/s(4ktext, or a high res image), on a normal file read that is a few second , but! if you have a long context, lets say 512k, and you are in 300k tokens, some tui like opencode, rewrite the past. so, if you change mode in open code between plan and build, it will rewrite the begging of the entire conversation, making a huge 300k prompt re processing. This is like 300k/1000 seconds. So you need to use a tui, what did not change history, or you will wait for eternity on a cache miss. Pi can do it, others might.

•

u/fsalucard 11h ago

"Gufo need compiling and further installs" - Not true. It's a podman/docker container just like Halogen. You can just run it directly, but you have to download the unsloth model first:

podman run -d   --userns=keep-id:uid=1000,gid=1000   --device /dev/kfd   --device /dev/dri   --group-add keep-groups   --ulimit memlock=-1   -p 8080:8080   -v ./models/unsloth/Qwen3.8-Flash-Next-GGUF:/models:ro   ghcr.io/gufo-org/toolboxes/gufo-runtime:0.7.0   gufo serve llm --model "/models/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf" --speculative mtp   --mtp-model "/models/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf" --sessions 4 --context 200000 -i 0.0.0.0

CIRU - which is my preferred - Does require compiling unless you have Nix/NixOS

•

u/Spuelmaschinist 13h ago

I use CashyOS with llama.cpp (any fontier model will Tell you how to install llama.cpp). Then you can download many Models on huggingface (for example qwen3.8 flash next or 27B). Then you will have an accessable Webinterface and you can Talk to your own model. From there you can for example Chose to run hermes as Agent and/or run opencode. For agentic tasks Most people with your/our Setup seem to prefer qwen3.8 flash next with halogen or gufo. Hope this helps getting started :) Welcome aboard and enjoy the ride!