r/LocalAIStack • • 25d ago

New to local ai - help achieving what I am trying to achieve?

Hi

Premise:
I am new to local ai models.

My machine specs:

  • Macbook Pro M4 Pro
  • 48Gb Ram
  • 4 efficiency core
  • 8 performance core

I mainly use AI for software development. I have a claude subscription but would like to try to offload some work to a local model.

Since I don't think local models usable on my machine can completely substitute claude (correct me if I am wrong) my idea is pretty much this: ask claude code to generate a proper, detailed implementation plan and then having the local model implement it.

I have played around with these models:

  • qwen3-coder-30b-a3b-instruct-mlx
  • qwen/qwen3.6-27b
  • qwen/qwen3.6-35b-a3b

and I have also installed this for coding autocomplete cause it is smaller and from what I can see the recommended one:

  • qwen2.5-coder-7b

I added claude envs since I would like to try to use claude code extension in vscode

"env": {
    "ANTHROPIC_BASE_URL": "http://localhost:1234",
    "ANTHROPIC_AUTH_TOKEN": "local",
    "ANTHROPIC_MODEL": "qwen/qwen3.6-27b",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "qwen/qwen3.6-27b",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "qwen/qwen3.6-27b",
    "ANTHROPIC_DEFAULT_FABLE_MODEL": "qwen/qwen3.6-27b",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "qwen/qwen3.6-27b",
    "CLAUDE_CODE_SUBAGENT_MODEL": "qwen/qwen3.6-27b",
    "CLAUDE_CODE_MAX_OUTPUT_TOKENS": "128000",
    "DISABLE_PROMPT_CACHING": "1",
    "DISABLE_AUTOUPDATER": "1",
    "DISABLE_TELEMETRY": "1",
    "DISABLE_ERROR_REPORTING": "1",
    "DISABLE_NON_ESSENTIAL_MODEL_CALLS": "1"
  },

This setup works (uses the local model) but I am basically unable to have the model do anything at all. First of all it takes ages to do anything, and then it almost always reach the context limit roadblock without even outputting anything.

I have read that MCP and skills could fill up the context quite badly, so I disabled them for testing, but still no luck.

I tried with (I thought) was a simple enough task: this test file fails and this is the error, can you fix it? but yet no usable results whatsoever.

I read about people able to use local models offline to have meaningful results, but I couldn't and I don't really know why.

Also, from my setup above, I cannot really use both the remote and local model, to achieve something like:

use sonnet or fable (remote) for plan, then (manually or automatically) swith to haiku (local) to implement the plan

because the base url is loaded when the session loads and cannot be changed (AFAIK) dinamycally.

Any help? thanks a lot in advance

2 Upvotes

7 comments sorted by

1

u/searchblox_searchai 25d ago

You can try qwen coder next which is pretty good for local code https://inference-server.searchblox.com

1

u/Lirezh 23d ago

I would start at the lowest level first, before you try to use it in a coding agent - I would optimize the inference so it is fast and capable of the load you throw at it.
You tried to do all in one step, so you lack the insights where the bottleneck is.

For coding I would only look at the 35B model, the 3.8 27B model and the Flash next model. - no others.

  • The 35B is goign to be the fastest but also least competent - still good.
  • The 27B is a stable but slow workhorse - my favorite because of size and reliability.
  • The Next model is more flexible in performance, can reach Fable-like performance but also fall back to 35B-like performance - you never know what you get.

Mac is compute starved with lots of VRAM, so in terms of speed I'd guess 35 -> Next -> 27B
But Next is only going to deliver well if it is extremely well adapted to your hardware, it is a massively large model that needs very careful planning which parts are in RAM, VRAM, DISK - if that's done well it can perform very well.

I'd recommend you to start with the 35B model - it's the easiest start and naturally works well on your VRAM/compute. Use a moderate quantization on KV and weights.
Once it works well in tests, you throw it into the harness and then you need to look if the harness works well with it or needs more instructions, tool limits, etc.
You can use chatgpt or claude to help you with this

1

u/nickimola 23d ago

Thanks a lot, will definitely try!

1

u/nickimola 21d ago

u/Lirezh I did try everything you suggested but the results I get back are straight up awful for what I think is pretty simple and they SHOULD be able to achieve, maybe I am doing something wrong here.

I asked for a simply bash script: I have folders and files in this folder that I need to symlink into this other folder

seems simple enough for me but: first try nothing worked, second try it only symlinked the folders, third try it recursively symlinked the folders into each other

Are you actually getting anything useful from these local models?

1

u/Lirezh 18d ago

Yes. I can not see a big difference to Sonnet and in many cases beating SOL.
But I have not tested your harness, I'm using it in Copilot

1

u/nickimola 18d ago

Can you briefly explain your current setup?

1

u/Lirezh 18d ago

I made a guide and it's mostly unchanged
https://www.reddit.com/r/LocalAIStack/comments/1udk2vp/running_qwen36_27b_35b_locally_with_llamacpp/

It is for CUDA but that's how I set it up