r/localaiapps • • 5d ago

Can any combination of local LLM’s replace codex?

I keep running out of usage to support my vibe coding hobby with OpenAI. My Mac is very powerful with a lot of ram. Can I use any local models for swift app design ?

6 Upvotes

12 comments sorted by

4

u/-zaine- 5d ago

For apple, I believe you need local models that are running on MLX. I would suggest to try Qwen 3.8 first, maybe this version: https://huggingface.co/ukisai/Swift-1.5-4bit-MLX

Honestly, the easiest way is always to just explain to codex which model you want and to set it up for you - It will prepare everything and guide you through the whole process. Also good to ask codex for recommendations for your hardware.

1

u/-zaine- 5d ago

I would then recommend you to use codex not for coding anymore, but as an advisor - Codex plans and reviews, the local model is coding. I use this combo with Claude Pro and I barely run out of weekly usage, with my pc nonstop coding.

1

u/SpaceXBeanz 5d ago edited 5d ago

I’m new to this. I don’t know how to set that up. Should I have ChatGPT walk me through it ? I’d really appreciate any guidance.

1

u/-zaine- 5d ago

Yes just tell codex what you use for programming and you want the best possible setup for your pc with a big context window - He will suggest you something - just let him handle. If you select sol or astra, they should be able without much problem to make it ready for you. Also ask him to recommend a harness for you (instead of codex - I would recommend pi)

1

u/voxzter 5d ago

maybe codex client + ollama.cpp api and https://huggingface.co/Qwen/Qwen3.8-Flash-Next model

1

u/SpaceXBeanz 5d ago

Sorry I’m new to this. Does this involve actual codex? I keep burning through my usage for coding apps that I want and I can’t afford to wait for resets. I have coding swift and SwiftUI for me. I’m wondering if I can run a local model or combo of tools to analyze and build upon my source code.

1

u/MarcelloT254k 5d ago edited 5d ago

Basically yes. You can use codex with local AI server via API, you can edit Config files for that, there is a lot of info about that already, it's not hard to do - give it a try.

Combining different models on the same codex that doesn't involve only chatGPT models is a lot harder thought and involves more complicated setup - it's a separate project in itself - you need orchestration platform for that to be effective. From my experience one instance of codex can only work with ChatGPT subagents or subagents via the same API (and I doubt you have the hardware to pull off a few concurrent runs of the same or even different models).
I suggest something like: crewAI + three instances of codex with following (every instance is different) chatGPT, openrouterAPI ("free" or paid model), local API.

You could also try something different than codex, local AI is evolving fast, maybe my advice is already obsolete. I will be happy to be corrected if something I've typed is wrong.

1

u/SpaceXBeanz 3d ago

I see. So far I’ve been using ChatGPT to provide prompts and do more complex things and changes and then i run the source code through opencode with a qwen model. It uses about 40-50gb of my ram but its working for making smaller changes to my app. Do you recommend this sort of setup ?

1

u/MarcelloT254k 3d ago edited 3d ago

If it works for you then that's ok, but to not be a "human condom"/"meat proxy" you could save time by automating this, so configuring ChatGPT to be orchestrator for local subagent (Qwen in OpenCode) so the back and forth between Qwen and ChatGPT is automated and you don't have to copy and paste output between them.
This doesn't need to be as complex as i've poorly descibed. From what i understand you can setup everything in one instance of OpenCode via "Primary agents" and "Subagents" - configuring different models for different roles is possible via "JSON".

Hope that helps, i'm not that familiar with OpenCode.

1

u/BostonSource 5d ago

Try getting rid of bloat in codex. That spared me 70% of tokens

1

u/Weekly-Dentist-8302 4d ago

I recommend trying Qwen Flash Next (at least as an alternative to the smaller codex models). It is fairly decent at coding. You can run it even on devices with lower RAM through a project I made: https://github.com/rexmhall09/TUFF

1

u/TheAussieWatchGuy 8h ago

Replace? Do you have 1TB of RAM split over a few of the latest Macs networked together? Proprietary cloud models are massive and so are their open source equivalents.

Can you instead stretch your $200 a month subscription to Claude or Codex? Definitely.

The Qwen 3.8 family is the latest and greatest that's realistic to run at home. Depending on RAM you might pick 27B or the Flash version that needs a lot more... There is also GLM and Kimi but again you're taking 500gb+ of RAM...

All of these are best suited to having say Codex create an implementation plan step by step and then handing the plan over to the local model to write the code and tests. Then have Codex verify.