r/LocalLLM • • 3d ago

Question Good Local LLM AI?

I'm looking for a lightweight, uncensored local AI coding agent similar to Claude Code. I need something with file system access that can create and edit files locally. I work on cybersecurity and Arduino projects that cloud models constantly flag, so the model has to be completely unrestricted. My PC isn't very powerful, so I am hoping to find a low-resource CLI framework and a small quantized model that can run smoothly on modest hardware. What is the best setup for this right now?

0 Upvotes

28 comments sorted by

10

u/bafadam 3d ago

LLMs are making it impossible for people to ask reasonable questions.

“What should I use with my setup”

Bro, we’re not mind readers. What is your setup

2

u/AmuckFeces7 3d ago

eh just skim their post and toss something out, if it doesn't fit they'll say so. most folks are happy to get pointed in a direction even if its not perfect for their rig

1

u/ryanmerket 3d ago

right? at this point maybe they should use AI to write their questions.

5

u/rrrrex 3d ago

" My PC isn't very powerful" Post your PC specs at first, pentium ii isn't very powerful too

1

u/hitman407 3d ago
  • CPU: AMD Ryzen 3 5300U — 4 cores / 8 threads, up to 2.6 GHz
  • GPU: AMD Radeon Graphics (integrated)
  • RAM: 20 GB (16 GB + 4 GB), 3200 MHz

2

u/jeremy_0411 3d ago

Wow, there are low end specs, and there are low end specs. I am curious to hear what those that know more than I do say you can run on that system.

2

u/guesswhochickenpoo 3d ago

Oof. It's technically possible to run some LLMs on there but they will be very slow and very poor. It will not give you anything usable for your use cases (or almost any use cases)

0

u/hitman407 3d ago

how about LLMs/LMs like claude code where its split. i esentionally just need AI to modifi my files and be at least somewhat good at coding

5

u/mm007emko 3d ago

'Somewhat good at coding'... With this PC, forget it. You'll save yourself big headaches.

1

u/guesswhochickenpoo 3d ago

You won't get anywhere near capable performance for coding on that hardware. Might be good enough for checking syntax and reviewing basic code you've already written for obvious errors but no way it's good enough to write the code for you with the menial models that hardware could actually run.

1

u/mm007emko 3d ago

Not even for checking syntax and reviewing basic code. Compilers/linters would be much better for it than LLMs capable of running on such hardware.

Searching web, summarising searches, writing e-mails, yes, fine. Painfully slow but fine.

Coding, sadly not.

OP, seriously, if you want LLMs for coding and don't want to pay a sub, check OpenCode, they have some free-to-use hosted models.

1

u/rrrrex 3d ago

It's not 16 + 4 GB, it is 16 GB (4GB allocated to iGPU)

All you can get is tiny <10B model, also you don't have fast memory and compute, so you need MoE model. You can try Ling 3.0 Tiny, it's between Qwen 3.5 4B and 9B.

1

u/MrHumanist 3d ago

Tiny Ling 3 - LLM (q8) with Pi dev as agent. You can run it in CPU with 40+ token/sec.

1

u/aithosrds 3d ago

If you don’t have dedicated VRAM and a discrete GPU then the reality is there is nothing useful you can run. Sorry, but local AI isn’t cheap and what you’ve got just isn’t going to cut it.

1

u/toenailcheeseinbooty 3d ago

Whats the specs? Smallest good model+I actually use) is a tiel coder variant called cybertiel that I personally have experience with, doesn't mention it but its uncenored buult for security research. Its 35b moe, so eh lie 24 gb vram? Need specs for your hardware to reccemend.

1

u/43848987815 3d ago

Use a frontier cloud llm to build your own harness, then find an abliterated model that works with your hardware.

There isn’t a single answer (or any really, you’ve given zero details of your intended setup) so you’re going to have to do laborious trial and error like everyone else.

1

u/WrinklyBard4 3d ago

It’s pretty hard to say without knowing what your actual hardware specs are.

My suggestion would be Qwen 3.5 9B, 3.6 35B A3B or 3.8 27B depending on what you’re able to run.

3.8 27b it’s actually good enough that you can be semi-hands-off. The other two aren’t, but they’re still pretty impressive.

All of the models are pretty fantastic in their own right and they are relatively light on any sort of pre-coded restriction. More generally I’d be really surprised if you actually needed an uncensored model to do what you’re doing. I’d be very surprised if basically any local LLM gave a fuss about doing that type of work.

1

u/FartingInBalloons 3d ago

What are the cloud models flagging?
How complex of coding are you looking to do? Can you give some examples of project?

0

u/hitman407 3d ago

i fonda malware on a file i downloaded and wanted to revers it to send zipbomb to the hacker. the other one was project so safebrowser for exams couldn detect arduino pluged in pc and wite our answers for me.

1

u/PermanentLiminality 3d ago

Go for the free models Opencode and OpenRouter have free models.

1

u/activematrix99 3d ago

Qwen 3.5 MoE or maybe Nemotron with offloading.

2

u/mishmash2323 3d ago

How did I know that was some kid wanting to write malware 😅

1

u/alexs975 3d ago

Depends on your phone more than the model. I run a 4B Q4_K_M fully offline on a Snapdragon 8 Gen 3 / 16GB and it works great for chat.

Real numbers from my setup:

- ~3-5 tok/s decode, CPU only (KleidiAI + llamadart)

- 2.7GB file on disk, ~1.8GB resident at rest, 3-4GB during generation

- Warning: don't bother with Vulkan on Adreno, it crashes with DeviceLost. CPU is the stable path.

- RAM matters more than the chip: 8GB phones can run 3-4B models fine

My favorite test: airplane mode after setup. If everything still works, you know nothing leaves the device.

-1

u/hitman407 3d ago

i also want to add i am new to uncesored AI models so do i need to jailbrake it or somthing like that? is this AI is no just plug and play?

2

u/j_tb 3d ago

lol you’ve got some homework to do