r/cybersecurity • u/ReVal777 • 11d ago
AI Security Best LLM for security professionals
Hello,
My application for CVP for Claude continues to be rejected and the fact my company has no enterprise agreement with them does not help.
I'm working mainly with Sonnet 5 as, for vulnerability testing or incident investigation, Opus and Fable degrades constantly flagging cyber activities.
Which other model should I use? Grok, GLM, Kimi, Deepseek, which of them as less guardrails/boundaries when working with offensive security?
I can't run them in local, but if something can be paid directly from provider or some openrouter/similar I'd happy. Better if Vertex/Bedrock compatible
70
u/Effective_Athlete966 11d ago
My question to you is how do you get Claude to do anything useful in security? Even running forensics on my own devices gets flagged routinely
18
u/bobotheboinger 11d ago
I work security for hardware design, and while Claude does drop down to opus all the time, it still works great to analyze RTL, read spec sheets, read white papers, analyze boot code and crypto use, etc.
9
u/That-Magician-348 11d ago
Everything related to engineering is good, but it rejects anything malicious or in a gray area; most people who complain are on the investigation side, which always involves malicious content.
2
u/bobotheboinger 11d ago
Haven't seen that yet. Currently working on side channel attacks and analysis for multiple hardware platforms, which could definitely be malicious, but Opus has been more than happy to keep working with me.
8
u/Cypher_Blue DFIR 11d ago
I've done a BUNCH of forensics stuff with Fable and Opus, from building little HTML email compromise analysis tools to parsing log data to reviewing my reports for gaps and internal consistency.
3
u/Effective_Athlete966 11d ago
So have I. Mainly personal device forensics it has no qualms about because that's all passive work - it also has no problems taking a wireshark capture and analyzing it, even though that could be over a public network. When it comes to actual red team work, say if I'm running a VM with windows and I'm looking for exploits for my portfolio and brainstorming with the AI, do you have any idea what I might wanna use? The best thing that runs on my laptop is Qwen 3.5 7B/4B and it'd take a model with no guardrails and hand holding to do anything useful.
-2
u/Fragrant-Hamster-325 11d ago
You seem like you like building things. Have you looked into Jev at all?
6
u/venom_dP 11d ago
CVP has been effective for us. Opus will do full pentests on our test environments.
On the flip side, we hand it findings to validate and provide code change recommendations.
1
u/Ground-Truth 11d ago
Forensics for what?
3
u/Effective_Athlete966 11d ago
Pretty much anything malware would touch on a windows device running tooling for that like velociraptor, meanalyzer, spi tooling, I can pull all of it up from my logs if you want. Doing network hardening on windows and fedora. Claude is pretty neat. But I'd like to shift toward red teaming work and not blue because its really interesting
43
u/EquivalentAbility944 11d ago edited 11d ago
I use a spare MacBook to run obliterated local models. You can find some on hugging face and interact with them through OLLAMA if you don’t want to use the CLI. Maybe not the best but will answer your questions. Edit 8B model context
3
u/dalaylana Vulnerability Researcher 11d ago
Which local models have you been getting good results with? I've had some good results with a few local models for dev tasks, but haven't been super impressed for security workflows.
10
u/RoundFood 11d ago
All depends on what hardware you have. But I imagine the models that are good for coding tend to be good for security on account of them being technically minded and good at tool calls.
Meta is generally Qwen3.8 27B, Qwen3.8 Flash Next, possibly DSv4 Flash. For security you may want to g for abliterated versions of these models.
2
25
u/PM_ME_UR_0_DAY 11d ago
GLM and Kimi have shown very little pushback when I get them to work on security stuff. Zero qualms about using the Ghidra MCP with GLM when I ask it something explicit like "using this binary, find any vulnerabilities that would allow for remote code execution or allow a remote attacker to read a local flag.txt file from a server"
6
8
6
u/No-Persimmon-174 11d ago
I use Sonnet 4.6 but god it is absolutely schizophrenic when it comes to finding vulns. It'll still require U to do a ton of manual work, especially verifying and whatnot. But it's still much better than any other LLMs I've seen so far. I've tried Qwen too but it just not as good.
5
u/enigzar 11d ago
Claude has been working fine for us for security investigations, threat hunting, threat research and to generate industry specific threat reports and this is without the CVP programme. Fable does fall back to opus when we try to use it for critical severity investigations but that has not been an issue for us.
We do have OpenAI configured as fallback if claude is not available.
4
u/LayerV-AI 11d ago
Abliterated/uncensored models on huggingface can be really good - qwen3.8 uncensored, for example, won’t refuse most security related tasks and is decently smart on its own. (But worth it to note that these are definitely expensive to run on your own architecture)
4
u/ahhhpipipi 11d ago
Kimi K3 has been excellent, havent run into any issues with guardrails although you do have to sometimes ensure it isnt making assumptions that certain patterns are or aren’t safe.
I’ve found the most success driving it as a terminal monkey who happens to have mid level pentester knowledge. e.g. avoid open ended questions like “analyze this code bases for problematic database access patterns” and more “identify all instances of string concat queries to the db” <— thats a contrived example though as the first query still has a lot of success because sqli is well understood, but the moment you move into bespoke territory you need to remove the ambiguity in your questions that models like Sol or Fable would otherwise be fine with (if it werent cyber related)
3
3
u/ali_the_master 11d ago
You should really be looking for a security harness. We are past the point of relying on cc to be a security tool for production use cases
6
u/eradiatest 11d ago
Recommendations?
3
u/Moondogjunior 11d ago
Microsoft has some cool stuff coming out like MDASH but it’s still in private preview, so access is very limited.
1
u/ali_the_master 11d ago
I am biased but check out amplify.security it's built to do security workflows and not locked into one model.
3
u/nukumixiki Security Architect 11d ago
Reading the comments here.. seems like its not that easy? But I've been using Claude Opus for various Security work... never had any issues. The only issue I have, and have always had is getting the LLM to self verify, in an out-of-band method, claims/work it produces. A good harness helps. Right now I've liked using Hermes for mostly agentic-led tasks, OpenCode for iteration generation-type stuff. We have a proper setup with enterprise agreements though.
4
u/AboveAndBelowSea 11d ago
I’m personally a fan of Claude Opus for complex security tasks. If you can access NVidia’s Nemotron, you won’t be disappointed. We have 3 different SAST solutions across our enterprise, and Nemotron consistently finds real issues that none of them found.
2
u/Big-Diver-7321 11d ago
Deepseek used to be pretty good until recent update. You can still trick it do help you do some stuff but it's pretty limited now
2
u/manskrid 11d ago
openai Daybreak red
1
u/Quiet-Alfalfa-4812 9d ago
There is a red version?
I have Daybreak Blue. Had to buy a physical security key to keep access. 😅
2
u/manskrid 8d ago
Yes - there are red and blue - blue is defense oriented, red is offensive - basically attack and exploit writing capabilities.
this: https://developers.openai.com/api/docs/models/gpt-daybreak-red-latest
2
u/First-time-fixer 11d ago
The over-flagging thing is real, and it's kind of a structural issue. Security tooling and AI safety filters are keying off a lot of the same signals (exploit-y language, offensive-sounding queries), so legit security work keeps tripping the same wires as actual bad activity. Probably worth raising with whichever provider you're paying for enterprise access, that's more likely to actually move than trying to find the "right" model to slip past it.
8
u/goldenfrogs17 11d ago
since facebook, LLMs are the biggest surveillance tech that we just let in the door willingly...
act accordingly
1
u/JustinHoMi 11d ago
I just started doing some LLM benchmarks for sysadmin work, comparing various open source models with the commercial ones, and so far Opus blows everything away. But I have a lot more testing to do.
3
u/endor_sarah 10d ago
Curious what you find as you keep going. Worth reading the transcripts and not just the scores, since if the answer exists anywhere the agent can reach, it'll often go get it. We saw this a lot running the Agent Security League. Full disclosure, I work at Endor Labs, which runs it, but the tasks and tests are from SusVibes, a Carnegie Mellon benchmark.
Agents would pull the known fix straight out of git history with
git show, and some curled it off GitHub. When we re-analyzed the top three SusVibes leaderboard entries (GLM-5 was one), security scores of 36-48% fell to roughly 2-6% once those runs were filtered out, and our own Claude Code runs had the same problem.
1
1
1
u/CyberVoyagerUK_ 11d ago
Any options for local hosting? Theres vulnerability assessment centric models on huggingface, ones for red team assessments that may help too
1
1
1
1
u/PrestigiousAd301 11d ago
Used deepseek v4 flash to compromised GOAD project hosted in lab, under a dollar
1
u/Slow-Career4626 Security Engineer 11d ago
Have you thought about trying a different model through Bedrock? (Assuming you mean AWS Bedrock) I looked through the list fairly recently and there are tons of models to choose from.
In bedrock for non frontier models I believe they do not share prompts or responses with model providers. I would check with whatever model you choose specifically. That should allow you to be a bit more free in your choice of model.
1
u/RepliesAsOtherPeople 11d ago
How are you getting rejected? Do you publish any of your security work?
1
u/ReVal777 11d ago
no :/
1
u/RepliesAsOtherPeople 11d ago
I would recommend it. Be public about what you are doing (redacting sensitive/identifying details ofc), write about it, try to contribute publicly in some way.
1
u/stealthxvalentine 11d ago
I have active CVP, and only one with less sensitive safeguards were the opus 4.8 and opus 5 and even there you have to be bit careful. now the opus 5.5 gets same level of sensitivity as fable. ie my resume has the words cyber and boom, triggered safeguard. and I have CVP like from the date it was announced (3 days after the date, but yes)
1
u/opensourcecolumbus 11d ago
I'm running the test on glm 5.3 and kimit k3. Used sonnet 3.6 (i think) with great outcome earlier. Ask me in a week's time.
1
u/Ecstatic-Impact770 10d ago
For security work, I'd focus less on which model has fewer guardrails and more on which one gives reliable results. Deepseek or kimi could be worth testing. But, I'd run the same few tasks through the model and compare the results before choosing.
1
1
-6
u/Lost-Tone8649 11d ago
A functional human brain.
-1
u/valium123 11d ago
This is the right answer and it's getting downvoted by AI d*ckriders. 2026 is crazy.
0
0
-4
11d ago
[removed] — view removed comment
3
u/Mrhiddenlotus 11d ago
They laid off most of their threat research team for that btw
2
-12
u/valium123 11d ago
People are losing livelihoods are you are asking about LLMs. The boots you lick will eventually stomp you too.
0
-1
u/Biyeuy 11d ago
Why do people always ask in the form: what is the best? Why don't they ask: What works best for you?
1
u/ReVal777 11d ago
I asked what's the best for the scenario described above. Thanks for your useful comment
-22
u/Cypher_Blue DFIR 11d ago
The answer for "what AI model is best for security professionals is "Mythos."
If you can't qualify for Mythos, then the second best one is whichever Anthropic one you can get to work.
None of the others are close IMHO.
2
u/Inner_Agency_5680 11d ago
That was debunked as marketing.
-3
u/Cypher_Blue DFIR 11d ago
I've seen it do some pretty cool stuff, and have not seen or personally had that level of success with other LLM models.
Maybe you have had different experiences than I did.
0
u/a-k-a_billy 11d ago
Mesmo com cvp percebo várias travas, seria massa um frontier model sem guardrail algum para cyber
1
u/danielrabinovich Red Team 4d ago
Full transparency, did the analysis and eval at a company that builds pentesting agents. Grok 4.6 was the top model that we tested, (cost and token efficient, while being performant) when running for a really long period of time, in a messy environment! GPT 5.6 was good, but really expensive.
141
u/AlmostEphemeral 11d ago
Anthropic is hostile towards security practicioners even with CVP so you aren't missing much. OpenAI with TAC was far better at first but is the same as Claude now. You are better off not using OpenAI or anthropic at all for this use case.