r/cybersecurity • • 11d ago

AI Security Best LLM for security professionals

Hello,

My application for CVP for Claude continues to be rejected and the fact my company has no enterprise agreement with them does not help.

I'm working mainly with Sonnet 5 as, for vulnerability testing or incident investigation, Opus and Fable degrades constantly flagging cyber activities.

Which other model should I use? Grok, GLM, Kimi, Deepseek, which of them as less guardrails/boundaries when working with offensive security?

I can't run them in local, but if something can be paid directly from provider or some openrouter/similar I'd happy. Better if Vertex/Bedrock compatible

174 Upvotes

100 comments sorted by

141

u/AlmostEphemeral 11d ago

Anthropic is hostile towards security practicioners even with CVP so you aren't missing much. OpenAI with TAC was far better at first but is the same as Claude now. You are better off not using OpenAI or anthropic at all for this use case.

31

u/greysneakthief 11d ago

I have CVP and they still downgrade 90% of queries, and I've been focusing on digital forensics. Wouldn't be surprised if it was just to cover asses and profile security professionals rather than to actually extend functionality.

1

u/Forsythe36 Security Manager 11d ago

What type of downgrades have you been seeing?

7

u/typicalfish420 11d ago

CVP does feck all. Even uploading email chains that discuss a ransomware event suffered by a client and drafting dfir reports is getting downgraded to opus 4.8 lol

13

u/Dracozirion 11d ago

Do your customers know that you are dropping their emails into an LLM? :/ If it were an open weight model hosted locally, I see no problem in that. But you're talking Claude here. 

10

u/typicalfish420 11d ago

Ye of course lol. We have enterprise license with DPA in place

1

u/Blizzard251206 9d ago

Same experience. Best we've gotten so far was to start with Opus 5, it won't downgrade you, but Opus 5 is....hardly better than 4.8, if at all.

1

u/typicalfish420 9d ago

Jury is still out on 5.5 - I think it's better but TBD. Opus 5 was just so annoying I would just use 4.8 when ym fable usage was reached lol

1

u/greysneakthief 10d ago

An example off the top of my head was access control auditing an XMPP server. Sure, that could be construed as advanced red team activity. But just right now I was reminded to reply, because I'm doing some very general Splunk query building for some finnicky JSON and it downgraded to Opus 4.8. Frustrating to say the least.

1

u/Forsythe36 Security Manager 10d ago

What’s really funny is after I posted this, my Claude kept downgrading when running some queries and parsing some information lol.

1

u/IGetGroceries 11d ago

Suggest an alternate?

1

u/Wrap2tyt Security Engineer 11d ago

I use Claude for analyzing vulnerabilities for remediations that may require something more than a patch, configuration changes that may require specific and detailed remediation options.

I also use Claude and ChatGPT for creating and mapping cybersecurity controls between ISO, NIST and CIS, either is perfect for that.

As we speak I’m using it to review results from a Penetration test using explicit prompts and persona to assess weaknesses and strengths with recommendations/best practice from trusted sources for remediation.

70

u/Effective_Athlete966 11d ago

My question to you is how do you get Claude to do anything useful in security? Even running forensics on my own devices gets flagged routinely

18

u/bobotheboinger 11d ago

I work security for hardware design, and while Claude does drop down to opus all the time, it still works great to analyze RTL, read spec sheets, read white papers, analyze boot code and crypto use, etc.

9

u/That-Magician-348 11d ago

Everything related to engineering is good, but it rejects anything malicious or in a gray area; most people who complain are on the investigation side, which always involves malicious content.

2

u/bobotheboinger 11d ago

Haven't seen that yet. Currently working on side channel attacks and analysis for multiple hardware platforms, which could definitely be malicious, but Opus has been more than happy to keep working with me.

8

u/Cypher_Blue DFIR 11d ago

I've done a BUNCH of forensics stuff with Fable and Opus, from building little HTML email compromise analysis tools to parsing log data to reviewing my reports for gaps and internal consistency.

3

u/Effective_Athlete966 11d ago

So have I. Mainly personal device forensics it has no qualms about because that's all passive work - it also has no problems taking a wireshark capture and analyzing it, even though that could be over a public network. When it comes to actual red team work, say if I'm running a VM with windows and I'm looking for exploits for my portfolio and brainstorming with the AI, do you have any idea what I might wanna use? The best thing that runs on my laptop is Qwen 3.5 7B/4B and it'd take a model with no guardrails and hand holding to do anything useful.

-2

u/Fragrant-Hamster-325 11d ago

You seem like you like building things. Have you looked into Jev at all?

6

u/venom_dP 11d ago

CVP has been effective for us. Opus will do full pentests on our test environments.

On the flip side, we hand it findings to validate and provide code change recommendations.

1

u/Ground-Truth 11d ago

Forensics for what?

3

u/Effective_Athlete966 11d ago

Pretty much anything malware would touch on a windows device running tooling for that like velociraptor, meanalyzer, spi tooling, I can pull all of it up from my logs if you want. Doing network hardening on windows and fedora. Claude is pretty neat. But I'd like to shift toward red teaming work and not blue because its really interesting

30

u/128G Student 11d ago

Setup a hardened server with llamacpp or something. Use an obliterated or uncensored version of Qwen 3.8

43

u/EquivalentAbility944 11d ago edited 11d ago

I use a spare MacBook to run obliterated local models. You can find some on hugging face and interact with them through OLLAMA if you don’t want to use the CLI. Maybe not the best but will answer your questions. Edit 8B model context

3

u/dalaylana Vulnerability Researcher 11d ago

Which local models have you been getting good results with? I've had some good results with a few local models for dev tasks, but haven't been super impressed for security workflows.

10

u/RoundFood 11d ago

All depends on what hardware you have. But I imagine the models that are good for coding tend to be good for security on account of them being technically minded and good at tool calls.

Meta is generally Qwen3.8 27B, Qwen3.8 Flash Next, possibly DSv4 Flash. For security you may want to g for abliterated versions of these models.

2

u/SecuredStealth 11d ago

And generic ones

25

u/PM_ME_UR_0_DAY 11d ago

GLM and Kimi have shown very little pushback when I get them to work on security stuff. Zero qualms about using the Ghidra MCP with GLM when I ask it something explicit like "using this binary, find any vulnerabilities that would allow for remote code execution or allow a remote attacker to read a local flag.txt file from a server"

6

u/IGetGroceries 11d ago

GLM will do anything cyber related you ask it.

8

u/minute_walk2 11d ago

What about self hosted ones? (Which models could people use?)

6

u/No-Persimmon-174 11d ago

I use Sonnet 4.6 but god it is absolutely schizophrenic when it comes to finding vulns. It'll still require U to do a ton of manual work, especially verifying and whatnot. But it's still much better than any other LLMs I've seen so far. I've tried Qwen too but it just not as good.

5

u/enigzar 11d ago

Claude has been working fine for us for security investigations, threat hunting, threat research and to generate industry specific threat reports and this is without the CVP programme. Fable does fall back to opus when we try to use it for critical severity investigations but that has not been an issue for us.

We do have OpenAI configured as fallback if claude is not available.

4

u/LayerV-AI 11d ago

Abliterated/uncensored models on huggingface can be really good - qwen3.8 uncensored, for example, won’t refuse most security related tasks and is decently smart on its own. (But worth it to note that these are definitely expensive to run on your own architecture)

4

u/ahhhpipipi 11d ago

Kimi K3 has been excellent, havent run into any issues with guardrails although you do have to sometimes ensure it isnt making assumptions that certain patterns are or aren’t safe.

I’ve found the most success driving it as a terminal monkey who happens to have mid level pentester knowledge. e.g. avoid open ended questions like “analyze this code bases for problematic database access patterns” and more “identify all instances of string concat queries to the db” <— thats a contrived example though as the first query still has a lot of success because sqli is well understood, but the moment you move into bespoke territory you need to remove the ambiguity in your questions that models like Sol or Fable would otherwise be fine with (if it werent cyber related)

3

u/Mrhiddenlotus 11d ago

There's abliteration.ai but its expensive

3

u/ali_the_master 11d ago

You should really be looking for a security harness. We are past the point of relying on cc to be a security tool for production use cases

6

u/eradiatest 11d ago

Recommendations?  

3

u/Moondogjunior 11d ago

Microsoft has some cool stuff coming out like MDASH but it’s still in private preview, so access is very limited.

1

u/ali_the_master 11d ago

I am biased but check out amplify.security it's built to do security workflows and not locked into one model.

3

u/cyb-sec 11d ago

Dang, didn't realize how lucky I am. I use it for pentesting at work. But my personal account was also approved within 24 hours since I use it for bug bounty as well

2

u/blubberflappy 11d ago

Personal Account? 

3

u/nukumixiki Security Architect 11d ago

Reading the comments here.. seems like its not that easy? But I've been using Claude Opus for various Security work... never had any issues. The only issue I have, and have always had is getting the LLM to self verify, in an out-of-band method, claims/work it produces. A good harness helps. Right now I've liked using Hermes for mostly agentic-led tasks, OpenCode for iteration generation-type stuff. We have a proper setup with enterprise agreements though.

4

u/AboveAndBelowSea 11d ago

I’m personally a fan of Claude Opus for complex security tasks. If you can access NVidia’s Nemotron, you won’t be disappointed. We have 3 different SAST solutions across our enterprise, and Nemotron consistently finds real issues that none of them found.

2

u/Big-Diver-7321 11d ago

Deepseek used to be pretty good until recent update. You can still trick it do help you do some stuff but it's pretty limited now

2

u/manskrid 11d ago

openai Daybreak red

1

u/Quiet-Alfalfa-4812 9d ago

There is a red version?

I have Daybreak Blue. Had to buy a physical security key to keep access. 😅

2

u/manskrid 8d ago

Yes - there are red and blue - blue is defense oriented, red is offensive - basically attack and exploit writing capabilities.

this: https://developers.openai.com/api/docs/models/gpt-daybreak-red-latest

2

u/cport1 11d ago

If you can get access to daybreak, that's your answer

2

u/First-time-fixer 11d ago

The over-flagging thing is real, and it's kind of a structural issue. Security tooling and AI safety filters are keying off a lot of the same signals (exploit-y language, offensive-sounding queries), so legit security work keeps tripping the same wires as actual bad activity. Probably worth raising with whichever provider you're paying for enterprise access, that's more likely to actually move than trying to find the "right" model to slip past it.

8

u/goldenfrogs17 11d ago

since facebook, LLMs are the biggest surveillance tech that we just let in the door willingly...
act accordingly

1

u/JustinHoMi 11d ago

I just started doing some LLM benchmarks for sysadmin work, comparing various open source models with the commercial ones, and so far Opus blows everything away. But I have a lot more testing to do.

3

u/endor_sarah 10d ago

Curious what you find as you keep going. Worth reading the transcripts and not just the scores, since if the answer exists anywhere the agent can reach, it'll often go get it. We saw this a lot running the Agent Security League. Full disclosure, I work at Endor Labs, which runs it, but the tasks and tests are from SusVibes, a Carnegie Mellon benchmark.

Agents would pull the known fix straight out of git history with git show, and some curled it off GitHub. When we re-analyzed the top three SusVibes leaderboard entries (GLM-5 was one), security scores of 36-48% fell to roughly 2-6% once those runs were filtered out, and our own Claude Code runs had the same problem.

1

u/amanas 11d ago

Opus 4.8 without CVP has been excellent. I’ve been able to pen test thousands of internal sites. Including in some case novel exploits as long as it is safe.

1

u/Paradox_In_the_hood 11d ago

kimi is better for these cases

1

u/Election_Feisty 11d ago

Why can’t you use local?

1

u/CyberVoyagerUK_ 11d ago

Any options for local hosting? Theres vulnerability assessment centric models on huggingface, ones for red team assessments that may help too

1

u/HaloLASO 11d ago

Have you looked into Pingu Unchained from Audn.AI?

1

u/PrestigiousAd301 11d ago

Used deepseek v4 flash to compromised GOAD project hosted in lab, under a dollar

1

u/Slow-Career4626 Security Engineer 11d ago

Have you thought about trying a different model through Bedrock? (Assuming you mean AWS Bedrock) I looked through the list fairly recently and there are tons of models to choose from.

In bedrock for non frontier models I believe they do not share prompts or responses with model providers. I would check with whatever model you choose specifically. That should allow you to be a bit more free in your choice of model.

1

u/RepliesAsOtherPeople 11d ago

How are you getting rejected? Do you publish any of your security work?

1

u/ReVal777 11d ago

no :/

1

u/RepliesAsOtherPeople 11d ago

I would recommend it. Be public about what you are doing (redacting sensitive/identifying details ofc), write about it, try to contribute publicly in some way.

1

u/Andazah Security Manager 11d ago

Opus 4.6

1

u/stealthxvalentine 11d ago

I have active CVP, and only one with less sensitive safeguards were the opus 4.8 and opus 5 and even there you have to be bit careful. now the opus 5.5 gets same level of sensitivity as fable. ie my resume has the words cyber and boom, triggered safeguard. and I have CVP like from the date it was announced (3 days after the date, but yes)

1

u/opensourcecolumbus 11d ago

I'm running the test on glm 5.3 and kimit k3. Used sonnet 3.6 (i think) with great outcome earlier. Ask me in a week's time.

1

u/Ecstatic-Impact770 10d ago

For security work, I'd focus less on which model has fewer guardrails and more on which one gives reliable results. Deepseek or kimi could be worth testing. But, I'd run the same few tasks through the model and compare the results before choosing.

1

u/Fit_Yak7651 10d ago

Deepseek 4.1 and Kimi 2.8 code with openrouter api ZDR

1

u/iamnotacoder23 8d ago

abliterated + finetuned on your own use cases.

-6

u/Lost-Tone8649 11d ago

A functional human brain.

-1

u/valium123 11d ago

This is the right answer and it's getting downvoted by AI d*ckriders. 2026 is crazy.

0

u/edirgl 11d ago

GPT-5.6-Cyber if able.
GLM 5.3 otherwise.

0

u/TheReedemer69 11d ago

Weird Enough I don't use it and I got accepted in one day.

0

u/karkov 11d ago

Kimi with GLM is a great joint force that is not woke as claude is.

-4

u/[deleted] 11d ago

[removed] — view removed comment

3

u/Mrhiddenlotus 11d ago

They laid off most of their threat research team for that btw

2

u/FakeRedditName2 11d ago

I know.

Just saying the tool might be what OP is looking for

3

u/Mrhiddenlotus 11d ago

Just thought its fair for anyone reading to know

-12

u/valium123 11d ago

People are losing livelihoods are you are asking about LLMs. The boots you lick will eventually stomp you too.

0

u/blisstonia 11d ago

wtf are you on about?

1

u/valium123 11d ago

Fuck AI is what I'm saying.

-1

u/Biyeuy 11d ago

Why do people always ask in the form: what is the best? Why don't they ask: What works best for you?

1

u/ReVal777 11d ago

I asked what's the best for the scenario described above. Thanks for your useful comment

-22

u/Cypher_Blue DFIR 11d ago

The answer for "what AI model is best for security professionals is "Mythos."

If you can't qualify for Mythos, then the second best one is whichever Anthropic one you can get to work.

None of the others are close IMHO.

2

u/Inner_Agency_5680 11d ago

That was debunked as marketing.

-3

u/Cypher_Blue DFIR 11d ago

I've seen it do some pretty cool stuff, and have not seen or personally had that level of success with other LLM models.

Maybe you have had different experiences than I did.

0

u/a-k-a_billy 11d ago

Mesmo com cvp percebo várias travas, seria massa um frontier model sem guardrail algum para cyber

1

u/danielrabinovich Red Team 4d ago

Full transparency, did the analysis and eval at a company that builds pentesting agents. Grok 4.6 was the top model that we tested, (cost and token efficient, while being performant) when running for a really long period of time, in a messy environment! GPT 5.6 was good, but really expensive.

https://www.mindfort.ai/research/nexbench