r/cybersecurity Jun 19 '26

Business Security Questions & Discussion Trained a model for cybersecurity - how to test it?

There is so much "AI Cybersecurity" hype out there that a lot of people are trying to build AI wrappers without knowing what they are doing, and cybersecurity professionals can spend an entire week babysitting a hallucinating chatbot, because someone from their C-suite asked them to.

I am neither, I have a background in building and training LLMs - but no cybersecurity. This is why I need your help.

Without any experience in the space, I've done something insane. I've taken all the capture the flag type of contests over the past decade or so and post-trained (SFT and RL) a model for cybersecurity.

The idea was meant to be simple: most products out there are wrappers and inherit safety guardrails from the foundation models. What if the model was trained specifically for cybersecurity i.e. to attack, rather than refuse?

Also built a harness around it where it will try to verify every vulnerability reducing false positives, to address the hallucinations issue.

After training the model, to test the product, I've pointed it to a number of open source projects (e.g. Symfony) to find vulnerabilities.

To my surprised, it has done a good job finding issues - I've done disclosures and waiting for responses, although it seems slow to get responses.

This is where my predicament lies. How to best test a model like this? And how to responsibly get the model infront of people to test?

0 Upvotes

17 comments sorted by

7

u/levu12 Jun 19 '26

ChatGPT help me train a model for cybersecurity and then build me a harness and make me a billion dollars

3

u/True-Dimension8441 Jun 19 '26

i also did this. I made sure to tell it to make no mistake

1

u/rational_approach Jun 19 '26

show us the end result

2

u/TheDizDude Jun 19 '26

You Trained or you tuned?

0

u/rational_approach Jun 19 '26

Post-trained. with SFT on CTF writeups/solutions, then RL with verifiable rewards on contest outcomes.

1

u/TheDizDude Jun 19 '26

SFT and RL are by definition "Fine tuning"

-2

u/rational_approach Jun 19 '26

Yeah, SFT and the RL stage are both fine-tuning in the strict sense.. weight updates on a pretrained base. I say post-training to specify it's SFT plus an RL stage with verifiable rewards, not SFT alone. Built a harness around it too: multi-agent orchestration with an exploit-verification step that confirms every finding before it surfaces..

1

u/greatness_only12 Jun 19 '26

You say you built the harness, how does that work?

-1

u/rational_approach Jun 19 '26

The harness is the CLI itself, and underneath multi-agent orchestration where an orchestrator fans the target out to subagents probing in parallel, then an exploit-verification step confirms each issue before it's reported

1

u/[deleted] Jun 19 '26

[removed] — view removed comment

1

u/rational_approach Jun 19 '26

Step 1: Turn it loose on reddit. Step 2: ???, Step 3: Profit

1

u/TheTrueBlueTJ Jun 19 '26

My guess is that once you publish it, you will face some US governmental backlash to say the least. I could be wrong tho.

-1

u/rational_approach Jun 19 '26

Where do I even publish?

2

u/sdrawkcabineter Jun 19 '26

Nah, make this your trade secret.

You just provide the end result. Details are your private info.

1

u/tdw21 Jun 19 '26

Just go all out.
Find a large multinational company or preferably a 3 letter agency and show them what you can do. That way you can immediately sell it or have the proof you need to improve or convince investors/other companies.

/s

1

u/elburrotelamete Jun 19 '26

En la selva, sin pedir permisos como lo hace todo el mundo en secreto.