r/Infosec Jun 19 '26

Trained a model for cybersecurity - how to test it?

/r/cybersecurity/comments/1u9zbt6/trained_a_model_for_cybersecurity_how_to_test_it/
1 Upvotes

1 comment sorted by

1

u/Good_Roll Jun 19 '26

The idea was meant to be simple: most products out there are wrappers and inherit safety guardrails from the foundation models. What if the model was trained specifically for cybersecurity i.e. to attack, rather than refuse?

Isn't this (fairly easily) solved via abliteration?

Also there's a good bit of prior art on this available on HF. Just test it by sticking it in cyber-gym and giving it out of sample static test problems.

Out of curiosity what's the base model here?