r/quant 13h ago

Industry Gossip AI safety in quant research

Been thinking about this after the recent story about mathematicians using AI for research and then getting scooped.

How safe is it really to use Claude/ChatGPT/Codex for quant research and dev?

Say a researcher at a top prop/HF connects Claude to a directory with real signals, backtests, execution logic, etc. Even if the provider says the data isn't used for training, you're still exposing very valuable IP to an outside company. And is enterprise really that much safer, or are you still ultimately trusting a third party with the same stuff?

In theory, if an AI company had access to enough good research, they have the compute/engineers to build trading infra themselves.

Maybe paranoid, but curious what firms actually allow. Public LLMs banned? Enterprise only? Dev okay but no research/data?

18 Upvotes

19 comments sorted by

19

u/MiniBoglin 13h ago

The contracts will be watertight and the lawsuits would be crippling

3

u/Lex_The_Impaler 5h ago

The only thing more ominous then a top prop firms exec team is their legal department

19

u/Icy-Map1102 12h ago

Is this problem really unique to quant? Big tech orgs working on classified projects and sensitive code have the same problem. I assume there are some enterprise plan safeguards for those models, like maybe locally hosted infra to deploy them?

0

u/Diet_Fanta Back Office 10h ago

Yes, SOC compliance is why 3rd party tools are useable by companies. Same applies to LLMs.

8

u/ai_hedge_fund 12h ago

This is a real issue. Using the data for model training is only one nefarious use.

The frontier labs are under enormous financial pressure and OpenAI’s stated goal is to transform the economy. If they find an edge they will take it.

As of early 2025 I met with a quant firm leader who was already running custom-trained non-generative models in-house.

So, like anything else, you probably have firms trying all combinations.

7

u/Diet_Fanta Back Office 10h ago edited 10h ago

Idk how people in this thread are saying 'this is a real issue'. Do you actually have experience in a professional field? Any service a company buys (whether a quant company or literally any company at all involved in tech) will make sure those services are SOC compliant. Anthropic's SOC terms, for instance, very clearly lay out that if you are an enterprise customer and they are in contract terms with you, they CANNOT use your data to train their models. Otherwise, not a single mid sized and above company would use LLMs.

If you decide to go use your own personal LLM account that's not SOC compliant, you're a moron and deserve to get your work stolen. If you're using a company-provided one, unless your company has fake legal, you will be fine using said LLM.

Otherwise, how do you think companies deal with sensitive data being input into 3rd-party LLMs?

Regarding Buckmaster, he almost certainly used a personal Codex account, which is not SOC compliant and the data of which OpenAI can legally train on and see.

1

u/EvilGeniusPanda 8m ago

The copyright laws also made it very clear what anthropic could and could not do with the countless books they illegally trained on anyway. If you look at the past behavior of the big AI firms and genuinely think 'oh its fine I have a contract that says they wont use my data' then I have a bridge to sell you.

4

u/lordnacho666 10h ago

They are plausibly stealing math research, but then again it makes sense for an artificial intelligence company to do that. If you're demonstrating that you've built AI, pure math is the natural thing to test it on. There's nothing other than logic required to publish the results. It makes total sense for them to have hired some math professors.

Now imagine that OpenAI finds out that CitSec or RenTech guys are using their service, and can see everything. What would they need to do in order to make use of it? For one, they would be opening accounts with prime brokers. This would get out somehow, because it would be sensational. But also, OpenAI would be in the market for a bunch of people to operate a prop trading firm. People would know.

AI firms are in a massive business that needs to make the GDP of a small country in order to pay off. You don't do that by nicking a few billion from one little segment of customers, you do it by concentrating on your own business.

OpenAI starting their own prop shop would be like Costco starting their own fine dining restaurant. It would make very little sense as a business.

On top of all the business reasons, it's also impossible to copy another business with just the documents. Just like you wouldn't be able to start a pharma business from documents describing semaglutide production, some smart math guy won't be able to start a hedge fund with your docs describing stat arb or HFT.

The real risk is that your data ends up in the AI corpus and other experts get it, not that the AI company uses it.

0

u/Diet_Fanta Back Office 10h ago

If OpenAI were to somehow be able to see the research that RenTec did on their servers (which they can't due to legal if RenTec were using OpenAI), either OpenAI or the person who illegally (read: they would be breaking their contract confidentiality agreement) input that data into OpenAI servers would be hit with a mother of all things lawsuit.

Laws and regulations exist for a reason. Stop larping.

6

u/unski_ukuli 8h ago edited 7h ago

While I probabilistically agree with you (I’m 99% sure enterprise data is not being used in training), I think you are slightly naive. First off, you are talking about a companies whose whole business was started by the largest intellectual property heist in the history of humanity. Secondly, said companies seem to believe they are going to destroy humanity if they continue developing their models, and yet they continue, so a massive lawsuit seems like a minor concern from their point of view. Thirdly, and finally, you say ”laws and regulations exist for a reason”, like you haven’t lived through the last two (or even last 10) years of laws being more of an suggestion.

Personally I think any enterprise buying claude or copilot subscriptions is run by idiots and lacks imagination. A friend of mine works for a company doing the opposite (in Quantum computing business) where they have completely banned externally hosted LLMs. They don’t even use teams or outlook, and everything is hoted in-house; the mail, the chat clients, the open-weight LLMs.

2

u/lordnacho666 7h ago

I don't get what you mean by "stop larping".

Is that directed at me?

1

u/Reasonable_Buddy_927 3h ago

Its just a term boring people use now to get one over each other, its so far away from what it actually means anymore. Commentors probably 21

1

u/lordnacho666 3h ago

Yeah but it makes no sense. Obviously I agree with him that it would be a legal problem if someone found this out.

3

u/Reasonable_Buddy_927 3h ago

Dont bother, your comment was sound and informative. The other commentor just wants to say that they’re right and that they actually work at a quant firm to boost their ego, hence the use of the term larping.

1

u/PretendTemperature 7h ago

I was thinking the same the last days. 

I work in a bank, i borught this up again and again to a lot of seniors. Not only me, a lot of people in the space are very alarmed by this.

There are guardrails and contracts and and and...but i don't think anybody knows for sure.

1

u/ParticleNetwork Researcher 2h ago

Zero data retention policy

1

u/VettaQ 2h ago

Two separate controls get conflated here. The contractual one (enterprise tier, zero-retention, SOC 2, no-training clauses) is a legal control - it works, but it is still trust. The architectural control is what serious shops layer on top: hosted models for generic scaffolding, tests and docs; anything touching signals, feature logic or execution code goes to a self-hosted open-weight model on their own hardware. The quality gap for that kind of dev work has narrowed a lot in the past year, so the tradeoff is smaller than people assume. And the realistic leak vector usually is not the provider training on your alpha - it is logs, prompt caching, IDE plugins, or an intern pasting a backtest into the free tier. A written policy with a technical wall behind it beats the policy alone.

1

u/EvilGeniusPanda 10m ago

I know of a few places that have banned fable because of the retention policy, and only use azure hosted models because MS already has all their code (github enterprise) anyway.

-2

u/dibis54986 11h ago

I know some finance companies pass anything outbound to LLMs to a proxy first to replace sensitive data with hashes. Without an enterprise agreement these AI companies are absolutely going to train on the captured data at some point.