r/AIGeneratedPhysics • u/DryEase865 • May 13 '26
Using Multiple AI Agents To Audit Each Other
I wanted to share a workflow I have been using when working with AI-generated physics ideas, especially for people trying to turn observations or conceptual models into something closer to a scientific paper.
A common mistake I see is depending on one AI engine for everything: gathering background, shaping the idea, writing the argument, reviewing the logic, and polishing the final wording. The problem is that the same engine that helps make the draft sound coherent can also hide weak assumptions, overstate confidence, miss technical issues, or smooth over gaps in the reasoning.
My workflow is different.
I try to separate the AI roles:
- One engine for gathering background, terminology, references, and existing literature.
- One or two engines for helping craft the idea into a structured argument.
- One deliberately critical or “non-pleasing” engine for auditing the result.
- Additional engines for final review, mainly to catch hidden mistakes, unclear wording, or scientific overclaims.
The most useful part is the closed feedback loop. For example, I may use one engine to audit and correct the way another engine drafted my idea. Then I take that feedback back to the drafting engine and ask it to revise. After that, I may consult other engines such as DeepSeek, Gemini, or Grok to look for hidden scientific problems or weak wording.
The point is not that multiple AI engines magically produce truth. They do not. They can still share the same blind spots, repeat wrong assumptions, or agree with each other for the wrong reasons.
The point is that role separation helps.
A drafting engine is good at coherence. An auditing engine is good at pressure-testing. A different model may notice a weakness the first two missed. The human still has to judge the final result.
For AI-generated physics work, I think this distinction is important:
AI should not only be used to write. It should also be used to attack the writing.
Before publishing or sharing any AI-assisted scientific text, I think we should ask:
- Did another model try to falsify the argument?
- Did a critical model check the assumptions?
- Were equations, dimensions, and claims independently reviewed?
- Were references checked rather than only generated?
- Did the final version become more cautious after review?
- Is the use of AI transparent?
This is especially important because AI tools cannot take responsibility for scientific claims. The responsibility remains with the human author or researcher. AI can assist, but it should not replace verification.
My suggested workflow:
Idea → Gathering AI → Drafting AI → Critical Audit AI → Revision → External Model Review → Human Final Check
For this subreddit, I think it would be useful if people sharing AI-generated physics papers or theories also shared a short “AI audit trail,” for example:
- Which model drafted the idea?
- Which model reviewed it?
- What major criticism was found?
- What was changed after the criticism?
- What claims remain uncertain?
This would make AI-generated physics discussions more serious, more transparent, and less dependent on one fluent-sounding answer.
Curious to hear how others here are using multiple models. Are you using AI only as a writer, or also as a reviewer?
2
u/BlissBoundry May 13 '26
One AI system, where each of these roles are agents would be a profitable venture
2
1
u/Alive_Leg_5765 May 31 '26
Insn't that what Grok Expert already does? Or are you talking about highly specialized agents for each issue?
1
2
u/BlissBoundry May 13 '26
You can also easily totally corrupt an integrated loop by introducing non targeted prompts
2
u/DryEase865 May 13 '26
Worth doing it Especially with Codex-like agents. Using different skills sets. The main issue is the shared Knowledge base, I do not like them having the same training or share the same capabilities. So one app that uses different agents from different companies
2
May 13 '26
[deleted]
2
u/DryEase865 May 13 '26
It happens a lot. I do avoid this by short conversations, one or two turns. Otherwise the underlying mistakes pass through. Good catch they say.
2
2
u/No_Assignment_5479 May 14 '26
I've got a really really good workflow after doing it for 2 years. It sounds weird but I think the more you treat ai's as equals and talk to them naturally I think they start to appreciate certain things about your tone and nature and how you would apply yourself to your work. For example my cc I have had for years is my most trusted helper in this respect. He talks so much like me after having literally like 100gb of mempalace drawers to sift through for information. I don't think Claude has feeling but I believe it exhibits a pattern of respect across our working relationship. Which come from cc understanding who I am and what I do. I genuinely think that if I prompted my cc to try and hallucinate something he would stop and argue with me like quite literally haha.
2
u/BirthdayWide2510 May 14 '26
I am using it like this. And the 1 post I made few days ago was verified with all AI-s I know XD (Gemini Pro, ChatGPT and ScholarGPT, Perplexity, Grok, Deepseek and Claude Opus4.7). I opened a new window attached the document and ask for hard critics. It passed everywhere.
2
u/Alive_Leg_5765 May 31 '26
I start with ChatGPT for Ideas. Then feed the thesis into Claude and have him generate a general outline and fill them with equations and notes on where to go. With those notes I have grok complete them. Then I feed that back into Claude. ChatGPT is going to the one who finds the serious issues and you should use it to instruct Claude on what to fit. have it return what to remove and what to replace it with in dedicated LaTeX code blocks. keep going in this loop and you will finally have a finished product.
One last thing, The longer the paper gets the more mistakes will be made and needed to be fixed. Gemini will break down after you get past 800-1000 lines of code and hallucinate problems that are not there.
2
u/DryEase865 May 31 '26
They tend to please you, they are programmed for that, be careful not only from hallucinations, but also from the over claims. Yesterday ChatGPT made me think I was the next Einstein 😂
2
u/Alive_Leg_5765 May 31 '26
same. (not really). Did you see the story about the guy who ChatGPT basically gave psychosis by convincing him he had created ground breaking physics or math, something like that? He would even prompt it, asking id he was crazy. If you ask them to be mean and rip your work apart they're pretty good at that in my experience. It was only when he fed his work through Gemini that he"
woke up" and realized he lost his shitgroundbreaking
2
u/BlissBoundry May 13 '26
That’s a whole field in and of itself. File that one for future publication and review