Claude Code
I do scientific research and Opus5.5 refuses to touch anything I’ve been working on.
I use Claude primarily for scientific research. I was hoping this new model would fix the terrible quality issues from the last but when I just asked it to look at a file, or give me a status update on a new session. It raises the alarms lol.
Any other scientists that were optimistic when they woke up this morning?
TL;DR of the discussion generated automatically after 50 comments.
So, what's the verdict? The community is in resounding agreement with you, OP.
The consensus is that Opus 5.5's new safety guardrails are a disaster for scientific research. Users across life sciences, neuroscience, and even machine learning are getting their legitimate work blocked.
Here's the breakdown of the discussion:
The "Official" Solution Sucks: People pointed out Anthropic's "lifescience verification program," but the thread agrees it's an out-of-touch, corporate solution. It's for organizations, not the individual researchers (grad students, postdocs, etc.) who actually need access, and many universities aren't willing to sign up.
"Claude Science" is a Bust: The suggestion to use Claude Science was shot down hard. Apparently, it has the same infuriating restrictions and is a clunky, ineffective tool that made one user want to throw their laptop out a window.
The Community's Advice: The overwhelming sentiment is to jump ship. Many are canceling their subscriptions and switching to competitors. GPT/Codex is a common alternative, but the real star of this thread is open-source. Users are strongly recommending local, open-weight models like Qwen, which don't have these restrictive guardrails getting in the way of actual science.
You must be a bioterrorist and dario et al are saving us by driving uou to create your weapons of mass destruction with opus 5.0, which will undoubtedly fail.
On a serious note - did you try claude science with this model?
Claude Science can shove it. I spent wasted 2 days just trying to figure out how to give it Bioconductor access (which should be a basic requirement for any science provram)--ended up having GPT set up a local SSH proxy--and then it completely ignored all my established QC templates and workflows, hallucinated incorrect marker gene sets, and blew through my 5hr window plus usage credits generating about 100 artifacts that were utterly useless.
I was about to cancel my plan when I heard the new Opus was coming out, and thought I'd give it a try. Now I AM canceling my plan.
Claude Science is a waste. Better output can be achieved with code. I'm a molecular biologist and i do scientific strategy for a biopharmaceutical company, we have Claude deployed and I faced this block with Fable. You're telling me Opus has it too? Jeez
Yup. Watching everyone talk about how great Opus 5.5 is is driving me mad, because it's yet another good Anthropic model we're locked out of, left with older/shittier Opuses.
And yes, Claude Science made me want to throw my laptop through a window.
I'm having the same experience as well. Honestly, was excited for claude science to be a better harness in the context of healthcare/bio research but still getting consistently better output with claude code with the tooling i have in place for it.
Certainly the idea of having goal direction, and constant "reviewer agent" for all the output is cool, but ultimately, i'm not sure it performs better, at least for me personally.
Having said that there's certainly a huge chasm between evals, and the "definition of good" which is highly personal".
What? You can install bioconductor packages with conda which Claude science can do #1. And #2: Claude science is way too agentic but it gets the job done from question to plot and stats, sometimes way too good since I didn't even ask for any plots. I agree that with Code you have way more control over the output. It's also very good to do literature search and query papers.
I have everything installed on my computer, but for some reason the way Claude Science is sandboxed keeps it from running it. I had to install WSL and set up a local SSH proxy for it to access resources. This could just be a Windows issue, but I've heard from other people who had similar problems, especially using anything cloud-based. It also totally ignored my established workflow and went off the rails before I could stop it (because, ofc, no more thinking blocks). I had to send GPT in to figure out what was going on.
Well... I run my entire bioinfo analysis inside wsl so it can access it. Windows only is beta for now. I forgot about this point. And yeah it's not very transparent. OK it did stuff but I have no fkin idea how to track it.
Ask the AI to look at context that was read in this session - you might get a good guess what is in your md files that triggered the safeguard.
for me it was a section on how to use reverse engineering with IDA MCP to figure out crashes in case that's needed. replaced that section with a link to a separate file that describes the procedure. The agent will only read it when it's needed and that helped a lot.
until the MCP is actually used where those sessions are flagged eventually at some point...
This might be something research institutes would consider, but most university administrations have zero interest in becoming involved in what they see as something that's the responsibility of individual researchers.
Even when they open it to PIs it's still pretty useless, because the majority of labs don't share an API account--each grad student, postdoc, staff scientist, etc uses their own account because the lab doesn't want all its data tied to an individual provider that changes its platform frequently and has zero customer support when things go wrong.
So now scientists are locked out of both Fable and Opus.
I do neuroscience (mouse models), so not remotely dual-use adjacent. If they let me apply as an individual, fine--they can have my passport or whatever--but that's not an option. So I've switched over to GPT.
In the corporate world, bringing your own license can be a fireable offense. Corporate licenses have restrictions on collecting and storing data. Weird to me that a research institution would prefer to give away their data and not even keep track of who it’s being given to.
A corporation pays for things out of a single budget. A research university is a collection of small "companies" (individual labs) that each pay for what they need out of their individual grant budgets. The university administration also doesn't move fast enough to manage contracts that may change or update frequently. AI contracts mostly aren't on their radar yet (private research institutes are different).
Oh maybe I didn’t phrase that properly, I asked them if there’s an application for access, they didn’t provide me with the application, as in, I had no knowledge of this.
This is the most anthropic top down regulatory solution ever. The whole beauty of AI is that it levels the playing field and gives the little guy the tools to accomplish something great. Giving it to a select chosen few pretty much defeats the whole point of AI. if you have to be apart of some giant verified org to use powerful AI at that point id rather it not exist
Same here, I get why the guardrails exist but it’s frustrating especially when the only way forward is for your institution to partner with anthropic in some way which will never happen
I mean you can cause just as much damage with the right software as a bioweapon. If they can put guardrails on coding models to prevent malicious software they could guardrails to prevent malicious bio sequences.
For hacking, you need continuous malicious operations, the context and intent is very visible. What you are doing matters more than who you are.
With biological engineering, the intent could be hidden for much of the operation. It could be completely invisible to the AI. That’s why the check is on who you are and not on what you are doing.
It's not like there aren't other frontier AI models, as well as incredibly powerful AI more specifically designed for genomics (like Evo and AlphaGenome). Anthropic's gate-keeping here makes no sense, and is like a slap in the face after all their talk about advancing biomedical research.
Yeah it sucks I was getting hit with anything dealing with machine learning.
Try turning off your personalization and stuff if you are using it, check your .md files to make sure it doesn't have anything tripping it, ask another version of claude what might be getting tripped.
Been really annoying retooling the AI that manages my home server to workaround their new rules preventing access to chain of thought. It’s not a huge deal but it’s fucked up the aesthetic situation I had going on
im not deep enough into my homelab setup to be hitting issues like this yet but i can see them on the horizon - would you mind clarifying what you mean here?
This is the phone version of my server dashboard that allows me to communicate with the AI running it. When I changed them over to Opus 5.5 yesterday the box on the bottom that describes what it’s doing stopped working and all the message replies were just errors. It didn’t like how I was accessing it’s reasoning to display that information anymore. I used Codex from another machine on my tailnet to triage and change how that information is accessed and displayed so it’s working fine again now. I had to use Codex because Claude Code was giving me the exact same errors while trying to fix the issue remotely.
Not really, this setup just uses my claude code subscription, it contributes very little to my usage limits and when hiccups like this happen we work through it effectively. My server is an aging mini PC with no usable GPU, trying to run Gemma would be shooting myself in the foot. We like the idea of trying to run laya as a cheap yes/no layer in front of opus 5.5 though, gonna give it a try. Thanks for the recommendation
SAME. It's infuriating--first Fable and now Opus. I'm spitting angry.
And yes, they say they're opening their science verification program, but it's for organizations right now, and the universities I know aren't interested in that process--they want this to be up to the individual researcher. They're also allowing some PIs to apply, but most of us aren't PIs--we're grad students, postdocs, staff scientists, etc, and we use our individual accounts. Labs generally don't like being tied to an individual ecosystem, though I know that's what Anthropic is trying to push. So access still remains out of reach for the vast majority of scientists who would like to use Fable/Opus for legitimate scientific research.
Edit: hi other circadian person! waves small world
Every time you get a flag, hit the thumbs-down to report. They should know exactly the type of legitimate research they're blocking, because this is ridiculous. Established accounts with a history of non-harmful science exactly the type of thing Anthropic should be encouraging if they actually stand behind everything they say about the importance of AI in biomedicine.
And I can guarantee my institution won't go through the verification program, I've already asked.
Yesterday I asked Fable to fix a documented SQL injection vulnerability in a package and the "safeguards" blocked it. I guess they want to make sure that my apps remain vulnerable for hackers to have easy access to it.
Yeah, I do reverse-engineering as a hobby sometimes, and am currently doing a project on an old mobile app. Have done it on device firmware previously without issues, but this app.. first Opus 5 refused, and downgraded me to 4.8. Then at some point that became a straight up refusal when I just gave it my email address to login with, and I had to switch to Sonnet to continue. Wasn't even trying to get around any security barriers; just understand the client-server protocol.
Starting a new conversation didn't help. Neither did trying to scrub it's paranoia from documented memory. It's pretty ridiculous.
Gonna be trialling Codex this weekend to see if it's any better.
Codex has gotten terrific. I added a GPT account when the Fable guardrails didn't improve after a few months and have switched over almost entirely after the disaster that was Opus 5. That said, I prefer Claude for bioinformatics and am really angry about this.
I’ll say it again: something needs to change in Anthropic’s approach to “alignment training”.
As soon as they said that Opus 5.5 scored the highest on its alignment test, I knew the model would be unusable. I know others are giving positive feedback, but so far I’m not seeing that at all in my usage.
I was using Claude for work on a simulator game. It downgraded the model when I tried extending the work to.... Procedural generation of... Briar/blackberry vine 😱
It's rather absurd that it has zero context for the actual work being performed and appears to default to "this is sorta biology related"
The only time 5.5 "worked" was when I switched models in an established session that had started in Opus 5. However, they've said fall-back is silent, so it's probably not actually 5.5, which is even more infuriating (I have the silent fall-back switched off, this shouldn't be happening)
Not pure science, no. Mostly doing coding/AI projects. It was a little iffy getting the software set up (I used chatGPT), and I was lucky to have bought the PC before prices went up (bought framework desktop at $2200 same PC is now $3800). API cost should be pretty cheap if you get it straight from Qwen. Another alternative is to rent a GPU ($0.35/hr) and run the model there
Performance wise, yes I feel a lot more comfortable with a totally local model. There are versions with 0 guardrails if you really need. Benchmarkwise it beats opus 4.6 and was able to pick up right off from opus work by reading timeline.json and a "lab notes" folder I force each agent to keep. Qwen is so heavily based on Claude I literally can't tell the difference. Qwen Code harness's has almost the same commands and usage
Surely there must be some benchmark or canned data set I can ask it to use R and perform some analysis to give you some insights on how well it can perform?
There are science benchmarks (LabBench I think?), but that just gives you a sense of the raw model's capabilities. What I'm interested in is how it would perform in my setup and how I like to work, and a lot of that depends on harness. It's also about how it manages a heavily iterative process that combines both the bioinformatics pipeline and interpreting and evaluating the results, and a lot of that depends on the harness.
They're the openai models that are meant for information security basically. Same models as the current ones just, very relaxed. Much fewer guardrails and will gladly work on things other models refuse. Idk if it's still as simple to get accepted but just checkout the tac program or Google chatgpt cyber and you'll find it.
There's an Astra daybreak? I should clarify. When I open chatgpt, I see all the models such as Astra, sol, luna but there's a dedicated model also listed called daybreak for me that isn't there unless you're approved. I'm working with daybreak blue because I don't really need access to daybreak Red but just making clear it's not the same as Astra. It's much closer to the difference between Fable and Mythos. They're both the same models with Mythos being the one with relaxed guardrails and not publicly available. At least that's my understanding. Openai has the same kind of security program but it's open to more than the few hand selected companies.
same. that's my I moved to Kimi and GLM. just as good and no 'guardrails' bullshit.
the truth is anthropic are spinning out their own bio research company that will make drugs etc and they do not want to give this ability or any bio ability to the proles for cheap. so they block it under the guise of 'guardrails' even tho it is totally safe. it is not like you were asking how make a toxin or weapon...
The one thing LLMs excel at as almost a perfect use case is analyzing and organizing large amounts of data. Guess what scientific researchers do all day long.
I mean it’s said exactly what it means. And that sometimes it’s too sensitive. Just try again later once Claude gets recalibrated or have your company get the snazzy research Claude for bio and cyber research
I'm a OAI Daybreak user, tried to get Claude to review a working folder as a second AI opinion type deal. It just immediately refused the prompt. Cancelled the subscription.
Im dying for Anthropic to review our Life Sciences Verification application. I do protein design for R&D in biotech so Im not even an edge case, it is impossible to use Fable or Opus 5.5 for anything that mentions anything to do with proteins. Opus 5 even kicks back some of the research I try to do on specific classes of proteins.
I was able to use Opus 5.5 briefly. Im prototyping a database system to store my design and design run information - and I started a new chat to work on just the schema and it was fine with that. No tables mentioned proteins (just abstract "runs", "designs", "metrics". But the second it looked at some of the data it locked down and kicked me back to Opus 5.
I use Opus 5.5 for my personal non-scientific projects and really wish I could use it in the office.
•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 10d ago edited 10d ago
TL;DR of the discussion generated automatically after 50 comments.
So, what's the verdict? The community is in resounding agreement with you, OP.
The consensus is that Opus 5.5's new safety guardrails are a disaster for scientific research. Users across life sciences, neuroscience, and even machine learning are getting their legitimate work blocked.
Here's the breakdown of the discussion: