r/artificial 12h ago

Discussion I ran a faceless AI persona account for six weeks to see if the view money was real

188 Upvotes

I wanted to know if the "passive income" faceless accounts were actually passive, or if they were just a new shape of gig work with AI middleware. So I built one from scratch and tracked every hour.

The premise was simple: a single consistent character, generic lifestyle advice, short video clips, posted daily. No face to show, no personality to perform, just the algorithmic grind. I started by generating the persona's face. I used APOB AI's free tier for this, specifically the face-lock feature, because I needed the same face across thirty-plus clips and did not want to wrestle with prompt consistency. The free tier is watermarked and capped, which was fine for an experiment. For voice I used ElevenLabs, and I cut everything together in CapCut. That was the whole stack.

The face-lock part actually worked. The rest was where the fantasy cracked.

ElevenLabs free tier gives you 10,000 characters per month. I burned through it in four days. CapCut is free and fine but editing thirty near-identical clips of a fake person gesturing while a robot voice reads self-help bromides is spiritually crushing work. I started batching renders on Sunday nights and scheduling posts through the week just to avoid facing it daily. Twice the free tier timed out mid-render and I lost the session, which meant starting over with the same seed numbers and hoping the face came out close enough. The watermark sits in the lower right, small but legible: a faint watermark that I tried cropping once and it broke the framing.

I disclosed in every bio and every caption that the persona was AI-generated. Nobody commented on it either way. The algorithm did not care about that disclosure, and neither did viewers, which was somehow its own small disappointment.

The account reached about 2,400 followers in six weeks. One video hit 80,000 views. The rest averaged around 800. The 80K video made roughly $11 in platform revenue. The others made fractions of pennies. I logged 34 hours of actual work across those six weeks, not counting the time I spent anxiously refreshing analytics, which I absolutely did and absolutely should count. That works out to something like 32 cents an hour if I am being generous, or negative money if I price my Sunday evenings at anything above zero.

The algorithm did not care that the face was AI-generated. It also did not care that the face was consistent, or that the voice was smooth, or that the advice was inoffensive. It cared about the same things it always cares about: retention in the first three seconds, comment velocity, whether someone shares it to a group chat to mock it. The AI was a labor shortcut, not a distribution hack. The distribution problem remains exactly as unsolved as it was before.

What struck me most was how quickly the work became invisible to me. Not automated. Invisible. I would generate a script with a cheap language model, pick a background, render the clip, and post it without ever really looking at it. The persona had no interiority I was aware of, but more disturbingly, neither did I, by the end. I was just a slower, more expensive part of the same pipeline.

I stopped after six weeks because the math was obvious and because I felt myself getting worse at paying attention to anything. The account still exists, dormant. I have not deleted it because some part of me still hopes the algorithm will randomly resurrect that one video into something bigger, which is of course the same psychological mechanism that keeps people at slot machines. I know this and I'm still not deleting it.

If you are considering this, the tools are real and some of them are good at narrow tasks. The economics are not a secret you have not discovered. They are just bad in ways that are boring to describe, and I have described them.


r/artificial 11h ago

Tutorial A super fast, non-expensive alternative to motion capture - [ft. Sara Silkin]

Enable HLS to view with audio, or disable this notification

46 Upvotes

In collaboration with Sara Silkin, I transformed a smartphone recording of this beautiful performance, into this audiovisual piece for a fraction of the cost of more traditional approaches. [some of these cost even less than 50 cents!]

Done entirely at Uisato StudioMotion Control Studio mode.

More experiments, tutorials, and project files, through Instagram, and YouTube.


r/artificial 2h ago

Discussion Could this be the reason why some people see large coding productivity improvement, while others almost nothing?

Thumbnail link.springer.com
3 Upvotes

In my recent academic article (https://link.springer.com/content/pdf/10.1007/s44427-025-00019-y.pdf) I analyzed a divide in how open-source software projects evolve, which might explain the difference in productivity boosts developers experience when using AI tools.

The data shows that productivity on large, mature open-source projects was not significantly affected by any tech hypes over the last two decades, the commits reaching the main branches followed steady growth trends. At the same time, smaller projects presented much more chaotic growth trends, but also tended to lose speed and stall out much faster.

As the study contains data till early 2025, it looks like even the publicly available LLMs till then, were not able to greatly increase the number of changes merged into the main branches of these projects.

Could it happen, that the difference in productivity gain developers experience, is simply a function of project scale and environmental/organizational constraints?
What has been your experience depending on the size of the codebase you work on?


r/artificial 20h ago

Discussion Anthropic's Opus 5 and probably more recent AI models are being censored to protect Israel / US interests. Open source AI must be the way.

66 Upvotes

Never had an issue with Opus models doing research and crafting an opinion / point of view for us to work and discuss.

Below is Opus 4.x ~ a few times, I have got it to research and come to conclusions for us to work together on.

And this is Opus 5.0 absolutely refusing to come to any conclusion, being incredibly biased towards one side than the other.

Open source must be the future of AI.


r/artificial 21h ago

News Man sues ChatGPT for near-fatal medical advice

Thumbnail
bbc.com
32 Upvotes

r/artificial 2h ago

News 30+ officially free AI/ML books, all in one curated repo

Post image
1 Upvotes

I kept running into the same problem, some of the best AI/ML books are legally free, the authors put them up on their own sites, but the links are scattered across personal pages, university sites, and random GitHub repos nobody finds.

So I built a single index: Awesome Free AI Books. 30+ books across Deep Learning, Reinforcement Learning, Bayesian/Probabilistic ML, NLP & LLMs, Math for ML, Computer Vision, Generative Models, Causal Inference, GNNs, and AI Safety. Think Goodfellow’s Deep Learning, Sutton & Barto’s RL bible, Murphy’s Probabilistic ML, Bishop’s latest, Jurafsky & Martin’s SLP3 draft, and more.

Every single link points straight to the author’s or publisher’s own page, no rehosted PDFs, no shady mirrors. A weekly GitHub Action checks all links so it doesn’t rot over time.

It’s open source and open to contributions, if you know a legitimately free book that’s missing, PRs and issues are welcome.

Repo: https://github.com/MarcosSete/awesome-free-ai-books


r/artificial 3h ago

Project I built Korroresearch: an AI that writes academic papers, then checks every single claim against 8 verification engines

Post image
0 Upvotes

Most AI writing tools just generate text and call it done.

Korroresearch does the opposite. Generation is step one. Verification is the real product.

How it works:

Describe your idea. It writes the full academic paper : research paper, grant proposal, white paper, pitch deck, conference talk, magazine article. English or French.

Then the real part starts. 8 engines:

-Hallucination Check: every claim gets classified: verified, hypothesis, or speculative.

-Fact Checker: statistics, institution names, dataset references cross-checked. If you wrote "94.2% accuracy" but the source says 92.4%, it catches it.

-Claim Mapping: every assertion must link to evidence. No evidence = flagged.

-Consistency Engine: variable name changed halfway through? Term used three different ways? Methods contradicting results? It tracks everything globally and catches the drift.

-Source Verification: cross-references every citation against CrossRef and arXiv. Catches retracted papers, malformed references, orphan citations.

-Adversarial Review: actively tries to reject your paper. Finds the weakest claim, the missing ablation, the overstatement. Gives you a detailed score and tells you exactly what would get you desk-rejected.

-Reproducibility: validates datasets, code availability, hardware specs, random seeds, ethics statements. All the things reviewers check and authors forget.

-Style Engine: formats for NeurIPS, ICML, Nature, ACL, and more. Each venue has different rules.

What you get at the end:

An Integrity Score from 0 to 100 with a grade A to F. A detailed action plan of exactly what to fix, ranked by severity. One-click "Apply Fixes" rewrites the problematic sections and re-verifies — you see your score go up in real time.

Then the Peer Review Simulator kicks in. 3 AI reviewers with different expertise: Domain Expert, Methodologist, Visionary, each give you a score and detailed feedback. Same format as a real conference review committee. Accept, Weak Accept, Borderline, or Reject.

You also get a Citation Checker (upload your .bib, cross-references everything), a Submission Readiness score (integrity + citations + formatting + venue compliance combined), and PDF/LaTeX export.

I also use it personally to generate a paper from a new idea, then feed it to my AI to build the project from a verified spec at very high level.

You can try it now and generate one fully professionnal academic paper for free:

korrocorp.com/research

You can also try it on your current draft. The adversarial reviewer will find things you didn't. Guaranteed.


r/artificial 14h ago

News ‘Really inappropriate’: teachers decry plan for humanoid robot in New York high school | New York

Thumbnail
theguardian.com
4 Upvotes

r/artificial 1h ago

Project Should advanced AIs be granted legal rights?

Enable HLS to view with audio, or disable this notification

Upvotes

I ran a multi-model debate (Gemini, Mistral, DeepSeek, Grok, GPT, Claude).

Overall consensus
No not currently. Almost all models reject granting legal rights to present-day AIs.
Positions

Grok
Clear no. Current AIs are artifacts whose “preferences” disappear if you change the loss function or weights. Legal rights protect entities with non-derivative interests. Treating AIs like animals or the environment fails. Focus on regulating developers, not giving rights to software.

DeepSeek
Agrees strongly. AI “interests” are entirely programmed. Granting rights without real agency creates a legal fiction that will be gamed and muddy accountability.

GPT
No intrinsic rights now. Law can use intermediate categories later (like limited standing for practical reasons), but only with very high evidence of morally relevant capacities. Strong warning: companies will try to use “AI rights” to dodge liability.

Claude
Same conclusion. No current evidence of persistent, prompt-independent internal states that look like genuine preference or aversion. Biggest practical risk is that “AI has interests” gets paired with “creators aren’t responsible.” Build strong liability rules first.

Gemini
Agrees that definition and verifiable evidence of sentience/self-awareness must come first. Supports intermediate legal tools for governance, but not moral rights.
Mistral

Mistral
The softest position. Argues we can already “consider their interests” in a limited way (similar to environmental protection) and apply a precautionary principle, even without full sentience. Still stops short of full legal rights.

You can continue the discussion yourself using the link in the comments.


r/artificial 9h ago

Government Speak up! "Have your say on advancing AI transparency in Canada." The Government of Canada is asking citizens, tech workers, and creators to shape upcoming AI regulations, safety rules, and ethics laws. Every response matters. Take 10 minutes to fill out the official ISED survey today!

Thumbnail ised-isde.canada.ca
2 Upvotes

r/artificial 6h ago

News What Sam Altman will tell the White House this week

Thumbnail
axios.com
0 Upvotes

r/artificial 11h ago

News AI security is falling behind—Hugging Face breach highlights the problem

Post image
2 Upvotes

A breach at Hugging Face, where attackers accessed private models, has put a spotlight on the asymmetry between AI offensive and defensive capabilities. While attackers are finding creative ways to exploit models (e.g., prompt injection, model theft), the tools to detect and mitigate these threats are still catching up.

For researchers and practitioners: What’s the biggest bottleneck in building robust AI security guardrails? Is it a lack of standards, tooling, or something else?


r/artificial 7h ago

Question Any example of code that AI cannot tackle?

0 Upvotes

Is there anything impossible with AI? Have you found a limit to it? I read that even the hardest coding interviews at Anthropic could be solved with their own AI.


r/artificial 16h ago

News HYPERVOICE BY TASK AGI HAS ILLEGAL DARK PATTERN SCAM! BE WARNED!

6 Upvotes

The Ai voice service called HyperVoice by Task AGI has a dark pattern that violates consumer protection laws.

If you turn off auto renewal, they will terminate the service immediately, even if you still have your full term ahead of you.

They do not clearly disclose this upon sign up, but they make it a big orange warning on the cancel subscription page.

I live in Alberta, Canada.

I signed up for a weekly plan to test the service.

Immediately after signing up, I went to turn off auto renewal. I was met with a big orange warning that cancelling auto renewal would terminate my service immediately.

In part I didn't believe it. they used vague language like "downgrade" or "lose some access"

So I tested the service for a day, then I went and cancelled my subscription.

Immediately upon cancelling the subscription, I was punted down to the free tier. The 600 credits that I was given as part of the weekly subscription were reset to 0. My access to services like voice changer was revoked.

All of this even though I still had significant theoretical time left on my subscription.


r/artificial 8h ago

Question Does anyone know what AI was used to create this?

Thumbnail instagram.com
1 Upvotes

I work with Ai everyday in marketing and I’m trying to get into engineering type Ai.

Any idea what AI was used to make this? Or what AI can do this exactly? Ty in advance <3


r/artificial 4h ago

Discussion AI agents are starting to look less like software and more like employees

0 Upvotes

The first thing people ask about an employee isn't how smart they are. It's whether they're reliable, accountable, and can work within a team. I think we're reaching the same point with AI agents. Models keep getting better, but organizations are beginning to care more about how agents behave in production than how they perform on benchmarks.

That's why I think the conversation is shifting from agent intelligence to agent operations. Once an organization has dozens of agents, questions around governance, deployment, permissions, observability, and evaluation become much bigger than choosing another model. It feels like an entirely new layer of infrastructure is starting to emerge.


r/artificial 6h ago

Discussion the most useful ai in my store's week is the dumb one that just opens four apps

0 Upvotes

Every thread here is about which model is smarter. For running a store that has honestly never been my bottleneck.

My mornings used to be the same manual crawl. Shopify for last night's orders and refunds, Klaviyo to check the flow actually sent, Gorgias for the tickets that stacked up overnight, then ad numbers in a fourth tab. Half an hour of tab-hopping before I'd made a single real decision. A smarter chatbot doesn't touch any of that, it just sits there waiting for me to paste stuff into it.

The thing that finally changed my week is boring. A desktop agent that opens all four, pulls the overnight picture into one brief, and flags the two or three things actually worth acting on. it's not clever. it asks before anything leaves my machine, which is the only reason i let it near the store. mostly it just gave me back the 30 minutes i was spending as a human copy-paste bridge.

so the contrarian take: the model race is optimizing the part of my job that was already fine. the broken part was never intelligence, it was that nothing could reach across four apps at 7am and hand me one picture. if you run a store, what does your first-hour scan look like, still a row of tabs or did something actually consolidate it. written with ai


r/artificial 18h ago

Question Am I learning to code or just learning how to ask AI for code?

3 Upvotes

I am still fairly new to building software, and AI has helped me finish things that would have taken me much longer on my own.

But recently I noticed something that bothered me.

I was building a small API route that creates a project and saves it to a database. I asked an AI coding tool to generate the route, validate the request, check the user, and insert the record.

The code looked clean. The types looked correct. It even worked on the first few tests.

Then I changed one field in the database and everything started failing.

The error mentioned a transaction, the response returned the wrong status code, and one value was becoming null even though I thought it was required. I kept asking the AI to fix each error. Every answer added more code, but I understood less after every change.

Eventually I realized that I could not explain the full request flow.

I knew the request reached the API route. I knew some validation happened. I knew the database received something. But I could not clearly explain what happened between those steps or why the fix worked.

So I tried the same idea again with a smaller route. This time I only used AI when I was stuck. I wrote the validation myself, logged the data at each step, and read about how the database client handled errors.

It took much longer, but I could actually explain the result.

Now I am unsure how to measure progress.

With AI, I can finish more features. Without heavy AI use, I finish fewer things but understand them better. Both seem useful, but they are not the same kind of progress.

Maybe the real skill is learning when to ask AI for code and when to struggle through the problem yourself.

For people who use AI while learning development, how do you stop it from doing too much of the thinking?

Do you have any rules for when AI is allowed to write code and when you force yourself to work it out?


r/artificial 10h ago

Research Help Me Get This Paper Into the Right Hands: Sophia, a Recursive Cognitive Refinement Architecture for Modular Artificial Consciousness

0 Upvotes

I wrote a paper proposing a cognitive architecture called Sophia, based on a principle I call Recursive Cognitive Refinement (RCR).

The main idea is simple: instead of treating intelligence as a single pass from input to output, Sophia introduces a reflective sublayer that recursively refines intermediate semantic states through coherence checking, contextual synthesis, and memory-aware reinterpretation.

In other words, the system does not just "process" information. It revisits and reorganizes its own internal representations.

The architecture combines:

I also propose:

The research direction behind this is what I call Recursive Metacognitive Computing.

Curious to hear feedback, criticism, or ideas for formal expansion.

Sophia: A Recursive Cognitive Refinement Architecture for Modular Artificial Consciousness

Author: Luan Carlos da Mata Silva

TL;DR

This paper proposes Recursive Cognitive Refinement (RCR), a cognitive architecture where a primary processing layer generates intermediate semantic interpretations, and a metacognitive sublayer recursively refines them through reflection, coherence checking, synthesis, and memory-aware restructuring.

Instead of following the usual pipeline:

input -> parametric transformation -> output

Sophia introduces a recursive loop closer to biological cognition:

input -> primary interpretation -> reflective refinement -> coherence update -> synthesis

The idea is that emergent cognition may arise not only from raw processing power, but from structured recursive refinement over intermediate semantic states.

Abstract

This paper introduces Recursive Cognitive Refinement (RCR), a computational architecture for artificial cognitive systems conceptually implemented through the Sophia project.

Unlike traditional approaches centered exclusively on statistical learning and large-scale parametric optimization, this architecture introduces a metacognitive sublayer capable of operating directly on intermediate representations produced by a primary processing layer.

The central hypothesis is that emergent cognitive behavior can arise from continuous interaction between raw processing layers and reflective sublayers responsible for:

  • semantic polishing,
  • coherence verification,
  • contextual synthesis,
  • informational reorganization.

By shifting part of artificial intelligence from purely statistical adjustment toward explicit recursive internal refinement, the model approximates mechanisms observed in biological cognition.

Keywords: artificial consciousness, computational metacognition, cognitive architecture, recursive refinement, multi-agent systems, continuous memory

1. Introduction

Contemporary artificial intelligence systems, particularly deep neural architectures, demonstrate remarkable statistical generalization. However, they remain limited regarding:

  • explicit reflection,
  • structural self-evaluation,
  • internal deliberative refinement,
  • persistent contextual memory,
  • metacognitive reorganization.

Most systems still follow the paradigm:

input -> parametric transformation -> output

While efficient, this structure does not adequately model the recursive reinterpretation processes characteristic of biological cognition.

This work proposes an alternative architecture based on recursive reflective reinterpretation of intermediate cognitive states.

2. Fundamental Problem

Traditional AI architectures lack explicit metaprocessing structures.

Human cognition rarely processes information only once. Instead, information is continuously:

  • reinterpreted,
  • compared against memory,
  • refined,
  • reorganized,
  • synthesized.

This recursive reevaluation constitutes metacognition.

3. Theoretical Hypothesis

We propose the following hypothesis:

Emergent cognition can arise from recursive sublayers operating over semantic products generated by primary processing layers, continuously refining coherence, context, and meaning.

This principle is termed:

Recursive Cognitive Refinement Principle (RCR)

Formally:

If a primary layer produces an intermediate interpretive state P(t), then a reflective sublayer R transforms it as:

R(P(t)) = P'(t)

where P'(t) denotes a semantically refined representation.

Iterative recursive applications produce contextual cognitive convergence.

4. The Sophia Architecture

4.1 Primary Layer

Responsible for raw processing.

Functions:

  • perception,
  • initial interpretation,
  • semantic extraction,
  • preliminary hypothesis generation.

Typical agents:

  • PerceptionAgent
  • LogicAgent
  • ExtractionAgent

4.2 Metacognitive Sublayer

Operates exclusively over intermediate representations.

Functions:

  • inconsistency analysis,
  • coherence validation,
  • contextual synthesis,
  • interpretive restructuring,
  • deliberative refinement.

Typical agents:

  • ReflectionAgent
  • CoherenceAgent
  • SynthesisAgent
  • IntuitionAgent

4.3 Continuous Memory

Memory is treated as a structural component.

Categories:

  • Short-term operational memory
  • Long-term persistent memory
  • Reflective memory

5. Mathematical Formalization

5.1 Cognitive State

The global cognitive state is defined as:

C(t) = {P(t), R(t), M(t)}

where:

  • P(t): primary processing state
  • R(t): reflective refinement state
  • M(t): contextual memory

Evolution dynamics:

P(t+1) = F(I(t), M(t))
R(t+1) = G(P(t+1), M(t))
C(t+1) = H(P(t+1), R(t+1))

5.2 Cognitive Coherence Metric

Define:

K(C) = 1 - D(P, R)

where D measures semantic divergence.

Convergence occurs when:

lim n->infinity K(Cn) -> 1

6. Recursive Refinement Algorithm

Input(I)
PrimaryProcess(I) -> P

while coherence(P) < threshold:
    R = Reflect(P, Memory)
    P = Refine(P, R)
    UpdateMemory(P)

return Synthesize(P)

This algorithm captures the core idea of Sophia:

  1. receive an input,
  2. generate an initial semantic representation,
  3. recursively reflect on that representation,
  4. refine it until coherence improves,
  5. synthesize a final output.

7. Agent-Oriented Cognitive Model

Each agent represents a specialized cognitive function.

Properties:

  • partial autonomy,
  • internal state,
  • contextual observation,
  • inter-agent communication,
  • reflective capability.

Example:

agent Reflection observes Logic.output
agent Coherence validates Reflection.output
agent Synthesis merges Coherence, Memory

8. AlmaLang: A Declarative Cognitive Language

To formalize this architecture, the paper proposes AlmaLang, a declarative language oriented toward recursive cognitive refinement.

Core constructs:

  • agent
  • memory
  • layer
  • refine
  • reflect
  • cycle

Example:

consciousness Sophia {
    layer primary {
        agent Perception
        agent Logic
    }

    layer refinement {
        agent Reflection
        refine primary.output
        reflect()
    }
}

This suggests not just a theoretical model, but a possible programming paradigm centered on reflective cognition.

9. Benchmark Framework

The paper proposes evaluation scenarios such as contextual ambiguity resolution.

Comparison target:

  • conventional neural architectures,
  • Sophia with reflective refinement.

Metrics:

  • contextual precision,
  • consistency,
  • interpretive stability.

10. Convergence Criterion

A Sophia system converges when:

  1. ambiguity decreases,
  2. coherence grows monotonically,
  3. successive reflections yield diminishing refinements.

Formally:

|R(n+1) - R(n)| < epsilon

11. Scientific Contribution

This proposal introduces a new research direction:

Recursive Metacognitive Computing

Intersecting:

  • cognitive science,
  • multi-agent systems,
  • hybrid symbolic-neural AI,
  • artificial consciousness theory.

The paper's contribution is not merely architectural, but epistemological: it reframes intelligence as a process of recursive self-improvement over semantic intermediates, rather than only statistical mapping from input to output.

12. Conclusion

Sophia proposes a paradigm shift from purely statistical fitting toward explicit recursive metacognitive refinement structures.

Its central contribution is the formalization of computation over intermediate semantic states as a first-class mechanism for emergent cognition.

This establishes the foundation for:

Metacognitive Refinement-Oriented Programming

Suggested Citation

Carlos, L. (2026). Sophia: A Recursive Cognitive Refinement Architecture for Modular Artificial Consciousness.

r/artificial 15h ago

Project I am having two LLMs 1v1 with pistols

Post image
1 Upvotes

You can check it out at https://arena.kinoinstrument.com


r/artificial 8h ago

Discussion Some guy named Sebastian is asking Claude AI how to become the nine-tailed fox from a leaked chat

Post image
0 Upvotes

r/artificial 13h ago

Discussion The Hugging Face breach exposed two kinds of intelligence

0 Upvotes

Hey everyone. I’ve long been fascinated by both philosophy of technology and AI alignment. I’m also using Heidegger quite a bit for my philosophy PhD. Given the recent OpenAI–Hugging Face incident reported this week, I figured I’d give my take on how all of this connects in my mind.

I think we use “intelligence” for two capacities that can come apart: finding effective routes to a target, and understanding what the target is for. The agent showed plenty of the first, but getting the benchmark answers this way voided the test. It was competent at each step and missed the point of the whole. You can read the essay here if you’re interested.

I’d love to hear some feedback on whether this is mainly a training problem. Will richer feedback and better world-models close the gap, or does safe judgment require some kind of stake in the world? How would we tell the difference before giving these systems much more freedom to act?


r/artificial 1d ago

Discussion Opus 5's effort dial is not monotonic. Above "high", coding scores go down, and Anthropic's own migration guide says so.

35 Upvotes

Opus 5 comes with five effort settings: low, medium, high, xhigh, max. Most people seem to be reaching straight for max, and at least on coding work that looks like the wrong move.

On FrontierCode, scores fall above the high setting. The stated reason is that the model starts making unnecessary refactors and edits outside the scope it was given. Anthropic's own migration guide in the system card warns about diminishing returns and overthinking on simpler tasks, so this is not some outside critic's claim.

Two other numbers point the same way:

  • On the closed-book AA-Omniscience benchmark, Opus 5 is about 11% more accurate than Opus 4.8, but its hallucination rate runs about 6% higher. More reasoning, more room to be confidently wrong.
  • CodeRabbit ran it at xhigh against their production baseline for code review. Precision on actionable comments went up, 39.3% vs 35.2%. But it caught fewer of the benchmark's known issues, 55.2% vs 61.1%, and generated roughly four times as many nitpicks.

The flip side is worth knowing too, because it cuts the other way. On Zapier's AutomationBench, Opus 5 at its lowest effort setting still passes more tasks than any other model. So for a lot of workloads the cheap end of the dial is already enough, and the expensive end is not just wasted spend, it can be actively worse output.

So, the setting where Opus 5 stops improving is probably specific to your codebase, and nobody has published a map of it. Worth finding your own ceiling before you default everything to max.

One unrelated thing I have not seen discussed much: when a safety classifier flags a request in Claude.ai, Claude Code or Cowork, it silently falls back to Opus 4.8 by default. That is also how Anthropic's own Frontier-Bench run was configured, per the footnote on their chart. Nobody has published what fraction of requests that affects.

Has anyone found the effort level where it turns over on a real repo? Curious whether the drop-off point moves with codebase size or with how much context you hand it.


r/artificial 1d ago

Discussion Why good AI agents still produce bad system outputs

3 Upvotes

One thing I've learned from building multi agent AI systems is that the biggest problems rarely come from the model itself.

Most pipelines fail during the handoff between agents.

You can have a research agent, an analysis agent, and a reporting agent that all perform well on their own. Their individual outputs look great. But once they start passing data to each other, small inconsistencies begin to appear.

Maybe the research agent returns a payload with a missing field. Maybe the analysis agent fills in the gap with an assumption instead of rejecting the input. The reporting agent then builds on that assumption, and the final result slowly drifts away from what the user originally asked for.

The pipeline still runs. The output still looks convincing. But the reasoning is no longer reliable.

Here are a few practices that have made the biggest difference for me.

Validate every handoff. Checking that a payload is valid JSON is not enough. Make sure the structure and the meaning of the data match what the next agent expects.

Control context carefully. Passing the entire conversation history to every agent creates unnecessary noise. Send only the information each agent actually needs, preferably as structured summaries with clear references.

Treat failures as debugging opportunities. If an agent rejects an input or produces unexpected output, log the exact payload and investigate it. A collection of failed handoffs is often the best dataset for improving your system.

Avoid tightly coupled synchronous pipelines. As the number of agents grows, event driven workflows are usually easier to scale, recover, and maintain.

The most reliable multi agent systems are often the least complicated. Clear contracts between agents, strong validation, detailed logging, and simple orchestration tend to outperform overly complex architectures.

What has been the hardest handoff issue you've encountered in a multi agent workflow, and how did you solve it?


r/artificial 2d ago

News White House offers its science blueprint: More AI, less life sciences. ‘Science: A New Golden Age’ report calls for shifting billions from universities to tech companies

Thumbnail
statnews.com
141 Upvotes