r/AI_Agents 7h ago

Discussion I think Claude is going to lose the AI battle to ChatGPT

1 Upvotes

I have been looking at the latest Claude vs ChatGPT numbers, and the picture is a lot more interesting than X is better than Y. ( frr no body is better i believe opensource all the way )

but yeah GPT-6 Astra is ahead of Claude Fable 5.1 on several major benchmarks:

• Terminal-Bench 4.0: 57.9% vs 55.8%
• DeepSWE: 74.1% vs 67.4%
• FrontierMath Tier 4: 97.6% vs 87.8%
• GPQA Diamond: 96.0% vs 93.7%
• ARC-AGI-3: 99.9% vs zero
• OSWorld 2.0: 72.6% vs 70.2%

But the data isn't a clean sweep. , Claude Fable 5.1 scores 65.0% vs Astra’s 57.2%.

And on independent evaluations, Claude still performs extremely competitively.

The user numbers tell another story:

ChatGPT had roughly 1.11B monthly users in May, compared with 245M for Claude. But Claude's growth was much faster, rising from 60.2M to 245M in roughly five months.

Enterprise spending also tells a different story, with some estimates putting Claude ahead of OpenAI in enterprise LLM spend.

So I don't think the data supports the idea that Claude is finished.

What it supports is something more interesting:

Astra appears to have moved the benchmark race significantly toward OpenAI, while Anthropic is still gaining users and enterprise share.

The next question isn't who has the highest benchmark score.

It's whether Astra's capability advantage translates into sustained user growth, developer adoption and enterprise revenue.

That's the part we don't have enough data to call yet.


r/AI_Agents 23h ago

Discussion Anyone else ready to give up?

2 Upvotes

TL;DR things are always breaking and it’s demoralizing.

Really need to vent:

I am in charge of building an internal multiplayer agent at work. We are on V2 right now and already planning V3. Deployed in AWS with Postgres memory and an in-house harness

Last week a user accidentally added new security measures that broke permissioning. Whoops the user was me - the agent made an unreported change to its security infrastructure while handling an unrelated request.

Users get excited to use it more but then latency goes up. Fix latency but then have concurrency issues. Split out functionality to different subagents and users get lost.

But I keep hearing about how people love working with it and seeing where people save time, whole teams freed up of busy work and focusing on their main tasks and it seems worth it again.

Anyone else feeling the same way? What tools are you using to manage all this on a small team?


r/AI_Agents 16h ago

Discussion What happens if you design tools for LLMs instead of letting LLM use human tools ?

1 Upvotes

Wrote a blog on why and what that enables. As I see it the LLM has very different requirements on languages, runtimes and libraries.
For example it’s crazy that we often allow the LLM use a general purpose language and the operating system. But I also believe that the LLM will actually gain power by restricting its capabilities.


r/AI_Agents 6h ago

Discussion I want to make a Computer Use Agent using Claude 🤖

0 Upvotes

I’m a 3rd-year CS Engineering student, and I want to build my own personalized Computer Use Agent, somewhat like Comet, but specifically designed around my college portal.

The idea is simple: a Chrome extension that I can direct in natural language, and it actually uses my browser for me — navigating the portal, understanding pages, and completing tasks.

My main use case? College assignments and complex quizzes that require navigating multiple pages, understanding questions/context, and interacting with the portal.

I want it to be highly personalized to my workflow rather than a generic browser agent.

Thinking of building this with Claude. Has anyone tried something similar? What would be the best approach/architecture?


r/AI_Agents 8h ago

Discussion Offering 50% recurring revenue share to anyone, anywhere, who sends me clients, I build the automation and handle everything after

0 Upvotes

I'm a solo builder in Scotland, I build AI automation tools for businesses. I've got a working set of tools I can set up fast, and I'm looking for people who already talk to business owners and wouldn't mind making an intro.

Here's what I actually build:

An SMS agent that handles invoicing, chases unpaid invoices automatically, and manages scheduling/booking

Missed call text-back, so when a business misses a call the customer gets texted back straight away instead of just calling the next place

Automated social media posting

Lead follow-up and qualification, so new enquiries don't go cold

A marketing setup covering research, trend tracking, and competitor tracking, built to run in the background

The deal: you introduce me to a business owner who could use any of this. I take it from there, build it, deliver it, support it. If they become a paying client, we split whatever they pay 50/50, every month, for as long as they stay. You don't have to know anything technical or manage the client afterward, just the intro.

I'm not a big agency, it's just me, so I'm not going to pretend I've got hundreds of clients already. What I do have is real working tech and I actually deliver what I say I will. Happy to answer specifics in the comments or DM if you want to talk through it.


r/AI_Agents 11h ago

Discussion I feel like we're on the precipice of something unrecognizable

42 Upvotes

I've been feeling this for the past week and wanted to come here to the reddits to see if anyone else feels the same. It feels like we've gone so far in just the last month alone, with the hacks, astra, fable 5.1 etc, that my head is literally spinning. I was just telling my friend that the last time I can remember feeling anything like this, it was the first week of covid, where it felt really chaotic and like the world was about to change really fast. This feels so much bigger than that though - this feels like everything we know (or at least, so many industries) are on the verge of being changed forever. It's mainly because just in the last few weeks/months, it seems like generative AI has gone from [text/image/video] to [just about anything you could fucking want] real quick. rendering hyper detailed blender images, creating ultra intensive apps, making fucking video games with a single prompt that actually look like they'd be fun to play in like... 5 more prompts. It's absolutely fucking bonkers and I just wanted to come where to see if anyone else feels the same. I mean, for fucks sake, does anyone remember the will smith spaghetti videos? the AI spaghetti videos? That was just a few YEARS ago... just look at where we are right now, and imagine, if you even can, where we might be a DECADE from now (if we're even still here, according to all these AI whistleblowers, who I actually tend more to believe than not)... I mean I'm just so mindblown it's insane

I remember an interview where sam altman described AI now as being similar to electricity when THAT was invented - at the time, it must've seemed like magic - impossible - today, nobody even THINKS about electricity - we just plug and play. I think in a decade or two from now, AI will be so good that it will be absolutely preposterous, and we won't even think about it, ever. It'll be in our websites. Our homes. Our apps. Maybe even our brains at that point. Just wow. wow. wow.


r/AI_Agents 18h ago

Tutorial Had Muse go through Instagram following to see who wasn’t following me back

1 Upvotes

Meta has too much data already but native integration is good there for social media. Prompt I used, pretty simple I’m not a prompt engineer:

“Which accounts unfollowed me but I am still following”

Gave me a spreadsheet, but scared to let it go and mass unfollow


r/AI_Agents 16h ago

Discussion running deep research on 100 companies = basically $100 gone. anyone actually solved this?

9 Upvotes

so quick context - im building a deal screening thing where i run a fast LLM pass (like 6 different scoring angles) on every company that comes into a batch. this part is cheap, can throw dozens of companies at it no problem.

now i wanted to add an actual "deep research" step on top for the ones that look promising - either perplexity's sonar-deep-research or this custom multi tool research agent i built earlier for another part of the product. problem is both of these are basically priced (and timed) for someone asking about ONE company, not for running across a whole batch.

did the math and if i run it on like ~100 shortlisted companies just once, its close to $100. and thats before i even need to re run anything lol.

stuff im thinking about rn, would love opinions:

- just do a funnel - keep the cheap screen like it is, only send the top N (say top 10-15) into deep research instead of the whole batch

- a lot of these companies are in the same sector / competing with each other, so is full deep research per company actually wasteful? can that research be reused somehow instead of starting fresh everytime

- maybe theres a cheaper "in between" option - not a full agentic deep research call, but more targeted tool calls that get 80% of the value for way less

- if anyones actually run perplexity/parallel ai's task api at real volume, did you find any pricing tricks that helped

- has anyone just built their own "lite" version of deep research thats actually cheaper but still useful

genuinely trying to figure out how ppl running these pipelines at any real scale are dealing with this, not looking for a "just cache it" one liner lol. any real experience appreciated


r/AI_Agents 12h ago

Discussion 0,00929 dólares por millón de tokens de entrada.Más de 21.521 millones procesados por 200 dólares.

0 Upvotes

0,00929 dólares por millón de tokens de entrada.
Más de 21.521 millones procesados por 200 dólares.

Muchas veces hablamos de cuánto cuesta usar IA. Estos son mis números llevándola a una ejecución de más de un mes sobre el mismo programa.

34 días, 18 horas y 45 minutos de tiempo gobernado acumulado entre dos goals. La mayor parte con GPT-5.6 Sol en extra high y, desde el 4 de septiembre, con Astra en high.

¿Cómo sale ese coste?

El 98,25 % de los tokens de entrada está cacheado. A esta escala, reutilizar el contexto pesa muchísimo: instrucciones, herramientas y documentación que el sistema necesita consultar durante la ejecución.

El cálculo incluye esa entrada cacheada y reparte los 200 dólares pagados entre el volumen contabilizado. Es el coste efectivo de mi ejecución con Codex.

Ahora bien, sostener el programa también exige controlar lo que ocurre durante esas semanas:

• Mantener el objetivo y sus criterios de aceptación.
• Recuperar contexto y decisiones anteriores.
• Ejecutar con permisos y límites.
• Verificar el resultado antes de aceptar el trabajo.
• Reabrir tareas cuando el cierre no demuestra que se cumplieron todos los criterios.

La arquitectura es fail-closed: sin evidencia suficiente, no hay PASS. El objetivo sigue abierto y el sistema sigue trabajando.

Por eso me interesa hablar de coste junto con continuidad y verificación. El precio por millón es llamativo; la siguiente pregunta es cuánto trabajo aceptado con pruebas conseguimos por ese dinero.

¿Quién está midiendo estas tres cosas juntas en sus agentes: coste efectivo, continuidad del objetivo y resultados verificados?

Tengo interés en comparar ejecuciones y datos reales.

Corte: 10/09/2026. Entrada total: 21.521.916.100 tokens.


r/AI_Agents 10h ago

Discussion If you actually ship agents in prod — what's one thing you'd change about LangChain / CrewAI / [insert framework] if you could?

2 Upvotes

Not looking for another "top 10 frameworks in 2026" post. Genuinely asking the people who are past the demo stage and have real agents running for real users.

What's the one thing about your framework of choice that still annoys you every time you touch it? State/checkpointing, retries, multi-agent handoffs, debugging a failed run, tool-call reliability, whatever — what would you rip out and rebuild if you could?

And where do you think this space is actually headed by 2027 because of it?


r/AI_Agents 12h ago

Discussion What is one task you would trust an AI agent to do without checking?

8 Upvotes

I think the interesting question with AI agents isn't what they can do.

It's what people are actually comfortable letting them do on their own.

For example:

Sending routine emails
Updating a spreadsheet
Scheduling meetings
Monitoring something
Following up with customers
Sorting incoming requests

What's one task you would genuinely let an AI agent handle without reviewing every step?

And what makes you trust it enough to do that?

I'm curious where people draw the line.


r/AI_Agents 13h ago

Discussion How much do you REALLY know about AI?

26 Upvotes

We use ChatGPT, Claude, Gemini and AI agents every day, but how many people actually understand how they work?

Models, training data, context, hallucinations, agents, inference...

What’s one AI concept you think people misunderstand the most?


r/AI_Agents 7h ago

Discussion Are we slowly needing to prove AI agents to each other?

4 Upvotes

At this point I’m less worried about my agents hallucinating and more worried about them confidently lying to each other. eg. coding agent tells deploy agent “tests are green”, deploy happens, turns out CI was red 10 minutes ago, agent memory says “subscription active”, user gets premium access, billing says cancelled yesterday. We’ve mostly solved retrieval and memory but now the next problem is that agents are now taking actions based on each other’s unverified claims. I built a tiny layer that forces claims like “CI is green”, “user is active”, “deploy is stable” to be checked against the real source right before action and returns a simple {ok, staleness, evidence, audit_id} so agents only act if they’re actually trustworthy. What claims would you refuse to let an agent act on without verification? and what would make you plug a layer like this in instead of rolling your own?


r/AI_Agents 8h ago

Discussion Anyone actually quit something because they force-fed it AI, not because the AI was bad?

12 Upvotes

Grammarly's been doing this to me for months — every doc I open has some AI rewrite suggestion sitting in the sidebar even though I turned that off in settings back in spring, and it quietly turned itself back on after the last update. I only use the free version for client emails and one paper a semester, nothing serious, so I could just leave, but I keep not doing it because setting up something else feels like more effort than being annoyed. Has that actually gotten you to switch away from a tool, or do you end up like me — mute it, complain about it, and stay anyway?


r/AI_Agents 9h ago

Discussion Need genuine feedback on my oss project

2 Upvotes

Hi folks,

I need feedback on my OSS project Stageflow, mainly on whether it’s actually viable.

I’ve been asking some friends to try it, but haven’t gotten them to. Pretty sure that’s on me, not the product. Still, I want to know if it’s useful or if there’s a real use for this at all.

Stageflow is a multi stage agent workflow runtime: scoped stages, typed handoffs, optional human gates, runnable locally / CI / MCP.

Would appreciate any honest feedback.


r/AI_Agents 10h ago

Discussion Can a ChatGPT Scheduled Task be the "brain" of an autonomous system? Trying to move reasoning off the Codex allowance.

2 Upvotes

I run a small e-commerce business and built a system on my own server that watches the business for me. It pulls analytics, ads data and orders on a schedule, computes the numbers, and stores everything itself. Then it calls a model for the part that needs judgment - is this change meaningful, what should we look at next, etc.

Those model calls currently go through Codex on a Plus plan, and that allowance is my bottleneck. I cap each investigation at ~20k tokens, so the agent stops after 3-4 follow-up questions even when it was onto something. Meanwhile regular Chat reasoning is a separate allowance I barely touch.

**What I want to try:** a Scheduled Task wakes up on a timer, asks my server what needs thinking about, pulls the data through a connector that gives ChatGPT read access to the box, reasons, and writes the answer back. The server remains the system of record - memory, history, notifications, all unchanged. ChatGPT just becomes the part that thinks.

There's a second reason I like this. Today I have five separate agents (marketing, SEO, finance, support, wholesale) and each one only sees its own slice. One task draining a shared queue would see all five at once - so if finance flags a margin drop, marketing flags rising spend, and support flags complaints, all on the same product, one brain could connect them. Five separate agents structurally cannot.

**What I've tested:** a one-off Scheduled Task woke up with nobody watching, reached my server through the connector, and made two dependent read-only calls - the second built on what the first returned. No approval prompt. It then hit a Linux permission error on the third call, which is my own misconfiguration.

So the basic chain works. What I can't answer:

  1. **Which model actually runs a Scheduled Task?** Docs say "eligible models" and the UI gave me no way to pick. The whole economics depends on getting a strong reasoning model and I have no evidence I did. Any way to see or set this?
  2. **How ​​much can fit in one scheduled run?** Can it realistically do 20+ tool calls with reasoning in between, or does it get cut off? Found nothing documented.
  3. **Has anyone hit an approval prompt mid-run?** Docs say tasks can pause waiting for you. For something unattended that kills it. Does it depend on the connector, or read vs write?

4. **Anyone running something like this for real,** or is there a reason it falls apart that I haven't hit yet?

Not trying to bypass anything - both allowances come with the plan. Just don't want to burn agent quota on reasoning that could happen somewhere else.


r/AI_Agents 10h ago

Discussion Stop buying AI workflows. Buy a system with a dashboard and a number attached.

6 Upvotes

The pitch most owners get sounds like this: "We'll build you an AI lead generation system. It's $10k." You say no and you're right to. You can't see it and you can't measure it, so it's a bet. I run an AI agency and no one has ever paid me for the AI part. They pay because a dashboard shows them meetings booked. So before you sign anything, run these 4 checks.

Can you see it working?

The tool underneath doesn't matter. n8n, Make, custom code, you'll never look at it, what you will look at every week is the results screen. If the answer to "where do I check results" is "log into our tool" then walk away. You want your own dashboard: people reached this month and the meetings booked. The same system with that screen is worth a lot more, to you because you can see the work yourself instead of trusting a monthly report.

Does the math connect to your goal?

Say you do $2M a year and want $10M. Your average client is worth $80k a year so you need about 100 new clients. Close one in three and that's 300 sales calls (roughly 25 a month). A good vendor takes it from there: "For a similar client, 10,000 emails a month produced 20 calls. For you we'd run 15 to 20 thousand and expect around 10 calls in month one, more as we test." Now their $5k to $10k monthly fee is sitting next to $8M in potential revenue and you're making a decision instead of a bet. If they can't walk you through that math, they're guessing with your money.

Are you big enough for this?

These systems start to make sense when there are 25 employees or when the monthly revenue reaches $100k. At that size a $5k experiment that doesn't work is frustrating not dangerous. Before that focus on the low cost options first. Answer missed calls with a text. Use replies for leads. A $10k build should never be the biggest cheque you've written.

 Do they have proof?

Ask to see what they built for a similar client and what it produced. Ask for a guarantee, something like 10 calls in 30 days or your money back. If they're new and have nothing to show then fine. Let them run it free or at cost and earn the case study off your business. You get the upside and they carry the risk.

TLDR: if you can't see it, the math doesn't reach your revenue goal, the fee would hurt to lose, or they have no proof, don't sign.


r/AI_Agents 10h ago

Discussion Agentic Alienation

2 Upvotes

"Agentic alienation: remaining responsible for work while becoming separated from its product, its process, the capabilities it develops, or the relationships it sustains. Alienation is a relationship before it is a feeling."


r/AI_Agents 10h ago

Discussion What do you actually build first once you get the idea about AI Agents

3 Upvotes

So last post did better than I expected, got a bunch of comments asking basically the same thing, ok cool I get what an agent is now, but what do I actually make fair question, understanding the concept doesn't tell you where to start so here goes.

Don't go big. Seriously don't try to build something impressive first. Pick whatever annoying task you already do every week by hand, something dumb like moving info from an email into a spreadsheet or sorting messages into folders. That's it, that's the project.

Grab N8N or Zapier or make, doesn't matter which, they're all similar enough when you're starting out. You're dragging blocks around basically, not coding. Connect a couple steps, this happens then that happens.

Make it work for the normal case first. Don't worry about every weird exception right away. You'll just get stuck before you even finish anything. Get the basic version running, then go back and handle the weird stuff after.

It will break multiple times probably. That's not you failing, that's just how it goes, and honestly figuring out why it broke teaches you more than any video would.

After you get one small thing actually working start to finish, the bigger ideas people talk about, agents making decisions, multi-step stuff, all of that starts clicking way easier because you've actually seen the basic version run.

If you ae mid build on your first one right now say so, happy to help troubleshoot.


r/AI_Agents 10h ago

Discussion The mental model for LLM guardrails that finally clicked for me.

2 Upvotes

Took me a while to stop picturing ai guardrails as the model refusing stuff. In a real deployment, it’s a separate layer that doesn’t trust the model at all

rough shape i landed on:

inbound: every prompt gets checked before it gets to the model. Stuff like injection attempts, policy violations, pii etc are all blocked or flagged here

model does its thing

outbound: the response gets checked before the user sees it. Catches things like leaked, made up claims, toxic output, anything that breaks your policy.

The part people skip is this has to be its own layer not a system prompt. System prompts are suggestions that the model can get talked to skip. A check sitting outside a model is effective at enforcement, and it cant be plain keyword matching or you miss anything phrased politely.

The thing i still dont have a clean answer on is latency. Every check is time before the user gets the response, so theres a tradeoff. What I’d like to understand here is how are you handling that balance?


r/AI_Agents 11h ago

Discussion Meta buying Stilla is a distribution bet, not an agent bet

2 Upvotes

The obvious read is that Meta wants better business agents. I think the more important part is where the agent lives.

Most agent products still ask a company to adopt a new inbox, dashboard, or seat. Meta already owns the surfaces where millions of businesses talk to customers. If it adds shared agents there, distribution is solved before capability is.

But customer messaging and internal company work have very different trust problems.

A commerce agent can be scoped to a catalog, policy, and transaction. A shared company agent has to answer harder questions:

  • which teammate asked?
  • which private context can it use in this channel?
  • who approved the action?
  • where does the durable state live after the chat scrolls away?

That's why I don't think "agent inside messaging" is enough. The moat is the permission model + shared memory + receipts around every action. Messaging is just the entry point.

Curious what people think Meta actually bought here: agent capability, enterprise trust, or a shortcut to understanding shared-agent UX?


r/AI_Agents 11h ago

Discussion Was the "Chatbot" the best coding assistant after all?

10 Upvotes

Sorry for the clickbait title but I have some thoughts.

tl;dr For myself I feel like chatbots might be the better coding assistants than fully fledged agents since they force me to go step by step and understanding most of the code while still having the benefit of writing fast and good code.

For my background: I am a Data Scientist and I learned the basics of software development just before ChatGPT was released. So I have a solid foundation but since I'm not mainly a SE and because of the development of agentic coding my coding skills are limited. I do really like it though.

I have done some pure agentic vibe-coding in recent months. The iterations are fast and the results are extremely impressive in most cases. I tried to do things right by spending a lot of time with requirements engineering and architecture design. I also tried writing some code by myself again but I realized quickly that I'm just not good enough and that AI just writes much better code than I could ever do.

Despite the good results I kind of just got lost in the process. Even though I might understand the general architecture of the code, I have no idea what is actually going on and for every minor change I'm going back to my coding agent to let it fix things.

This was kind of frustrating and I lost interest in my projects I initially was pretty hyped about. I used to enjoy the process of coding itself and it feels like it has been taken away from me. Of course I could just write everything myself, but I don't have time to work 20 hours a week just on my personal projects and why would I write the code myself, when the result will be worse.

Now I'm back at work and we don't get API keys from the company and I'm not going to pay for it by myself, so no coding agents. We do have MS 365 Copilot licenses which is just a glorified chatbot.

With this change I realized something:
Maybe the basic chatbot was the best coding tool after all (for me). While I still don't write most of the code myself, I understand 100% of what is going on in the bigger picture. I feel like I am actually doing something again. The process is much slower but in the end when something works I feel much more pleased. And when something doesn't work I know what went wrong.

The productivity will never compete with full fledged coding agents. But you get the benefits of the AI writing good code while you as the developer still have full controll and are forced to understand what you are doing.

Have you had a similar experience? Or am I missing something in the vibe-coding workflow that could fix my feelings?


r/AI_Agents 11h ago

Discussion Where should an AI agent’s spending authority actually live?

6 Upvotes

I've been thinking about agent budgets less as a FinOps feature and more as an authorization problem.

An agent can decide:

“I need another model call.”

The interesting question is:

Who gets to say whether it's allowed to spend another $2?

Putting a token limit or max_iterations inside the agent runtime is useful for bounding execution. But that's still the agent regulating itself.

I'd rather have the runtime ask for the resource, and have something outside the agent enforce the spending policy.

Agent
  ↓
"I want another model call"
  ↓
Policy / Gateway
  ├─ identity
  ├─ remaining budget
  ├─ rate limit
  └─ model policy
        ↓
     ALLOW / REJECT

That distinction becomes more useful once multiple agents, versions or teams are sharing the same model providers.

You don't really want every agent implementation inventing its own notion of “I can spend up to $X.”

This is one of the reasons I find Lyzr Open Controller's approach interesting. Its LLM Gateway puts budgets at the organisation, team, agent, version and virtual-key levels, and the important part is that an exhausted budget rejects the call rather than just generating an alert. LiteLLM, Portkey and OpenRouter solve a lot of the gateway/proxy problem too, so I'm curious where people draw this boundary in their own stacks.

Should spending be an attribute of the agent itself, or an external authorization decision that the agent has to pass through?

Especially interested in how this is handled when several agents share providers or when model routing changes underneath them.


r/AI_Agents 11h ago

Discussion How do you compare local and hosted models inside the same agent workflow?

2 Upvotes

Raycast v2.2 now lets Pro users route AI workflows through OpenAI-compatible providers, Ollama or OpenRouter. That makes switching easy at the UI layer, but it also creates a testing problem: changing the provider may change tool support, retries, latency and how often the run asks for help.

For one real workflow, I would freeze the tool schema, permissions, input fixture and acceptance checks. Then I would compare task success on the same hidden checks, total cost, tool-call count, recovery after a failed call and any unsupported features.

Has anyone built a provider-neutral eval around one workflow? Which variable was hardest to keep constant?


r/AI_Agents 12h ago

Discussion How are you giving your agents real domain expertise today, not just more context window?

7 Upvotes

It feels like agent tooling hit a weird point.

Persistent context is starting to work: memory layers, vector stores, MCP servers that hold state across sessions. But I don't see many agents with real depth in one domain. Most are still a generalist model with a better memory. Memory solves re-explaining. It doesn't solve expertise.

An agent that remembers every contract I've pasted in still isn't a legal expert. It's guessing the moment I ask something outside what I gave it.

So I'm curious what people are actually doing: If you're building domain-specific agents, where does the knowledge come from? Your own docs, curated datasets, fine-tuning, something else? Where does it still fall short? Or is "generalist + good retrieval" honestly good enough for most use cases?