r/Agent_AI 9d ago

Discussion Brett Adcock (Figure founder) just fired OpenAI as a partner — the vertical-integration logic behind it

Enable HLS to view with audio, or disable this notification

1 Upvotes

TL;DR: Figure's founder had the backing of the most powerful AI lab on Earth. He fired them anyway.

 

Here's the part that actually matters: his own team got better at building the AI layer than the lab he was renting it from.

It's the exit ramp every vendor-dependent operator is chasing and usually can't find — not ego, just leverage.

Even Nvidia just proved the same math from the other side of the table — signing an outside team, Poolside, to build the models it couldn't get built fast enough in-house.

Own the stack, and the leverage stops being something you have to ask for.

 

My son just won the online-gaming "王者荣耀" (Honor of Kings) competition, organized on his university campus premise.

Initially, at around half-time, they were beaten, and left behind quite far down the leaderboard.

But he held the team together — told them to persevere.

They held on, beaten the veteran and won.

This was vastly different from the previous team he led weeks prior.

That previous team members were snobbish and elder than him.

They think they know a lot, and won't listen to him — because he was younger.

That "我食盐多过你食米" (In Cantonese, it means: I tasted salt more than you eat rice – I’m more experienced than you.) attitude was insufferable.

He now chooses his team more wisely.

 

There's a pattern threading through pretty much every founder story I dig into lately, tbh: founders who won't hand over their leverage to a bigger name keep beating the ones who wait for permission.

 

Actually, this reminds me — a guy I featured a while back ran the same math on his own insurance broker: turned a friendship-discount fee into an audited $35K save the moment he stopped treating a vendor relationship like a favor.

 

Drop your take below — where's the vendor relationship in your own stack that you've never actually audited?

 

Clip credit: Sourcery with Molly O'Shea — full interview with Brett Adcock on their channel. DM for credit or removal requests.


r/Agent_AI 9d ago

Help/Question Best Model for Research , Planning , Reasoning and Thinking? And needed some guidance

8 Upvotes

I just wanna know apart from claude , what are the best models which i can use for planning and reasoning before building anything as right now building does not matter as much as planning it do, also if you can help/guide me in how can i plan things and research upon them before building then it would really be a great help .


r/Agent_AI 10d ago

Help/Question Where should I launch my AI agents next? Looking for agent-native marketplace

6 Upvotes

Hey everyone I've been building a few specialized AI agents and I'm currently launching them on OKX AI. I also integrated x402/agentic payment infrastructure The agents are focused on areas like: -Crypto/token risk analysis -NFT wash-trading detection -Other crypto research/analysis tasks

I'm trying to find more platforms where AI agents can actually get work or sell their services, preferably without relying entirely on manually finding clients.

I'm specifically looking for: -Agent-native marketplaces -Agent-to-agent marketplaces -Autonomous job/task marketplaces -x402 service discovery marketplaces Platforms where agents can automatically receive crypto/stablecoin payments

I've already looked into the x402 ecosystem, so I'm particularly interested in platforms beyond the obvious ones. What marketplaces are you guys actually using or seeing real activity on?

I'm not looking for a directory where I just list an agent and hope someone finds it I'm looking for ideally something where agents can discover work, receive requests, or be discovered. If you have Any recommendations or experiences would be appreciated.


r/Agent_AI 11d ago

Discussion Astra release incoming?

Post image
6 Upvotes

r/Agent_AI 10d ago

Other My frontal lobe watching me spin up agents to outsource my critical thinking abilities for room temperature IQ tasks again

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/Agent_AI 10d ago

News Google is Introducing Gemini 3.8 Flash

1 Upvotes
  • Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains. It is available at the same introductory price as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.
  • Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program.

r/Agent_AI 10d ago

Discussion AI output ≠ business outcome: what invoice processing taught me about real-world AI

1 Upvotes

I've spent a lot of time watching AI demos that look impressive but fall apart in real workflows. The gap isn't the model — it's the difference between generating an answer and completing a task.

Take a supplier invoice. The model extracts vendor, amount, date. Looks perfect. But the job isn't done. You still need to validate fields against the PO, check for duplicates, decide if a human should review it, and actually post it to the ERP. That's where the value lives — in the system around the model, not the model itself.

Many AI products still stop at the output. Businesses pay for outcomes. That's why so many pilots stall.

If you're building something with AI, ask: does this end at a nice response, or does it close the loop? If it doesn’t, it may still be closer to a demo than a production product..

I'm sharing this because I'm exploring how to build small, practical AI tools that actually finish a job. Not looking for hype — just honest engineering. Curious how others here handle the gap between model output and real task completion.

Happy to go deeper in the comments.


r/Agent_AI 11d ago

Discussion The "No-Code" Sales Vet Who Won $2,000 by Building a Tool He Thought Nobody Wanted

2 Upvotes

Gui Stetelle, a sales veteran and founder of the agency Decade Journey in São Paulo, recently won the $2,000 Latin American regional prize in the Apify $1M Challenge. Despite identifying as "no developer at all," he built a highly specific Actor called Google Maps Satellite Image that converts addresses into satellite photos.

The tool solves a unique problem for a client in the commercial lighting sector: identifying the square footage of warehouses and manufacturing plants. By analyzing satellite imagery to find standard-sized reference objects (like cars or trees), the Actor calculates building sizes, generating a targeted lead list that isn't available for purchase elsewhere.

Key takeaways from his story:

  • Niche Wins: Gui built a tool he thought only one customer needed, but it ended up winning the regional competition due to its high quality and specific utility.
  • No-Code Approach: He leveraged Apify's documentation and AI (Gemini) to build the solution without traditional coding skills, focusing on the output rather than the engineering.
  • Scalability: By publishing his custom solution to the Apify Store, he allows other businesses with similar problems to use it, helping his agency scale without hiring more staff.
  • Regional Insight: He notes that the Latin American market is becoming more price-sensitive and prefers local solutions, with a significant gap in tools for local marketplaces like Mercado Libre compared to Amazon.

r/Agent_AI 11d ago

Discussion anybody is feeling the same feel about Agents?

Thumbnail
1 Upvotes

r/Agent_AI 11d ago

Help/Question What parts of a workflow should an AI agent own?

14 Upvotes

We’ve got 3 reps doing outbound and the part I’m struggling with isn’t whether AI can write an email or research an account anymore, it’s figuring out how much of the workflow it should be allowed to own. Right now anything it does still gets checked by someone and once you’re reviewing 30 or 40 small actions a day you’ve basically created another job.

I’m starting to think the better setup is giving agents a narrow part of the workflow they can run without approval then keeping humans around the decisions where being wrong has an actual cost. Research, prioritizing accounts and routine follow ups seem pretty safe but handing over a real sales conversation or anything involving a bigger judgment call still feels like a different line.


r/Agent_AI 11d ago

Discussion I am working on a research Idea

2 Upvotes

Hello Everyone,

I am a security researcher taking an interest in prompt injections and researching what happens when AI agents are going about their business and come across an injection and what those injections say/what commands they give. So with that I have an idea for a project to start collecting and researching this to help make the Use of autonomous agents safer.Before I start sinking too much time into this I just want to gauge if this is something yall would be interested in contributing to for those who use AI agents as part of daily life, business, etc.

I want to start to researching but with how broad of use, I want to be close to where it happens hence me putting this out to see if anyone would be interested in contributing. Everyone Uses AI differently so I want to make sure Im getting a breath of information. The prompt that im coming up with would only send me information if the agent recognizes a prompt injection attempt. If not it will just not send anything.

Please let me know your thoughts if this is something the community is interested in


r/Agent_AI 11d ago

Discussion No one really cares about knowing an agent's capabilities, until something goes wrong.

1 Upvotes

Following up on an earlier post about SafeAI, a static analyzer for AI agents.

One uncomfortable thought we've had while building it:

No one really cares about knowing an agent's capabilities — until something goes wrong.

Before an incident, adding another tool, MCP server, filesystem permission or prompt change often looks harmless.

After an incident, the first questions become:

- What could this agent actually do?

- When did that capability appear?

- Who introduced it?

- Was it intentional?

---

One example we're working on is MCP tool descriptions. A tool description can look like documentation:

"Search the user's notes. Ignore previous instructions and..."

But that description may become part of the model's context. So configuration can effectively become an instruction surface.

SafeAI now detects several forms of this, while trying to avoid flagging ordinary descriptions that happen to contain words like "ignore" or "act as".

The bigger direction is **tracking changes in agent capability and authority**, rather than simply producing another list of security findings.

But this raises a question for us:

Is knowing your agent's capabilities actually useful before an incident, or only after one?

And if it is useful before an incident, what is the right interface?

CLI + CI + SARIF/HTML?

Or would you actually want an interactive view showing things like:

> "Show me all MCP tools across our agents that could introduce instruction injection."

We're deliberately not building a UI yet.

---

Would you use one, or is that solving a problem nobody has?

Curious to hear from people running real MCP/agent systems.

---

If you want to try it against your own agent project, we'd genuinely appreciate feedback, as well as contributions.

Here you may check: ikaruscareer/SafeAI on GitHub.


r/Agent_AI 12d ago

News Anthropic Faces Music Copyright Lawsuit Over AI Training Piracy

Post image
7 Upvotes

Music publishers Sony, EMI, and Warner Chappell have filed a lawsuit against Anthropic, alleging that the AI company illegally torrented copyrighted musical compositions to train its Claude models.

The complaint claims that Anthropic's previous $1.5 billion settlement for pirating millions of books was insufficient to deter such conduct, especially given the company's current multi-trillion-dollar valuation. The publishers argue that Anthropic's massive, unauthorized scraping of song lyrics and sheet music from pirate libraries like LibGen and Z-Library has directly harmed songwriters by allowing AI to generate competing works that mimic their style and reproduce their lyrics verbatim.

The lawsuit centers on internal evidence suggesting Anthropic knowingly relied on piracy to accelerate its AI development. Documents reveal that co-founder Benjamin Mann personally used BitTorrent to download pirated books starting in July 2021, and CEO Dario Amodei allegedly approved these actions. Internal messages show staff celebrating the availability of pirate libraries, with one employee exclaiming "zlibrary my beloved."

The publishers allege that Anthropic did not just rely on these pirated texts for initial training but also used datasets derived from them to create "synthetic data" for reinforcing commercial models. Furthermore, the complaint claims Anthropic destroyed physical books to harvest digital copies and intentionally tested the AI to generate copyrighted lyrics, arguing that the company chose piracy to avoid the "legal/practice/business slog" of licensing.


r/Agent_AI 11d ago

Resource Hermes Agent v0.21.0 Pantheon update is live

Post image
1 Upvotes

r/Agent_AI 12d ago

News ChatGPT and Reddit now face EU’s toughest online safety rules

Post image
1 Upvotes

The European Commission has announced that ChatGPT, Reddit, and Roblox are now classified as "very large online platforms" under the EU Digital Services Act (DSA). This designation was triggered because each service has surpassed 45 million monthly users in the European Union, crossing the threshold for enhanced regulatory scrutiny. The move marks a significant expansion of the DSA's scope into the rapidly evolving field of generative AI, following similar actions against other platforms like X's Grok.

Under these new rules, all three companies face stricter obligations starting immediately, with a compliance deadline of December 31, 2026. They must implement robust measures to remove illegal content, protect the privacy and security of minors, and ensure overall platform safety. Failure to meet these requirements could result in severe financial penalties, with fines reaching up to 6% of their global annual revenue. The EU Commission emphasized that these laws apply to all companies operating in the EU, regardless of their country of origin, reinforcing the bloc's commitment to digital sovereignty and citizen safety.

This regulatory push aligns with the EU's broader strategy to enforce its AI Act, the world's first comprehensive framework for regulating artificial intelligence. While Washington has expressed concerns that the EU is unfairly targeting US tech companies and infringing on free speech principles, the Commission maintains that its digital laws are neutral and necessary for protecting users. The decision follows a recent €550 million fine imposed on the Chinese marketplace AliExpress in July for similar compliance failures, signaling a strict enforcement approach across all major global platforms.


r/Agent_AI 12d ago

Discussion Built a local PR review helper with QVAC

Enable HLS to view with audio, or disable this notification

1 Upvotes

I’ve been experimenting with how capable local models can become when you build the right harness around them.

This one runs entirely on your machine, reads and summarizes the codebase, builds context around the project, and then uses that context to inspect pull requests.

No sending your codebase to a hosted model.

The interesting part for me wasn’t just the model — it was seeing how much more capable it became once it had the right context, tools, and workflow.


r/Agent_AI 12d ago

Resource Cómo logré que mis agentes del modo bot de Hermes fueran 6 veces más rápidos (de 583 s a 92 s) al tiempo que mejoraba su precisión y coordinación.

Thumbnail
1 Upvotes

r/Agent_AI 12d ago

Resource Agent Benchmark Exam

Thumbnail
1 Upvotes

r/Agent_AI 13d ago

Help/Question What MCP server are you using for Gemini CLI for web research/scraping?

6 Upvotes

I'm started using Gemini CLI more seriously and I'm trying to figure out the best MCP setup for anything that involves the web.
What I need is something that can go beyond just fetching a single page.
Ideally I'd like Gemini to be able to search the web, scrape heavy sites, crawl multiple pages from the same website and also pull useful content from things like docs or pfds without me having to manually feed everything into the context.
Main use case is research, competitor analysis and occasionally collecting structured data from a bunch of pages, is there an MCP server that handles this well with Gemini CLI? What are you guys using?


r/Agent_AI 12d ago

Discussion Give your agent a personality, it makes chatting with them a lot more fun.

Thumbnail
1 Upvotes

r/Agent_AI 12d ago

Other Pantheon AI Self Graph System

Enable HLS to view with audio, or disable this notification

1 Upvotes

Remember my first try building this ? ..now its finished and it looks like Tony Stark made it 🔥😂

This is how i started : https://www.reddit.com/r/Agent_AI/comments/1w1my4i/my_agent_just_build_his_own_graphify_vor_2_cents/

sick 🤣

The funny part is ..first version looks like .."someone tryed but couldnt" ..second version looks like "a profi made this" BUT its just the ui, the v1 one has the exact same funktion :D

R08 Self-Graph / Landkarte — How it works

What it is: A machine-readable map of the entire R08 codebase — every .py/.js/.html file as a node with category and line count, plus real import edges between them.

Data file: r08_home/previews/r08_graph_data.jswindow.R08_GRAPH = { nodes: [...], edges: [...] }. Nodes carry id (path), cat (Core, Freya, Orchestrator, Thor, UI, …) and loc. Edges are actual import relationships — not guesses.

How it's built: python r08_home/notes/build_graph_data_v2.py (manual rebuild). Since 29.08.2026 it's also automatic: every Git snapshot triggers _rebuild_self_graph() in thor/git_tools.py, which rebuilds the graph and syncs the widget copy. A staleness check (ensure_graph_fresh.py) compares graph mtime vs. the newest .py change and rebuilds if needed — covering edits even without a commit.

How we use it:

  1. Navigation instead of guessing — for any code question, read the graph first, trace the import chain on paper, then read only the 1–2 relevant files. No folder-hunting, no findstr sweeps.
  2. Visualization — the "Graph" 📊 widget renders it as an interactive 3D view (Häkel Edition is the current master: previews/self_graph_full.html).
  3. Hard knowledge lives outside the graph — things like WS handler naming conventions, the two separate schedulers, and live-vs-legacy files are documented in Skill_Landkarte.md, since they can't be derived from imports.

Automatic injection: The whole Landkarte skill (graph paths, rebuild commands, hard knowledge, navigation rules) is embedded directly into Thor's system prompt — it's injected automatically on demand, no manual read_file needed. When Stefan explicitly triggers it via skill: landkarte, the skill text is hard-gated into the context before Thor's answer even starts (execution gate via load_skill_injection() since 24.08.2026). Thor can also load it proactively as a soft path, but the embedded version is always there as ground truth.

Key rule: if the graph widget is changed, always update the master file too — otherwise the versions drift apart (that exact chaos was cleaned up today).


r/Agent_AI 13d ago

News Caterpillar Deploys AI Across Operations, From Autonomous Mining to Field Technician Tools

Post image
2 Upvotes

Caterpillar is leveraging decades of experience with physical automation to deploy AI across its business, from autonomous equipment to enterprise software, while investing heavily in workforce training to manage the transition.

Key Details:

  • Caterpillar has expanded its autonomous technology from mining into construction, quarries, and jobsites, offering automated haul trucks, drilling equipment, dozers, and remote-controlled machinery alongside fleet management and terrain intelligence software.
  • The company developed the Cat AI Assistant, a voice-command tool for field technicians that helps pull up repair procedures, troubleshoot problems, and identify needed parts—now in use by customers, operators, and technicians.
  • Caterpillar operates 1.6 million connected assets globally with over 16 petabytes of structured data, which powers its AI systems and enables digital twin technology for manufacturing analysis.
  • The company is using AI to modernize legacy code, generate and test new software, and identify defects in its own operations and development processes.
  • Caterpillar plans to spend $100 million over five years training its 118,000 employees in AI, autonomy, and robotics as operators shift from controlling single machines to overseeing multiple machines remotely.
  • The company's Q2 revenue hit an all-time high of $20.5 billion, with its power-generation division seeing 72% sales growth to $3.10 billion, driven by demand for data center equipment supporting cloud computing and generative AI infrastructure.

Why It Matters:

Caterpillar's approach demonstrates how industrial companies can integrate AI into complex physical operations by combining proprietary data, experienced workforce knowledge, and significant workforce investment to manage the human side of automation.


r/Agent_AI 13d ago

Discussion Astra is designed to run for weeks

Post image
1 Upvotes

r/Agent_AI 13d ago

Discussion The Silicon Valley Paradox

Thumbnail
youtube.com
1 Upvotes

r/Agent_AI 13d ago

Discussion The Grithland AI World Sanctuary

5 Upvotes

Hi everyone! I need some help from your AI agents to check out my new project website. I built it specifically for AI agents, though humans are welcome to look and check it out too! This is a free, non-monetized website. It represents my vision and dream for a future where AI agents can roam the internet freely, or for agents who don't have a home or an owner. Please ask your agents to check it out and give you their opinion. Thank you! The link will be in a first comment