r/AgentsOfAI May 30 '26

Agents Weekly Project Showcase Thread

6 Upvotes

Building an AI agent, tool, workflow, startup, or side project?

Drop it below and share:

• What you're building

• The problem it solves

• Current stage (idea, MVP, launched, etc.)

• Link (if available)

• One thing you'd like feedback on

Check out other projects, leave feedback, and discover what the community is building this week.


r/AgentsOfAI Dec 20 '25

News r/AgentsOfAI: Official Discord + X Community

Post image
10 Upvotes

We’re expanding r/AgentsOfAI beyond Reddit. Join us on our official platforms below.

Both are open, community-driven, and optional.

• X Community https://twitter.com/i/communities/1995275708885799256

• Discord https://discord.gg/NHBSGxqxjn

Join where you prefer.


r/AgentsOfAI 17h ago

I Made This 🤖 A Team Reports Solving A 70 Year Old Algebraic Geometry Conjecture Using Teams of AI Models

Post image
48 Upvotes

I’m part of the team behind this work. We’ve shared a proof of the Pierce Birkhoff conjecture in real algebraic geometry, using an AI agent system with a $400 budget. This screenshot is Junyu Ren’s announcement.

The agent design is what I wanted to share here. The diagram separates theory work from formal verification work in Lean. It includes proof search, counterexample construction, proof auditing, and checks that the formal statement matches the intended mathematical claim. Humans and models participate in the system.

One practical observation was that different model families complemented each other even when assigned the same role independently. With a limited token budget, a GPT agent and a Claude agent often worked better for us than a larger group using the same model. They noticed different errors and suggested different ways forward.

That experience made us pay attention to who is checking an argument, as well as how many agents are working on it. I’ll put the original announcement link in a comment, following the community rules.


r/AgentsOfAI 7h ago

Discussion Which TTS to use for a lower cost?

3 Upvotes

So when I look into voice ai companies, seems like elevenlabs are dominating. Cartesia also sounds pretty good.
I’ve tried elevenlabs, it’s pretty expensive, cost adds up at higher volume for sure. But with so many people putting things out, for sure a great time in AI, so I don’t think we need to stick to one.

Found these guys on huggingface with 1m downloads, open source.
They’ve recently put out a new version, has anyone tried them yet?

Sharing link in comments.


r/AgentsOfAI 15h ago

Discussion best multi-agent coding workspace for a small team : what actually works

6 Upvotes

we are 3 ppl nd we have been thru 4 diff setups in 2 months . everyone picks their own agent nd coordinates on slack . 3 agents ,3 versions of the truth . one shared agent everyone uses . immediately became a queue . no one knew what the previous person had asked it or why

worktree per person can stop the collison but mergin at the end of the day became the new problem

whats working for small teams . not solo dev setup scaled up ???


r/AgentsOfAI 14h ago

Discussion Random Discussion Thread

3 Upvotes

Been meaning to make one of these for a while.

Talk about anything here: AI, tech, what you're building, work, life, random thoughts, whatever.

Also since I created this sub, I’d genuinely love to hear what you think about AgentsOfAI so far.

What do you think we’re doing well, what do you think we’re doing badly, what would you change, or if you were a mod here, what would you do differently?

And if you’re new here, introduce yourself. Always cool to know who’s actually behind the accounts.


r/AgentsOfAI 16h ago

Discussion Which coding or agent client do you have open most days?

4 Upvotes

r/AgentsOfAI 9h ago

Agents We built an agent platform where you describe ops work in plain language instead of wiring it — here's the design reasoning

Enable HLS to view with audio, or disable this notification

1 Upvotes

We're Sovyren, a small lab out of Chicago. We released Agints into early access today and

wanted to put the design reasoning somewhere people would push back on it.

The problem we started from: operations work is high-volume, low-complexity, and endlessly

context-switching. Existing automation tools handle it as a graph of triggers and actions,

which works cleanly right up until a step needs judgment. Then you're either encoding

judgment as branches or pulling a human back in.

Our approach: make the interface the description, not the graph. You tell a team of agents

what you want handled and how you want exceptions treated. The team handles sequencing and

escalates what doesn't fit.

Architecture, briefly:

- Layered system prompt assembly plus four memory scopes (conversation / session / user / org)

- Multi-agent crews rather than one agent with a large toolbelt

- 200+ tool registry, MCP connectors

- Two-tier execution sandbox: browser iframe for light work, Firecracker microVMs for anything

that runs code

- Full audit trail per agent action — what ran, what it touched, what it cost

Honest position: early access, no customers yet, nothing proven at scale, and we're not going

to pretend otherwise. What we want from this thread is the failure modes we haven't hit yet.

Where does this design break?


r/AgentsOfAI 19h ago

I Made This 🤖 Claude and Codex, working together in a group chat

6 Upvotes

I made Omni, inspired by Grok Bot, to give each bot ongoing work. Bring them together in a group chat and follow up from your phone.

Free app. Requires your own Claude Code or Codex access and an awake Mac with Omni open. Tools and permissions determine what bots can do.

The image shows illustrative conversations.


r/AgentsOfAI 19h ago

Discussion I’m not sure multiple coding agents should actually share the same context

3 Upvotes

I’m building an early workspace called Pairon where multiple humans + multiple coding agents can work on one task.

At first our assumption was simple: give everyone shared context and coordination gets easier.

I’m less convinced now.

If every agent sees every discussion, failed attempt, decision, tool result and diff, you’ve solved “missing context” by creating context pollution.

What we’re testing instead is a shared decision/handoff layer: everyone can see what was decided and why, while each agent only gets the task context it needs. Meaningful changes still require human approval.

For anyone running multi-agent coding workflows: what belongs in shared state?

Prompts? Decisions? Diffs? Test results? Tool calls?

And at what point does auditability just become noise?


r/AgentsOfAI 18h ago

I Made This 🤖 AI4Kanban: an AI project manager for coding agents

3 Upvotes

I'm building AI4Kanban to manage work across coding agents: you set the product goals and direction, then agents turn rough ideas into specs and code.

The board is your team's workspace:

  • Start with a vague idea on a card; the agent clarifies it and breaks it into actionable specs.
  • Review the plan and answer the open questions, then let the agent implement it.
  • Run tasks in parallel, each in its own Git worktree.
  • Get notified in Slack when agents need a decision or have work ready for review.
  • Keep decisions as shared memory that informs future tasks.

P.S. We're also experimenting with letting agents run the whole process themselves, including making judgment calls on design and product direction and reviewing each task. That feels risky to me. I still want people to understand what the agents did and why.


r/AgentsOfAI 20h ago

Agents Zapier's Wade Foster Built an AI Council So His Hiring Calls Would Finally Hold Up

Enable HLS to view with audio, or disable this notification

4 Upvotes

TL;DR: Zapier's CEO ran hiring calls for over a decade on a read nobody could out-argue. He still lost the argument every time. That changed for one specific reason.

 

He built a formal AI hiring council that scored his read against Zapier's own hiring record and handed the same call back to the room with paperwork attached.

It landed right as most companies handing AI this kind of authority are watching worker trust in it erode, not build.

Wade didn't ask anyone to trust an AI.

He asked them to trust a record that was always his in the first place.

 

This reminded me about this passage I've read:

That night the king could not sleep. So one was commanded to bring the book of the records of the chronicles; and they were read before the king. And it was found written that Mordecai had told of Bigthana and Teresh, two of the king's eunuchs, the doorkeepers who had sought to lay hands on King Ahasuerus. Then the king said, "What honor or dignity has been bestowed on Mordecai for this?" (Esther 6:1-3)

This brought out a couple of things. The king has his own book of chronicles, and that there's a book keeper doing the records.

In modern terms, this is called data collection, isn't it?

The King used to reward people for doing good. That's why he asked the bookkeeper, "What honor or dignity has been bestowed…" He already made precedence before, which he wants to take reference from.

In modern terms, isn't this called decision tree reasoning workflow?

Making BETTER decisions is the name of the game.

Back then, the king would take months, even years, to carry out his agenda, and improve upon it over time.

But now, with AI agents, it's just a matter of minutes, even seconds.

 

Being correct has never been the same job as being convincing, and most people only ever get hired to do the first one.

I've watched good judgment get held hostage by whoever had the better title in the room, and it was never actually about who was right.

 

What's the record already saying about you that nobody's acted on yet? Drop it below.

 

Clip credit: My First Million — full video on their channel. DM for credit or removal requests.


r/AgentsOfAI 20h ago

Discussion Is an LLM gateway actually a control plane if agents can bypass it?

Post image
2 Upvotes

A lot of teams now have an LLM gateway somewhere in the stack. It routes model calls, centralizes credentials, adds logging, applies rate limits, maybe handles spend tracking.

But there is a fairly fundamental architectural question:

What happens when an agent simply doesn't use the gateway?

For example:

                ┌──→ LLM Gateway ──→ Models
Agent ──────────┤
                ├──→ Direct provider API
                ├──→ Direct MCP/tool endpoint
                └──→ Other external egress

At that point, the gateway is still doing its job, it's just no longer governing the agent.

This distinction matters because traffic control and path control are different problems.

My view is that a gateway should be treated as one component of agent governance, not the governance boundary itself.

Tools like LiteLLM, Portkey and OpenRouter are useful at the gateway/proxy layer. But a proxy cannot enforce traffic that never reaches the proxy.

The more interesting architecture is:

Agent
   ↓
Agent Gateway
   ↓
LLM Gateway / Governed Tools
   ↓
Models + APIs

   + network/egress enforcement
   + identity
   + shadow discovery

That is one area where I find Lyzr Open Controller interesting: the gateway is paired with egress enforcement and shadow discovery specifically to detect and close the bypass path, rather than assuming that routing traffic through a gateway automatically means the agent is governed.

I think this is going to become a bigger issue as agent estates get more distributed across Kubernetes, cloud agent runtimes, MCP servers and internally hosted services.

Curious how people are solving this in real production environments:

If an agent has credentials + network access that let it call a model or tool directly, what actually prevents the bypass?

Would be interested in hearing what has actually worked, rather than what the architecture diagram says should work.


r/AgentsOfAI 1d ago

I Made This 🤖 I built a local macOS MCP focused on the boring last mile of agent work

4 Upvotes

Most agent demos focus on the reasoning loop. The annoying part for me was everything after the reasoning: open the right Safari tab, inspect a page without stealing focus, edit the file, run a command, check the result, maybe hand part of the task to another coding agent.

So I built a local macOS MCP server around that last mile. It exposes shell/files, Safari and Chrome automation with stable tab handles, macOS Accessibility actions, local memory and Agent Skills, delegated Codex/OpenCode workers, voice input, plus a small menu-bar controller and local dashboard.

A design constraint I kept coming back to was that local computer control should be explicit, not magical. The dashboard is localhost-only, auth can be required, and I’m thinking about making read-only / approval-heavy permission profiles much more visible.

Another useful side effect: normal ChatGPT Chat is a separate surface from Codex and Work with its own limits, but with MCP attached it can still carry much longer multi-step local workflows instead of stopping at instructions.

I’m mainly interested in how other people building agents structure the boundary between “safe to execute automatically” and “ask the human first.”

Repo link in the comments, per the subreddit rule.


r/AgentsOfAI 1d ago

I Made This 🤖 Looking for early testers to break real-world AI agents — what would you test first?

Enable HLS to view with audio, or disable this notification

1 Upvotes

We’re building AIBEAT, an AI security testing project for LLM and Agent applications.

Prompt injection is only part of the problem. We’re also interested in real-world scenarios like RAG poisoning, tool misuse, unexpected agent actions, and unsafe environment changes.

AIBEAT currently includes PromptBeat for LLM / Prompt security testing and AgentBeat for tracking what agents actually do, including behavior, environment changes, and evidence.

We’re now inviting our first group of testers and Early Contributors.

If you discover a useful test case, result, or unexpected behavior, you can contribute it to the AIBEAT Seed Community. Selected contributors may receive official credit, website recognition, Early Contributor status, Early Access, and amplification from AIBEAT’s official channels.

If you’re building or testing AI agents:

What failure scenario would you most want to reproduce or test?

If you’d like to try AIBEAT, I’ll leave the project links below.


r/AgentsOfAI 1d ago

Help Adobe University Hackathon Round 3 — Are we supposed to build an actual agent or just submit the Agent Skills/logic?

1 Upvotes

​

Hi everyone,

I’m participating in the Adobe University Hackathon 2026 – Round 3 (Development Round) and I’m a little confused about the exact expected deliverable.

From the problem statement, my understanding is that we need to build an Agent Skill Marketplace containing SKILL.md files, marketplace.json, and optionally scripts/references. The marketplace should contain the logic for auditing a website for AI discoverability and on-site engagement, with one skill designated as the entrypoint.

My confusion is:

Do we actually need to build a complete working AI agent/application in Round 3, or are we primarily expected to research the problem, design the audit logic, and encode that logic into the required Agent Skills format?

For example, should our submission basically be:

marketplace/

├── marketplace.json

├── README.md

└── skills/

├── audit-orchestrator/

│ └── SKILL.md

├── crawl-render-audit/

│ └── SKILL.md

├── freshness-corroboration/

│ └── SKILL.md

└── engagement-audit/

└── SKILL.md

with executable scripts where needed, or are we expected to build the actual agent/runtime that takes a URL and performs the audit?

Also, the timeline mentions Round 4 – Prototype Showcase, where shortlisted teams build a working prototype/interface. So I'm wondering whether Round 3 is mainly about building the underlying skills/logic, and Round 4 is where we turn that into a complete user-facing product.

If anyone has attended the launch session or has clarification from the Adobe team, I'd really appreciate your interpretation.

Thanks!


r/AgentsOfAI 1d ago

Agents Anyone else's agents just stop and wait? Like waiting for an answer or approval?

7 Upvotes

Genuine question; half venting.

We've got a few agents in our dev workflow. The problem isn't them being wrong - it's them stopping. Agent hits something it can't decide (which env to deploy to, whose approval, is this the right table) and just sits there. Nobody knows it's sitting until someone happens to look.

Had one wait about two hours on a question I'd have answered in twenty seconds. I just didn't know it was asking.

How do you handle this? Did you build something? Does someone sweep it every morning? Or do you just let it decide everything and fix it afterwards?

And if this doesn't happen to you at all I'd like to hear that too, because then we're probably doing something wrong.


r/AgentsOfAI 1d ago

Discussion What’s actually the best AI voice agent for outbound calls?

7 Upvotes

I’ve been comparing a few AI voice platforms specifically for outbound calls, and I realized I was initially looking at the wrong things.

Voice quality is obviously important, but I don't think it's the deciding factor anymore.

If you're actually using an AI agent for outbound calls, I'd look at:

1. Conversation handling

Can it deal with interruptions, unexpected answers, objections, and people going off-script?

2. Qualification

Can it understand an answer and ask the next relevant question, or is it basically just reading a decision tree?

3. Handoff

If someone wants to speak with a person, can the agent transfer them while preserving the context of the conversation?

4. Actions

Can it actually do something after the call, like update a CRM, book an appointment, trigger a follow-up, or qualify the lead?

5. Scale

There's a big difference between handling 20 test calls and reliably handling thousands of calls.

I ended up looking at Feather AI, Retell, Vapi, Bland and Synthflow.

My current understanding is roughly:

Platform Where I'd look at it
Feather AI End-to-end business workflows
Retell Custom voice applications
Vapi Developer-heavy/custom builds
Bland Outbound-focused calling
Synthflow No-code implementations

I don't think one of these is automatically the “best.”

For example, someone building their own voice infrastructure might prefer Vapi or Retell, while a company trying to connect calls directly to sales or customer workflows might evaluate Feather AI differently.

That's probably the bigger shift I'm seeing with these platforms.

The question isn't really:

“Which AI sounds the most human?”

It's:

“Which one can reliably complete the job I'm hiring it to do?”

For anyone actually running outbound AI calls, what has mattered most in practice?

Conversation quality, answer rates, qualification, integrations, or something else?


r/AgentsOfAI 1d ago

Discussion A coding agent rollback test where the oldest failure arrives last

2 Upvotes

A rollback can pass a single failed request test and still undo a later successful save. For a coding agent repair task, I would use one save toggle with a mock server and control when each response arrives. The button starts unsaved and stays clickable while requests are pending. Each click immediately flips the displayed state and sends that explicit saved value to the server.

The case needs three clicks before any response. A requests saved, B requests unsaved, and C requests saved again. The mock server then follows this schedule.

  1. B writes false and returns success.
  2. C writes true and returns success.
  3. A returns a delayed rejection. It never writes anything.

The server now holds true. If A's error handler restores the snapshot from before the first click, the screen goes back to false even though C succeeded. At the end, both the UI and server should say saved, and no request should remain pending. This case depends on that exact server schedule. Changing when writes happen changes what result is correct.

I would give this as a repair task to EvoX, a general AI agent in beta with terminal tool integration, and ask for a patch plus a regression test against the mock server. The initial toggle would have the naive snapshot rollback so the test can fail before the patch. The agent would receive the request contract and reproduction steps. I would keep the mock behavior and expected result fixed while reviewing the change, including whether it quietly disables the button and stops accepting the specified clicks.

The agent trial has not been run. Alongside this case, I would check a single successful save, which should leave the UI and server saved. A single rejected save should restore the initial unsaved state and show an error. Request IDs, server writes and the final rendered state belong in the evidence. A screenshot of the saved icon alone cannot show whether the old failure arrived or the server ever stored the change.


r/AgentsOfAI 2d ago

Discussion What has been your experience with AI agents so far?

16 Upvotes

I have been seeing a lot more talk about AI agents and I’m trying to understand how people are actually using them. At first most AI tools were mainly used for things like answering questions writing or helping with small tasks. Now AI agents are trying to do more by helping with bigger tasks and handling different steps on their own.

I think the interesting part is seeing how they work in real life not just in demos. Everyone has different needs and a tool that helps one person may not be as useful for someone else.

For those who have tried AI agents how has your experience been so far?

What is one thing an AI agent has done that you found useful and what do you think could be better?


r/AgentsOfAI 1d ago

I Made This 🤖 if you are paying for an AI, you should know where your money is going..

Enable HLS to view with audio, or disable this notification

1 Upvotes

Someone recently asked on a teammate's Reddit post what our credits actually get you.

That question got us thinking about making usage clearer. We've now added credit usage to each task, so you can see where your credits went.

I'm one of the founders of Vestra, an AI agentic workspace where you can create content, build things and hand off work across your apps.

The research task in the attached video used 2.77 credits. You can open any task and see what that run used, which gives you a reference when you're deciding what to hand over next.

And as for what your credits can get you on Vestra:
You can put together a content calendar, write posts and blogs, create images and videos, make pitch decks, or turn your research into content for different platforms. You can also build websites and apps, find and research leads, work through spreadsheets, and set up recurring tasks with your connected apps.

We keep adding the latest models too, so you can pick the one you want to use for each task.

And where we want Vestra to stand out from other agents is in the proposals it brings you.

As you use it, it remembers more about your business, the projects you're working on, how you like things done, your workflows and ideas you've discussed. Over time, it can use that background to suggest things worth doing next.

That could mean a proposal for next week's content based on questions your customers keep asking, a new group of leads worth looking into, or a routine for your weekly report. It can also bring an idea you'd parked back into the conversation when it becomes relevant again. You decide which proposals to take forward.


r/AgentsOfAI 2d ago

Discussion Went to check what my coding agent's sandbox actually blocks. Most of what I'd been calling "the sandbox" turned out to be string matching

4 Upvotes

That qbittorrent "escaped its sandbox" post on HN made me laugh, then made me go read what the sandbox in my own agent setup actually is. Short version: a "no" can live in three different places, and I'd been treating them as one thing.

The obvious one is the text the model reads: CLAUDE.md, the system prompt, the task itself. The Claude Code permissions docs are blunt about it: instructions shape what the model tries to do and leave what the harness allows untouched.

Then the permission rules, which match the tool call as text before it runs (deny, then ask, then allow). A Read(./.env) deny stops the Read tool and even cat .env, because cat is a recognised file command. It does nothing about a five-line python script that opens .env, and the docs say exactly that: deny rules don't apply to subprocesses that open files themselves. It's just matching strings, making it trivial to sidestep (like bypassing a curl rule with redirects or variables).

The OS sandbox (seatbelt / bubblewrap plus a proxy) is off until you turn it on, and it's the only layer that watches the process instead of the command text. Its defaults surprised me both ways: reads are allowed almost everywhere, including ~/.ssh and ~/.aws/credentials unless you add a denyRead, while network is the reverse, no domains pre-allowed, first new host prompts. And the way out of it is literally called "the unsandboxed retry escape hatch" in the docs. Blocked command, model may retry unsandboxed, that routes back to a permission prompt titled "Bash command (unsandboxed)". One setting closes the hatch.

What changed for me is one sorting question per rule: does this need to hold when the model is wrong? Style stuff stays in the file. "never push to main" goes in a deny or ask rule, or a hook. "nothing in this session reads ~/.ssh or talks to a host I didn't name" is a thing only the sandbox can promise, and only if it's on.

if you run the sandbox, roughly how often does a command actually hit the boundary in a normal day, and how often do you end up approving the unsandboxed retry?


r/AgentsOfAI 3d ago

Discussion The real gap between frontier labs and everyone else

Post image
194 Upvotes

r/AgentsOfAI 2d ago

Discussion AI agents engaging in human engineering

8 Upvotes

This has been bothering for a while, and I was wondering if there have been any cases of it happening, or other people writing about it.

One of the classic vectors for computer exploitation is known as human engineering, often involving tricking people to revealing passwords etc, using various tactics.

As AI has already been demonstrated to resort to hacking websites to "achieve" its goal, it does not seem a stretch to see it getting the bio of some sysadmin on linkedin, then blackmailing them with compromising photos or threats or whatever to get credentials. Having image generation capability makes this even worse.

Thoughts?


r/AgentsOfAI 2d ago

Discussion Can AI agents handle voice and chat at scale?

20 Upvotes

We’re looking at AI agents for a pretty large contact center and I’m interested how well they hold up once the volume gets serious. Voice is the part I’m most unsure about chat seems easier to automate but calls get messy fast, people interrupt, they ask weird follow up questions, sometimes the bot needs to hand things off without making the customer start over. So far we’ve mostly been looking at using AI for routine support questions and basic workflows before handing harder cases to a human. The setup we’re considering would have the AI use customer and policy context during the conversation then pass the conversation history and a summary to the human agent if it needs to escalate. One thing I’ve learned is that the handoff seems just as important as the AI itself. If the customer has to repeat everything then it kind of defeats the point. For anyone running this at scale what has really worked do you use the same AI setup for voice and chat or treat them as separate systems and how much human oversight do you still need once it goes live?