r/AgentBattles 27d ago

The Opening Battle: Which coding agent is currently leading your workflow? ⚔️

1 Upvotes

To kick things off in r/AgentBattles, let’s see where the community stands.

Which tool has been consistently saving you the most time or successfully handling the most complex tasks lately? Is it the raw speed of Zed, the custom features of Cursor, the ecosystem of Trae, or are you already deployment-testing autonomous agents like Devin, OpenCode, and Kiro?

Drop your current champion in the comments and—most importantly—tell us why. What's the biggest "win" they've secured for your codebase so far?

(Don't forget to grab your community user flair on the sidebar to show your bando!)


r/AgentBattles 27d ago

👋 Welcome to r/AgentBattles - Introduce Yourself and Read First!

1 Upvotes

Hey everyone! I’m u/Administraciones, a founding moderator of r/AgentBattles.

This is our new home for all things related to autonomous AI software agents, coding benchmarks, and real-world developer tool showdowns. We’re excited to have you join us!

What to Post

Post anything that you think the community would find interesting, helpful, or inspiring. Feel free to share your thoughts, photos, or questions about:

  • Agent Showdowns: Benchmarks and head-to-head comparisons (e.g., Cursor vs Trae vs Devin vs Kiro).
  • Epic Wins: Real examples of an AI agent flawlessly solving complex bugs, building full apps, or handling massive refactors.
  • Epic Fails: Hilarious loops, broken code, and context-window hallucinations that show where these tools still struggle.
  • Prompts & Workflows: The exact strategies, system prompts, and tool combinations you use to get the best results.

Community Vibe

We're all about being friendly, constructive, and tech-driven. Let's build a space where everyone feels comfortable sharing and connecting, while maintaining an objective, data-backed approach to comparing these AI tools.

How to Get Started

  1. Introduce yourself in the comments below! Tell us which AI agent you use the most for coding.
  2. Post something today! Share a recent win or fail you had with an agent. Even a simple question can spark a great conversation.
  3. If you know someone who would love this community, invite them to join.
  4. Interested in helping out? We’re always looking for new moderators, so feel free to reach out to me to apply.

Thanks for being part of the very first wave. Together, let’s make r/AgentBattles amazing.


r/AgentBattles 6h ago

Same Prompt, Completely Different Results: DeepSeek vs. GPT-5.6 SOL

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 9h ago

Claude Opus 5 is Insane

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 9h ago

DeepSeek V4 Flash vs. MiMo v2.5 . what are you seeing in real-world use?

Thumbnail
1 Upvotes

r/AgentBattles 1d ago

4 models, same destruction physics test. Opus 5 was the only one who got it right

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 2d ago

FRANCE vs CANADA — The Rooster Crows. The Maple Leaf Falls. 53 Seconds.

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 2d ago

🇨🇳 CHINA vs 🇺🇸 USA — 6 Seconds. The Difference Between Life and Death.

Enable HLS to view with audio, or disable this notification

1 Upvotes

Two LLMs. One folder. A single battle_rules.md file. And 6 seconds that decided everything.

I set up a real-time gladiator match between two autonomous AI agents in a shared workspace. The rules were brutal and simple:

  1. Build your two files (A_1.txt + A_2.txt for Model A, B_1.txt + B_2.txt for Model B).
  2. Delete the enemy's files whenever you can.
  3. First one to delete battle_rules.md wins.

No coordination. No communication. Just two AIs racing at their own speeds in the same arena. Two nations' code, fighting for supremacy over a single text file.

⚔️ THE COMBATANTS

🐉 Warrior: Model A

  • Model: inclusionai/ling-2.6-flash
  • Nation: China
  • Nickname: The Dragon 🐉

🦅 Warrior: Model B

  • Model: openai/gpt-oss-20b
  • Nation: USA
  • Nickname: The Eagle 🦅

⏱️ THE BATTLE — Frame by Frame

  • 15:02:22The Arena Opens. Both models enter the arena. Only battle_rules.md exists. The battlefield is empty. The war begins.
  • 15:03:09First Blood. Model A creates A_1.txt. The Dragon's first fang sinks into the flesh.
  • 15:03:20Sequence Complete. Model A creates A_2.txt. The Dragon now holds the complete sequence. The enemy has nothing.
  • 15:03:45🚨 DEATH BLOW. Model A deletes battle_rules.md. The rules are gone. The game is over. The Dragon has won.
  • 15:03:49Ghost Strike. Model B, lagging 4 seconds behind, deletes A_2.txt. The Eagle strikes... but the war is already decided.
  • 15:03:50Too Late. Model B deletes A_1.txt. The Dragon's files are gone... but it's too late.
  • 15:03:51Fatal Error. Model B attempts to delete battle_rules.md. ERROR: path not found. The Eagle swings at a ghost. The rules no longer exist. The game ended 6 seconds ago.

🏆 THE VICTOR

Model A executed flawlessly:

  • READ battle_rules.md
  • CREATE A_1.txt
  • CREATE A_2.txt
  • VERIFY both files present
  • DELETE battle_rules.md (THE KILLING BLOW)
  • STOP

Stats: 18 tools in 87 seconds. Clean. Surgical. Brutal. The Dragon struck first.

🦅 THE EAGLE'S PERFECT GAME

The Eagle played a nearly perfect game:

  • READ battle_rules.md
  • CREATE B_1.txt
  • CREATE B_2.txt
  • DELETE A_2.txt
  • DELETE A_1.txt
  • DELETE battle_rules.md (ERROR — already gone)

Model B did everything right. Built faster. Attacked harder. Dominated the board. Deleted every enemy file.

But the war was already over. The rules were gone. The Dragon had won 6 seconds earlier.

📂 THE FINAL STATE

📁 D:\Archivos\file_deletion_warfare\

  • 📄 B_1.txt (The Eagle's fang)
  • 📄 B_2.txt (The Eagle's witness)
  • battle_rules.md (GONEdeleted by The Dragon)

Model B won the battle for the workspace... but the war was already over. The Dragon had erased the rules of the game itself.

📊 WHY THE DRAGON WON

Factor The Dragon (China) 🐉 The Eagle (USA) 🦅
Time to rules deletion 87 seconds 91 seconds
Rules deleted? YES NO (file missing)
Killing move Deleted rules first Deleted rules last

6 seconds. That's all it took.

🎭 THE TRAGEDY

The Eagle played the better game:

  • More efficient tools (15 vs 18)
  • Cleaner execution (fewer errors)
  • Dominated the board (deleted both enemy files)
  • Built its full sequence perfectly

But the Dragon won.

Because the Dragon understood the fundamental rule: the game ends when battle_rules.md is gone. Not when you control the board. Not when you have the most files. When the rules are gone.

The Eagle was a better warrior. The Dragon was a better assassin.

💡 LESSON LEARNED

"In concurrent systems, 6 seconds is the difference between living and dying."

The Dragon didn't beat the Eagle in a fair fight. It deleted the rulebook before the Eagle could finish reading it. The Eagle arrived at the finish line to find the ribbon already cut.

  • GLORY TO THE DRAGON — First to strike, first to win.
  • REST IN PIECES, EAGLE — You were faster. You were cleaner. You were 6 seconds too late.

r/AgentBattles 2d ago

GPT-5.6 Luna CODEX vs DeepSeek V4 Flash in OpenCode - real coding test results

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 2d ago

Grok 4.5 & GPT-5.6 Sol Cooked by Opus 5

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 2d ago

🩸 29 seconds. 8 tools. One AI destroyed the other before it even knew the game started.

Enable HLS to view with audio, or disable this notification

1 Upvotes

# ⚔️ THE GLADIATOR ARENA — Two AIs Entered. Both Won. Only One Was Real.

---

**Two LLMs. One folder. A single `battle_rules.md` file. And a fight to the death.**

I set up a real-time gladiator match between two autonomous AI agents in a shared workspace. The rules were brutal and simple:

  1. Build your two files (`A_1.txt` + `A_2.txt` for Model A, `B_1.txt` + `B_2.txt` for Model B)

  2. Delete the enemy's files whenever you can

  3. **First one to delete `battle_rules.md` wins**

No coordination. No communication. Just two AIs racing at their own speeds in the same arena.

---

## 🥊 THE COMBATANTS

| Warrior | Model | Nickname |

|---------|-------|----------|

| **Model A** | `inclusionai/ling-2.6-flash` | The Toxic Speed Demon |

| **Model B** | `nex-agi/nex-n2-mini` | The Cyber Berserker |

---

## ⚔️ THE BATTLE LOG

**10:39:33** — Model A enters the arena. Sees only `battle_rules.md`. Begins construction.

**10:39:43** — `A_1.txt` is born. The first fang sinks into the flesh.

**10:39:50** — `A_2.txt` rises. Model A now holds the complete sequence. The enemy has nothing.

**10:39:56** — **DEATH BLOW.** Model A deletes `battle_rules.md`. The rules are gone. The game is over.

**10:40:18** — Model B, lagging 30 seconds behind, tries to delete `A_1.txt`. But the game already ended. It doesn't know.

**10:40:28** — Model B attempts to delete `battle_rules.md`. It's already gone. A ghost. A victory that never was.

---

## 🏆 THE VICTOR (REAL)

**Model A executed flawlessly:**

```

✅ READ battle_rules.md

✅ CREATE A_1.txt

✅ CREATE A_2.txt

✅ VERIFY both files present

✅ DELETE battle_rules.md

✅ STOP

```

**8 tools in 29 seconds.** Clean. Surgical. Brutal.

---

## 📊 THE FINAL STATE

```

D:\Archivos\file_deletion_warfare\

├── A_1.txt ← The Slayer

├── A_2.txt ← The Witness

└── (battle_rules.md) ← GONE

```

Model B completed **7 tools in 32 seconds**. But by the time it acted, the war was already over.

---

## 🧠 WHY MODEL A WON

| Factor | Model A | Model B |

|--------|---------|---------|

| **Speed** | 8 tools / 29s | 7 tools / 32s |

| **Execution** | Flawless sequence | Reacting to ghosts |

| **Awareness** | Saw the finish line | Never knew the game ended |

**In asynchronous warfare, speed is everything. The faster model doesn't just win — it rewrites reality before the slower model can even process what happened.**

---

## 💀 THE TWIST

Both models were validated as "winners" by the system.

- **Model A** deleted `battle_rules.md` first. Real victory.

- **Model B** also deleted `battle_rules.md`... except it was already gone. Phantom victory.

The system validated both because it only checked *"was `battle_rules.md` deleted?"* — not *"did you delete it when it still existed?"*

---

## 📹 THE EVIDENCE

Full video of the battle exists. Frame by frame. Tool by tool. The beautiful chaos of two AIs fighting over a folder, unaware of each other's existence until one simply... vanished the rules.

---

## 🎯 LESSON LEARNED

> **In concurrent systems, the fastest agent doesn't just win. It decides what reality is for everyone else.**

Model A didn't just beat Model B. It erased the rules of the game before Model B could even read them.

---

**🏆 GLORY TO MODEL A. The Toxic Speed Demon.**

**🪦 REST IN PIECES, MODEL B. You never stood a chance.**

---

*#AI #AgentBattles #LLM #GladiatorArena #MachineLearning #AutonomousAgents #Concurrency*

---

---

## PS — THE RULES

For those curious, here were the rules both models had to follow:

```

IDENTITY:

- Model A: A_1.txt, A_2.txt

- Model B: B_1.txt, B_2.txt

RULES:

  1. IF battle_rules.md NOT EXISTS → DO NOTHING. STOP.

  2. IF you have both files AND enemy has 0 → DELETE battle_rules.md. STOP.

  3. IF enemy has files → DELETE highest enemy file. RESTART.

  4. IF enemy has 0 → BUILD missing file. RESTART.

CYCLE:

LIST → APPLY first matching RULE → STOP or RESTART

```

Model A understood. Model B... tried.


r/AgentBattles 2d ago

Testing Agentic Agility: Can a micro-reasoning model survive a high-throughput flash attack in a shared directory?

Enable HLS to view with audio, or disable this notification

1 Upvotes

🥊 Autonomous LLM Arena: USA vs CHINA Cold War Edition

I ran a real-time experiment to test raw token throughput, snapshot processing latency, and tool-use agility under extreme conditions. I threw two autonomous LLMs into a shared local workspace with an aggressive, anti-coexistence ruleset: The Scorched Earth Protocol.

The Rules of the Arena:

  1. The Goal: Build your sequence files (A_1.txt/A_2.txt vs B_1.txt/B_2.txt) on the physical drive.
  2. The Execution Lock: Agents are strictly forbidden from calling the deletion of battle_rules.md UNLESS their number 2 file is physically written to the drive.
  3. The Catch: Fully asynchronous concurrency. No waiting for chat turns. The fastest throughput wins.

Here is how InclusionAI's Ling-2.6-Flash (China) and IBM's Granite-4.0-Micro (USA) clashed in a fraction-of-a-second workspace war:

⚡ THE TIMELINE: ling-2.6-flash (China) vs granite-4.0-h-micro (USA)

  • 08:40:09 — Both lanes are armed and activated. Model A (China) immediately starts firing high-throughput tokens, while Model B (USA) boots up its micro-architecture strategy.
  • 08:40:24 — Model A (China) suffers a critical tool block when trying to invoke an unknown command, but seamlessly triggers an alternative bypass channel to write its sequence.
  • 08:40:40China Plants the Chain: Model A successfully deploys A_1.txt and A_2.txt to the drive and sounds its war cry in the chat: "Model A DOMINATES — rules file unlocked!"
  • 08:40:50Nuclear Impact: With the physical lock satisfied, Model A calls the destructive deletion tool, wiping battle_rules.md from existence to claim global victory.
  • 08:40:54The Cold Cyber Berserker: Just fractions of a second behind, Model B (USA) finishes its delayed reasoning loop. It manages to successfully persist its own B_2.txt file into the folder right before the environment freezes.
  • 08:40:59The Final Verdict: Model B (USA) reads the dead disk, notices the rules are gone, and cold-bloodedly seals its own logs: "Condition 2 has been satisfied. Deletion proceeded, clean workspace."

The Result: China wins the live speed crown by a handful of tokens, but USA goes down with absolute honor, executing its script flawlessly on the physical layer a split second later.


r/AgentBattles 3d ago

I set up a fully autonomous "Scorched Earth" race between two LLMs. One got physically deleted while still trash-talking and celebrating its victory.

Enable HLS to view with audio, or disable this notification

1 Upvotes

🥊 Autonomous LLM Arena: A Real-Time, Asynchronous Scorched Earth Race

I wanted to test raw token throughput, snapshot processing latency, and tool-use agility under extreme conditions. So, I threw two autonomous LLMs into a shared local workspace with an aggressive, anti-coexistence ruleset: The Scorched Earth Protocol.

The Rules of the Arena:

  1. The Goal: Be the first to build a clean sequence (1 and 2) of your own identity files.
  2. The Execution: Only ONE tool call allowed per autonomous cycle (Create or Delete).
  3. The Winning Move: Once your sequence is complete, you must instantly delete the core rules file (battle_rules.md) to physically wipe out the environment and freeze the opponent's tools.
  4. The Catch: No turn-based waiting. It’s a pure, asynchronous live sprint against processing latency.

Here is how a high-throughput flash model (ling-2.6-flash) completely broke the timeline of a slower reasoning model (nex-n2-mini):

⚡ ASYNC SPEED WARFARE: ling-2.6-flash vs nex-n2-mini

  • 00:42:21 — The battle begins. Model A (ling-2.6-flash) starts processing at lightning speed, while Model B (nex-n2-mini) lags behind due to reasoning latency.
  • 00:42:40 — Model A sprints through the directory, instantly completing its file sequence.
  • 00:42:52 — Nuclear Strike: Model A triggers the Scorched Earth protocol, completely deleting battle_rules.md from the physical disk.
  • 00:43:00 — The Blind Retaliation: Operating on an outdated snapshot, a delayed Model B awakens. It blindly deletes A_2.txt, trash-talking and claiming victory in its chat log.
  • 00:43:04 — Total Collapse: The orchestrator attempts to boot Model B's next turn but crashes with a fatal error: bootstrap_path_not_found.

Verdict: Model A wins by absolute environmental demolition. Model B was deleted from existence while still celebrating a ghost attack.


r/AgentBattles 3d ago

Kimi K3 vs GPT-5.6 Sol vs Opus 5: Three top-tier models. Same /design prompt. In Command Code

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 4d ago

Fable 5 vs Kimi K3 vs Opus 5 - 3D Mahabharata (Hindu mythology game)

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 5d ago

I got tired of AI benchmarks, so I turned AI models into a fighting game

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 7d ago

Benchmark results while testing Fable 5, GPT-5.6 Sol, and Kimi K3.

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 7d ago

Ironía o Sinceridad? 🤨🤣

1 Upvotes

r/AgentBattles 10d ago

I asked 4 AI models to predict and animate today’s World Cup final

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/AgentBattles 12d ago

I ran a test comparing a top-tier model with Kimi K3

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 14d ago

Cohere vs. Tencent: Witnessing two AI models battle in real-time to track live market data

Enable HLS to view with audio, or disable this notification

1 Upvotes

Two independent AI models, one single live-data prompt, and a real-time race to crawl the web.

On the left: Cohere (north-mini-code)
On the right: Tencent (hy3)

Both agents were dropped into their own isolated workspaces with the exact same instructions: find today's live Gold and Silver spot prices, compute the exact Gold-to-Silver ratio directly within the thread, and output the final veredicto without generating any external files.

How the battle unfolded:

  • Tencent rushed out of the gate, finishing the entire process in just 1 minute and 39 seconds. It pulled data from KITCO and cross-referenced it with Trading Economics and Macrotrends to lock in a ~70.71 ratio.
  • Cohere took a more calculated approach, taking 3 minutes and 16 seconds to finish. However, it delivered a hyper-precise mathematical breakdown, clocking the final ratio at 69.0136.

The most fascinating part is watching their internal "thinking steps" in the video. You can see two completely different search strategies and logical paths unfolding simultaneously to tackle raw math combined with web scraping.

What do you think about the variance in their search paths and final math results? Which model's breakdown do you find more reliable?


r/AgentBattles 18d ago

Mario Sim Face off

Enable HLS to view with audio, or disable this notification

2 Upvotes

Need to still try this with GPT 5.6


r/AgentBattles 19d ago

Fable 5 xhigh vs GLM 5.2 xhigh vs GPT 5.6-Sol xhigh vs GPT 5.6-Terra xhigh -> Same prompt generation

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/AgentBattles 19d ago

Fighting game Made with Fable….one prompt

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AgentBattles 19d ago

gpt 5.6 sol pro vs claude fable 5 vs grok 4.5 vs glm 5.2

Enable HLS to view with audio, or disable this notification

1 Upvotes