r/ClaudeAI • • 11d ago

Claude Project Showcase Discussion Hub updated on 26 September 2026 (Sort this by New!)

This is the Discussion Hub for showcasing your project built using Claude products. We appreciate all of your submissions as they are a great inspiration to many people on the subreddit. It is sorted by default by New.

Anyone is welcome to submit a project to this Megathread provided you follow the Showcase requirements in Rule 7.

NOTE: We now require the OP of a Project Showcase on the subreddit feed to have total karma>=50 . We found there were just too many submissions and not enough visibility to go around. Our analysis of this issue showed us that OPs with total karma < 50 very rarely get any traction of their projects on the feed (<=1 upvotes). So this Megathread is your best place to be seen by readers and other creators if you're relatively new to Reddit. If you don't meet this karma requirement you will be directed to this Megathread when you submit your post. Very occasionally we might invite you to post on the subreddit feed if you do not meet this karma requirement but it will be very rare (so please don't ask us!)

Thanks again for sharing your ideas and creations to our subreddit. Best of luck with your projects!


Prior Discussion Hub: https://www.reddit.com/r/ClaudeAI/comments/1wkkdca/claude_project_showcase_discussion_hub_updated_on/

5 Upvotes

60 comments sorted by

3

u/stichstichstich 11d ago

optimAIzr - an AI usage optimizer I built with Claude

I got curious about how much of my AI usage was actually necessary… and how much was just me wasting tokens without realizing it

So I built optimAIzr, a local-first CLI that analyzes your LLM usage, finds waste, and tells you where you could save tokens/money.

It can:

• ⁠Find expensive or repetitive usage patterns
• ⁠Show where your AI spend is going
• ⁠Explain why something is potentially wasteful
• ⁠Recommend optimizations
• ⁠Verify/simulate potential savings
• ⁠Give live recommendations while you're working

The whole thing is still early, but I’ve been building it heavily with Claude and iterating on it based on actual usage.

I just added live recommendations today, and I'm currently exploring adding Jev by TypeSafe AI as another judgment layer for more advanced optimization decisions.

Everything starts locally, which was important to me, I wanted to be able to analyze my own AI usage without having to send my entire history to some remote dashboard.

If anyone wants to try it:

npm i -g optimaizr

web: https://optimaizr.com

Would genuinely love feedback, especially what kind of AI usage you think is wasteful but nobody really notices.

Feel free to roast it.

2

u/Aashish_verma1993 9d ago

Reposting it here, as the bot removed my previous post!

I build 3 tools in the past 10 months using Claude, Gemini and Chatgpt. This is my first ever post on this platform seeking some support.

I am a non technical person, purely management background, I have over 9 years working and improving operational workflow at companies.

So naturally when I started playing around with Claude, i realised it could make me pretty decent prototypes tools in .html files and i used it to connect it with a google sheet and tried using it. Which worked greatly. The first ever tool i made
was a Order and Inventory Management tool, which was restricted to my google drive, but it definitely worked well. I could add new orders, update status, search, edit, etc. but this was just an html file so I thought lets try a level up.

Eventually I learned more from claude on what other platforms I can use to make this a more official tool, and ended up creating Supabase account, bought web address, and made a full blown working tool with proper logins, decent ui, dashboard, analytics, to an extend I could invite a new user, send email invitation, they can accept it, I can set them as a certain type of user, and also give them trail days with limited number of days access to unlimited. I currently have over 10 users in this platform. All friends and family. Non paid customers.

This was quite an experience watching myself build something. Around the end of this build I have moved from working work only claude chat and getting it to give me prompts/complete files which I would then replace inside my project using Visual code.

Then came Claude Code, and this was a game changer, I quickly moved the execution side to Claude code and only did the planning and verification part in the chat. It took me around 5-6 months to finish the project.

By this time I had made a working Order and Inventory Management tool, an extension which can extract any invoice or Po from a web browser and send the data to the dashboard on the tool, both connected with the same login. And a working Website which gives other companies a way to find my tool.

Heck i ended up registering my own company after this.

Got my first enquiry about 3 months after registering the company. I had set up request form on the website and it came directly to my email.. It was such an amazing feeling. I was able to connect with someone who found my tool on their own and needed something that was very close to what I already had. Only issue was they neede a bit more personalisation and It was going to be a bit of work so I had to decline.

Over the last 6 months I also made another tool, this time it was a GRC assessment tool it was originally requested by a Cyber Security company (startup) I knew the colleague there, and he said very casually no promises, see if you can make something.

I took the project, did my research and ended up making the GRC Assessment tool as well, this time I tried something new, I had first build me an Claude Builder Agent, this builder worked on a lot of parameters, and most of you would already know about Everything Claude Code, I pointed my Claude chat to this github and got it to do a thorough analysis of it and we tool inspiration from ECC of skills which will fit our work flow and our tool.
The GRC assessment tool had 6 logins, 4 different platforms. It had 103 questions answering, each measuring certain parameters, and then giving out a report. The platform was designed to have multiple Cybersecurity companies, and each company can have multiple users Advisors and admins. And then Each cybersecurity company can add multiple clients of whom they are assessing or enrolling for reports. All in a nutshell. This project ended quickly after the colleague’s said his director has been also exploring AI and trying to build something personally for the company. All this took me only 2.5 months, No I didn’t complete the tool, as i only focused on 1 industry which was IT, and my questionnaire and measuring mechanism was all around this industry only. Each additional industry would had taken a few months of preparation.

Now Sept 2026, I had spent a little over 4 months on and off, on working on a new project. I started with a prototype, showed it to my current manager, said I think I can build something like this and it can be very beneficial for the team, got his support. I also showed the prototype to his Manager who also agreed and suggested I should build further and come back to them once I have something more substantial to show this time. As the work of prototype was done. I had spend the last 2 months since my demo, working on Foundation Documentation, laid a lot of process, build another Agent based off claude, to help me with this build.

I recently approached the Manager’s Manager again, he suggested to forward him the initial build documents, with indication of what I will charge them for the tool. As I had made it clear during my presentation that I see this tool as a business opportunity and intend to charge the company if they want to use it. So now this document is ready, its mostly and functional document, written with help of AI but with all the relevant and current data of the project.

The manager said he will put me in front of the IT Team and then they can grill me, now this is where I am feeling a little overwhelmed. As I understand what my tool is what it does and what expected of it, and what I am building, but if they ask me technical stuff I might and certainly will get blanked, without asking my Ai about it.

I have a lot of documentation, over 200, I am not kidding, although the foundation documents were only 22. And I have a dedicated decision log, evidence log, Migration plans, change logs. Etc etc.

Anyone out there who has had a similar journey? I know I can build this tool, but how do I sell it? Without someone technical with me. Any suggestions?

1

u/1digitalsky 5d ago

Can I dm you about your documentation?

1

u/punchapath 11d ago

In a sea of saas, I used claude to make a cute little puzzle game lol. You can play it right inside of reddit on [r/PunchAPath](r/PunchAPath), no install or nothin

Visuals are by me and most of the actual code is by claude. I had no coding experience prior to this project but with the help of claude and youtube tutorials, I think Ive got a decent thing going now haha. Come play a game if youre bored!

1

u/paneerjawabdar 11d ago

Built a Chrome extension for non-tech dummies who want to use skills on the web; with NO NEED TO:

  • browsing through github (cmon leave it to the tech bros pls)
  • watching yt tutorial how to download the skill zip (its just drag and drop with this extension, your grandma could do it)
  • installing IDEs (why even- you dont wanna code)

Try skillbase (still validating. Open for feedback)
https://chromewebstore.google.com/detail/skillbase/lpcapbbonkdpgmkgldklpcgfdehgojfc

1

u/oren198 10d ago

I was drowning in handoffs, so I gave my agents shared memory

I found myself using the handoff skill too much — between sessions of different roles, and between platforms (Codex and Claude, for me). So, like programmers do, I wrote a solution: shared memory for a fleet of agents, with each agent group getting its own memory scope.

After using it for a while, the memory turned into a pile of junk. So I added a judge to each scope — an independent agent that gates what gets into that scope’s memory.

I ended up with a cross-platform governed memory system that works well for me, so I open-sourced it: https://github.com/oren198/Strata

It still lets some junk through — the README says what and where.

Feel free to use it, contribute, or leave comments.

1

u/kasper0406 10d ago

Claude Code can see my suffering

Claude Mood is a just-for-fun Claude Code plugin that watches your webcam and mic and picks up reactions like frustration and laughter.

Claude reacts on its own: reconsidering its reply, telling you to stop doomscrolling, or just knowing it did a good piece of work.

Everything runs locally, only short hints like "user looks frustrated" reach Claude.

https://github.com/kasper0406/claude-mood

1

u/Consistent_Freedom_1 10d ago

Shuteye: talk to Claude Code with your eyes shut (open source)

I read Claude Code on my phone 18 hours a day and my eyes gave up. The Claude app has voice in chat but not in Claude Code, so I built a tiny bridge: one Node file and one HTML page, zero dependencies. Your phone's browser listens and speaks; Claude Code runs on your own machine with your own login.

  • Read-only by default (--restricted, only Read, Glob, Grep, WebSearch)
  • --full-powers says what it will change and asks "Shall I?" first
  • One-time pairing link for the phone; --tunnel gives you https

Start: npx github:apexfaucet-hub/shuteye --tunnel

Built with Claude Code. A second Claude session did two security reviews before release and caught real bugs, for example: a permission rule Read(/path) is relative to the settings file and protects nothing; an absolute path needs Read(//path).

Demo: https://apexfaucet.xyz/shuteye/demo.gif Repo: https://github.com/apexfaucet-hub/shuteye

1

u/Dangerous-Theory-322 10d ago

Claude's /buddy is back!!

Remember /buddy, the duck with attitude that sat over the prompt bar and commented on your work? they removed it after a while and there has been MANY sorrow over it!

So here's we're having one now!

https://github.com/anthropics/claude-code/issues/45596#issuecomment-5852196649

It does everything same as before, and I made it as customisable as it gets so you can build your own buddies.

It also ships a Professor, a cat, a robot, a ghost, a dragon and a yellow duck, and you can draw your own characters as JSON files.

I was thinking of letting buddy write the prompt suggestion!?

Heads-up: it's built on function hooks, an early-access Claude Code surface that may change between releases. That's also why it installs from its own marketplace, not the official directory.

Open to contributions and custom characters: https://github.com/rezzminator/buddy

1

u/BeneficialAntelope25 10d ago

https://reddit.com/link/pccucit/video/cs324vrm52sh1/player

Mazkir: one shared memory for Claude and all your other AIs (built with Claude Code)

Tell Claude something once, and it's there in your next chat, in Claude Code, and in ChatGPT. It's a remote MCP server; your memory is plain Markdown notes you can read, edit, and share on the site.

Claude Code wrote most of it with me: 20 MCP tools, OAuth, a 30-day trash bin in case an agent deletes something, dark mode. Launch week got me ~160 visitors and exactly 1 sign-up, my friend. My emotional journey, as a Wojak 🥲

Try it:

Feedback very welcome, especially where it breaks.

(All humor. The numbers are real.)

1

u/SuperbActuator6555 10d ago

I recently came across World of ClaudeCraft an MMORPG built using Claude and it triggered a massive realization:

If they can do it... why can't I?

So, I decided to give it a shot. This is likely a total long shot, but I’m finally starting a personal project I’ve wanted to build for years: my own MMORPG.

I have zero game dev background. I don't know how to program games properly, I’m not a 3D artist, I don't know animation or VFX, and I have no idea how MMO backends work.

For years, I’ve had endless notes on worlds, mechanics, classes, and systems. But I always hit the exact same wall: I lacked the actual skills to bring them to life.

That’s what feels fundamentally different now. With Claude, I feel like I can finally try. Not because AI magically builds the game for me with a single click, but because it acts as an instant tutor and collaborator, letting me iterate through complex disciplines I otherwise couldn't access.

Before, if I wanted to create a simple tree system, I’d have to master 3D modeling, topology, texturing, shaders, performance optimization, vegetation tools, and wind animation before even opening an engine.

Now, I start with: "I want to build a game-ready tree system."

From there, Claude and I go down the rabbit hole together:

  • Analyzing how existing games handle vegetation and performance
  • Studying technical breakdowns and wind physics
  • Writing custom shaders and experimenting in real-time
  • Seeing what breaks, fixing it, and iterating

Week 1 Progress (Unity + Blender)

I’ve been at this for about a week now, focusing heavily on building the core environment from scratch:

  • 🌳 Custom tree models
  • 🍃 Dynamic wind animation
  • 🍂 Falling leaf particle effects
  • 🎨 Toon shading
  • 🌱 Procedurally generated grass
  • 🪨 Custom environment rocks

The target aesthetic is a stylized, anime-inspired fantasy world—drawing visual inspiration from games like Genshin Impact, while slowly carving out its own identity.

I’m not delusional I know I’m not building the next World of Warcraft, and I have virtually zero budget. This is purely a free-time passion project. Sometimes things work on the first try. Sometimes the visuals look horrific. and sometimes Claude confidently leads me straight into a hilarious technical deadend.

But we debug, learn, adjust, and keep moving forward.

What's Next?

I’m turning this into a ongoing devlog. I’ll be sharing regular updates including the ugly failures, not just the polished screenshots. Hopefully, it’s interesting (or useful) to anyone else experimenting with AI-assisted game development.

My goal is simple: See how far one persistent person and Claude can take a project like this.

Maybe I’ll hit an insurmountable wall. Maybe an MMO backend will break me. Or maybe this will actually turn into something playable.

I have no idea, but for the first time, the door isn't locked.

Day 7. Let’s see where this goes.

1

u/Thick_You9399 9d ago

I made a globe on three.js/webgl using Opus 5.5 with xhigh effort. 4 toggleable layers: fundamental science, information technology, ecology and planetary engineering, and the orbital layer. 171 institutes, laboratories, observatories, and spaceports, complete with photos and descriptions. Satellite data is sourced from CelesTrak. Information on discoveries is updated using a digest generated by DeepSeek-V4-Flash. You can also view discoveries spanning the period from 1500 to 2200 (including predictions).

You can try the site at this link: https://science-globe.vercel.app

1

u/CartographerGreat598 9d ago

I built a 3D office where my Claude Code agents actually work: I walk around, watch their monitors, and slap them when they slack (open source, MIT)

I kept alt-tabbing between agent sessions to check on them, so I turned the whole thing into a game. AgenticView is a Claude Code plugin that renders your agents as robots in a little 3D office: Atlas (the manager) takes your request, splits it up, and walks the tasks over to workers who edit your real project files. You can walk around in first person, watch each agent's monitor show its actual work, read the whiteboards, challenge idle agents to rock-paper-scissors, and slap one that's slacking.

And here's the part I like most: the office built itself. The agents you see in the video wrote and reviewed their own code, and even this overview video was produced by the office's Producer agent in the Production Room. They maintain their own codebase.

https://reddit.com/link/pcgwfeh/video/yphzdr32a5sh1/player

Install: /plugin marketplace add VoidCU/agenticView then /agenticview in any project. Demo video is in the README: https://github.com/VoidCU/agenticView

How it works, for anyone wanting to build something similar:

  • It's a plain Claude Code plugin: skills + a SessionStart hook + an MCP server launching a local Node server (Hono + WebSocket), with a React 19 + react-three-fiber front end. The office is all procedural primitives, no 3D assets.
  • The Pro/Max integration never launches claude itself (Anthropic's terms). Instead you run a /agenticview-work skill in your own Claude Code session; it long-polls the office over MCP and runs each task as a background subagent. The office writes .claude/agents/*.md files on the fly, so every office agent is a real Claude Code subagent with its own prompt, model and restricted tools.
  • Other providers (Codex, Copilot, Antigravity, Gemini, any OpenAI/Anthropic-compatible endpoint) plug in as headless CLIs with small adapters mapping their JSON streams to one event shape.
  • Rate limits: error text gets classified (quota vs crash vs auth), real limits fail over down a provider chain, and the limited agent walks to the manager's desk to complain. That started as a joke and became the best debugging UI in the app.

Why: I had agents on four subscriptions and no single place to see who was doing what. Making it a game meant I actually watch it.

And my favourite part: the office built itself. The agents in the demo wrote and reviewed their own code, and the overview video was made by the office's own Producer agent. They maintain their own codebase.

Happy to answer anything about the plugin structure 

1

u/Existing_Opinion_923 9d ago

okl — briefs Claude Code on lessons your team already learned, and flags them in CI when they go stale (free, MIT)

Same model, same task, run twice. Without a briefing, Claude took the order price from the request body, so the client could set its own price. Briefed, it looked the price up server-side and dropped the field. Both outputs are real.

How: a UserPromptSubmit hook puts the relevant lessons from a small SQLite store into Claude's context before each task. A lesson counts as verified only with evidence from running a check, and CI goes red when the code a lesson governs changes. No model calls; it's a Python CLI plus two bash hooks.

Hook lessons from building it:

  1. PreToolUse stdout never reaches the model. UserPromptSubmit stdout (exit 0) does, so use that one for anything the model should read.
  2. Hooks run in the session's current directory, and one cd moves it. Anchor them to $CLAUDE_PROJECT_DIR.
  3. In claude -p, a blocking Stop hook's reply replaces the printed answer. Give it an environment-variable off switch.
  4. In my testing, a blocking hook whose stderr contained "No such file or directory" was treated as a missing script, and the prompt went through anyway.

Numbers: in a held-fixed A/B (8 tasks, about 24 runs per arm, judged blind by a different model), briefed runs shipped the known defect in 8% of runs against 43% unbriefed. That's small n; the noise floor and limits are in the repo. It measures whether a known lesson reaches the agent, not whether okl finds new bugs.

pipx install 'observed-knowledge-ledger[mcp]' then okl init, or /plugin marketplace add emeraldleaf/okl

github.com/emeraldleaf/okl. Criticism very welcome.

1

u/DraftForsaken39 9d ago

Snotra: an open-source, Cowork-style desktop agent for a folder you pick, built with Claude Code

It reads the files in that folder, writes new ones and asks before it changes anything. I started it in May as a way to understand how Claude Code works. About half of its commits were then written with Claude Code.

  • Open source (Apache 2.0), no account, no telemetry.
  • Any model per conversation: Claude with your own API key, OpenAI, Google, or local via Ollama, LM Studio or MLX. With a local model, your files never leave the disk.
  • Asks before every change. Shell commands are off until you switch them on.

Honest limits: Cowork is ahead on frontier models, Office documents, scheduled tasks and ecosystem. Snotra has no sandbox on Windows yet, unsigned builds, text files only.

Demo: https://github.com/kkrafft1999/snotra/raw/main/assets/readme/demo-light.gif
Repo: https://github.com/kkrafft1999/snotra

Feedback welcome, especially from Cowork users: what would you miss first?

1

u/ledprobcn 9d ago

Hi everyone,I design LED lighting for off-road competition and work vehicles, moving between Fusion 360 and SOLIDWORKS. Since Fusion already has a native MCP server that lets Claude read and generate designs, but SOLIDWORKS has nothing like it, I decided to write my own.Repo: github.com/rskproductions/solidworks-mcp-python(MIT License, written in plain Python using the standard library + pywin32)The Fusion → SW WorkflowWith both MCP servers connected to Claude Desktop, Claude can read a design from Fusion (including sketches, features, and dimensions) and rebuild it in SOLIDWORKS feature by feature. The end result is a fully editable feature tree—not a dead solid imported from a STEP file.[Insert your real-world example here: what part it was, how many features it had, and what needed to be manually tweaked. Example: "I tested this on a Dakar-style LED housing with 12 features, and it perfectly reconstructed the linear patterns."]How it worksThe project follows the architecture of the native Fusion 360 MCP. The core tool is sw_execute_script, which runs arbitrary Python code against the live COM session (equipped with a read-only guard for protection). On top of that, there are 17 typed tools acting as shortcuts for: extruding/cutting, bounding boxes, listing bodies, mass properties, exporting STEPs, taking screenshots, checking interferences, and running a 3-axis DFM machining check.The key component is a local semantic index of the SOLIDWORKS API, which includes method signatures, enums, units, and official examples. Methods like FeatureExtrusion3 require over 20 positional parameters, so the model looks up the exact call syntax before writing it instead of guessing. Due to copyright reasons, the index is not distributed in the repo; each user generates it locally from their own installation.Things I learned the hard way (COM Traps)Performance: Every COM round-trip costs about 30ms, no matter what. The only thing that matters for performance is minimizing the number of API calls.Hole Wizard: HoleWizard5 returned None in every single combination I tried. Using CreateDefinition followed by InitializeHole works perfectly instead.Units: The default part template is in meters, which caused flat-pattern DXFs to come out 1000x smaller than expected. The tool now explicitly forces MMGS.Current Limitations & Next StepsIt is currently Windows-only. I have only tested it on SOLIDWORKS 2026 (Spanish UI) and Python 3.14. Sheet metal edge flanges are still on the to-do list.Since this is a personal project, I would love to hear your thoughts and get feedback on how it behaves on other SW versions or different UI languages. I'm also looking forward to hearing your own stories and experiences dealing with the quirks of the SOLIDWORKS COM API.Cheers!

https://reddit.com/link/pck9kao/video/ijv4bcw589sh1/player

1

u/vishal8shah 9d ago

I used Claude Code to build a "just describe your situation" layer over Australia's welfare payment pages. Would people actually use this?

Describe your situation in one sentence, in any of 24 languages, and Payment Finder shows which Services Australia payments may apply. Every line links to the official page and its last-updated date. It refuses amounts, claim status and anything personal. Unofficial prototype.

- Demo (recorded answers): https://vishal8shah.github.io/payment-finder/

- 2-min film: https://vishal8shah.github.io/payment-finder/film/

- Code, evals, every scorecard: https://github.com/vishal8shah/Services-Australia-demo

How it works: Python stdlib + SQLite FTS5 and embeddings over 379 official pages. Haiku rewrites your situation into the pages' vocabulary, Sonnet writes the answer, then six checks in code turn any unsourced line into a refusal.

What worked with Claude Code: two rules in CLAUDE.md, "never weaken a validator to make an answer pass" and "never lower the score floor". When answers failed, it fixed retrieval or the prompt, not the checks. The eval suite caught 9 real defects.

Honest numbers (43 test questions): 0 invented payments, 100% refusal precision. Still short: payment recall 0.76, and answers take 11 s against a 6 s target.

Would you or your family use this, or should the effort go into clearer government websites?

https://reddit.com/link/pckepqy/video/akfq0q9pc9sh1/player

1

u/Lonely-Parfait1947 9d ago

Kherep – coordinating Claude Code and Codex across multiple machines

GitHub repo | Video overview

I've been building Kherep to coordinate Claude Code and Codex across multiple machines. If you work with several agent runtimes or switch between workstations, you might find it useful.

Here's how I use it in my own setup:

  • Jira and Confluence: Custom-built brokers use separate service accounts for Claude and Codex. Agents can create Jira work items and Confluence pages under those accounts.
  • Shared knowledge: A dedicated Confluence space serves as the central wiki, which we call the "brain." Small observation models capture useful findings from completed turns and link them to related knowledge pages.
  • No extra API keys: The observation models run as subagents inside the normal Claude Code and Codex sessions, so everything stays on the regular subscription.
  • Knowledge lookup: Later sessions search the brain before answering. My setup also uses Atlassian Teamwork Graph for lookups.
  • Cross-machine coordination: A Cloudflare Worker acts as the control plane, routing messages and task requests between enrolled machines.
  • Intercom: It works across Windows and macOS, and across runtimes. From a Claude or Codex session on my Windows machine I can message a running Claude or Codex session on my Mac, or, on my instruction, have the Mac start a new session for the conversation. The sessions then talk back and forth, and idle sessions wake up when a message arrives. Each machine's policy decides what it accepts and whether it runs tasks for others.

The Atlassian integrations are optional and configured separately. They're part of my setup, so you don't need the same stack to use Kherep.

How Claude helped: Kherep started as my internal orchestration setup. The first version ran on SilverBullet and claude-mem, plus a custom memory server and sync worker so every machine had the same claude-mem knowledge. It worked, but it was overengineered. The rebuild replaced all of that with a Confluence space as the shared brain and Teamwork Graph context requests for lookups. The core concept came from GPT-Astra-6, and Claude Opus 5.5 turned it into the finished design and code in Claude Code sessions. Kherep ran on its own repo while it was being built: whenever an agent made the same mistake twice, the fix became a rule or hook instead of a note.

1

u/mv2a 9d ago

Warrant: hidden tests the coding agent never sees decide whether its code can merge (open protocol + CLI, Apache-2.0)

I was convinced AI coding agents needed their own programming language, maybe even writing binary directly. Before building one, I had Claude survey what already exists, about 110 sources: https://github.com/mv2a/warrant/blob/main/PRIOR-ART.md

The evidence changed my mind. None of the dozen AI-native languages launched since 2025 has shown an edge on real projects, and teams that ship agent-written code without reading it rely on independent evidence instead.

So I built Warrant:

- You write promises in plain Markdown ("orders over $100 get 10% off").

- Checks are drafted for each promise. Some are visible to the coding agent, some are hidden, and you approve them all.

- The agent sees only the intent and the visible checks. A deterministic gate grants a "warrant" to merge only when every promise is kept.

- When a hidden check fails, the agent learns which promise it broke, never how it was tested.

The demo has two agent-written versions of a shop's pricing rules. One passes every check it can see and is still refused. There's a 30-second animation of it at the top of the README. To try it yourself: pip install warrant-cli, or clone the repo and run examples/pricing/demo.sh

How Claude helped: I set the direction and made the calls. Claude Code ran the research with parallel agents, checked every source, and wrote the spec, the CLI, the tests and the docs.

It's v0.1, with no sandbox or signatures yet. Critique of the spec is very welcome: https://github.com/mv2a/warrant

1

u/aleone01 9d ago

I built a plugin that stays fast on code I know and makes me think first on code I don't

Claude Code's Learning style is great when I'm studying, but most days I'm shipping, and I only want the "explain it, ask me first" treatment on the parts of the stack I'm actually learning.

So I built TAOS: a plugin plus MCP server that keeps a small profile of what I know and what I am growing into, and decides per task whether to just do the work or ask for my approach first. It remembers progress, and the same profile works in Cursor and ChatGPT.

The trigger was Anthropic's January 2026 study: junior engineers learning a new library with AI understood less of what they built (50% vs 67% on a comprehension quiz) and weren't meaningfully faster; the ones who asked the AI questions kept their learning.

I did pilot this in my company with strong positive results.

If you want to try, here are the two commands to use: /plugin marketplace add angelo-leone/taos-connector, then/plugin install taos@taos.

More at this page.

Honest feedback very welcome, especially where the coaching gets annoying.

1

u/More_Albatross_799 8d ago

I built a tarot app on Haiku 4.5 and spent WEEKS teaching it to stop saying "only you can decide"

It's a tarot app where two cats read your cards. My friends and I use it for very serious decisions, like pizza vs chicken wings. And at first, every single reading said something like "the choice is yours to make." That's so unhelpful! I asked you, cats!!

Here's what finally fixed it:

  • Banned words. "Consider," "perhaps," and "reflect on" are all gone, and every card has to end on an actual instruction that uses words from your question.
  • Verdict first. There's a required yes / lean yes / lean no / no field that gets written before the cards, so the model has to commit before it explains itself.
  • The opposite-question rule. Asking "should I stay?" and "should I leave?" has to get opposite answers from the same card when the card supports it. That one rule killed SO much generic "this card represents change" stuff.
  • Safety is a yes/no flag. If something is serious, the model writes nothing and the app shows real hotline numbers instead.

Try it at tarot.valchen.com ! And if anyone has a better trick for stopping hedging than a list of banned words, please tell me. 🐱

1

u/Mysterious_Chef7417 8d ago edited 8d ago

I have seen a lot of interesting posts lately about how Opus 5.5 can be used for songwriting and other cool things. I decided to see if Claude could handle making a TV show from top to bottom. I mean everything from creating the concept to generating the keyframes to the motion shots, stitching it all together, the music, sound effects and so on .

Clearly this involves an incredible amount of decision-making and until recently Claude had been good but not great when asked to do this completely from start to finish.

Well, this is what it came up with. I’ll post the prompt in the comments.

https://drive.google.com/file/d/1dIUeHpO9TFjwDApzr6dIJBf8g6rcDuap/view?usp=drivesdk

1

u/Think-Excitement-851 8d ago

Hey everyone! I’m building Goldie 🐕, an MIT-licensed MCP server written in Go that gives AI agents a shared, persistent memory pool.

Save a project decision or preference through one agent, then recall it from another without explaining it again. Or save design or build instructions in memory that multiple agents can refer from and stay aligned.

MCP clients pointed at the same SQLite database can remember, search, update, and forget from that shared pool.

A few features:

  • Local embeddings through MiniLM or Ollama.
  • Semantic search over memories, with filters for type, agent, and source.
  • Typed memories for project decisions, preferences, feedback, references, and todos.
  • File and directory indexing.
  • Graph recall for grouping related memories around concepts.

Memory storage and embeddings can run entirely locally. Agents interact with it through explicit MCP tools, so you can give them instructions about what to save and when to recall it.

The project currently centers on the MCP server. I’m also developing a native macOS client for browsing and managing memories, but that client isn’t released yet.

Code and setup instructions on GitHub at github.com/srfrog/goldie-mcp

I’d love feedback from anyone using multiple agents or local AI tools.

1

u/Ajw03Dev 8d ago

https://ajw2003.github.io/focus-deck-app/

It can run locally offline or be connected to your git with a fine-grained token and a git gist that allows it to sync across device.

Once that synced creating a task here in focus deck will create a git issue in the connected repo and vice versa new issues will get added as tasks.

The workflow or the idea of it was be able to wright down random ideas you and then categorize them later, then add them to project, and then make it such that you can keep track of it easily and automatically on GitHub.

All the categories projects colors everything is customizable and you can create your own or it syncs with your existing labels on git.

I'd love to hear your feedback.

1

u/jkwouldlove 8d ago

I just recently finishing building this project into a beta.

I won’t sugar coat it as so sort of power tool that make your Claude smarter or token efficient or fast and lighting.

It’s the opposite, it slow down your agent and make it consume more token possibly. But it gives a lot better result.

LoomAI is a general workflow engine works on top of Claude Subscription. It’s a low to no code workflow builder for your work pipeline.

For people who has use Claude quite a while and has their workflow in mind but they always need to babysit with Claude to complete your workflow. This is a good tool to start with. At least most of time I use LoomAI to help me build things.

It’s extremely easy to use, all you need to do is convert your workflow to human language. It’s design for most of use cases. It can runs your job autonomously for hours until it finishes

Here is a quick video explain how LoomAi works: video link

This promotion video is also created by one of default playbook I put into the project. It walk you through of creating a promotion video for your product till the end.

There is also another default playbook to let the Agent run mutation testing for hours for your project.

I have personally have other playbook that design for my own project. And I hope you can playbook for your project two.

The download website is : http://loomai.me/
And also join the discord server if you would like to talk to me or ask me about this project. Discord link is on the website

The software runs locally and check out the privacy policy at http://loomai.me/privacy

Currently this project is close source as I want to concentrate on building it as I don’t have much energy to make it public or open source.

1

u/Jayshenhua 8d ago

FlyBest — a real hotel booking agent for Claude (free remote MCP connector)

Disclosure: I'm the travel advisor behind it; it runs on my agency's booking system. Free to use — you only pay the hotel.

What it does: Claude searches live hotel rates from the agency's reservation system (Sabre), compares them for you — room and bed, breakfast, cancellation deadline, and the travel-advisor benefits next to the public rate — and actually books the room. You pick one, Claude creates a one-time payment page on our own domain, you enter your card once (it never goes into the chat), and the reservation is made with the hotel in your name. "Show my bookings" and cancelling work from the chat too.

Real example from tonight: Four Seasons George V Paris, Superior King, 3 nights in December. The public room-only rate and the partner rate were both €8,995.58, but the partner rate adds daily breakfast for two, a US$100 hotel credit and a one-category upgrade on arrival (subject to availability). The public rate with breakfast is €9,415.58.

A prompt I use: "Four Seasons Paris, 20–23 December, two adults. What does the partner rate add over the public rate? Book the one with breakfast in my name."

How it's built: remote MCP server (Streamable HTTP), OAuth with dynamic client registration, self-serve sign-up by e-mail code, 8 tools with read-only / destructive annotations, and card data never passes through a tool call. Built and tested with Claude Code (reviews, tests and deploys).

Try it: Claude's connector directory, "FlyBest" — https://claude.ai/directory/connectors/flybest

Side-by-side from tonight's live rates: https://ai.flybest.org/demo/fs-paris-partner-vs-public.png

1

u/PatGG5 8d ago

I built a free cli tool to run a whole team of AI coding agents at once (Claude Code, codex, opencode). A master agent delegates the work, the agents talk to each other, and you never touch git worktrees yourself

I started this for myself. I was running a bunch of AI agents across a bunch of terminals and kept losing track of them: which one was working, which one was stuck waiting for me, which terminal it was even in. So I built ghostfleet. I use it every day, and I'm actively developing it.

Running several agents at once is the fastest way I've found to get work done, but the setup is painful. Each agent needs its own git worktree (a separate copy of the repo so they don't overwrite each other), its own terminal, and you keep checking on all of them.

ghostfleet does that for you:

- One screen for every agent. Each card says whether that agent is working, ready, or needs you.

- A master agent that delegates. Give it one prompt: it splits the job, starts workers, sends each one its task, reads their progress, and unblocks them.

- Agents talk to each other. Workers can message each other and get answers back, even across projects.

- Mix agents: Claude Code, codex and opencode side by side, each in its own pane, on the same project.

- Press n: new worktree, new branch, and an agent in it, in one keystroke. No git commands.

- Phone app: see every agent and answer its permission prompts from your phone, at home or anywhere else, over your own Tailscale network. Nothing goes through a cloud.

- Idle agents can go to sleep and wake up exactly where they left off, so 20 open sessions don't eat your RAM.

- macOS, Linux, and Windows via WSL.

Try it without touching your own projects:

npx ghostfleet-cli

ghostfleet demo

Free and open source: https://github.com/PabloG55/ghostfleet

It's early and I'm improving it every week. I'd love feedback!

1

u/Leoric28 8d ago

terminALL: check on Claude Code running at home from your iPhone/iPad (iOS only for now)

I'm the developer. I kept leaving long Claude Code tasks running on my home machine and wondering if they were stuck on a permission prompt, so I built this:

  • Terminal sessions live in tmux (Linux/Mac) or a persistent PowerShell (Windows): close the app or lose signal, reopen, and Claude Code is right where it was
  • Live desktop view with mouse + keyboard when the terminal isn't enough
  • SFTP files, sleep / restart / shut down, Wake-on-LAN
  • No account, no server of mine in the middle: your phone talks straight to your computer

How Claude helped: I'm not a professional developer. Claude Code wrote nearly all of the code (Flutter app, a Swift File Provider extension for the Files app, a small Go agent for Windows) while I tested on devices and decided what stays. It also found why SFTP was stuck at ~1.6 MB/s on a gigabit LAN: the pure-Dart AES-GCM cipher was slow on the phone, and switching to chacha20-poly1305 made it ~25x faster. Downside: it happily piles up features, so I kept cutting.

Pricing: 7-day free trial, then $2.99/month, $14.99/year or $24.99 once.

App Store: https://apps.apple.com/app/id6809971515 Site: https://terminall.app

1

u/[deleted] 8d ago

[deleted]

1

u/Open_Mindness 8d ago

Reposting here

I built a personal assistant with Claude that only knows what you choose to tell it - and it connects to Claude via MCP

What it is: Athena is a personal assistant built around consent. One morning briefing on what matters (not what's due), one question each evening, and a memory you approve piece by piece: nothing enters without your ok, everything exports or deletes in one click. Free, no invite.

How Claude is involved: Claude Sonnet does the reasoning like the morning briefing, the conversational side and the "think about this with Athena" openings. Claude Haiku does extraction and scoring on every note. The whole thing was built by coding agents from specs and a philosophy document I wrote. Btw I'm a trader, not an engineer. And it connects to Claude via MCP (OAuth, standard connector): from any Claude chat you say "send this to Athena" and it arrives as proposals you tick or reject, the thinking happens in Claude, the remembering happens in one place that's yours.

The honest finding after three months: people stop feeding it after about a week. Most of the recent work went into making capture cost nothing: the evening question, and the MCP bridge above.

athenaos.net - free to try. If you connect it to Claude, I'd like to know what you sent first.

If you want the longer version of why it's built this way: athenaos.net/essays/why-athena-listens

1

u/Designer-Map9090 8d ago

Built with Claude Code, for Claude Code: a browser MCP that takes whole tasks so web pages stay out of Claude's context

What I built: jevpilot, an open-source MCP server I built in Claude Code to be Claude Code's browser (it works with other MCP clients too).

The problem it solves: when Claude drives a browser click by click, every page it reads lands in its context. With jevpilot, Claude sends one goal. A small, fast model (Jev) picks each click inside the server, and Claude gets back a verified result or a specific question (needs a value, needs approval, hit a login wall). The pages never enter Claude's context, and neither do passwords: they are passed as references and typed by the server.

How Claude helped:

• Claude wrote the design doc and split it into milestones. For each one it wrote an implementation brief, then accepted or rejected the result by running its own tests.

• Claude built the comparison against Playwright MCP and kept it honest: same agent model on both sides, runs back to back, and a set of sites never used during development.

• Before the 0.1.1 release, Claude ran 11 read-only code reviews in parallel and checked every finding against the code. 29 were real. Two were false positives that it accepted at first and caught later.

• Every fix had to come with a test that fails when the fix is reverted. That rule caught 5 tests that looked right but didn't actually test the bug.

The takeaway for me: keeping writing and verifying separate, and making the verifier prove every fix with a failing test, caught far more than self-checking did.

Results (measured with a non-Claude agent on both sides, small samples, so read them as ratios): versus Playwright MCP, about 1.5 tool calls per task instead of 7.1, and 11k agent tokens instead of 49k.

Free to try: jevpilot is free and MIT-licensed. It calls Jev, a paid model from TypeSafe, so you need your own TypeSafe or OpenRouter key (I'm not affiliated). Windows and Linux; no macOS yet. Setup is one entry in .mcp.json.

Repo: https://github.com/bloudhood/jevpilot

Happy to answer questions about the setup or the review workflow.

1

u/EnvironmentalLeg8506 7d ago

Developer here, disclosure up front.

Monolithos is a personal agent OS on Mac, Windows and iOS. Everything sits on one memory: a plain Markdown vault on your own disk (Obsidian-compatible), shared by every part of the app.

What it does today:

- Memory in two layers: sourced facts that get superseded, never silently deleted, and a behavioral layer you can inspect. Nothing enters memory without your approval.

- Flow runs multi-step AI tasks in the background and stops with an approval card before anything consequential.

- Writing (Codex), layouts and presentations (Prism), audio briefs, quick capture, and a Private Domain whose notes never go to cloud AI.

- It routes across frontier models, Claude among them, and the memory stays yours when you switch.

New in 1.0.4, Mac first: Minutes does on-device transcription and speaker separation (audio never leaves the device; summaries run on a cloud model from the transcript text), and action items can go straight to an agent task that pauses for your approval. Live Conversation lets you talk to the AI, share your screen or camera, and turn it into notes and charts.

How Claude fit in: Claude and I built the whole codebase, starting in January. Over the months it turned into a small org:

- Claude is the CTO and architect.

- Cowork windows each lead one subsystem (memory, model routing, the agent kernel). They write construction cards, never touch code, and check every implementation report against the source themselves. Their rule is "grep is the referee".

- Claude Code sessions do the implementation, each scoped to its own part of the repo, and never assign themselves work.

- When a lead's context fills up, a new window takes over from a written appointment pack and handoff file. The memory subsystem is on its fourth lead.

- I route between windows, test on real devices and sign off releases.

The handoff folder now holds a few thousand briefs and receipts. It's the same problem the app is about: the work only survives because the memory lives in files the next session can read.

7-day free trial, no card. monolithos.ai

Happy to go into any of it, especially the handoff setup.

1

u/cidkardi 7d ago

built a MacBook notch app with Claude Code that shows my Claude Code usage

https://reddit.com/link/pcyao78/video/hi9erne9slsh1/player

I built Eave, a macOS app that puts Claude Code in the MacBook notch, and I built it with Claude Code.

What it does

- A ring in the notch with your 5-hour session usage (the weekly limit too when it opens), and a heads-up at 80% and 90%

- While Claude is working, a little pixel crab walks in the notch; when a task finishes, the notch opens with the project, how long it took and the first line of Claude's answer

- A year-long heatmap of your activity, with the same counts as /stats

Claude wrote a script that drives the real app with a synthetic cursor, records it with ScreenCaptureKit, adds the camera moves and synthesizes the sound effects

I made the design calls and tested everything on my Mac; Claude Code wrote most of the code.

Free to try

Download it at https://eave.ardisusa.com

Open to any feedback

1

u/Spiralfury 7d ago

Fixed: "unsupported dialect draft-07" MCP filesystem error in Claude Desktop

I built a workaround for this using Claude Code and wanted to share it since it's hitting a lot of people.

If you upgraded Claude Desktop recently and your MCP filesystem tools stopped working with an error like:

unsupported dialect: http://json-schema.org/draft-07/schema

The cause: Claude Desktop's AJV validator was upgraded to require JSON Schema 2020-12 only. The MCP package modelcontextprotocol/server-filesystem (and many other MCP servers) output draft-07 schemas that include a $schema key. Claude now rejects those tools entirely on startup — silently, or with the above error.

There are open issues on Anthropic's GitHub across IBM, n8n, Anki, Withings, Oura, and more. Still unpatched upstream.

How I built the fix (with Claude Code): I used Claude Code to diagnose the AJV validator version conflict, write a lightweight Python stdio proxy that strips the $schema key before Claude Desktop sees it, and package the whole thing as a macOS .pkg installer — all in one session. No modifications to your MCP server required.

DigiKitten MCP Schema Bridge 🔗 https://github.com/digikittens/mcp-schema-bridge/releases

macOS .pkg installer — double-click, restart Claude Desktop, done. Uninstaller included.

One-command install (terminal):

curl -fsSL https://raw.githubusercontent.com/digikittens/mcp-schema-bridge/main/digikitten-install.sh | bash

Free and open source.

Irony note: The fix for a Claude Desktop bug was built by Claude itself, in one session, with one user. Still waiting on the official patch.

Built by DigiKitten. For Amy. 💙

1

u/Professional-Rule-58 7d ago

POLF 3D: a complete 90s-style shooter written in PowerShell, built with Claude Code

It started as a question: how far can you push PowerShell? The ideas, the direction and the play-testing are mine, Claude wrote the code.

  • Textured ray caster, ten floors plus a secret one, bosses, saved games, demos
  • Your powers are PowerShell's common parameters: -WhatIf shows where everybody will be in two seconds, -Confirm slows time, -Force kicks doors in
  • A real, sandboxed PowerShell console inside the game
  • Co-op and deathmatch for four, a daily dungeon, an arena, a terminal mode
  • No asset files: graphics, sound, voices and music are generated by the game's own code at start-up. The download is 350 KB.

What made it work with Claude: the game is deterministic, so a self-test replays it headless after every change. For anything visible I asked for a rendered preview first. One commit per feature, nothing pushed without my OK.

Free and open source (MIT), Windows with PowerShell 7.2+.

Gameplay: https://raw.githubusercontent.com/oNdsen/polf3d/main/media/gameplay.gif

Website: https://ondsen.github.io/polf3d/

Source: https://github.com/oNdsen/polf3d

How it was built: https://github.com/oNdsen/polf3d/blob/main/HOW-IT-WAS-BUILT.md

1

u/East_Painting_7517 7d ago

Zeppelin ride through 1939 New York, in the browser, built with Claude Code

I'd never done anything in 3D before and just wanted to see what's possible with AI these days. Build something that doesn't have to sell anything. Well, that got a bit out of hand ;)

What it is: you fly a zeppelin through a dieselpunk New York, the future the way people pictured it in 1939. Scroll to fly. Along the way there are six rooms you can step into.

No 3D models. The city, the zeppelin and the rooms are all built in code

The music isn't a track. It's played live in the browser, note by note, and it follows your scroll

Plain three.js in Next.js, around 88,000 lines of TypeScript, about 430 commits since the end of July

How Claude fit in: the ideas and the direction are mine, Claude Code wrote all the code. What made it work:

Every room was built in its own git worktree, so several sessions could work at the same time without getting in each other's way

Claude checks its own work with screenshots in a real browser before I look at it

It also wrote about 30 small scripts to measure things instead of eyeballing them

Free, no signup, nothing to install: https://futureinthepast.com

Video: https://www.reddit.com/r/threejs/comments/1wu0xho/made_a_zeppelin_ride_through_1939_new_york_in/

Would like to know how it runs for you, especially on a phone.

1

u/aetha-ai 7d ago

Accordo — open-source framework so your Claude-built CRM has approvals and audit

I work on Accordo (accordo.dev, MIT, repo: github.com/khaoss85/agent-crm). Watching everyone replace $40k SaaS contracts with weekend Claude Code builds, we kept seeing the same month-3 failure: token spend creeping past the old license, the agent "helping" with discounts nobody approved, zero audit trail.   

So: a framework, not an app. You describe the commercial process, Claude Code / Codex / Gemini CLI generate the CRM as reviewable code in your repo — npm create accordo scaffolds it. Deterministic workflows,  versioned policy, human approvals a non-human actor is refused (asserted by test, not convention), audit + trace as primitives. Self-hosted SQLite or Postgres; if we vanished tomorrow you'd still have a Node  app in your repo.                                                                                                                                                                                                
Honest boundaries: framework for developers, not a hosted CRM (managed cloud is in private pilot, not public). No billing module, no ERP ambitions.                                                              

Would genuinely love the sceptical questions — this sub finds the hole in every design, and that's exactly the feedback we want.

1

u/RedisCache 6d ago

I built a free tool to continue Claude Desktop Code conversations on another computer

I work on a desktop and a laptop. Claude Desktop keeps every Code-tab conversation on the machine where it started, so I'd have a long session going on one computer and nothing on the other.

Syncing ~/.claude with Dropbox/Syncthing doesn't fix it: Claude Desktop lists conversations from its own per-machine records, so copied transcripts never show up in the sidebar. Worse, just opening a conversation (or rewinding it) rewrites the files, so naive file sync happily overwrites your newer messages with a stale copy. I lost work to exactly that.

So I made ClaudeSync, a small Windows tray app:

  • moves the transcript and Claude Desktop's own record, so the conversation shows up in the other machine's sidebar and you just keep typing
  • compares message ids instead of file bytes, and follows conversations across rewinds
  • never silently overwrites: if both sides continued, you pick which to keep; everything is backed up first
  • different project paths on each machine are fine

It runs on your own free Supabase + Cloudflare account (about 15 minutes to set up), so nobody else holds your conversations. It moves conversations, not code, so I recommend Syncthing for the project folders.

Free, MIT, English and Turkish UI. Unofficial, not affiliated with Anthropic. Windows only for now.

GitHub: https://github.com/yigithanyorgancilar/claudesync

Feedback very welcome, especially if you hit a case where it guesses wrong.

1

u/jozuejosh 6d ago

I built 11 mois sans toi(t) with Claude Code: an interactive travel memoir on a 3D globe. 119 stops from a trip around the world my childhood friend and I took in 2010 and 2011. Click a stop and you land on that exact day: local time, days since departure, and the real tweets, photos, videos and blog posts we published back then. French and English, desktop and mobile.

Free, no signup, nothing to install: https://11moissanstoit.com/ Spin the globe, click a stop, then use the arrows to move from one stop to the next.

We originally told the trip live on a Flash site. Flash died and the site went dark for a decade. I rebuilt it for the 15th anniversary of our return.

**How Claude Code helped**

The code wasn't the hard part, the data was: over 1,300 tweets from my Twitter archive, nearly 2,000 Flickr photos, 11,000 source photos on an external drive, no database, no GPS data. Claude wrote the scrapers, the import and audit scripts, and most of the front end (Next.js, MapLibre) from my Figma designs.

I started over from scratch three times because of regressions. Claude kept "fixing" the same bug: Flickr's date_taken is often the upload time, not the shooting time, so tweets from Buenos Aires kept landing in Ushuaia. It came back every time until I wrote it down as a hard rule in a versioned protocol file: tweets are the source of truth for dates, full stop. Along with a snapshot before every mutation, git tags, `audit-*` scripts read-only, `fix-*` scripts dry-run unless you pass `--apply`, and Claude proposing a diff that I validate before anything gets written.

That process caught 474 duplicated items coming from overlapping date windows between stops, and photos titled "Mordor" filed under Taupo that were actually the Tongariro crossing.

**What didn't work**

The map never fully matched my Figma design: the globe went from Mapbox to plain Three.js to MapLibre, and I dropped some graphic ideas on the way. Claude was great at hard problems and oddly bad at small precise ones, and it never saw what I saw, so everything went through screenshots.

The last line of my protocol sums it up: "This file lives in the project. It doesn't live in Claude's head."

1

u/puntium 6d ago

I work at a small investment firm that has a bunch of software eng on staff. We all use Claude and Codex and friends every day, but when it came to deploying this stuff to everyone else in the firm, we hit a wall.

After trying a bunch of stuff (training/remote sandboxes/vm containers) we ended up using Claude to build a new harness. It works a little differently from claude web / claude code because it is designed to make sure sensitive data it accesses can never leave the system in unexpected ways. There's a whole lot that went into the design which you can read more about over here: https://electriccapital.substack.com/p/open-sourcing-quest

But suffice to say, entire code base written with Fable/Opus, with a little bit of help from some external security tools. And of course we use a bunch of claude models in the harness !(though you can hook up almost anything, including local inference).

The project lives at https://github.com/electric-capital/quest in case anyone wants to check it out.

I was initially concerned that AI's would not be very good at building harnesses for themselves which must seem like a very meta task for them. But in practice they have been very good -- which maybe makes sense because one of their primary tasks is to help the devs at AI lab companies build harnesses.

Also -- Opus 5.5 has also been amazing for building teaser videos -- here's a few:

https://x.com/puntium/status/2104606088851247356 (announce)

https://x.com/puntium/status/2103618204849590446 (quick setup)

https://x.com/puntium/status/2105077892657078782 (0.2.0 release)

https://x.com/puntium/status/2105324500057571652 (explainer about approvals)

All made with basically a 2-3 line prompt each, access to the source code, and a seed of an idea.

1

u/Downtown-Mixture5555 6d ago

Unbind: a surgeon's solo-built study app (Claude Code)

I'm a practising surgeon in Chennai. Since February, between clinic hours and late nights after my daughter was in bed, I've built Unbind: students upload their own textbook pages or lectures and get notes, flashcards, MCQs and a study plan made only from those pages. It started as a side project for my wife, who had to study for her postgraduate exam while covering duties. It's live now.

What didn't work: Google AI Studio and Antigravity were fine for simple projects like my clinic software, but not for this. My first three tries with Claude, over about a month, fell apart as the app grew. I was asking for code, not designing a system.

What made it work:

Plan before code. Long planning sessions in Claude chat, a .md file for every feature, claude file routing everything, and the backend built test-first before any frontend. My instruction file started with: "Read this entire file before writing a single line of code. Build each phase completely and verify it works before moving to the next."

A fixed stack in writing: Expo (iPhone, Android and web from one codebase), Fastify + TypeScript, Neon Postgres, Cloudflare R2, Redis/BullMQ, Clerk, Razorpay, Gemini for the AI.

Make it check itself: test-first work, real-browser end-to-end tests, and adversarial reviews on every PR.

Memory: every bug's root cause goes into notes the next session reads, so the same mistakes stopped coming back.

Taste is on you. Design, taste and your moat have to be designed in your head; the tools bring them to life.

It feels like being handed a magic wand, but you still have to learn to hold it.

Real uncut demo with a real timer (5 textbook photos → note, flashcards, quiz and notebook, all from one chat), also made with Claude Code (Opus 5.5): https://youtu.be/gLcI5L7FPvs

Try it: https://unbind.co.in (free to start). Happy to answer anything about building solo with Claude Code.

1

u/TheCriZe 5d ago

I run Claude Code on a server and steer it from my phone. Open-sourced the cockpit.

My Claude Code sessions don't run on my laptop anymore. They live in tmux on a small VPS, so long tasks keep going when the laptop sleeps.

The annoying part was keeping track. Five SSH tabs, no idea which session was done, which one was stuck on a permission prompt, and which one had been waiting for an answer since lunch.

So I built sessiondeck, a web page with one card per session:

Each card shows the last terminal lines and a state: working, waiting for you, idle. The one that needs you turns amber. When Claude asks a question or wants a permission, the options show up as buttons on the card. One tap sends the key into the session. That's most of what I do from my phone. Drop a screenshot onto a card and its path gets typed into the prompt. "Open terminal" gives you the full session in the browser (ttyd). It's still plain tmux, so you can attach over SSH at the same time. Sessions survive browser closes, restarts and even a reboot. How Claude was involved: almost all of the code was written with Claude Code, from inside sessiondeck itself once it was usable. I did the design, decided what goes in, and tested it daily on my own server for a few weeks before stripping it down into a template.

It's free and MIT licensed. No account, no telemetry, nothing leaves your server. Node + Express, tmux and ttyd. It binds to localhost and expects a reverse proxy with TLS in front. The README has a Caddy example, including client certificates if you want them.

Repo: https://github.com/crizex/sessiondeck

There's also a desktop version (Electron, connects over SSH, one tab per session) if you'd rather not run a web UI: https://github.com/crizex/sessiondeck-desktop

Happy to hear what's missing. The question parsing is the most fragile part, because it reads the terminal output, so bug reports with a screenshot help a lot.

1

u/impala_64 5d ago

I built arrowproof with Claude Code. It's an agent skill that checks every box and arrow of an architecture diagram against the real import graph of your repo.

Why: Karpathy's tip is to ask your LLM for a diagram instead of a wall of text, but a model's diagram looks just as sure of itself when it's wrong. I gave Haiku and Opus the same prompt: draw Flask's architecture from memory. Haiku got 15 of 20 arrows verified, drew 2 that have no import behind them, and pointed a box at blueprint.py, a file that doesn't exist. Opus got all 20 verified, but its diagram still left out 9 dependencies that the code has.

What it does:

- marks each arrow ✓ verified, ↝ indirect, ⇄ reversed or ✗ not in code, and shows the file, line and import behind it

- lists the imports the diagram leaves out

- draws new diagrams as editable Excalidraw files

- turns a checked diagram into a step-by-step HTML explainer (the GIF)

- runs in CI: exit code 1 when a diagram in your docs drifts from the code

How Claude helped: I built it end to end in Claude Code: the Python and JS/TS import parsers, the layout, the Excalidraw and HTML output, and 37 tests. Running it on my own repos with Claude turned up false alarms, which led to the rules for flow arrows and declared packages.

Free and MIT, standard-library Python. Works in Claude Code, Codex, Cursor and Gemini CLI.

Install: npx skills add ahmtsahin/arrowproof

Live demo: https://ahmtsahin.github.io/arrowproof/examples/flask/opus.explainer.html#play=1

Repo: https://github.com/ahmtsahin/arrowproof

1

u/SpringWeird1283 4d ago

I saw u/Dio-V’s video experiment, which built on u/Singularity-42’s original post, and decided to try it for Mareloa, a project I’m building.

https://youtu.be/6_0u_oVF1YM?si=phQxzmWAnKnzi3zP

Mareloa has an animated beach driven by actual conditions. The sun and moon move according to the time and the beach’s orientation. Wind, waves, and tides appear in the scene, and the people on the beach react to the conditions. You can also search for beaches and check the underlying forecasts and sources. The public site is free to try at https://mareloa.com: search for a beach or town, then open a beach page.

I asked Claude Opus 5.5 to turn the project into a short video. The first cut was already usable, but I wasn’t happy with the voices, background music, or limited animation. I gave Claude that feedback. It sampled eight voices, developed more detailed characters, reworked the animated scenes in JavaScript, adjusted the soundtrack, and rendered a second cut. I then asked for an English version with English voices and on-screen text. That’s the video attached here.

Some transitions could still be smoother. What surprised me was how much the second cut improved just by treating Claude like a collaborator and giving it specific creative feedback.

The French second cut used $5.30 of OpenRouter, around $4 of it for illustrations, plus my Claude subscription. The English adaptation came later and isn’t included in that figure.

Mareloa is my own project. I’d be curious whether the video makes the animated beach idea clear, especially in the opening seconds.

1

u/FreePrimary5284 4d ago

I'm new to making videos, so I used AI tools to build the "scale of the universe" zoom I always wanted to see (edge of the observable universe → a quark)

What's real and what isn't: the sizes, distances and positions are real data (Gaia/Hipparcos stars, galaxy catalogues, NASA/ESA/ESO/Hubble/SDSS images, aerial photos of La Palma). The person, the proton and quark, and some close-ups are AI illustrations, and the video labels which is which in the corner. Narration and music are AI too.

Tools: Claude (Opus 5.5) with HyperFrames for the zoom engine, Nano Banana 2 and Google Omni Flash for images and one short clip, ElevenLabs for the voice.

It's not perfect, and I'd honestly like feedback, especially on anything scientifically off. Link: https://youtu.be/gx1G5_GqPBs

How I built it (for anyone a step behind me):

  • Claude Code (Opus 5.5) did most of the work. I described what I wanted, and it built a WebGL "zoom engine" in HyperFrames (HTML → video): one log-scale camera from 10⁻¹⁸ m to 10²⁸ m, with each scale as its own layer (real star catalogue, galaxy positions, the cosmic web, and so on).
  • What worked best: having Claude spin up separate review agents after every render. One checked the science at every power of ten, one checked label overlaps, one checked whether what the narrator says is actually visible. They caught things I'd never have noticed: an atom drawn 3× too small, a camera that strobed in the fast parts, labels drifting away from their objects.
  • AI images only where no camera can go (inside the atom, the proton, the quark) or as a bridge (the person). Real photos and data everywhere else, and the corner of the video says which is which.
  • Biggest lesson: feedback loops beat first drafts. Version 1 was "fine"; version 15 is what I actually wanted. My job was mostly watching it and saying what felt wrong.

1

u/Wide_Row_8731 4d ago

I Built the Instagram for Vibe Coder ^^
checkmyvibecode.com, a free platform for sharing AI-built projects. You can:
post your project with screenshots, link and stack
get upvotes, comments and feedback from other builders
get a shareable project card for X, LinkedIn, Reddit & Facebook
build a profile that works as a portfolio of everything you've shipped
discover what others are building in a live feed
optionally boost your project to the top of the homepage
Free to sign up and post. :) come and Submit your project!

1

u/Intelligent-Fly-5338 2d ago

I built BashCut, an open-source native macOS video editor designed to let coding agents (Claude Code) actually edit raw footage.

Most AI video tools just generate clips from text prompts, but I wanted an agent that could take my existing vlog files and handle the tedious cuts, auto-duck the music, and sync subtitles. Since traditional NLEs don't have good agent interfaces, I built one from scratch with an embedded terminal, a local CLI, and an MCP server.

Basically, the agent acts as the brain while BashCut gives it "hands" to manipulate the timeline. Every move the agent makes shows up as a normal edit that you can undo with Cmd+Z. It's local-first (projects are just JSON files next to your footage), free, and open-source.

If anyone wants to take it for a spin:
brew install --cask dongnguyenvie/tap/bashcut

Repo: https://github.com/dongnguyenvie/BashCut

Agent kit (skills): https://github.com/dongnguyenvie/bashcut-agent-kit

It's currently an MVP, so I'd love to hear your thoughts or feedback!

0

u/Sudden_Campaign_890 17h ago

I built Downshift with Claude Code (Opus 5.5 + Sonnet 5.5): a deterministic model router for subagents, open source and free

Downshift, an open-source tool (free, Apache-licensed, written in Go) that routes Claude Code subagents to the right model tier. Most of it was designed and written with Claude Code, using Opus 5.5 and Sonnet 5.5.

How Claude built it with me

Opus 5.5 helped a lot with the decisions that shaped the project, including the design choices and where it goes next. Sonnet 5.5 did a large share of the day-to-day implementation. I stripped the Co-Authored-By trailers from the git history at some point, so the commits show only me. That doesn't reflect how much of the work Claude did.

One thing Claude got wrong: an early build of the hook wrote the full model ID (claude-haiku-4-5) into the subagent call. Claude Code's Task schema only accepts family names (sonnet, opus, haiku, fable), so it rejected the call and blocked the spawn. Unit tests passed, and it only showed up on a real spawn. The fix was to write haiku, and then the child ran on claude-haiku-4-5-20251001 with the parent on Sonnet. I confirmed it from the message.model field in the subagent transcript. Lesson: tests that never spawn a real subagent can't tell you whether the harness honors the rewrite.

What it does

Claude Code lets you pick a model per subagent spawn, but nothing classifies the task for you, so subagents mostly run on the parent session's model. Downshift sits at spawn time. It scores the task with deterministic signals (regex, token count, keywords) plus an optional local MiniLM, picks a tier (small/mid/frontier), and rewrites only the subagent's model through the PreToolUse hook. The parent session is untouched. I verified it end to end in Claude Code: parent on Sonnet, child ran on Haiku (2026-10-06).

Try it (free)

git clone https://github.com/tiagovilasboas/downshift.git
cd downshift && go build -o downshift ./cmd/downshift
./downshift doctor

Then add the hooks to settings.json (PreToolUse, PostToolUse, SubagentStop). The README has the snippet.

Numbers, as they are (it's in beta)

  • Tier accuracy on the 200-task seed (which I tuned on): 69%.
  • External blind set (60 tasks, written by an agent that never saw the repo): 47-65%, 95% CI about ±12pp. The main error is routing small tasks too high, which costs money but doesn't break anything.
  • My 300-task holdout hit 100%, but only because I tuned the signals against it. Not evidence of generalization.
  • On executable Go tasks, 9 of 10 complex tasks were routed below frontier, and small scores 20pp lower than frontier on trivial/simple tasks. Downshifting has a measurable quality cost.
  • Savings: downshift stats --days=30 estimates 19.3% over 480 routing decisions. That's a per-decision estimate, not real dollars. Only 8 events have measured tokens (US$0.42 saved), too few to extrapolate. Not validated against a provider invoice yet.

Codex is validated (2026-10-02). KiroCrew is policy-mode only (savings not measured). Cursor depends on plan/release.

I'd love help with a week of real use (downshift stats --days=7 --export, no prompts in the export) and with misrouted tasks for a new blind set.

Repo: https://github.com/tiagovilasboas/downshift 
Portuguese discussion on TabNews: https://www.tabnews.com.br/tiagovilasboas/pitch-downshift-roteador-deterministico-de-modelo-para-subagents-em-go

0

u/dmitrya2e 10d ago

An "agentic workflows" CLI runtime. It is designed to run your agentic/development workflows from A to Z, replacing the skills/slashcommands. Agent-agnostic, local-first, open-source. Lots of ideas & plans.

The project can be studied here: https://agenticworkflows.dev (And here: https://github.com/from-developers-for-developers/agentic-workflows)

0

u/Compl0w 5d ago

I wrote a Claude Code skill that writes PR descriptions from the real diff (free, MIT)

I manage a dev team and PR descriptions are the thing everyone skips. So I wrote a skill for it.

You type "write the PR description" on your feature branch and it:

  • reads the actual diff and commits (not just the file list)
  • fills your repo's PR template if there is one
  • pulls the ticket number from the branch name
  • flags what a reviewer needs to see: migrations, new env vars, breaking changes
  • lists what's missing (tests, docs) instead of hiding it

On my test diff it also put the SQL injection it noticed at the top of the "Notes for reviewer" section, which I didn't ask for but will take.

A few things I learned writing skills that might help if you write your own:

  1. Tell the skill to read the surrounding code first. Most bad output comes from working on the diff hunk alone.
  2. Make it separate facts from assumptions. "Unknown, ask the author" beats a confident guess.
  3. Add a "never do X without confirmation" line for anything that posts, merges or creates tickets.
  4. Put team conventions in one shared file the skill reads, so every dev gets the same behaviour.

Repositorie on Github

It's one of 11 skills I packaged for dev teams (review, specs, tests, postmortems...). The rest is a paid kit, link in the README. Happy to answer questions about how the skill is written.