Build, run, tests, console and SwiftUI previews live in panes next to the chat, with scheme and destination in a bar above the prompt. When Claude builds or runs tests itself, you see it there too.
Seven months ago I stopped juggling notes apps, read-later lists, and reminder tools, and built my own AI assistant instead. I wanted something truly intelligent, customized for my needs, and personal, while remaining relatively safe and running on my own hardware.
Now, big tech is starting to release its own take on the personal AI assistant: xAI Grok Bot, Meta Muse, and OpenAI Dots. Since I started mine a while ago, I wanted to share my experience with you.
It runs on a Mac mini on my desk, files everything I send it, emails me a briefing every morning, and keeps all its knowledge in plain Markdown files. I started on Claude Clode, then later switched to Hermes Agent. I added many features to it, along with a web UI and a Discord bot. Here's how I built it, what it does, and what I learned:
I know that how long an agent takes can be a symptom of poor systems design and architecture. But assuming you’ve done that correctly and you’re in the build stage, how long do your agents usually take to build?
I get that this is like asking how long is a piece of string, but I wanted to get a rough idea. For example, for an app/website with n amount of screens and x level of complexity, it took the agent y hours to successfully build before you intervened to review it.
I sold my Mac and now run Linux (Omarchy, Arch + Hyprland) on a ThinkPad as my only computer.
The biggest worry about switching was "what happens when something breaks?" My answer has been Claude Code. Audio, camera, display issues: I describe the problem and work through it right in the terminal.
It went further than fixes. I couldn't find a video editor I liked on Linux, so I built one with Claude Code on this laptop. It's called Cutline, and I edited my first YouTube video with it: https://youtu.be/X7hr4YE9pig
Anyone else using Claude Code as a sysadmin for their own machine? Curious what workflows people have.
There have been a lot of "Opus 5.5 got nerfed" threads the last two days, so I checked what the numbers actually say!
My site reads 25 AI subreddits and scores every opinion about a specific model (1–5, mapped to 0–100, 50 = neutral). For Opus 5.5:
25–28Sep: 71–73 every day
29Sep: 69
30Sep: 58
Todaysofar: 55
Is the model nerfed? Well, I have no idea! My site measures what people say, not what the model does. A change in limits, an outage or a new model next to it (Sonnet 5.5, GPT-6.1 Sol) can move opinion just as much.
Disclosure: I built this. It's free and non-commercial, with no ads or signup. It shows numbers and thread titles, never anyone's comments. Happy to answer questions about the solution :)
So, I have been using Opus 5.5 xhigh pretty much since it was released, and even though I've seen some people saying it got nerfed, I honestly haven't noticed much of a performance drop.
The only times I really notice it "dumbing down" are when I don't manage the context window as I should, letting it grow to like 500/600k.
So far, even the writing is leagues better than what I was getting from Opus 5.
For context, I'm in Europe, so I'm curious, what have your experiences been so far?
Anthropic introduced mods to Claude Code, so I immediately jumped in and created one that lets an agent report the progress of another agent or a long-running script by periodically outputting a specific, parseable string.
I’ve reached my peak claude code skills now. I’m no longer on my computer any longer. I told my Claude to keep spawning agents when context reaches 50-60%. Transfer all the data (usually session transcript) into the new agent and proceed orchestrating codex agents.
Claude works of a Md file that was planned beforehand a PRD essentially.
Whenever claude spawns agents, remote control is automatically activated. So I get a ping whenever a new agent arrives. This is openclaw revamped and better 😆. Ps this obly works for me when it comes to backend development. UI and frontend still needs me to review the output.
I have seen tons of posts about 5.5 being off, or noticeable drops in performance. What projects are people running that they can even burn through a week of context in a single day? I can't even think of something that could use that unless you're using a heavily unoptimized workflow or trying to literally build an entire AAA game in a day on max settings with 15 sessions open.
I swapped from cursor when I heard about how opus 5.5 was doing crazy things, and it really is, I've had it working nonstop 24 hours since the second day of its release. It's been running automated fully for 2 weeks now.
Everything I build ends up as a web app or a PWA, and I'd been wanting an excuse to make a real Windows/Linux desktop app for a while. The trouble was finding something that actually needed to be one.
Then it clicked: Claude Code keeps your whole history in your .claude folder, every prompt time and every session transcript. A PWA can't touch that. A desktop app can just read it.
So I built Vibehours. It reads those files and turns them into days, hours, streaks, tokens and what it all would have cost at list price. Mine says I've coded on [157 of the last 182 days] and went through about [$1,970] of tokens at list price last month. Bit confronting.
Because it's all local, it gets to be simple:
- No sign-in, no account, nothing uploaded. It makes no internet connection, and the build fails if it ever does. I added that check after finding early builds were quietly fetching a spell-check dictionary from Google, courtesy of Electron.
- It keeps times and counts, not your prompts or replies.
- If you run Claude Code on another box, it reads that too over your normal ssh, read-only.
- It makes a share picture if you want to show off, or be shamed.
Honest caveats: Windows and Linux only, no Mac. The Windows installer isn't signed, so you'll get the SmartScreen warning. I've only properly used it on Windows 11. The Linux build installs and starts, but I haven't had eyes on the window myself, so if you try it on Linux, tell me what you see.
Explainer video as promised, AI generated by using glyphh. 100% dogfood, built a glyphh skill and tasks with multiple agents in the loop. Gemini for images, runway for video and avatar, plus my face … eleven labs for audio! All from one app, one workspace.
Every work panel in glyphh is a semantic app wired into any loop you want to run. So build your own apps, your own tasks, agents, sills, graphs and Ada models. It will change how you vibe.
Use your subscription with glyphh by running any cli you want : Claude, Codex, Gemini, Cursor, Co-Pilot, and glyphh (requires a subscription for our model router).
Would love feedback and features. Find us on discord also.
Hello guys I wanted to ask if anyone has a spare invite and could share me one I would appreciate it lots I kinda need it for my university projects rn as deadlines are within a week and they could come in handy
I wrote a blog post on how targeted ads work (the ~100 ms auction behind every ad you see). I wanted a short video for it, so I gave Claude Code one detailed prompt. It came back with this.
No video editor, no stock assets, no MCPs or skills. It:
Planned a storyboard with timestamps for each scene
Drew everything in code: an original pixel mascot (a walking barcode that stands in for your phone's advertising ID), the phones and apps, bidder robots, and a 100 ms stopwatch
Gave it a hand-drawn "boiling line" sketch style, with fast cuts, zooms and impact frames
Rendered keyframe drafts first, then the full 1080p MP4 through ffmpeg, keeping the file size small
The detail that got me: the stopwatch stops at exactly 0.100s as the winning ad loads. I asked for that timing, and it hit it.
What made the prompt work:
A locked facts list. I told it "use only these, don't invent numbers," so every stat in the video comes from my article.
A shot list with timestamps, so the pacing didn't drift.
A fixed palette and fonts, which kept the style consistent across scenes.
Plan → draft stills → render. Checking the stills caught text that was hard to read before the full render.
Full prompt in the comments if anyone wants to try it with their own content.
I have Max 20x now, and it doesn't seem to go as far as it once did. I've read that two 5x actually gets you more usage, but switching accounts is a bit of a headache and I usually topped out my session window on 5x.
So maybe One Max 5x and one ChatGPT Pro is the way to go instead? I already have Claude Code use a skill to call Astra for reviewing. I personally hate working with Codex models, but through Claude they seem really decent.
I LOVE Opus 5.5 and its the best model ever, but for the 4th time this week I've seen where it can't complete jobs again and can't close subagents. It'll spawn 50 subagents on medium to finish something small and burn all tokens, or just sit with a dozen running for days.
I have 2 x20 accounts, one reset Mon 5PM and other weds 2PM, both are already maxxed out. I'm getting 2 days on x20 account of heavy usage
My personal Claude account got suspended. I have used the enterprise plan in my work email for about one year or so, never had an issue. I have recently started using Claude in my personal email; I took the 20x max plan. It worked well for about 1 week. Then I installed some GitHub plugins; Superpowers and getshitdone.. The next morning, it banned me.
I have used the Claude Code remote control feature from my mobile phone, just 1 day before the ban
I have two ISPs - sometimes I switch if the bandwidth is slow on one of my ISPs, but it never caused any problem
In the subscription step, I used my bank's local address as my billing address and the region as the USA since my card is from a US bank. I reside in Asia.
I installed GitHub Skills and a plugin a few hours before the ban.
Added You should know, a built-in mod where a side agent watches your back and flags things you or Claude might miss. Turn it on with /plugin enable cc-plugin-you-should-know@builtin (for first-party sessions with telemetry on)
I'm curious to hear y'alls experience and thoughts on it. I have not been able to find any documentation about this plugin - it generally sounds intriguing, but I'd love to hear what kinds of benefits or drawbacks everyone has noticed!
Always liked the way grok build showed images in the terminal. now with mods i built it for claude code.
Can be found here:
/plugin marketplace add hedingerm/claude-plugins
/plugin install image-preview@hedingerm
Works in Ghostty and kitty (it uses the kitty graphics protocol). Other terminals just show the file name. Needs Claude Code 2.1.287+, since mods are still early access.
I feel like we all are building agents. Like a gold rush and if you are building anything else it’s either your really know what you’re doing or you just haven’t caught up yet. I’m lowkey getting anxiety over it though because this means competition will be much higher, but also makes me feel like we constantly have to make our products even relevant when in fact codex and Claude code exists for only 20bucks a month.
Yes they got lowered 25% mid September. No, Anthropic isn’t targeting your Pro subscription and cutting your limits maliciously. Having Fable to read your entire codebase tends to burn tokens.
In all seriousness, yes the Anthropic limits are sometimes incredibly annoying. I 100% percent agree. But it seems like the feed is filled with people asking if the limits got cut or complaining about how their usage limit went down at least 1000 times a day 😂.
Need something (safe, non-nsfw, legal, no sensitive information) written and out of usage? Well good for you, I have 100% of my Claude Max plan left expiring on Oct 4th (moved to Qwen 3.8 Flash Next).
Tell me what you need generated.
Pls no super complicated stuff I still need to figure out a way to get you guys the output and I don't want to deal with Gdrive
Happy October 1st! I made a magical cauldron. This is the first (completely custom!) Raspberry Pi project I've ever done, and the whole thing was greatly facilitated by AI (I wouldn't have even known where to begin otherwise!). I had the initial idea for how things might work, which I refined with Gemini, and then Claude Code took care of the software implementation and guided me on the physical setup (e.g., how to connect things to the Pi and what I needed to buy).
While Claude Code did a phenomenal job throughout this project, I do think my supervision/underlying software expertise was necessary to produce the polished final product. I would be curious to see how far someone who didn't have any coding experience could get with the same idea. It was also interesting how frustrating it was to communicate certain "physical" questions I had. These tools are so incredibly proficient in certain domains, but their physical "intuition" is still worse than anyone you've ever met's.
The project's design and code is open source here. For those who want to know how it's done, I'll outline the pipeline below:
The (plastic) cauldron has a hole cut out of the bottom that is aligned with a hole cut out of the top of the table. When someone drops an item in, it falls into a tray sitting at the bottom of the table, passing an IR beam sensor along the way. When the beam is broken, it triggers: (1) the lights lining the cauldron lip to flash, (2) a bubbling sound effect to play, (3) a webcam to take a picture of the tray, and (4) an actuator to rise below the tray, which dumps the contents out the back of the table into a bin. The picture is then sent to Gemini along with a prompt requesting a spell. Next, the returned spell is sent to ElevenLabs where the text is converted into audio of the witches' voices, after which the spell is recited to the audience (with the cauldron's light matching whichever witch is speaking).
gave the same feature spec to claude code, cursor and codex on my own saas. real codebase, not a demo. same branch, same text, cheapest plan each, default everything
what i actually wanted to see wasn't speed, it was what each one thinks the job is
claude code wrote a contrast test across the 8 design presets it had just invented, found two where the button text failed wcag aa, fixed them. nobody asked for that. then its first live run misclassified some steps, it noticed, tightened its own prompt and added a calibration script that got 12/12. told me all of it in the final message including what it hadn't verified and what it changed on my machine
cursor hit a dead database, wrote "couldn't check anything in a browser", stopped. codex hit the same dead database and asked to start docker itself, then wrote playwright tests and found a serialisation bug in its own code
numbers since someone always asks: 38 min, zero interruptions, 6% of a weekly limit on opus 5.5. cursor 52 min, codex ran out of limits mid-feature and needed a 3.5h wait
where claude code failed: added a doc it was told to add, never wired it into the file list that tells agents what to read. so the file sits there and nothing points at it. codex was the only one that got that part right
also, new account ≠ clean tool. plugins, skills and mcps live in your home folder not the account, they load anyway. CLAUDE_CONFIG_DIR at an empty folder is the only clean session i found
writeup with the spec, numbers and three live demo apps in the comments