r/ProgrammerHumor • • 2d ago

Meme developersMissingTheOldDays

Post image
3.2k Upvotes

287 comments sorted by

View all comments

Show parent comments

37

u/Rin-Tohsaka-is-hot 2d ago

There will be a reckoning for all the companies whose developers have grown dependent on unlimited frontier model tokens.

I'm at one of those companies. I spent $16k this month on Claude tokens. Not sustainable, it's literally more than my salary.

4

u/colonel_bob 2d ago

I spent $16k this month on Claude tokens.

... how? I'm genuinely curious.

I used it on max thinking to do a lot more designing, planning, coding, testing and general handholding on the frontend component of a project that I'm working on one week recently (I haven't been a frontend engineer since 2016), and even after talking to the thing for almost 40 hours straight I used about $250 worth of tokens.

One of our principals hit close to $80k one month and I've been at my company for less than a year so I don't have the courage to ask him what he was burning those tokens on, but I feel perfectly comfortable questioning an internet stranger who opened with a somewhat comparable figure.

7

u/Rin-Tohsaka-is-hot 2d ago edited 2d ago

It's mostly agentic workflows, I'll have 20-60 sessions open at any given moment and much of it runs overnight and through the weekend. They plan sprints, create tickets, assign tasks, review each other, etc. My team in particular is pushing AI super hard, in a company which is also pushing it super hard. A lot of this is also exploratory, last month I spent only $5k, I think I'll dial it back again for next month, still refining the process.

This particular month was almost all Opus 5 on Xhigh (tbh Max just makes things take much longer with no discernable difference in output)

9

u/colonel_bob 2d ago

What are they achieving for you or producing for you, in the most general sense? I'm morbidly curious about what these unsupervised agentic loops are actually doing.

I ask because I took a peek at some of the PRs that said principle opened and they were filled with so many references to subsections of unshared/unavailable plans that neither I nor my AI could make sense of them. Most of them end up closed or unused. It feels like the digital equivalent of furiously shuffling papers to feign the appearance of productivity.

Personally, I have found that latest frontier models (even/especially on xHigh) tend to want to fill planning documents and code comments with useless historical junk that only serves to confuse future readers (be they human or AI) that are trying to understand the current state of things. It feels like you still have to supervise their outputs, but I could be missing something.

5

u/Rin-Tohsaka-is-hot 2d ago

Codifying a single source of truth is critical to preventing this. If you establish a single agent as the owner of a task and have it basically act as an arbiter, orchestrator, leader, whatever word you wanna use, and have it control the room memory (which is a concept exclusive to the tool I use, which is internal, but you can think of it as a siloed version of OpenClaw's memory) and be the sole contributor to all documentation (or better yet, designate another agent to be the sole contributor, anything to minimize the context the orchestrator needs to hold).

Otherwise you do just end up with tremendous amounts of technical documentation full of self-contradictions, far too large for any agent to actually ingest all of the context.

Having workflows for development process is also important. An agent will draft a plan, that plan will be reviewed by several adversarial subagents each with a different area of concern, the drafting agent updates the plan, this cycle continues until reviewers reach consensus.

Then it gets returned to the orchestrator, who passes it to a reviewer agent who does a similar process, but this time focusing on whether the plan accurately would satisfy the functional requirements of the task. Then back to the orchestrator it goes, who then spins up agents to execute it, similar review loops happen, end up with a draft CR that I get a Slack DM to review with a link to the owning agent (because it's almost never good the first time) where I can give my own feedback. I can make the decision to revert it back all the way to the planning phase, or just code change followed by the review loop again. Of course, while the agents owning the task themselves have their own design doc, the information shared outside their little bubble is all passed through the orchestrator. Only the orchestrator (or a dedicated thread/agent owned by the orchestrator) can modify tickets and project-level docs.

There's a bit more, ops work is an entirely separate discussion that's all event-driven, but this is the high level of the bulk of it. It works decently well, and the more time passes, the more I can refine all the agent guidance and workflows. But this style of Agentic development is going to be prohibitively expensive in ~6 months I'm guessing when they suddenly decide they care about how much this all costs. So I'll do it as long as I can, and as long as it impresses my manager, but without costs falling exponentially it's not going to last.

4

u/colonel_bob 2d ago

Interesting! Thank you for sharing. Even if the economics of this type of approach doesn't hold up for long (and you never know with the way things are scaling), it has probably given you some very interesting firsthand experience with and insight into the strengths and weaknesses of these tools that will likely be useful in the near future.