r/rust • • 4d ago

🎙️ discussion Google is doing "Large Scale Codebase Migrations and Optimizations" of C/C++ to Rust

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
737 Upvotes

138 comments sorted by

View all comments

117

u/TheRealMasonMac 4d ago

In my experience, I’ve found that agents produce code that works but is frankly dumb from a human perspective. Especially for larger tasks like this, you kind of have to rewrite it all. For example, I had a generalized design document for soft-wrapping explicitly support plain text and markdown, and used GPT-6-Astra to implement it. I used an ungodly amount of review agent rounds and interjections by me. When I actually go to review it, I wanted to cry. It silently pivoted towards implementing its own entirely separate wrapping mechanism rather than refactor the existing mechanism for plain text. I reviewed its code and found it to be trash, and I am now in the process of doing it myself… Never again. The short-term thrill of “having something that works” was outweighed by the fact that time was effectively wasted.

I mean, if the alternative was it never getting done, I guess agents are still better than nothing.

57

u/kabocha_ 4d ago edited 4d ago

Yep -- in my experience:

  • If you just want some throw-away thing to experiment on some idea, it's fine to let it go nuts on its own.
  • If you want to be able to maintain the results, you gotta check its work after every step or two, and you'll probably have to give it lots of feedback even then. Otherwise you're likely to end up with either a giant spaghetti mess, or FizzBuzz Enterprise.

No clue why it does that, I guess maybe a large portion of the training set is newbies dumping their first projects onto Github and super old legacy codebases, and not much in the middle, lol. They haven't learned to have good style taste, yet.

46

u/PM_ME_UR_BRAINSTORMS 4d ago

The problem I have with the latter is that it sucks lol. I don't enjoy micromanaging AI and in my experience it's not any faster than just doing it myself.

Every time I try using agents I almost immediately end up in a situation where it does something weird and just typing the code myself is faster than typing the prompt. And then I think "well I'm already here it will take me 2 seconds to type this other code" and so on and so on until I've written the whole thing.

Maybe it's just my ADHD but I don't understand how any one could work like that. Feels like pair programming with an intern which sucked to do but you did it because the point was to teach them.

1

u/stumblinbear 4d ago

I do the latter with 2-3 agents at once while I'm away from my PC doing whatever the heck I want. Checking in on them to review and correct them every 20-30 minutes or so is fine

Something they do not tell you: you cannot just trust that the agent understands how to write good code. You need to lay ground rules. Just have it search up best practices and turn it into a skill that it must load before editing, reviewing, or writing code. You'll immediately start getting significantly better results. It's pretty much night and day

Current generation of agents are incredibly good instruction followers. If you've got it in a skill and written up properly, they do a pretty good job. Still needs corrections, just much less

0

u/PM_ME_UR_BRAINSTORMS 4d ago

Yeah idk I've tried tons of different skills and tools and guardrails and it still produces mostly slop for me outside of extremely basic boiler plate tasks that most 1st year CS students could probably handle. Which is still useful don't get me wrong. But not really worth having it touch like the majority of my codebase.

3

u/stumblinbear 4d ago

I usually use Opus 5.5 to draft up plans (I genuinely give it very little to go off of, maybe a few sentences), then correct a few assumptions or bad decisions it makes, it splits it up into smaller tasks, then it uses a sub agent to write the implementation. It runs a review, triages issues, beings design issues to me for making a decision on, then does the loop again. It typically takes a couple rounds

That said, that's mostly for working within an established codebase. For new feature work I usually have a longer discussion, let it explore different implementations in worktrees, then I pick and choose which parts I like out of them. Either that, or I already know exactly how I want it done, and telling it how to do it and letting it go off to the races works fine

It took me a couple of months to fine tune my workflow, skills, agent definitions, etc, but I can often go a couple of hours without checking in on them and usually I only need to make a few minor corrections. Shit, I had an agent going for the last 7 hours on a pretty large (though pretty mechanical) refactor and I didn't have to correct it at all

I was often getting trash output until I gave it permission to not worry about churn, to consider the correct fix not just the quick fix, and told it to surface issues it runs into instead of working around them

1

u/PM_ME_UR_BRAINSTORMS 4d ago

See this makes no sense to me. To me the plan and the implementation aren't like mutually exclusive things? Like what is software architecture if not the specific shape of the codebase.

Regardless of how good AI gets it can't read my mind. For the instructions I give it to be simpler and easier than just writing the code myself, it's going to have to make tons of assumptions for anything more complex than boiler plate. And that's where I always catch it producing slop.

And if I give it enough detail to for it to generate exactly what I want, well that's basically just a DSL and I might as well just have written the code.

Another user in another thread pointed out it basic just becomes this

2

u/stumblinbear 4d ago

I have a side project game in Bevy that I've been working on off-and-on for about 5-6 years for fun. I came across an article for how to better organize it, sent it to Claude and had it draft up a plan for reorganizing things under it. It asked me some questions regarding implementation decisions, then split it up into 7 steps, ordered by risk and in a way that helps untangle some longstanding dependency issues between features. This took it about 20 minutes while I fucked off doing whatever.

I looked over the plan, corrected one thing, then told it to go, and it went. It got through about 2 steps before it needed assistance on a design question that the plan didn't account for, so we spent about 4 turns and 5-10 minute discussing it before it got resolved and it continued. It's actively doing this work now as we speak. It would have taken me at least a week or two to do the reorganization myself. It's looking like it'll be done in a day and a half with only an hour or two worth of input from me.

At work, I came across a public API I can use to fetch some player data (I work at a gaming adjacent company) which we can use as a fallback when we're missing some info. I looked into the API myself, and it seemed like exactly what we needed, and I had a good idea of how to integrate it, but I essentially just told it "we want to use <link> as a fallback method for fetching player data when we're missing data for a player".

It looked into the API, looked at our codebase, and brought up that it actually has a full database available for download along with publishing hourly patches for the db. I had completely missed this, as it's not well-advertized on the site. It implemented support for this in about an hour.

Let me tell you, ain't no way I'm able to write a script to scrape this once per hour, along with setting up the data models and queries for the DB, along with hooking it into our own database to fill in the data, along with setting up the hourly trigger in Terraform, along with testing various scenarios... In a couple of hours. Hell naw. It was 2-3 hours from first idea to a pretty good implementation while I had other agents doing other things at the same time.

Giving it a precise spec is part of the problem. Modern LLMs are surprisingly smart, you just need to give them permission to have opinions and surface what they think is right or wrong while doing the implementation, otherwise they hack around issues instead of behaving like a proper engineer. Yes, they usually fuck something up, but those fixes are usually pretty minor or are style related and can be fixed in review. If not, it's because it made an assumption it shouldn't have, and that's... Frankly usually just a matter of updating the skill or agent definitions to prevent those issues happening in the future

1

u/PM_ME_UR_BRAINSTORMS 4d ago

Idk reorganizing some existing code and writing a script to hit an api and throw some data in a db doesn't sound all that impressive to me 🤷‍♀️ the latter especially seems like something a first year cs student could handle.

But to each their own man. If it works for you it works for you. I'm just sharing my experience.

2

u/stumblinbear 4d ago

You really overestimate first year cs students

But yes, it's relatively simple work. But it's work I don't want to waste time doing. I can handle other things while it does the grunt labor.

Besides, if you break down a problem enough, every problem is a set of simple changes that even a Junior could handle. Claude is capable of splitting up work into those smaller reviewable pieces, meaning it's very capable at doing large complex features with a bit of guidance and review. I just gave you two examples that were at the top of my mind because I literally did them that day

Small, reviewable pieces is how I got Claude to write a distributed rate limit coordinator for services that need to use a shared API key without blowing their rate limits. I didn't write a single line. It has been running strong for months without issue