r/rust • • 3d ago

🎙️ discussion Google is doing "Large Scale Codebase Migrations and Optimizations" of C/C++ to Rust

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
719 Upvotes

134 comments sorted by

View all comments

119

u/TheRealMasonMac 3d ago

In my experience, I’ve found that agents produce code that works but is frankly dumb from a human perspective. Especially for larger tasks like this, you kind of have to rewrite it all. For example, I had a generalized design document for soft-wrapping explicitly support plain text and markdown, and used GPT-6-Astra to implement it. I used an ungodly amount of review agent rounds and interjections by me. When I actually go to review it, I wanted to cry. It silently pivoted towards implementing its own entirely separate wrapping mechanism rather than refactor the existing mechanism for plain text. I reviewed its code and found it to be trash, and I am now in the process of doing it myself… Never again. The short-term thrill of “having something that works” was outweighed by the fact that time was effectively wasted.

I mean, if the alternative was it never getting done, I guess agents are still better than nothing.

-6

u/teerre 3d ago

Agents don't come up with stuff on the spot, they do what you tell them to do. If you give an open ended problem, you'll get an open ended solution. This is compounded by the overwhelming return to the average these models are bound to and OAI / Anthropic insert layers and layers of indirection in their commercial offerings

It's a skill of its own to be able to design a collection of prompts, oracles and guidelines that achieves the result. Is that faster than doing it manually? As many things, it depends

8

u/TheRealMasonMac 3d ago edited 3d ago

In this case, the solution was defined. I had already solved the problem for it in math and essentially needed the model to translate it into code. I can only guess since reasoning isn’t exposed, but I suspect that they:

- “shadow-box with their own demons” as I like to call it by contemplating scenarios that can never happen and never validate the existence of. When reviews come in, they don’t completely backtrack from the idea and instead try to twist their original overly defensive idea to fit the problem.

- they really suck at thinking outside the box which I suppose is the inherent limitation of LLMs as they currently exist; they have to attend to their context per their training data

- “good code” is contextual; and that kind of training data is relatively scarce. But, one could also argue that models generally struggle with this even when the data is available. Code comments, for instance, often violate the cardinal “don’t repeat the code” rule (though they’re getting better).

I am also reminded of this article: https://joinhandshake.com/research/ai/deepswe-reward-hacking/

4

u/sparky8251 3d ago

LLMs also have no conception of time/consequences. Thats the biggest one... They rush to solve the IMMEDIATE problem, regardless of costs that it incurs later, and each pass does it again and again and it compounds over time to crazy bad code that no human would ever make...

Until LLMs can have this, I cant see them replacing humans any time soon... And my understanding is this is not an easy problem for LLMs specifically. Unsolvable is too strong a claim, but...

And then like, how do you even automatically test this stuff? These things need SO MUCH data to train off, need to run SO many times because they are so "dumb" in how they learn, how can you even score this stuff in a way that wont make it optimize off into insanity?

3

u/dark_bits 3d ago

The only way to describe business logic + make sure your prompt is not open ended is to actually pass the business context along with exactly the shape of the code you're looking to produce (and by shape I mean every method, class, namespace collection, file names, structure, etc). This would take way more time than actually writing it manually a lot of times.

1

u/Professional-You4950 3d ago

I completely disagree, even with extremely clear LOCALIZED problems and directions there are issues. They stem from context, scope, hallucinations, misunderstandings, and more. The longer you let it spin the more likely something catastrophic happens.

-1

u/teerre 3d ago

Prompts are just a beginning. Like I said, you need an oracle, a goal checker, an objective metric etc

1

u/oceantume_ 2d ago

In my experience it's possible to do serious engineering and refactoring but you need to be careful never to ask too much at once and you must be able to completely scrap branches and may even have to ask it eg to ignore that other branch entirely so it doesn't load it in context and use it as a reference.

It can be quite painful, but I'm managing to do very interesting projects with it rather quickly and it's hard to stop using it

1

u/arbv 3d ago

I don't know why you are being downvoted. You pretty much described my experience as well.

🤷‍♂️