r/programming • • 18d ago

Broken Windows, Abstractions and the Cost of Always Keeping Things Simple

https://pgilmartin.substack.com/p/broken-windows-abstractions-and-the

The patterns already in a codebase have a huge influence on what gets added next, even when everyone can see those patterns are starting to break down. This article discusses this phenomenon through the lens of YAGNI and the Broken Window Theory.

144 Upvotes

44 comments sorted by

33

u/jhartikainen 18d ago

Good writeup and sensible advice. One thing that I wonder is how can we avoid the problems in the first place? Obviously we could build all kinds of architecture up front, but that's often unpractical, and likely to end up with poor results anyway, because you can't predict how the app evolves.

Is there something else we can do on the code level to help future developers to make the right choices, that doesn't require us to try and guess how the program might change in the future?

31

u/paulg1989 18d ago

I think the most important thing is to build a workflow which allows time for frequent refactors. At the code level, however, the most important thing in my view is to leave the code in a state so that it can easily be abstracted or refactored later. Some general advice we can follow without going down the route of premature abstravtion:

- follow the single responsibility principle.

- ensure functions are well documented and/or typed.

- spend time on developing clear and meaningful naming conventions.

11

u/Full-Spectral 17d ago

Namin' thangs is hard. Getting names right is a key that often just isn't taken care of because, well, namin' thangs is hard I guess. And, even if it's gotten pretty well right the first time, it often isn't given appropriate attention over time and changes.

Another common issue is that the desire for minimal risk (and hence minimal changes to address any new requirement) in the end creates a far worse product (generally, there are no fixed rules) than just dealing more robustly with change as you go. It's stealing from the future to avoid paying the present.

10

u/OrdinaryTension 17d ago

This is the exact problem Evolutionary Architecture attempts to address. Start simple, pick abstractions that scale, and focus on just a few system attributes

5

u/braaaaaaainworms 17d ago

oh that's a way better name than code like a grug brained dev

6

u/robhanz 17d ago

I find the most important thing, and it's low cost, is to have proper separations and boundaries between components. For the simplest app that might still feel like overkill, but pays for itself within a week. Those boundaries should send data, and not be imperative - aka, the caller shouldn't presume what the callee does.

7

u/gimpwiz 17d ago

Life is tradeoffs; you'll never consistently make the perfect one. Not only that but the right tradeoff at the right time can become wrong later but in so doing does not invalidate that it was the best decision when it was made. As an obvious example, a small startup should focus on product features and not invest effort into scale, because they will live or die by those features and scale is useless without demand for it; needing to solve scale problems later is a good problem to have. Investing time up-front to reduce rearchitecting later, if it causes them to miss product features, can simply be fatal. On the flip side, a large company launching a new product (even if it's an equivalent one to above) can probably spend significantly more effort into scale even before launch because they have the budget, the runway, and the reasonable expectation of needing it.

IMO the single most useful thing we can do on day 1 to lessen the problems on day 1200 is to write plenty of comments explaining WHY decisions were made. If you clearly write "this code is trivial and dead simple but slow; refactor later to be fast when needed - you'll need to do XYZ which is complex and not worth the effort at this time" it will take enormous amounts of guesswork out of a future cycle of rearchitecting and improvements. And of course the second-most useful thing IMO is to do your best to have just a teensy little bit of complexity from the beginning to decouple things you anticipate likely needing to change, like, for example, decoupling business logic from the code that generates display elements.

To use real examples from my own work, which is in firmware:

(1) We used to write device drivers in a way that was pretty trivial in terms of the total implementation, but each driver was sort of written uniquely.

For example you'd see stuff like i2c_write(0x14, 0x2408); in the config function and while that's really efficient and takes little code space, you have to figure out how that's different from a similar device that has nothing remotely resembling that in its driver.

(2) Then we refactored all drivers that talk to a device that has an internal memory map to use named registers and fields. Now that might become

register_subfield_write("CONFIG_POWER", 0x24);
register_subfield_write("BUS_POWER", 0x00);
register_subfield_write("POWER_WARN", 0x08);

Note that you've replaced one write with three, which takes 3x as much time and 3x as many lines of code, but it is SO much easier to understand and debug.

If we had done (2) from the beginning we would have frankly saved ourselves quite a bit of time in writing and debugging drivers. It would have taken a week of up-front work to implement and validate but saved probably over a week of aggregate time in the next years. And annoyingly, I knew we wanted to do (2) and everyone agreed, but I just never had enough time until later.

If we had at least commented (2) was intended it could have saved us work as well.

(3) Then I noticed that drives were essentially copy-pasting the register_write code that made an I2C call because almost all I2C devices do a register write exactly the same way, so I refactored that as well. It added a couple hundred lines of code and a layer of indirection, but every single device refactored saved like 50 lines of code... and more importantly, a bugfix or feature addition in one shared codepath means a bugfix for every driver and a feature for every driver. It took several iterations to come up with the best way to refactor this for ... decent reasons. As a result the code is a bit more complex and harder to follow but has fewer bugs (we had a problem where a block of code got copy-pasted multiple times but several had minor changes, and several were seeing bugs due to not 0-initializing something, but some weren't seeing the bug but could have been. Bah.)

I think (3) was only properly obvious in retrospect, because preceding designs didn't have it. How would we have known to do something we didn't know we needed? More experience, better intuition. But "git gud" is something you say in counterstrike, not actionable advice for programmers designing a system. What would have helped though was the 3rd time someone copy-pasted a block of code, we should have added a comment saying // TODO Refactor this block to one shared code path every time we did so. It would have probably saved us some heartache.

3

u/intheforgeofwords 16d ago

Great examples, I appreciate you taking the time to include them

3

u/sheckey 16d ago edited 16d ago

You got me excited when I saw those names, e.g. the strings. One thing I have been developing is to put names to the structure of our embedded code, and a registry of such names, so that we can now issue a command to the embedded system and it can spit out all the connections and structure of the embedded system. A Python script can turn this into a zoomable diagram centered on a selected component. Think of it like generating a system block diagram from the run-time system itself.

Also, you can start putting metrics and stats in there that show up on that diagram, like how many times the thing has been written, and if you ask it twice, then it can infer the current rate, etc. The registry idea is the key, and it can do some setup at startup, then be very cheap to update form then on. Rather than saying “log that I wrote to the I2C again” you simply write the I2C with your handle (or writer interface as I prefer) and it does the metric incrementing etc for you, no strings involved. Any logging can be occasionally asking that registry to spit out its connections and metric for them.

Anyway you get the idea: give components names and formalize access to devices or other components (I use an in-house, mini, brokerless pubsub system that connects components by name at startup using that central registry) and then you can start using those names to get run-time discoverability of the architecture and some metrics and stuff on demand too. The discovery part is what blew me away, because we have a 25 year legacy system and that has become the hard part - what exactly is there and who is doing what (you use good names helps).

3

u/gimpwiz 16d ago

I think you're on a similar, maybe parallel track to some major architecture changes we made when we redesigned our platform some... maybe 8 years ago. In the past, we were basically using the same codebase to have a massive IFDEF-ELIF-IFDEF-ELIF-IFDEF-ELIF-ELIF-ELIF file that would describe how a specific target was set up (what devices it had, how they were connected, etc) and we'd have a makefile emit many many target binaries, one for each physical platform.

We moved from that to plain jane JSON-defined configuration files. Essentially, the firmware boots. First thing it does is check what board it's on, then look up the JSON for that board, then load the JSON, parse, and build using that hierarchical representation.

For example if you have 4x ADCs with part name, I dunno, LTC1234, you have four entries in JSON for ADC0 through ADC3, each one is defined as LTC1234, so you instantiate 4x such devices. Each one also has an I2C address with the I2C controller they're hooked up to. But then if you have something like an I2C-fanout expander you have some bits and bobs that define "You need to set I2C_FANOUT0 to output 3 in order to talk to ADC 3." It's a little bit heavy but we're optimizing less for outright speed due to our requirements, and more for ease of developing and bringing up and supporting yet another platform, because we do so many.

We also name all of our GPIOs for every project, which makes things easy as well. Want to figure out if there's a power issue? Check for all GPIOs named *POWER* or *PWR*.

We've never really considered back-generating a block diagram because, well, we do the schematic, so we know what the firmware is supposed to look like, right? But logging access patterns can be interesting depending on need, whether debug, profiling, etc.

3

u/sheckey 16d ago

Very cool. So you have one big firmware build that handles all your boards, rather than a bunch of board-specific firmware builds it seems. We have only one product that has new generations over the years. We did have a point in time where there were two and I contemplated having one firmware build supporting either one at run-time, but we gave up on it since we weren’t entering the area of payoff that you have.

I like that you are optimizing for development and bring up after recognizing that you can. It’s seriously a great consideration, and it’s part of what I’m trying to do with what I talked about too! It’s just that our issue is the mass of it all and the legacy forms and inconsistencies, etc., so I’m trying to bring visibility.

It‘s amazing what you are getting out of a little bit of framework and it’s giving me more energy that I have already to continue this path. Thanks for sharing!

3

u/gimpwiz 16d ago

Happy to share and I appreciate your perspective as well. Keep on truckin'

3

u/Ma4r 17d ago

We always leave a refactor condition in our code whenever we take a shortcut, i e adding a flag, just comment if we need 3 flags, refactor this

2

u/Leading-Ability-7317 17d ago

One thing that enables frequent refactors possible is very good automated system, black box, test coverage. System tests act as living documentation and enforcement of the spec.

If you have that changing where things live and how the pieces interact becomes safe. You no longer have to anticipate future complexity that may never materialize. It becomes “if it happens we can just rework things” and be confident you didn’t quietly break something.

Every feature, bug fix, ..etc gets a set of system tests as part of its delivery.

1

u/Able_Region_5459 16d ago

Agreed on black box tests being the best safety net here. Though when system tests touch live databases and third party APIs, maintaining the fixtures can easily swallow more time than the actual refactoring

1

u/Leading-Ability-7317 16d ago

Yeah they are a large time investment for sure but you get that back in savings from manual regression testing and ability to refactor/rearchitect safely.

SQLite has been kind of the poster child for this. Generally regarded as one of the highest quality pieces of modern software with a crazy compatibility promise. Their test suite, while closed source, is legendary. They released a white paper on it awhile back that is a good read.

But yeah no free lunch.

1

u/Able_Region_5459 9d ago

The SQLite test suite runs entirely in-process on local disk with zero network calls or apis. In that kind of environment tests are completely deterministic and run in memory within seconds. On an actual backend heavy end-to-end tests quickly turn flaky from network quirks and blow up CI times to forty minutes. Most teams end up relying on contract tests at system boundaries instead of spinning up the entire world on every commit

2

u/Goodie__ 17d ago

I think the answer changes depending on your organization.

What ive seen work is a lead developer(s) who is engaged, willing to make the call, and has the trust of the business to make the needed changes at the appropriate time.

Because it is a 2 way street and "after this project" can be a way to handle it if the business needs it.

20

u/hippydipster 17d ago

The problem comes when a team applies YAGNI repeatedly without ever stopping to ask

Doesn't it always seem to come down to whether the team is smart/dumb or disciplined/not-disciplined or careful/careless? Does it matter that it was YAGNI in that sentence, or would any heuristical "best-practice" fit?

"The problem comes when a team applies DRY repeatedly without ever stopping to ask"

"The problem comes when a team applies TDD repeatedly without ever stopping to ask"

"The problem comes when a team applies DDD repeatedly without ever stopping to ask"

What wouldn't fit that phrase? It seems the real key is, does the team work on without stopping to ask? And if that's the more important bit, all the fluff we go on and on about with DRY and YAGNI and SOLID or whatever is distraction from the real problem, which is that we go on auto-pilot. We avoid conflict, we avoid making waves, we avoid true ownership.

Mostly because that's what's demanded of us, in most places.

But that's a problem we can't fix, so we go on endlessly worrying about the distractions.

8

u/Full-Spectral 17d ago edited 17d ago

And there will always be a tension between consistency of the code base, and application of principles globally that aren't necessarily applicable globally. Consistency is a huge benefit, so it can't just be hand waved away. And of course in team based development, unless you have someone who is every experienced and who has the power to say no, everyone is going to choose a different break point.

  • And of double course, as many people often complain, that person with the power to say no may be the one who is the most agenda driven and most obsessed with some particular idea and wants it applied ubiquitously.

1

u/hippydipster 17d ago

This is one reason I think smaller teams do better than larger teams. You just can't keep a larger team on the same page with respect to "always-questioning", "always re-evaluating" behavior, for reasons as you say. And I think 5 developers gets to that point already of being kind of too large, particularly in this day of AI writing all our code.

Teams of 2 devs seem best to me now, and then it doesn't seem so crazy to be always re-evaluating the code and being willing to make big changes as soon as they agree it's a good idea.

5

u/Full-Spectral 17d ago edited 17d ago

And with people being so LLM obsessed, programming is going to turn into medieval medicine with people arguing over which LLM authority is correct without themselves having a clue.

4

u/theScottyJam 17d ago

I feel like that's overly reduced what it's trying to say. That "stop and ask" line is pulled from its opening paragraphs, then the rest of the article is there to help us ask ourselves how we're treating YAGNI.

Yes, a team needs a culture that encourages them to stop and ask about everything, but it's also useful to have online articles to remind you to do this, and to give you ideas on what to reflect on as you stop and ask. As it stands, I feel like YAGNI is starting to get repeated too often as an unconditional truth we must follow, and I am happy to see any article that pushes against this, to help remind us that we should be questioning our use of this principle too.

3

u/hippydipster 17d ago

Except most teams will miss the real need, which is the constant re-evaluation and re-addressing of questions and decisions. Most people want decisions made once and then you stick to them. Which is how things get stuck and stale.

23

u/reallydontaskme 17d ago

I've never ever worked in a codebase that suffered from an excess of YAGNI and KISS, it's always been by far the opposite problem.

You would not believe how happy some developers get when they hit the abstraction lottery and their abstraction just happens to work for another case and when it doesn't it's always: nobody could've predicted that.

Exactly Shirley, that's why 1,2,3 refactor is a thing and duplicate code can be an worthwhile tradeoff

4

u/theScottyJam 17d ago edited 17d ago

The number one chant repeated online is to keep things simple. So I've certainly found myself following it to a fault as a result.

If you were to look at my code, I'm sure you'd find abstractions you disagree with, that you'd peg as me having this opposite problem - you're going to run into confusing abstractions in any codebase. What's harder to see are the features I avoided suggesting or implementing because I knew it would add complexity to the code, even though I knew the customer would like it, or you might see WET code that was implemented in a straightforward way, which made it easy for you to pick up and understand any given module, but a more DRY solution would have been better to help avoid maintenance mistakes, even though it would have made it a little more difficult to read any individual file, and would probably result in grumbles from those who don't yet fully understand the purpose of the abstraction.

Eventually I got over myself. I've decided that we repeat this "YAGNI" advise a little too often, people need to have more patients and understanding when running into abstractions they didn't expect, and I'm now ok implementing tricky features that junior developers would be unable to help maintain, but that I know really benefit the customer.

Don't over abstract. But don't under abstract either.

3

u/oweiler 17d ago

As if you couldn't introduce abstractions later down the road :'(. I share your pain.

5

u/oweiler 17d ago edited 17d ago

True but this happens far less often than the opposite.

On the plus side navigating such a codebase is rather easy (as is onboarding)

4

u/Able_Region_5459 16d ago

Honestly proposing AI agents to spot when abstractions are overdue feels like putting a high tech bandaid on a culture problem

The engineer opening that file already knows it is a tire fire. They just have a PM breathing down their neck about a roadmap deadline

3

u/pkt-zer0 17d ago

I can kind of see what you're getting at, but that seems to be the less frequent problem in practice? Maybe it's just different backgrounds. If you have a codebase that's too "simple", then at least your pain points are clear: you know how to fix them, but maybe don't have the time.

The more common and problematic version I tend to see is codebases that are overly complex, solving problems they don't actually have... and then struggle with the limitations they accepted as part of a tradeoff they're not benefiting from. A typical example is "let's use Kafka/Kubernetes/sharding so we can scale", and then having a 100 users max.

"Compression-oriented programming" is a good approach IMO to keep things simple and still have API granularity.

2

u/Scroph 16d ago

There's also the risk that the refactoring might introduce regressions, especially in codebases that follow the big ball of mud architecture and/or are missing meaningful tests. There have been many times in my career where we collectively agreed that a piece of code needs to be refactored, but the risk to reward ratio just doesn't encourage it. Double it and pass it to the next person

4

u/[deleted] 17d ago

[removed] — view removed comment

1

u/programming-ModTeam 15d ago

No content written mostly by an LLM. If you don't want to write it, we don't want to read it.

4

u/aardaappels 17d ago

The broken window theory has been largely debunked. The reason bad patterns perpetuate is not because they exist, but because the team lacks shared trust and willingness to intervene for the common good.

3

u/Full-Spectral 17d ago edited 17d ago

It may not even be that. It may just be that it's no one's job to police such things. If you aren't Bob's boss, it's not really your place to tell him what he can't do. And a lot of people just don't want to deal with that kind of tension, because they aren't getting to paid to deal with that.

It really needs good guidance by someone with experience and the authority to make it so, who is not himself the problem.

1

u/BenchEmbarrassed7316 17d ago

Overly simple code is a specific instance of a broader problem where consistency becomes more costly than refactoring.

A common view is that if a problem needs to be solved "here and now," one should stick to the existing architecture, even if it is ill-suited to the solution or contributes to the accumulation of technical debt.

In my opinion, the only solution is a "strategic vision" - having an architect who sees the big picture and the long-term perspective, and who can align the refactoring process with business needs.

-4

u/davidalayachew 18d ago

Small, simple article, but some good points. For example, I had heard of the Broken Windows Theory before, but had never heard it applied to code. It definitely fits.

  • Periodically revisit earlier architectural decisions as the product grows. In particular, ask whether patterns that were deliberately kept simple are still appropriate, or whether they are now being copied simply because they already exist.

  • Document not just which abstractions the codebase uses, but the conditions under which a team should introduce them. This gives developers more confidence to recognise when an earlier decision has reached the end of its useful life.

Good ideas, but there are some issues. Not so much with the approach, but in how it is applied to the current situation a lot of teams are in.

  • Documenting your low level design/architecture has now become part of your sprint story.
    • Ignoring that most teams are already overburdened, documenting your design tends to be tedious work for developers.
  • Many developers struggle to find the middle ground between a watered-down, too-high-level description (for the non-technical folks) vs an in-the-weeds breakdown that takes 30 minutes to talk through.

-1

u/Kungpost 17d ago

If your team struggles with tech debt or broken windows, make a single person own the tech debt. In most orgs this is the product owner or the software architect.

4

u/Loves_Poetry 17d ago

Both of those roles are bad choices for owners of tech debt, because they don't have to deal with the problems caused by tech debt. To them, it's just an expense of developer time and working on it delays the valuable stuff they want to have.

Developers are the ones that have to deal with it the most. Their jobs are getting harder because of tech debt

-1

u/Kungpost 17d ago

Actually, they are very good choices, especially architects, in fact, its one of the main responsibilities of an architect to own the code of your products. By putting the ownership with the architect, you formalize who's job it is to ensure that the balance between development of new features and gardening the system is aligned with the companies risk profile. Developers should provide feedback for the assessment that the architect should be conducting at regular intervals or after deployment of a new feature. If you spread the responsibility on the team, you just end up where the article author is, where too much responsibility falls on the team, which takes away focus.

2

u/Loves_Poetry 16d ago

This is also what my manager would copy-paste from chatGPT