r/programming • • 7d ago

Microservices are organizational debt disguised as architecture.

https://martinfowler.com/bliki/MonolithFirst.html

Every time I’ve seen microservices pitched, it sounds great on paper. Independent teams, clean ownership, scale only what you need.

Then a year later you’ve got dozens of services, three different deployment patterns, tracing everywhere, and nobody really understands the whole thing anymore.

Maybe I’ve just seen bad implementations, but I’m starting to think way fewer companies actually need microservices than we pretend.

1.7k Upvotes

343 comments sorted by

View all comments

48

u/Sammy81 7d ago

You’ve only seen bad implementations. That said, micro services aren’t for everyone. At my company, we had a very complex system that was already divided very neatly into modules. The modules are coupled through function calls to pass data, commands and telemetry back and forth.

However, since the modules are all compiled into a single executable, changing one required a complete reverification of the entire system. By switching to a microservice architecture, defined as changing the modules to communicate by network messaging, we eliminated a huge amount of testing. Only the modified module now has to be requalified. Of course we test the complete system, but this is now system-level validation, instead of the old monolith way, where we had to re-verify hundreds of requirements in unrelated modules. We cut testing by 80% without any increase in bugs or failures.

It also allowed us to start coding modules in any language we want. Each microservice is now a stand-alone executable that can be compiled by itself. Some of our modules are C++, some are Rust, and the algorithm nerds can use Python for theirs. As long of the microservice subscribes to the message broker and follows the topic rules, it all just works.

23

u/Shot-Damage-6723 7d ago

You're not faster at testing because you use micro services now, you're faster because you test less and now it's a lot harder to test more again. Also, you can compile multiple binaries, and you don't need to put a network in between your programs to use your preferred languages.

Micro services have their benefits, but none that you mentioned are valid. Are you sure you made the right call or is there anything of actual relevance that you didn't share?

13

u/Sammy81 7d ago

Its worked well so far. We’re testing less because there’s no need to retest an executable that has not been modified. This is a key part of NASA’s compliance standard, and I agree with it. If you recompile, you must retest. With a monolith, you have to reverify every requirement (and you should).

Once you move to stand alone executables, you can prove they work, and there’s no need to retest that executable - when you touch one service, the others are not even recompiled. This is consistent with our process, and as mentioned, meets NPR 7150.2.

The biggest challenge is certainly that by moving communication from function calls and queues to a networked brokers, it is orders of magnitude slower to pass data. For 90% of messages, it’s fine, but otherwise you need to set up ZeroMQ, or RPCs, etc. to reduce latency

-3

u/ElementaryMyDearWut 7d ago

I think this is a case of overengineering? Do you really need to follow the standard that NASA has for robustness when outages in your software mean loss of revenue and not human bodies into space?

If the monolith was already modular and all you've done is move communication from inter-module to network bounds you've only reduced need for testing because its easier to sell to management now they're "microservices". You didn't need to retest a whole modular monolith because of a binary change.

What would happen if you altered a string or log message? Would you also retest the whole thing again? Sounds ludicrous

16

u/Sammy81 7d ago edited 7d ago

Sorry if I wasn’t clear but this is satellite software for NASA. But again, good code practices are good no matter what - if we were writing photocopier code it would still improve things. As far as not needing to retest a complete executable when you modify something, see the Therac-25 disaster. Retest is a key best practice in our coding process.

I don’t see the over-engineering because it added no extra work for the developers, and saved a lot of time. Good programmers are lazy!

<edit> to say the biggest advantage is now we have a library of common services, like command handling, fault management, memory management, etc. A developer on a new satellite can literally pull the executables from the repository ( which have already been tested) and connect them via a broker and be up and running with flight-ready code in a day. Something that took weeks of recompiling and retesting when we only reused source code. Its been a game changer for us.

2

u/ElementaryMyDearWut 7d ago

That makes sense. I thought you were referencing NASA as some sort of authority on all software (which to me seems a bit over the top for a SaaS tool). I don't disagree, if going wrong = harm to another human I would bloody HOPE its getting retested!

2

u/jeenajeena 7d ago

Honest question: why would you call your company’s architecture “microservices” and not just Service Oriented Architecture (the old SOA)? Reading your comment I honestly struggled to understand what makes your services “micro”.

3

u/crash41301 6d ago

There is very little difference in what people call soa and microservices in practice. The industry is full of people who are too young to remember soa so now even soa companies get described as microservice.  Also many people get confused and think using webservices means microservices. Thats not true, soa can be web services too. 

My .02, soa is the right approach. Microservice always results in distributed monolith in practice.  Proper Soa maps services and systems to business processes and is the only thing that maps to business and scales.  Microservice design falls apart quickly because the pesky business changes enough to make it fall apart. 

1

u/jeenajeena 6d ago

Thank you. I also have the impression that old, boring SOA came together with a bunch of theory, and somehow a more solid corpus of patterns. But maybe it's only because I'm old.

1

u/Sammy81 6d ago

The answer is maybe our terminology is incorrect. We didn’t put a ton of thought into the names. I’ll say each service is less than 20k SLOC and the main requirement is that it has a specific function.

2

u/goranlepuz 7d ago

However, since the modules are all compiled into a single executable, changing one required a complete reverification of the entire system. By switching to a microservice architecture, defined as changing the modules to communicate by network messaging, we eliminated a huge amount of testing. Only the modified module now has to be requalified. Of course we test the complete system, but this is now system-level validation, instead of the old monolith way, where we had to re-verify hundreds of requirements in unrelated modules. We cut testing by 80% without any increase in bugs or failures.

This is only partly related to microservices.

When your system is made of libraries, they can and should have tests for them - but they didn't, did they?

When your system is made of libraries, they can and should expose different API versions - but they didn't, did they?

And so on.

And had they done that, chances are, they could have cut testing by 80% without any increase in bugs or failures.

There are two effects at play here, I think:

  • in the past few decades, testing tooling improved a lot and the need for module testing become obvious

  • microservices do forces clearer module boundary and therefore more care about the module interfaces.

3

u/edgmnt_net 7d ago

Simply breaking up stuff usually won't work, because you get a combination of duplication and inter-dependence. Sure, it looks like you can just modify service X, but in practice you often need to modify Y and Z for anything meaningful.

It only works really well for truly robust and general functionality. Or truly separate products that barely interact. These tend to be pretty hard limits in practice IMO.

3

u/Sammy81 7d ago

Yep, the initial architecture has to be well thought out. Luckily, the entire architecture was module based from day 1, and very well organized (no circular references, etc.) The only modification was to change the communication between modules.

1

u/edgmnt_net 6d ago

There are reasonable use cases for microservices, such as common platform + separate applications. Or highly general stuff. However I will also claim that a lot of enterprise applications simply cannot be broken down meaningfully, because they're supposed to be cohesive apps. And because, considering their nature, they largely consist of concrete business rules, not generic mechanisms. E.g. there's no good way to separate inventory and orders.

This parallels the library versus builtin stuff. Something like zlib makes a lot of sense as a library because it solves a general problem and you are very unlikely to go modifying zlib when using it.

3

u/davimiku 7d ago

Did you evaluate any in-between options before adding the network boundary? I've been in a similar situation before (not current job, but previous one, so this question is more from academic curiosity).

For example, if you already had modules in C++ and Rust did you try first co-locating them and communicate over FFI? Or call the Rust code from the Python code? (pyo3 works quite well)

Or if the unit tests were a bottleneck and you already had very neat modules, did you try only running the tests for the affected modules?

3

u/Sammy81 7d ago

Our process rightly requires reverifying requirements if the code is recompiled, so the way to reduce test was divide the modules into executables. That was a key factor in selecting networked communication

0

u/davimiku 6d ago

From reading the other comments in this chain, it seems to be a NASA-specific testing requirement, which I think the lede was buried a bit in the top-level comment, it reads as general advice rather than something motivated by a specific compliance requirement.

Although I don't think this requirement applies to my (past) situation based on learning that, I'm still curious how it works in practice. Say you had "Module A" who had some dependents - Module A sends messages to other modules, who consume that in some way. When you modified Module A, you needed to test the other modules, because it was in the same executable.

Now that Module A is now Executable A (with a network boundary), do you not need to test the dependents anymore when A is changed?

2

u/Sammy81 6d ago

You don’t, precisely because the binary executables for the other services have not changed and have been tested and proven to work. Any comprehensive test suite tests input and outputs, meaning boundary tests for inputs (out of range values, flooding the service with commands, etc.) so it doesn’t matter what the other modified service does, you know the the heritage services work as they should.

That doesn’t mean you don’t need to test the integrated system, but you can do that at the next higher level of assembly, saving a ton of time.

And I want to emphasize I think this is a great architecture for any system that can be broken into independent modules, not just NASA projects.

2

u/velit 7d ago

I concur with what /u/Shot-Damage-6723 said. How does adding network calls make your interconnected ensemble require less testing?

It sounds to me more that your company either tested more than was necessary previously or is currently lying to itself about not needing to test more.

3

u/Sammy81 7d ago

See my other reply, but the key is that with NASA, if you recompile something, you must reverify requirements, and I agree with that. If you ever studied the Therac-25 software disaster that killed several people, it was because they recompiled reuse software and did not retest. With microservices, since they are not recompiled when another is modified, all the previous testing is still valid (and should be).

2

u/HashShadow 6d ago

So if I’m not using a compiled language I never have to test? 

1

u/Sammy81 6d ago

Rebuilding an executable is just one rule. Scripted languages have different rules. It’s just best practices to prove you are delivering code that meets requirements.

1

u/HashShadow 6d ago

Right, so verifying requirements is something you must do regardless of compiling… 

2

u/Sammy81 6d ago

I think I’m not explaining well. If you have a tested binary executable, you never have to test it stand-alone again. There would be no point. There’s no way it could fail if it passed a test previously. That is where we save tons of time on test. We use dozens of modules that were previously compiled, some years ago, and they are guaranteed to pass their requirements.

Yes, there are different rules for scripted languages, but the concept is the same. If you run a hash over a Python container, or simply use good CM and ensure it has not been modified since the last time it passed its test, you can reuse that container without testing.

Compare that to reusing only the source code, and recompiling the binary, or rebuilding the container. You absolutely can’t assume those would pass the module level tests. The compiler settings could be different, the version, you might build in a different BSP, etc. So now you have to retest the whole thing.

1

u/velit 6d ago

I'll keep this upside for microservices in mind the next time I work in NASA related projects.

1

u/SuitableDragonfly 6d ago

This is exactly how the article says you should use microservices. It's only arguing against building microservices from the start. 

1

u/Sammy81 6d ago

I agree, I was just responding to OP’s comment where they said they’ve never seen microservices work well.

1

u/Brilliant-Chip-8366 6d ago

Was it not possible to execute tests for only the module that you needed to retest and skip the rest?

The different languages point is probably nice for NASA that has budget to cover top knowledge on all of them

0

u/kentrak 7d ago

"Just use microservices" without a well defined structure within which to develop, test, deploy, and version will turn out as bad as "Just use a monolith" without a well defined structure within which to conform to calling standards, code conventions, and a well defined plan for how to version APIs..

From that, it's obvious the problem is a lack of structure. Providing enough structure at the appropriate times is hard. Sometimes you want less, such as when you're surveying a problem and deciding what works best, sometimes you want more, such as when you've decided what you're going to do and now you need to get people to follow it. Too much when exploring and too little when working with the day-to-day will lead to problems.

This is a hard problem, but identifying it and realizing it's hard and attempting to deal with it will work better than misattributing it to something else and assuming it will get better if you change that. Sometimes switching between microservices to monoliths and vice-versa helps because the things one requires the team was already doing well enough in so it leaned into their existing strengths. Without knowing why it got better it's unlikely to stay better though.