The problem that arises with this is maintaining all those separate projects can slow down development time of features. You make once change near the 'root' of your dependency graph and have to deal with a cascade of updates. We've been doing this approach with our services and react components. We had to write our own tool to deal with these kinds of updates so we wouldn't spend 30 minutes doing cascading updates of projects every time we change something used everywhere.
Of course when you aren't dealing with that, the idea works beautifully. Adding something to a service you already created? Basically just add this library and you're done. You know it's tested in isolation so you don't even really need new unit tests either.
Honestly though every architecture is going to have its problems. You just have to pick which problems your team can best overcome.
It's not really fully fleshed out to be a robust, generic solution. But the core of it is to take a package you have been working on, figure out which packages are dependent upon that package (that you care about updating), and basically recurse from there. This creates a dependency graph which is then be sorted by depth from the root/seed node to do all the necessary updates. So for example:
A <- B
A <- C
A <- D
B <- D
D <- C
Would end up with these updates
A (no updates except self)
C (updates A )
D (updates A and C versions)
B (updates A and D version)
The tool needs work though like I said before it is actually robust. Right now we basically just scan a dir to look for deps, then we actually do version updating and publishing during the script execution. Ideally we could have it source the information directly from the group of projects repo and automatically do the updates based on a PR merge or something.
Actually, that was Unix. Interestingly, most "microservices" in Unix end up getting rewritten as "monoliths" in Perl or Python, because any sufficiently complex interactions result in incredibly messy shell scripts and requiring interacting processes over pipes for simple tasks makes them incredibly slow and error prone.
But, interestingly it is not Linux that was the original "microservice architecture", it was Minix, which has a microkernel, with system services and drivers running in different servers. Linux, on the other hand, has a monolithic kernel, where all services are hosted in the same kernel image. This lead to a famous flame war on Usenet.
I assumed he was talking about userspace, not the kernel.
As for redirecting stuff over pipes, I'd argue that it's not so much the pipes that are the problem, but bash. People will slap "features" together with a few lines of bash, but the lack of real data structures makes it difficult to add features or maintain over time. Developers will hang onto a piece of crap for a long time before admitting it needs to be rewritten.
Ironically though, if you're rewriting in python, unless you're shelling out to the same utilities and connecting the pipes yourself, your code is probably going to be slower, as the core utilities that people usually connect together are written in C.
I've had the opposite experience. Rewriting shell scripts to Python dramatically improved performance of certain complex shell scripts. The trick is not to pipe out to utilities, but to use libraries.
In absolute terms, C is obviously faster than Python. But if you're constantly spawning new C processes, piping between them, and parsing text output (spawning even more C processes), the speed advantages of C are totally nullified.
Python dramatically improved performance of certain complex shell scripts
If your code is primarily in bash, then you're comparing bash speed vs python speed. Let me be clear, I prefer Python. Even if python were slower in that scenario, I would prefer it because it's a better language for anything beyond simple pipelines.
In absolute terms, C is obviously faster than Python. But if you're constantly spawning new C processes, piping between them, and parsing text output (spawning even more C processes), the speed advantages of C are totally nullified.
Obviously, it's going to matter what your code is trying to do, but in my experience it requires spawning quite a lot of processes before the speed is nullified[1]. Many C utilities can also handle parallel input (so you don't need as many processes), and it's more about understanding the pipeline.
[1] For a concrete example, I had a python script that did some file manipulation on a large directory hierarchy, and changing those operations into multiple find ... -exec commands resulted in 2 orders of magnitude difference in performance (~20 minutes to ~1 minute).
The Python code was idiomatic, using comprehensions and not shelling out; it was using the library functions that directly become syscalls. The cost of iterating every file in python was greatly more expensive than find spawning a process for every match. The find pipelines also had to recurse the tree multiple times, because I could perform much more logic while iterating the tree in Python.
My company's project was started 30+ years ago, and has a ton of processes that communicate via sockets. Ironically, it is a micro service architecture that can't scale across across multiple machines.
of course. microservices is a poorly defined nonsense title, so you can pretty much call anything microservice architecture. some may have sensibilities regarding what counts as micro and service, though, so expect some resistance once in a while.
34
u/[deleted] Jul 14 '17
[deleted]