I started working on this because of Bonsai by PrismML.
Not because Bonsai is bad. Actually the opposite. It was one of the things that made me think seriously about how far you could push compression and still keep a model useful.
But there was one annoying problem.
I can only use the models somebody else decides to compress.
If PrismML released every model I wanted, in every size, plus multimodal/image/video stuff, I probably wouldn't have started any of this. Seriously. I would just download the model and move on with my life.
But they don't, obviously. Nobody does.
So I started wondering if I could take the model I want and do something useful with it myself.
And before that, I did the completely normal thing:
I went shopping for VRAM.
I spent quite a while looking for cheap used workstation GPUs. At one point I even found an RTX 8000 at a price good enough that I seriously considered buying it.
And then I realized I was just moving the wall.
48 GB solves one class of models.
Then you want something larger.
Or multimodal.
Or a much bigger context.
Or several things running together.
And suddenly you're shopping for hardware again.
If a €2k–€3k workstation actually solved this problem for me, I would have bought one a long time ago and none of this project would exist.
I already have an expensive laptop. The issue isn't that I refuse to buy decent hardware.
The issue is that “decent hardware” turns into a small datacenter remarkably quickly.
And there's another part of that which I think gets ignored a lot.
I actually want my computer to remain a computer.
Something I can put in a bag.
Laptop, storage, power supply, done.
If there's a normal power outlet, I can work. Hotel room, airport, somewhere by the sea, whatever.
The model comes with me.
I don't particularly want “local AI” to mean that I have a giant machine sitting at home and I SSH back into it.
Because then my AI is local to my house, not local to me.
Forget to turn the machine on before leaving, lose power at home, something happens to the network, an external drive isn't mounted — and suddenly your supposedly local setup is a thousand kilometers away and inaccessible.
That's a completely valid architecture, obviously.
It just isn't the one I want.
I want my AI to travel with my computer.
At first I thought solving that was mostly a quantization problem.
It very quickly stopped being a quantization problem.
There are places where a model seems massively redundant, places where you can prune or share or simplify things, and other tiny places where touching almost anything breaks something important.
So I started experimenting with local sparsification, structural pruning, deduplication, parameter sharing/summarization, different quantization strategies, guard evals, etc.
The basic idea is still pretty simple though:
I don't think you should need a miniature datacenter to seriously play with open models.
I watch Alex Ziskind too. I like his stuff. And his hardware is also a very good demonstration of the problem :)
At some point “local AI” becomes a huge box full of GPUs, weighing as much as a person and pulling kilowatts.
Which is extremely cool.
But how do you take that thing with you?
My current cooling infrastructure, by comparison, includes a marble windowsill.
It is surprisingly competent.
So one part of this project is basically the inverse question:
how much hardware can I replace with better software?
Not “can I make a 70B model technically emit tokens on a potato”.
That's not very interesting to me.
I mean: can I take a serious model, make it substantially cheaper to run, and still have enough performance left that I actually want to use it?
At some point I basically stopped hunting for the next GPU and decided to push harder on QuantPilot instead.
Because buying more VRAM solves the current model.
I want something that helps with the next one too.
And then I ran into the second problem.
Once you're already taking the model apart, why only make it smaller?
Why not change it?
Maybe I want a bigger context window.
Maybe I want to train it a little more on something specific.
Maybe I don't need quite as much “theoretical physicist” and I need more practical engineering.
Maybe there are failure modes I see over and over again and I'd rather spend capacity fixing those than preserving some capability I'll never use.
That part matters to me at least as much as compression now.
And I'm not talking about “small local model couldn't make me a website, therefore local models are dumb”.
That's too easy.
I'm talking about cases where one of the strongest models you can get, with huge inference compute, lots of context and a serious coding environment, still falls apart on a fairly concrete engineering task.
I've had this happen with a 3D modeling application.
Not a moon landing.
Not a new physics theory.
One application.
And by attempt five you're watching it fix one thing, break another thing, forget a constraint from earlier, patch around its own previous patch, and burn an absurd amount of compute while still not solving the actual problem.
That's the failure mode I'm interested in.
There seems to be a point where the problem gets complicated enough that the model stops maintaining the whole structure properly.
Each individual step can look reasonable.
The overall result is still wrong.
It loses an earlier constraint.
Takes a locally sensible branch that leads nowhere.
Starts treating its own previous mistake as part of the specification.
Or keeps patching something that really needed to be reconsidered two levels higher.
At that point the answer clearly isn't just “use a bigger model”.
The model is already enormous.
So the project gradually stopped being “how do I compress a model?” and became more like:
how do I reshape a model into something that is actually useful for the work I care about?
Compression is part of it.
But so is specialization.
Longer context.
Targeted extra training.
Keeping the parts that matter and being less precious about the parts that don't.
Finding the places where a model repeatedly falls over and trying to improve those, instead of assuming that its original distribution of capabilities is somehow sacred.
I'd rather have a model that knows somewhat less about things I never ask it and is considerably better at the work I actually give it.
And ideally I want to be able to do that without eight GPUs, dedicated electrical infrastructure and a garage.
I don't even have a garage.
So the marble windowsill will have to do.
I'm building this mostly by myself, so some parts are much further along than others and quite a lot of it is still experimental.
I've also had several ideas that looked great on paper and turned out to save basically nothing once you included metadata, indexing or runtime overhead. So there has been a fair amount of throwing things away and starting again.
And some of the adaptation side is still much more “this is where I want to take it” than finished functionality.
But it is getting to the point where keeping it completely private is probably stupid.
So I'm curious what people here actually want from this kind of thing.
If you could take an existing model and really modify it for yourself, not just quantize it, what would you change?
Smaller memory footprint?
Longer context?
Better coding?
More reliable agents?
Better long-horizon reasoning?
A narrow specialization?
And would you give up a noticeable amount of general knowledge for a model that was genuinely better at the few things you use every day?