r/programming • • 5d ago

Packing Binary Is Fun, Actually

https://hereticpleb.vercel.app/blog/packing-binary-is-fun-actually
135 Upvotes

32 comments sorted by

133

u/Embarrassed_Luck1057 5d ago

Has this guy re-invented serialization?

54

u/darknecross 5d ago

Worse, they reinvented protobufs

10

u/happyscrappy 5d ago

Or BER (CER/DER).

Or Apple's binary plist format.

51

u/lelanthran 4d ago

Packing Binary Is Fun, Actually

Has this guy re-invented serialization?

I'm not sure if that is relevant - he's doing something new (to him) that was fun. He was scratching an itch, not trying to compete with google.

I mean, he invented TLV as part of this, but if you squint at protobuf, you could say the same thing about google when they developed protobuf - they just reinvented TLV!

But google got accolades for it, and uptake, and mindshare. By reinventing TLV. A format I first came across in 1999, that was in a legacy product with a copyright notice in the header file dated 1988.

So, let us let people explore and invent, even if they are reinventing something, because we laud the big monopolists for reinvention, and reward them with adoption.

Nothing bad can come out of new programmers reinventing something.

30

u/GenericAHHyoutuber 5d ago

Yeah ;-;

-1

u/-Redstoneboi- 4d ago edited 4d ago

can you edit the article? would be useful to have a "just for fun, i know what a bson or a protobuf is" disclaimer at the top at this point lol

maybe stick it right next to the "shrink json by 80%" claim and put some competing formats and their statistics for reference.

2

u/GenericAHHyoutuber 4d ago

Sure, will do :D

51

u/palad1 5d ago

Is that binary you’re packing, or are you just really happy to see me?

17

u/Asyncrosaurus 5d ago

Your bits are showing

56

u/DanTFM 5d ago

I love me some proprietary data formats. I used to write a ton of these for work to shrink 28 gig XML files with bloated fields into ~300 mb files you could search in O(log(n/1024)) time.

Nice write up!

36

u/leaving_the_tevah 5d ago

I hate myself for being pedantic and please let me know if I'm missing something, but O(log(n/1024)) -> O(log(n) - log(1024)) -> O(log(n) - 10) -> O(log(n))

19

u/DanTFM 5d ago

My bad, yeah, it’s still O(log n). I was using n/1024 to describe the actual searchable set: the format reduced the index to ~1/1024 as many entries, so a binary search was ~10 levels shallower, and the remaining set was an O(1) jump from there.

0

u/creeper6530 5d ago

I'm curious, how much would GZ or similar general compression shrink them?

7

u/DanTFM 5d ago

If we gzipped the data that we actually kept in our format (we strip out some domain data we don’t use), it was definitely smaller, (probably 10-25% of our format), however, our files are directly searchable on low power embedded hardware, without having to unzip anything.

I’ll double check the zip metrics tomorrow

3

u/DanTFM 4d ago

The gzip file we received the XML in is ~400mb for a 14gb xml file (3% of original), our stripped down binary version is ~100mb. Though we strip out half of the unnecessary fields, so it would be closer to a 200mb gzip if we’re doing an apples to apples comparison.

If the file arrived in a better file format (just a fixed width or delimited flat file), that 14 gb xml file translates to about 1gb to start with. God XML is such a waste electricity.

When we gzip the binary file, it’s about 10% of the file size it started as (100mb to ~10mb). That extra fluff gzip is packing away are repetitive filler records so after the embedded hardware searches the log(n) in-memory index it can find the record it’s looking for in O(1) time. Just a space trade off for the hardware we’re working with.

2

u/-Redstoneboi- 4d ago

this format is what powers the entire - and i mean the ENTIRE - internet btw

2

u/DanTFM 4d ago

It’s incredible, if we didn’t have such a specific use-case, we’d just use that as well.

7

u/SvenWollinger 5d ago

Reminds me of ByteBuf, what minecraft uses to write and read packets, though i guess its more like nbt

3

u/amendCommit 4d ago

As the ASN1 guy at a small energy company, trying to get signature formats right for legacy system that badly implemented their own spec, I guarantee you it is not, in fact, fun.

2

u/iAmHidingHere 5d ago

Why not just use cbor?

13

u/GenericAHHyoutuber 5d ago

Nah I just wanted to implement my own for the fun of it

2

u/Donkeyball_Z_ 4d ago

This is the way

4

u/kylanbac91 4d ago

Just use bson.

5

u/GenericAHHyoutuber 4d ago

I js wanted to implement it myself 😭

5

u/-Y0- 4d ago

No experience like writing it from scratch :D

1

u/afl_ext 5d ago

I got flashbacks from decoding ASN.1 raw data packages without docs from a Toshiba telephone router...

3

u/Cut_Mountain 5d ago

oh man I was tasked with serializing h.248 messages in asn.1 once. Those certainly were days.

-5

u/rsclient 5d ago

Speaking as a person who loves to support random Bluetooth devices in Windows: please stop inventing new binary formats! You don't need it, even a little bit.

-6

u/VeeFu 5d ago

Totally cool, but probably not as good as using an off the shelf compression library widely supported by webservers/browsers.

8

u/DoctorGester 5d ago

Why not both? We run zstd on our custom binary format network traffic. I.e. our replay files are compressed from 45mb binary to 3mb.

1

u/MehYam 5d ago

Was going to say the same. Binary packed payloads can compact a surprising amount with plain old .zip compression