Sam Rose @samwho.dev · May 6

I have the exact same payload encoded in JSON and as a protocol buffer. The JSON is 34kb, the protobuf is 15kb. I compressed it in a variety of different ways and was surprised to see that the JSON compresses to a smaller file more often than the protocol buffer does. Does this surprise you?

19 likes 9 replies

?

Replies

Mastro.{js,ts} · May 6

I'm guessing protobufs omit the keys (since they are in the schema), but they are already compressed. Compressing an already compressed artifact is never optimal.

Andri Óskarsson · May 6

Interesting. What is the trade-off in terms of serialization cost though?

Swizec Teller · May 6

JSON compresses really well because it’s a super repetitive format. Sounds like protobuf is naturally smaller (I haven’t used it much) so it stands to reason it wouldn’t compress as well. Big improvement you can make to JSON as a wire format is JSON-LD so you can stream

GSV(E) A Suffusion Of Diggity · May 6

Not much. Protobuf already minimizes the frequently repeated parts of the json file that are low hanging fruit for a compression algorithm, right? So I imagine the remaining variation is due to how the different compression algos respond to the particulars of your dataset?

Erwan Martin · May 6

I ran the same experiment years ago and I had the same conclusions. We ended up ditching protobuffs for our API.

Jim Jonah · May 6

I thought about moving the payload of my card game engine away from json but didn’t for the same reason, it was basically a wash but would lose readability.

It’s Joe! · May 6

Protocol buffers have a few other designed-in virtues around memory layout and parsing, so does make some sense (but yeah, it was surprising 😀)

P-Y · May 6

Now that we understand, an example: Let's take 2 timestamps 1700000000000 1700086400000 JSON ASCII (13B, alphabet=10): 31 37 30 30 30 30 30 30 30 30 30 30 30 31 37 30 30 30 30 38 36 34 30 30 30 30 proto varint (6B, alphabet=256): 80 d0 95 ff bc 31 80 88 af a8 bd 31

P-Y · May 6

Absolutely surprising. Folks often forget they can compress protos, which is useful if the field values are repetitive (e.g. large strings) I wouldn't expect compressed json to ever yield a smaller payload than compressed protom cc @swank.ca any idea?