invariantify looks at the structure in your data before choosing how to compress it. Different data compresses differently, so we measure by domain and publish the results below, including the worst one.
| No. | Data domain | Measured result | Remark |
|---|---|---|---|
| 01 | Structured data (JSON) | 17.3× | Full 912 MB GitHub Archive event dump, not a sample |
| 02 | Backups and version history | 71.4× | 20-version incremental backup set, cross-file dedup |
| 03 | vs. zstd, xz, bzip2, brotli (max level) | 0 losses | 45 wins, 20 ties across 65 real files; up to 1.75× tighter on the best case |
| 04 | Documents (PDF) | 3.8–6.2× | Varies by how much of the PDF is already-compressed image data |
| 05 | Log files (Apache, HDFS, Linux, OpenSSH, Spark, Zookeeper, HealthApp) | 12.3–56.6× | 7 real production log formats from the LogHub benchmark set |
| 06 | Already-compressed binary media (JPEG, PNG) | 1.0× | The floor. Measured on all 24 Kodak reference images. Compressed data leaves nothing to gain. |
All values are measured, not projected, logged run by run. Row 06 is the worst case and it is printed on purpose.
| File | Original | invariantify | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| bib | 111,261 | 22,745 | 4.89× | 4.05× (bzip2) |
| book1 | 768,771 | 204,277 | 3.76× | 3.30× (bzip2) |
| book2 | 610,856 | 134,552 | 4.54× | 3.88× (bzip2) |
| geo | 102,400 | 48,149 | 2.13× | 2.12× (brotli) |
| news | 377,109 | 99,775 | 3.78× | 3.34× (brotli) |
| obj1 | 21,504 | 9,357 | 2.30× | 2.30× (brotli) |
| obj2 | 246,814 | 61,472 | 4.02× | 4.02× (xz) |
| paper1 | 53,161 | 13,161 | 4.04× | 3.44× (brotli) |
| paper2 | 82,199 | 20,322 | 4.04× | 3.31× (brotli) |
| paper3 | 46,526 | 12,214 | 3.81× | 3.17× (brotli) |
| paper4 | 13,286 | 3,655 | 3.64× | 3.09× (brotli) |
| paper5 | 11,954 | 3,661 | 3.27× | 2.92× (brotli) |
| paper6 | 38,105 | 9,669 | 3.94× | 3.42× (brotli) |
| pic | 513,216 | 33,909 | 15.14× | 15.14× (bzip2) |
| progc | 39,611 | 9,905 | 4.00× | 3.40× (brotli) |
| progl | 71,646 | 12,037 | 5.95× | 5.11× (brotli) |
| progp | 49,379 | 8,572 | 5.76× | 4.99× (brotli) |
| trans | 93,695 | 13,450 | 6.97× | 6.08× (brotli) |
| File | Original | invariantify | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| alice29.txt | 152,089 | 35,751 | 4.25× | 3.52× (bzip2) |
| asyoulik.txt | 125,179 | 33,369 | 3.75× | 3.16× (bzip2) |
| cp.html | 24,603 | 6,189 | 3.98× | 3.56× (brotli) |
| fields.c | 11,150 | 2,268 | 4.92× | 4.08× (brotli) |
| grammar.lsp | 3,721 | 1,012 | 3.68× | 3.26× (brotli) |
| kennedy.xls (spreadsheet) | 1,029,744 | 26,706 | 38.56× | 38.56× (bzip2) |
| lcet10.txt | 426,754 | 90,473 | 4.72× | 3.96× (bzip2) |
| plrabn12.txt | 481,861 | 129,378 | 3.72× | 3.31× (bzip2) |
| ptt5 (fax scan) | 513,216 | 33,909 | 15.14× | 15.14× (bzip2) |
| sum | 38,240 | 9,516 | 4.02× | 4.02× (xz) |
| xargs.1 | 4,227 | 1,307 | 3.23× | 2.86× (brotli) |
| File | Original | invariantify | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| reymont (Silesia corpus, Polish novel) | 6,627,202 | 1,068,830 | 6.20× | 5.34× (bzip2) |
| attention.pdf (public arXiv paper) | 2,215,244 | 584,284 | 3.79× | 3.79× (xz, tie) |
| File | Original | invariantify | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| Apache_2k.log | 171,239 | 3,025 | 56.61× | 34.14× (zstd) |
| OpenSSH_2k.log | 225,216 | 5,162 | 43.63× | 30.68× (xz) |
| Linux_2k.log | 216,485 | 6,632 | 32.64× | 26.77× (zstd) |
| Spark_2k.log | 196,268 | 6,611 | 29.69× | 27.93× (xz) |
| HealthApp_2k.log | 187,456 | 7,252 | 25.85× | 20.16× (xz) |
| Zookeeper_2k.log | 279,891 | 11,274 | 24.83× | 18.85× (bzip2) |
| HDFS_2k.log | 287,848 | 23,317 | 12.35× | 10.24× (zstd) |
| File | Original | invariantify | Ratio | Best of zstd/xz/bzip2/brotli |
|---|---|---|---|---|
| kodim01–24.png (all 24 Kodak reference images, identical result) | 736,501 avg | 736,517 avg | 1.0000× | ≤1.00× (none compress it either) |
| kodim01/05/12/18.jpg (4 derived JPEGs, identical result) | 182,274 avg | 182,290 avg | 0.9999× | ≤1.00× (none compress it either) |
Petabytes of user files and artifacts, most of them structured or repetitive. Across 65 real files tested at max settings, invariantify never lost to standalone zstd, xz, bzip2, or brotli, so switching is never a downside on the files those tools already handle well.
Version history is the most redundant data there is, which is why it is our best measured result at 71.4x on a 20-version backup set. Longer retention in the same footprint, and restores move less data over the wire.
JSON events are highly structured, measured at 17.3x on a full 912 MB dump, not a favorable sample. Application logs are even more repetitive: up to 56.6x on real Apache logs. Compressing them well cuts both the storage bill and the egress bill, at every hop where the data sits or moves.
On already-compressed binary media such as JPEG and PNG, invariantify measures 1.0x across all 24 images in the standard Kodak reference set. That is the honest floor, and it is on this page because a benchmark table without a worst case is an advertisement, not a measurement. PDFs turned out not to belong in this floor: two real documents we tested measured 3.8x and 6.2x, since a PDF's own internal structure varies far more than an already-DCT-compressed image does. If most of your data is genuinely already-compressed pixels, we will tell you the gains are modest before you spend a day integrating anything.
zstd and xz are excellent general-purpose compressors and we do not pretend otherwise. invariantify differs in one way we can state without disclosing the method: it looks at the structure in your data before choosing how to compress it, then still checks its choice against zstd, xz, bzip2, and brotli at their own maximum settings. Across 65 real files tested this way, it never lost: 45 outright wins, 20 ties, up to 1.75x tighter than the best single alternative on the best case. A full methodology will be published with the evaluation build so you can reproduce the comparison on your own data.
Not today. The core is proprietary. We plan a public evaluation build so you can verify these numbers on your own data before committing to anything, which we think matters more than reading our source.
On genuinely already-compressed binary media, such as JPEG and PNG, you get about 1.0x, the floor in the table above. PDFs are a partial exception, since they often carry their own compressible structure: two real PDFs we tested measured 3.8x and 6.2x. If your data is mostly already-compressed pixels rather than documents, the size gains alone probably do not justify a migration, and we will say so when you tell us what you store.
We are pre-launch. Private pilots come first, prioritized by data type, since the gains depend on what you store. Leave your details below and tell us what your data looks like; that is genuinely the fastest route in.
We reply with what the measured numbers say about your data type, including when the answer is that the gains would be modest. No drip campaign follows.