Compression
BAMFS stores already-compressed bytes and commits to both the stored and the logical bytes.
BAMFS stores already-compressed bytes. A file commits to its stored bytes
through the chunk hashes, and — when compressed — to its decompressed bytes
through a second commitment, outputHash.
Compress off-chain, upload, inflate on read, verify the inflated result.
The two commitments
function createFile(
bytes32[] calldata chunkHashes,
string calldata mimeType,
uint8 compression,
bytes32 outputHash
) external returns (bytes32 cid);outputHash is keccak256(logicalBytes). It must be bytes32(0) when
uncompressed and non-zero otherwise; either mistake reverts
InvalidOutputHash.
Both the codec id and outputHash feed the CID, so two files with the same
stored bytes but different codecs — or the same codec and a different claimed
output — have different CIDs. See the CID formulas.
Codec ids
The compression byte is a bit-field: the high bit picks the mode, the low
seven bits pick the algorithm.
[ bit 7 = mode ][ bits 0..6 = algorithm ]
0 → per-file compress the whole file, then chunk the output
1 → per-chunk compress each chunk independentlyA per-chunk variant is always 0x80 | <per-file id>, so the mode is derivable
from the id alone.
| Algorithm | Per-file | Per-chunk | Inflates on-chain |
|---|---|---|---|
none | 0x00 | — | n/a |
gzip | 0x01 | 0x81 | no |
brotli | 0x02 | 0x82 | no |
zstd | 0x03 | 0x83 | no (reserved) |
lz4 | 0x04 | 0x84 | no (reserved) |
deflate | 0x05 | 0x85 | no (reserved) |
fastlz | 0x06 | 0x86 | yes |
| app-defined | 0x40–0x7F | 0xC0–0xFF | if you register an IDecompressor |
0x07–0x3F are reserved for future well-known codecs. Pick from the
app-defined range and make sure every reader agrees on the mapping.
Why FastLZ is the default
FastLZ is the only codec with a Solidity decompressor, which makes it the only
one where the chain can inflate the stored bytes and derive the IPFS CID itself.
gzip and brotli compress harder but revert CompressionNotSupported on-chain —
their IPFS CID has to be computed off-chain instead.
| Codec | Ratio | Encode | Decode | Use for |
|---|---|---|---|---|
none | 1.0× | — | — | Already-compressed media. Maximum dedup. |
fastlz | ~1.5–3× | fast | fast | Compressible text with on-chain IPFS parity. |
gzip | ~2–4× | medium | fast | Generic text, when off-chain CID derivation is fine. |
brotli | ~3–5× | slow | medium | Long-lived static assets. |
Already-compressed binary assets should use none so their chunks deduplicate
across files.
Read and write semantics
Writers compress, then chunk — the chunk hash is over the compressed
bytes — and pass outputHash for the logical bytes.
Readers fetch chunks, concatenate, decompress, and verify against
outputHash. The SDK raises OutputHashMismatchError on a mismatch.
The SDK's readFile always inflates client-side, so it handles every codec.
The on-chain FileStore.readFile is for on-chain consumers and only inflates
codecs with a registered IDecompressor.
--compress must match between upload and diff. The CID commits to the
stored bytes, so a mismatch re-derives every file under the wrong codec and
reports the whole tree as changed.
