BAMFS
Use BAMFS

Compression

BAMFS stores already-compressed bytes and commits to both the stored and the logical bytes.

BAMFS stores already-compressed bytes. A file commits to its stored bytes through the chunk hashes, and — when compressed — to its decompressed bytes through a second commitment, outputHash.

Compress off-chain, upload, inflate on read, verify the inflated result.

The two commitments

function createFile(
    bytes32[] calldata chunkHashes,
    string calldata mimeType,
    uint8 compression,
    bytes32 outputHash
) external returns (bytes32 cid);

outputHash is keccak256(logicalBytes). It must be bytes32(0) when uncompressed and non-zero otherwise; either mistake reverts InvalidOutputHash.

Both the codec id and outputHash feed the CID, so two files with the same stored bytes but different codecs — or the same codec and a different claimed output — have different CIDs. See the CID formulas.

Codec ids

The compression byte is a bit-field: the high bit picks the mode, the low seven bits pick the algorithm.

[ bit 7 = mode ][ bits 0..6 = algorithm ]
  0 → per-file   compress the whole file, then chunk the output
  1 → per-chunk  compress each chunk independently

A per-chunk variant is always 0x80 | <per-file id>, so the mode is derivable from the id alone.

AlgorithmPer-filePer-chunkInflates on-chain
none0x00—n/a
gzip0x010x81no
brotli0x020x82no
zstd0x030x83no (reserved)
lz40x040x84no (reserved)
deflate0x050x85no (reserved)
fastlz0x060x86yes
app-defined0x40–0x7F0xC0–0xFFif you register an IDecompressor

0x07–0x3F are reserved for future well-known codecs. Pick from the app-defined range and make sure every reader agrees on the mapping.

Why FastLZ is the default

FastLZ is the only codec with a Solidity decompressor, which makes it the only one where the chain can inflate the stored bytes and derive the IPFS CID itself. gzip and brotli compress harder but revert CompressionNotSupported on-chain — their IPFS CID has to be computed off-chain instead.

CodecRatioEncodeDecodeUse for
none1.0×——Already-compressed media. Maximum dedup.
fastlz~1.5–3×fastfastCompressible text with on-chain IPFS parity.
gzip~2–4×mediumfastGeneric text, when off-chain CID derivation is fine.
brotli~3–5×slowmediumLong-lived static assets.

Already-compressed binary assets should use none so their chunks deduplicate across files.

Read and write semantics

Writers compress, then chunk — the chunk hash is over the compressed bytes — and pass outputHash for the logical bytes.

Readers fetch chunks, concatenate, decompress, and verify against outputHash. The SDK raises OutputHashMismatchError on a mismatch.

The SDK's readFile always inflates client-side, so it handles every codec. The on-chain FileStore.readFile is for on-chain consumers and only inflates codecs with a registered IDecompressor.

--compress must match between upload and diff. The CID commits to the stored bytes, so a mismatch re-derives every file under the wrong codec and reports the whole tree as changed.

On this page