Skip to content
Compress text

Measure how far text compresses

Measure how far text compresses

Read the code length trex’s context-mixing model gives a text: the size a coder driven by that model would write, in bits per byte. The model reads tokens, orbits and a prior learned from a corpus. No compressed stream is written (compression).

compress needs a build with the compress feature:

cargo build --release --features compress

Read the code length

$ trex compress --text 'the quick brown fox jumps over the lazy dog the quick brown fox jumps'
trex compress: 69 bytes -> 15 bytes  (1.626 bits/byte, 21.7% of original)
  coder: logistic mix + orbit + baked prior   0.00 MB/s   (--compare for the full table)

The size is the code length in bits rounded up to bytes.

Compare the models

--compare prints every model’s bits per byte and speed beside a byte-pair-encoding baseline, and --no-baked reads with no prior, which shows what the prior is worth on your text. --chunks N, --gpu and --hybrid split the work for speed at some cost in length, as choose a backend measures.