Skip to content
Segment without delimiters

Segment text with no delimiters

Segment text with no delimiters

Cut text that has no separators into units, where the text before a point stops predicting the text after it (seam).

English with the spaces gone

--english scores each cut against letter counts from a paragraph of English, so even a short string divides:

$ trex seam --english --text 'thebirdsflyoverthehouses'
trex seam: 24 bytes, order 3, 6 segments (bidirectional branching entropy)
  segments:
    [     0..3     ] "the"
    [     3..8     ] "birds"
    [     8..11    ] "fly"
    [    11..15    ] "over"
    [    15..18    ] "the"
    [    18..24    ] "houses"

A word whose letter clusters the paragraph does not hold can be cut inside. To split against a list of words you supply, parse with a dictionary.

Any other stream

Without --english the reader takes its counts from the input itself, so it needs enough text for its units to repeat. Score it against known boundaries with --recover, which removes a file’s spaces, segments what is left and compares the cuts with where the spaces were:

$ trex seam --recover --compare-bpe --max-bytes 16384 moby.txt
trex seam --recover: 16383 bytes, 3169 true boundaries, order 3, passes 1
  seam (no dictionary):   P=0.543 R=0.755 F1=0.632  (4405 cuts, 2393 hit)
  count-BPE (200 merges): P=0.365 R=0.939 F1=0.525  (8165 piece boundaries)
  -> seam recovers word boundaries +0.107 F1 over count-BPE: the bidirectional
     predictive signal (past<->future branching entropy) BPE has no access to.

On the first 131,072 bytes of the same text seam reaches F1 0.723. --order K sets the longest context it reads; the default is 3.

Cut records at the seams

--record seam makes each seam segment a record, so a query or a count reads the text in the units the reader found (record units).