Skip to content

Records

Records

trex reads an input as tokens, and cuts it into records: lines, paragraphs, blocks, the runs between two matches of a pattern, or units an axis finds with no delimiter at all. A record query asks of each record which of several patterns it holds. Templates and patterns built from records are on the building patterns page. The flags and parameters are on the CLI and PowerShell pages.

Tokens

Each token has a kind and a span; a typed token’s text parses as its kind. Whitespace tokens are kept in the stream, and a pattern never matches one.

let text = b"retry 3 times in 1500ms from 10.0.0.1";
let kinds: Vec<(&str, &[u8])> = trex::lexer::lex(text)
    .iter()
    .filter(|t| t.is_significant())
    .map(|t| (t.kind.name(), &text[t.span()]))
    .collect();
assert_eq!(kinds[4], ("duration", &b"1500ms"[..]));
assert_eq!(kinds[6], ("ip", &b"10.0.0.1"[..]));

Get-TrexToken leaves whitespace out unless -IncludeWhitespace asks for it. The kinds are the token atoms a pattern names.

Record units

UnitA record is
linea line, without its newline; the default
paragrapha run of lines holding something other than whitespace
filethe whole input
blocka balanced bracket group, from its opening bracket to its closing one; the one unit whose records nest
unit, unit:ROLEa supertoken, the statement, clause or argument list, of the role call, assign, kv, list, numeric or plain where one is named
periodthe stream’s own record period in supertokens; the whole input where it has none
seamthe seam axis’s segments, cut where the past stops predicting the future
bind, bind:Qthe bytes cut where they hold together least under the input’s pair field, as many cuts as seam makes, or a cut at every bond in the weakest Q per cent
autoseam or bind, whichever puts more of its cuts at the input’s line starts; seam where they tie
texturethe spectral texture regions
shapethe shape axis’s regions between its silhouette change-points

--record-start PATTERN makes a record of each run from one match of the pattern to the next, or to the end of the input, the bytes before the first match in no record; --record-span PATTERN makes each match a record.

use trex::records::RecordUnit;

let text = b"a 1\nb 2\n\nc 3";
assert_eq!(RecordUnit::Paragraph.records(text), [(0, 7), (9, 12)]);
assert_eq!(RecordUnit::parse("line").expect("a unit").records(text), [(0, 3), (4, 7), (8, 8), (9, 12)]);

The same units cut a window, the context printed around a match, and the records a count or a table counts.

Record queries

A query keeps the records holding every pattern (--all), at least one (--any), none (--none) or at least N (--at-least N), and drops a record holding a pattern --not names. A record definition or --not with no quantifier is --any. Every pattern is scanned once over the whole input, and a record holds a pattern where one of its matches overlaps the record, so ^ and $ read lines and \A and \z the input whatever the record is.

$ cat notes.txt
from 10.0.0.1
to bob@x.com

from 10.0.0.2
nothing

mail amy@y.org
$ trex scan --all -e '\I' -e '\E' --record paragraph notes.txt
from 10.0.0.1
to bob@x.com

$ trex scan --any -e '\E' notes.txt
to bob@x.com
--
mail amy@y.org

$ trex scan --any -e '\I' -e '\E' --record paragraph --json notes.txt
[{"line":1,"start":0,"end":26,"text":"from 10.0.0.1\nto bob@x.com","patterns":[0,1]},{"line":4,"start":28,"end":49,"text":"from 10.0.0.2\nnothing","patterns":[0]},{"line":7,"start":51,"end":65,"text":"mail amy@y.org","patterns":[1]}]

A record a query keeps prints as the lines it covers, each path:line:text over named inputs and bare over one, with -- between records that do not touch; --count counts the records, -l and -L name the files, -m N keeps the first N records, and --json gives one object a record with its first line, byte span, text and the patterns present, by index in the order given or by name under --patterns. A query takes no --format, --explain, context lines, -v, --passthru, -x or --single-match.

Entries of several lines

A timestamp at the head of a line starts a log entry, so --record-start '^ \T' makes each entry, however many lines it runs, one record:

$ cat events.log
2026-09-15T10:00:00Z login ok
  user bob@x.com from 10.0.0.1
2026-09-15T10:01:00Z login failed
  user amy@y.org from 10.0.0.2
2026-09-15T10:02:00Z logout
  user bob@x.com
$ trex scan --all -e '\E' -e '\I{in:10.0.0.0/24}' --record-start '^ \T' events.log
2026-09-15T10:00:00Z login ok
  user bob@x.com from 10.0.0.1
2026-09-15T10:01:00Z login failed
  user amy@y.org from 10.0.0.2

$ trex scan --none -e '\I' --record-start '^ \T' events.log
2026-09-15T10:02:00Z logout
  user bob@x.com

$ trex scan --any -e '\E' --record-start '^ \T' --count events.log
3

A supertoken is a record with no delimiter: unit:assign makes every binding one, whatever the language.

$ cat code.txt
x = 1
foo(3)
y = 3
$ trex scan --any -e '\N{>=3}' --record unit:assign code.txt
y = 3