Your first patterns
Your first patterns
The atoms, alternatives, repetition and the named structural shapes called lenses: the part of trex that overlaps with a regex.
The token atoms
Each atom matches one token of a kind:
| Atom | Matches | Atom | Matches |
|---|---|---|---|
\N | a number | \U | a URL |
\W | a word or identifier | \T | a timestamp or date |
\Q | a quoted string | \P | a punctuation token |
\I | an IP address | . | any one token |
\E | an email address | "lit" | a token equal to lit |
The pattern syntax lists versions, UUIDs, money,
durations and the rest. An email is one token, so \E matches it whole.
Alternatives and repetition
A | B matches either; + one or more, * zero or more, ? one or none, and {m,n} between m
and n:
$ trex scan '\N | \E' --text 'id 7 mail a@b.com'
[3..4] "7"
[10..17] "a@b.com"
$ trex scan '\N+' --text 'coords 1 2 3 stop'
[7..12] "1 2 3"
$ trex scan '\N{2,3}' --text '1 2 3 4 5'
[0..5] "1 2 3"
[6..9] "4 5"
+ takes as many as it can, so \N+ is one match of three numbers, and {2,3} cuts a longer run
into the largest pieces it allows.
Lenses
A lens names a structural shape found across languages. @call is an identifier followed by a
balanced parenthesis group:
$ trex scan '@call' --text 'foo(1) bar(2, 3)'
[0..6] "foo(1)"
[7..16] "bar(2, 3)"
@block, @kv, @flag, @list and @range name the others (lenses).
Where next
Binding and balance covers named registers with back-references and balanced brackets.