Extract fields
Extract fields
Pull typed values, key and value pairs, and the columns of a table out of text.
One typed value
Name the atom for the kind: \E an email address, \N a number, \I an IP address, \U a
URL, \T a timestamp, \V a version; the token atoms
lists them all.
$ trex scan '\E' --text 'ping bob@x.com please'
[5..14] "bob@x.com"
Key and value pairs
Bind each side to a register. A literal = is written "=", since a bare = is the
back-reference operator.
$ trex scan '\W:k "=" \Q:v' --text 'name = "bob" city = "nyc"'
[0..12] "name = \"bob\"" captures: k="name", v="\"bob\""
[13..25] "city = \"nyc\"" captures: k="city", v="\"nyc\""
--json writes each match’s registers under captures.
The columns of a table
Write one row as a pattern, naming the columns to keep. ^ starts it at a line’s first token,
and the header row does not match, since age is no number:
$ cat people.csv
name,age,score
alice,30,95
bob,25,88
$ trex scan '^ \W:name "," \N "," \N:score' people.csv --format '${name} ${score}'
alice 95
bob 88
For lines whose layouts differ, build the pattern from marked examples instead of writing it.