Skip to content
Extract fields

Extract fields

Extract fields

Pull typed values, key and value pairs, and the columns of a table out of text.

One typed value

Name the atom for the kind: \E an email address, \N a number, \I an IP address, \U a URL, \T a timestamp, \V a version; the token atoms lists them all.

$ trex scan '\E' --text 'ping bob@x.com please'
[5..14] "bob@x.com"

Key and value pairs

Bind each side to a register. A literal = is written "=", since a bare = is the back-reference operator.

$ trex scan '\W:k "=" \Q:v' --text 'name = "bob" city = "nyc"'
[0..12] "name = \"bob\""  captures: k="name", v="\"bob\""
[13..25] "city = \"nyc\""  captures: k="city", v="\"nyc\""

--json writes each match’s registers under captures.

The columns of a table

Write one row as a pattern, naming the columns to keep. ^ starts it at a line’s first token, and the header row does not match, since age is no number:

$ cat people.csv
name,age,score
alice,30,95
bob,25,88
$ trex scan '^ \W:name "," \N "," \N:score' people.csv --format '${name} ${score}'
alice 95
bob 88

For lines whose layouts differ, build the pattern from marked examples instead of writing it.