Aggregates
Aggregates
A table groups the matches of a pattern by a key rendered at each match, counts each group, and
can aggregate a register’s typed values in it. count-by orders the rows by key, top by
count with the key breaking ties, and uniq prints the keys alone. The flags and parameters
are on the CLI and
PowerShell pages.
The examples read these files:
$ cat access.log
10.0.0.5 GET /index.html 200 5120
10.0.0.7 GET /login 302 0
10.0.1.9 GET /missing 404 312
10.0.0.5 POST /login 200 88
$ cat sizes.txt
alpha 1000B 80ms
alpha 2000B 80ms
alpha 4000B 80ms
beta 8000B 300ms
Count by a key
The key is a report template, as scan --format takes one, so a register’s typed slice is
what the rows count: ${ip:octet1-3}, ${u:host}, ${e:domain}, or a composite such as
${m:upper}/${u:host}.
$ trex top '\I:ip \W:verb' '${verb}' access.log
GET 3
POST 1
$ trex count-by '\I:ip' '${ip:octet1-3}' access.log
10.0.0 3
10.0.1 1
$ trex uniq '\I:ip' '${ip}' access.log
10.0.0.5
10.0.0.7
10.0.1.9
Every key is printed. -n N cuts the table to N rows and says how many keys and matches there
were, so a short table never reads as a complete one; Python returns every row and
-MaxCount writes the first groups. A key an accessor leaves empty prints as - and is
counted, so the rows sum to the matches.
Aggregates
--sum, --avg, --min, --max, --p50 and --p95 each take one register and add a
column, computed in the value’s base unit through the parse a value predicate compares with.
--sum and --avg take the numeric kinds: number, byte size, duration, money and percent.
--min, --max and the percentiles also take timestamp, version and address. An aggregate a
kind cannot carry is refused before anything is scanned, naming the register, its kind and the
kinds the aggregate takes; a key whose matches bound no value prints -.
$ trex count-by '\W:h \Z:s' '${h}' sizes.txt --sum '${s}' --min '${s}' --max '${s}'
count --sum s --min s --max s
alpha 3 7000 1000 4000
beta 1 8000 8000 8000
$ trex count-by '\W:h \Z \R:d' '${h}' sizes.txt --sum '${d}' --duration-unit ms
count --sum d
alpha 3 240
beta 1 300
An aggregate names one plain register: ${u:host} is a slice of text and ${a}/${b} names two
values, so both are refused. In PowerShell each aggregate is a property of the group in the
.NET type that holds its kind: a number a long or a decimal, a duration a TimeSpan, a
timestamp a DateTimeOffset.
Averages and percentiles
--avg is exact: --avg-form repetend (the default) brackets the repeating digits and
rational writes the fraction in lowest terms. A percentile falling between two observed
values is decided by --percentile: nearest (the default) takes the value at
ceil(p x n) in sorted order, linear interpolates and takes numeric kinds only, lower
takes the observed value at or below, and hybrid interpolates where the kind allows it and
takes an observed value where it does not.
$ trex count-by '\W:h \Z:s' '${h}' sizes.txt --avg '${s}'
count --avg s
alpha 3 2333.(3)
beta 1 8000
$ trex count-by '\W:h \Z:s' '${h}' sizes.txt --avg '${s}' --avg-form rational
count --avg s
alpha 3 7000/3
beta 1 8000
$ trex count-by '\N:n' all --text '10 20 30 40' --p50 '${n}'
count --p50 n
all 4 20
$ trex count-by '\N:n' all --text '10 20 30 40' --p50 '${n}' --percentile linear
count --p50 n
all 4 25
Keys by place and by member
${path} counts the matches by input and ${line} by line. Over a set, from --patterns FILE
or several patterns, ${pattern} counts them by the member that made them.
$ cat rules.trex
# what a line may hold
let host = \I:addr
let mail = \E:e
\N{>=100}
$ cat notes.txt
from 10.0.0.1 at 500 to bob@x.com
x = 7 and 10.0.0.2
$ cat other.txt
mail amy@y.org
$ trex count-by --patterns rules.trex '${pattern}' notes.txt other.txt
4 1
host 2
mail 2
$ trex count-by '\E' '${path}' notes.txt other.txt
notes.txt 1
other.txt 1
JSON and windows
--json prints the rows as an array of objects, each aggregate under its flag’s name and a
value with no aggregate as null; --values and --duration-unit spell the values as scan --json spells them. --head N, --tail N and --lines A..B group the matches of one
window of each input, and --record UNIT counts that window in records.
$ trex top '\I:ip' '${ip}' access.log --tail 2
10.0.0.5 1
10.0.1.9 1