Skip to content
Aggregates

Aggregates

Aggregates

A table groups the matches of a pattern by a key rendered at each match, counts each group, and can aggregate a register’s typed values in it. count-by orders the rows by key, top by count with the key breaking ties, and uniq prints the keys alone. The flags and parameters are on the CLI and PowerShell pages.

The examples read these files:

$ cat access.log
10.0.0.5 GET /index.html 200 5120
10.0.0.7 GET /login 302 0
10.0.1.9 GET /missing 404 312
10.0.0.5 POST /login 200 88
$ cat sizes.txt
alpha 1000B 80ms
alpha 2000B 80ms
alpha 4000B 80ms
beta 8000B 300ms

Count by a key

The key is a report template, as scan --format takes one, so a register’s typed slice is what the rows count: ${ip:octet1-3}, ${u:host}, ${e:domain}, or a composite such as ${m:upper}/${u:host}.

$ trex top '\I:ip \W:verb' '${verb}' access.log
GET   3
POST  1

$ trex count-by '\I:ip' '${ip:octet1-3}' access.log
10.0.0  3
10.0.1  1

$ trex uniq '\I:ip' '${ip}' access.log
10.0.0.5
10.0.0.7
10.0.1.9

Every key is printed. -n N cuts the table to N rows and says how many keys and matches there were, so a short table never reads as a complete one; Python returns every row and -MaxCount writes the first groups. A key an accessor leaves empty prints as - and is counted, so the rows sum to the matches.

Aggregates

--sum, --avg, --min, --max, --p50 and --p95 each take one register and add a column, computed in the value’s base unit through the parse a value predicate compares with. --sum and --avg take the numeric kinds: number, byte size, duration, money and percent. --min, --max and the percentiles also take timestamp, version and address. An aggregate a kind cannot carry is refused before anything is scanned, naming the register, its kind and the kinds the aggregate takes; a key whose matches bound no value prints -.

$ trex count-by '\W:h \Z:s' '${h}' sizes.txt --sum '${s}' --min '${s}' --max '${s}'
       count  --sum s  --min s  --max s
alpha      3     7000     1000     4000
beta       1     8000     8000     8000

$ trex count-by '\W:h \Z \R:d' '${h}' sizes.txt --sum '${d}' --duration-unit ms
       count  --sum d
alpha      3      240
beta       1      300

An aggregate names one plain register: ${u:host} is a slice of text and ${a}/${b} names two values, so both are refused. In PowerShell each aggregate is a property of the group in the .NET type that holds its kind: a number a long or a decimal, a duration a TimeSpan, a timestamp a DateTimeOffset.

Averages and percentiles

--avg is exact: --avg-form repetend (the default) brackets the repeating digits and rational writes the fraction in lowest terms. A percentile falling between two observed values is decided by --percentile: nearest (the default) takes the value at ceil(p x n) in sorted order, linear interpolates and takes numeric kinds only, lower takes the observed value at or below, and hybrid interpolates where the kind allows it and takes an observed value where it does not.

$ trex count-by '\W:h \Z:s' '${h}' sizes.txt --avg '${s}'
       count   --avg s
alpha      3  2333.(3)
beta       1      8000

$ trex count-by '\W:h \Z:s' '${h}' sizes.txt --avg '${s}' --avg-form rational
       count  --avg s
alpha      3   7000/3
beta       1     8000

$ trex count-by '\N:n' all --text '10 20 30 40' --p50 '${n}'
     count  --p50 n
all      4       20

$ trex count-by '\N:n' all --text '10 20 30 40' --p50 '${n}' --percentile linear
     count  --p50 n
all      4       25

Keys by place and by member

${path} counts the matches by input and ${line} by line. Over a set, from --patterns FILE or several patterns, ${pattern} counts them by the member that made them.

$ cat rules.trex
# what a line may hold
let host = \I:addr
let mail = \E:e
\N{>=100}
$ cat notes.txt
from 10.0.0.1 at 500 to bob@x.com
x = 7 and 10.0.0.2
$ cat other.txt
mail amy@y.org
$ trex count-by --patterns rules.trex '${pattern}' notes.txt other.txt
4     1
host  2
mail  2

$ trex count-by '\E' '${path}' notes.txt other.txt
notes.txt  1
other.txt  1

JSON and windows

--json prints the rows as an array of objects, each aggregate under its flag’s name and a value with no aggregate as null; --values and --duration-unit spell the values as scan --json spells them. --head N, --tail N and --lines A..B group the matches of one window of each input, and --record UNIT counts that window in records.

$ trex top '\I:ip' '${ip}' access.log --tail 2
10.0.0.5  1
10.0.1.9  1