Summarize a log by its templates
Summarize a log by its templates
Read a log as the handful of line shapes it repeats, with how many lines each covers, then scan for the lines of one shape. A position whose text varies is a slot named by its kind (templates).
$ cat requests.log
10.0.0.1 GET /index.html status 200 12ms
10.0.0.2 GET /about.html status 200 8ms
10.0.0.1 POST /login status 302 40ms
10.0.0.3 GET /index.html status 200 11ms
10.0.0.9 GET /admin status 403 3ms
kernel: disk failure on /dev/sda
The shapes and their counts
$ trex templates requests.log
5 <ip> <word> <path> status <number> <duration>
1 kernel: disk failure on /dev/sda
Scan for the lines of one shape
Each template is also a pattern that finds exactly its lines:
$ trex templates requests.log --pattern
5 \I \W \L "status" \N \R
1 "kernel" ":" "disk" "failure" "on" "/dev/sda"
$ trex scan '\I \W \L "status" \N{>=300} \R' requests.log
[81..117] "10.0.0.1 POST /login status 302 40ms"
[159..193] "10.0.0.9 GET /admin status 403 3ms"
--record paragraph and the other units group records rather than lines, and --against FILE marks each template shared with another log or new to this one. The lines whose shape is
rare are a how-to of their own.