Skip to content
Find outliers

Find values that break their history

Find values that break their history

Find the values out of scale with what came before them, with the threshold read from the input at each token rather than written into the pattern. The baselines are on the context page.

Against the tokens before it

\N{>+1} matches a number an order of magnitude above the window of tokens before it, and \N{>+2s} one two of that window’s standard deviations above it:

$ trex scan '\N{>+1}' --text 'v 1 2 3 1 2 3 5000 2 3'
[14..18] "5000"

Against a key’s own history

The window mixes keys. Bind the key and compare against the values bound to it before: \N{>+1:k} reads the values that followed = or : after each earlier occurrence of what k holds, so a large size does not hide a slow latency:

$ trex scan '\N{>+1}' --text 'latency = 100 ; latency = 120 ; size = 5000000 ; latency = 90 ; latency = 12000 ;'
[39..46] "5000000"
[74..79] "12000"

$ trex scan '"latency":k "=" \N{>+1:k}' --text 'latency = 100 ; latency = 120 ; size = 5000000 ; latency = 90 ; latency = 12000 ;'
[64..79] "latency = 12000"  captures: k="latency"

The first occurrence of a key has no history and never matches.

Against its column

In a record that repeats, \N{>+2:phase} compares a value with the earlier values in its own column, with no delimiter named:

$ trex scan '\N{>+2:phase}' --text 'a , 10 , x ; b , 12 , y ; c , 9000 , z ; d , 11 , w ;'
[30..34] "9000"

Choose the delta

The delta is in orders of magnitude: +1 is ten times the baseline, +2 a hundred. With an s it is in the baseline’s standard deviations, which needs a baseline of more than one value. An empty baseline, such as the first token or a key’s first occurrence, matches nothing.