Skip to content
Parse with a token grammar

Parse with a token grammar

Parse with a token grammar

Parse input against named rules written over tokens rather than characters: number is a number token, ident a word, "+" a literal, and a left-recursive rule gives an operator its precedence. The grammar form is on grammars.

$ cat arith.grammar
expr   := <expr> "+" <term> | <expr> "-" <term> | <term>
term   := <term> "*" <factor> | <term> "/" <factor> | <factor>
factor := number | ident | "(" <expr> ")"

Parse and check

$ trex grammar arith.grammar --text '2 + 3 * 4'
(expr 2 + (term 3 * 4))

* binds tighter than + because term sits below expr.

Measure how ambiguous a grammar is

--count gives the number of derivations; a grammar with one rule for every sum reads 1 + 2 + 3 + 4 five ways:

$ trex grammar --grammar-text 'expr := <expr> "+" <expr> | number' --text '1 + 2 + 3 + 4' --count
5

@p after an alternative weights it, and --best and --prob give the probability of the most probable derivation and of all of them.