Results and Reducers
A parser does two things: it decides whether the input is valid, and it builds something out of it. The first part is the grammar. The second part is reducers.
What You Get Without Reducers
A grammar with no reducers still returns something - the tokens it kept, nested the way the rules were:
Sum : <T_DIGIT> (::T_PLUS:: <T_DIGIT>)* ;
1$parser->parse(StringSource::createFromString('2 + 3')); 2// [Token("2"), Token("3")]
That is occasionally enough. Usually you want numbers, or AST nodes, or a configuration array - and for that you attach a reducer.
Attaching A Reducer
A reducer is a block of PHP between -> and the rule body, and it runs when
the rule matches:
1Number -> { return (int) $children->value; } 2 : <T_DIGIT> 3 ;
Whatever it returns becomes the value of that rule, and gets handed to the rule above. The code is ordinary PHP - loops, conditionals, whatever you need:
1Name -> { 2 $name = $children->value; 3 4 if (\str_starts_with($name, '$')) { 5 return new \App\Ast\VariableNode($offset, \substr($name, 1)); 6 } 7 8 return new \App\Ast\ConstantNode($offset, $name); 9} 10 : <T_NAME> 11 ;
Braces inside strings are safe - the block is read by a real PHP lexer, so
"{" is a string, not the end of the block. An empty block (-> {}) is the
same as writing no reducer at all.
The same thing through the builder:
1use Phplrt\Parser\Context; 2 3$number->setReducer(static fn(Context $ctx, mixed $children): int 4 => (int) $children->value);
The Variables
Inside a code block, these are available:
| Variable | What it is |
|---|---|
$children |
What the rule matched. This is the important one. |
$ctx |
The full Context object |
$offset or $begin |
Where the rule starts, in bytes |
$length |
How many bytes the rule covers |
$end |
Where the rule ends - $offset + $length |
$source |
The source being parsed |
$content |
The whole content of that source |
$rule |
The id of the rule being reduced |
All except $children and $ctx are shorthands the compiler expands for
you - $offset becomes $ctx->begin, and so on. They are only declared if
you use them, so there is no cost to the ones you do not.
The position covers the tokens the rule kept, so a token dropped by
::T_NAME:: counts for nothing: the span of ::T_LP:: Expr() ::T_RP::
starts and ends where Expr does, parentheses aside. A rule that kept no
tokens at all - an optional that matched nothing, say - is empty at the
position the reading has reached, and its $length is zero.
What $children Contains
This is the part worth reading twice.
A rule that recognizes a sequence - a concatenation or a repetition - gets an array:
1Pair : <T_DIGIT> <T_DIGIT> ; // $children = [Token, Token] 2List : <T_DIGIT>+ ; // $children = [Token, Token, ...]
Any other rule gets the single value it recognized:
1Number : <T_DIGIT> ; // $children = Token 2Choice : Number() | Name() ; // $children = whatever matched 3Maybe : Number()? ; // $children = the Number, or nothing
Which of the two it is depends on the kind of rule, not on the input.
And the arrays are flattened into the parent. If a nested rule returns a list, its items are spliced into the list of the rule above rather than nested inside it:
1Root : <T_A> Pair() <T_B> ; 2Pair : <T_DIGIT> <T_DIGIT> ; 3 4// Root's $children = [Token(a), Token(1), Token(2), Token(b)] 5// not [Token(a), [Token(1), Token(2)], Token(b)]
This is deliberate - it keeps the result flat and predictable - but it means that a rule which should produce a group must say so by returning a value of its own. A reducer returning an object or a scalar is never flattened:
1Pair -> { return $children; /* an array */ } // still flattened 2Pair -> { return new PairNode($children); } // stays one value ✔
Because a rule can match one thing or a list depending on the input, reducers often start with a check:
1Expression -> { 2 // Just one operand: nothing to fold 3 if (!\is_array($children)) { 4 return $children; 5 } 6 7 // ... 8} 9 : Number() ((<T_PLUS> | <T_MINUS>) Number())* 10 ;
What A Token Is
Most of what a reducer is handed in $children is tokens:
1$token->id; // int - what the token is, compared against a constant 2$token->name; // string - what it is called, for messages; may be null 3$token->value; // string - the exact text it was read from 4$token->offset; // int - where that text starts, in bytes 5$token->size; // int - how many bytes it covers 6$token->channel; // Channel - the label it was read on
Compare by id, never by name. The identifier is the position of the
token's definition in the lexer; names exist for you and for error messages,
and the parser does not use them at all - which is why a token is allowed to
have none:
1if ($token->id === MyParser::T_DIGIT) { 2 // ... 3}
Generated parsers expose the ids as class constants, so you never have to write the numbers yourself.
offset and size are bytes, not characters, so the two of them address
the fragment directly:
1$content = $source->content; 2 3// The exact fragment the token was read from 4$text = \substr($content, $token->offset, $token->size); 5 6// Where the next token starts 7$next = $token->offset + $token->size;
For an ordinary token size is simply strlen($value). It differs only for a
token that entered a nested lexer, which is as
large as everything that lexer read.
Tokens print themselves in a form meant for error messages, with long values cut off and control characters escaped, so one can go straight into a message without worrying about what is in it:
echo $token;
1"23" (T_DIGIT) // a named token 2"@" (unknown token) // unrecognized input 3end of input // the terminal token
Captures
If a token's pattern has capturing groups, whatever they matched is on the token, which saves parsing the value twice:
1// "hi" from "([^"]*)" 3.14 from (\d++)\.(\d++) 2$string->captures; // ["hi"] 3$float->captures; // ["3", "14"] 4 5// Instead of trimming the quotes off $token->value by hand 6$content = $token->captures[0];
Captures are numbered per token, starting at zero - the first group of this
token is captures[0], no matter how many groups the tokens above it have. A
group that matched nothing still counts, so the numbering stays stable:
1"42" => captures: ["", "42"] from (\+|-)?(\d++) 2"-42" => captures: ["-", "42"]
The Interface
Write against the contract rather than the implementation:
1use Phplrt\Contracts\Lexer\TokenInterface; 2 3function describe(TokenInterface $token): string 4{ 5 return \sprintf('%s at %d', $token->name ?? 'anonymous', $token->offset); 6}
Phplrt\Lexer\Token\Token is the standard implementation, and
TokenEmbedding extends it for nested lexers.
Which tokens reach a reducer at all is a matter of
channels.
Building An AST
Here is the whole pattern. Define node classes:
1abstract class Node 2{ 3 public function __construct( 4 public readonly int $offset, 5 ) {} 6} 7 8final class NumberNode extends Node 9{ 10 public function __construct(int $offset, public readonly float $value) 11 { 12 parent::__construct($offset); 13 } 14} 15 16final class BinaryNode extends Node 17{ 18 public function __construct( 19 int $offset, 20 public readonly string $operator, 21 public readonly Node $left, 22 public readonly Node $right, 23 ) { 24 parent::__construct($offset); 25 } 26}
Nothing here knows about the parser: a node is handed exactly the values it needs, which is a little more typing in the grammar and leaves you with plain value objects.
Then build them in the grammar:
1%skip T_WHITESPACE \s++ 2%token T_NUMBER \d++(?:\.\d++)? 3%token T_PLUS \+ 4%token T_MINUS \- 5 6%pragma root Expression 7 8Expression -> { 9 if (!\is_array($children)) { 10 return $children; 11 } 12 13 $node = \array_shift($children); 14 15 // Fold left: 1 + 2 - 3 => ((1 + 2) - 3) 16 while ($children !== []) { 17 $operator = \array_shift($children); 18 $right = \array_shift($children); 19 20 $node = new \BinaryNode($node->offset, $operator->value, $node, $right); 21 } 22 23 return $node; 24} 25 : Number() ((<T_PLUS> | <T_MINUS>) Number())* 26 ; 27 28Number -> { return new \NumberNode($offset, (float) $children->value); } 29 : <T_NUMBER> 30 ;
Parsing 1 + 2 - 3 gives you a tree:
1BinaryNode(-) 2├── BinaryNode(+) 3│ ├── NumberNode(1) 4│ └── NumberNode(2) 5└── NumberNode(3)
Note $offset in the Number reducer - that is one of the
variables the compiler provides. Keeping an offset on
every node is what lets you point at the right place in the source when
something goes wrong later, during type-checking or evaluation.
The Context
Every reducer receives a Context as its first argument, describing where the
analysis is:
1static function (Context $ctx, mixed $children): mixed { 2 $ctx->rule; // int - the id of the rule being reduced 3 $ctx->begin; // int - where this rule starts, in bytes 4 $ctx->length; // int - how many bytes it covers 5 $ctx->source; // the source being parsed 6 7 return null; 8}
In a .pp3 reducer you rarely touch $ctx directly, because the common
fields have shorter names - $offset, $length, $end, $source, listed
above, and $end is worked out for you.
Returning Nothing
A reducer that returns null leaves the value alone: the children are passed
up as if there were no reducer at all. That makes null useful for reducers
that only observe:
1Debug -> { 2 \error_log('matched at ' . $offset); 3 4 return null; // do not touch the result 5} 6 : <T_NAME> 7 ;
Reducers Run After Parsing
One thing to be aware of: reducers do not run while the input is being read. The parser first recognizes the whole source, then walks what it recognized and reduces it bottom-up.
This means:
- a reducer never sees a rule that was tried and rejected - no wasted work, no side effects from a branch that did not win;
-
analyze()inMode::SyntaxChecknever runs reducers at all; - a reducer cannot influence parsing. It cannot look ahead, change what is matched next, or fail the parse to force a different alternative. If a decision depends on the input, express it in the grammar - that is what
In Generated Code
When you generate a parser, reducers become real methods, named after the rule they belong to:
1private static function reduceNumber(\Phplrt\Parser\Context $ctx, mixed $children): mixed 2{ 3 return (float) $children->value; 4}
Two practical consequences.
Your code appears verbatim in the generated file. It is debuggable and steppable, and a syntax error in a reducer is a syntax error in that file - so run the generator as part of your build, not at deploy time.
A grammar file has no use statements, so how a short class name resolves
depends on where the reducer ends up - the global namespace when the grammar
is read on the fly, the generated file's namespace when it is generated. The
safe answer is to write class names fully qualified:
1// ✔ works either way 2Number -> { return new \App\Ast\NumberNode($offset, $children->value); }
If the fully qualified names make a big grammar unreadable, you can declare the imports on the generated file instead:
1new Compiler() 2 ->load(FileSource::createFromPathname(__DIR__ . '/grammar.pp3')) 3 ->generate() 4 ->withNamespaceName('App\Parser') 5 ->withClassImport('App\Ast\NumberNode') 6 ->save(__DIR__ . '/Parser.php');
Number -> { return new NumberNode($offset, $children->value); }
The trade-off: that grammar now only works when generated. Pick one approach per project rather than mixing them.