phplrt 4.0

Results and Reducers

A parser does two things: it decides whether the input is valid, and it builds something out of it. The first part is the grammar. The second part is reducers.

What You Get Without Reducers

A grammar with no reducers still returns something - the tokens it kept, nested the way the rules were:

Sum : <T_DIGIT> (::T_PLUS:: <T_DIGIT>)* ;
1$parser->parse(StringSource::createFromString('2 + 3'));
2// [Token("2"), Token("3")]

That is occasionally enough. Usually you want numbers, or AST nodes, or a configuration array - and for that you attach a reducer.

Attaching A Reducer

A reducer is a block of PHP between -> and the rule body, and it runs when the rule matches:

1Number -> { return (int) $children->value; }
2  : <T_DIGIT>
3  ;

Whatever it returns becomes the value of that rule, and gets handed to the rule above. The code is ordinary PHP - loops, conditionals, whatever you need:

 1Name -> {
 2    $name = $children->value;
 3
 4    if (\str_starts_with($name, '$')) {
 5        return new \App\Ast\VariableNode($offset, \substr($name, 1));
 6    }
 7
 8    return new \App\Ast\ConstantNode($offset, $name);
 9}
10  : <T_NAME>
11  ;

Braces inside strings are safe - the block is read by a real PHP lexer, so "{" is a string, not the end of the block. An empty block (-> {}) is the same as writing no reducer at all.

The same thing through the builder:

1use Phplrt\Parser\Context;
2
3$number->setReducer(static fn(Context $ctx, mixed $children): int
4    => (int) $children->value);

The Variables

Inside a code block, these are available:

Variable What it is
$children What the rule matched. This is the important one.
$ctx The full Context object
$offset or $begin Where the rule starts, in bytes
$length How many bytes the rule covers
$end Where the rule ends - $offset + $length
$source The source being parsed
$content The whole content of that source
$rule The id of the rule being reduced

All except $children and $ctx are shorthands the compiler expands for you - $offset becomes $ctx->begin, and so on. They are only declared if you use them, so there is no cost to the ones you do not.

The position covers the tokens the rule kept, so a token dropped by ::T_NAME:: counts for nothing: the span of ::T_LP:: Expr() ::T_RP:: starts and ends where Expr does, parentheses aside. A rule that kept no tokens at all - an optional that matched nothing, say - is empty at the position the reading has reached, and its $length is zero.

What $children Contains

This is the part worth reading twice.

A rule that recognizes a sequence - a concatenation or a repetition - gets an array:

1Pair : <T_DIGIT> <T_DIGIT> ;      // $children = [Token, Token]
2List : <T_DIGIT>+ ;               // $children = [Token, Token, ...]

Any other rule gets the single value it recognized:

1Number : <T_DIGIT> ;              // $children = Token
2Choice : Number() | Name() ;      // $children = whatever matched
3Maybe  : Number()? ;              // $children = the Number, or nothing

Which of the two it is depends on the kind of rule, not on the input.

And the arrays are flattened into the parent. If a nested rule returns a list, its items are spliced into the list of the rule above rather than nested inside it:

1Root : <T_A> Pair() <T_B> ;
2Pair : <T_DIGIT> <T_DIGIT> ;
3
4// Root's $children = [Token(a), Token(1), Token(2), Token(b)]
5//               not  [Token(a), [Token(1), Token(2)], Token(b)]

This is deliberate - it keeps the result flat and predictable - but it means that a rule which should produce a group must say so by returning a value of its own. A reducer returning an object or a scalar is never flattened:

1Pair -> { return $children; /* an array */ }  // still flattened
2Pair -> { return new PairNode($children); }   // stays one value ✔

Because a rule can match one thing or a list depending on the input, reducers often start with a check:

 1Expression -> {
 2    // Just one operand: nothing to fold
 3    if (!\is_array($children)) {
 4        return $children;
 5    }
 6
 7    // ...
 8}
 9  : Number() ((<T_PLUS> | <T_MINUS>) Number())*
10  ;

What A Token Is

Most of what a reducer is handed in $children is tokens:

1$token->id;      // int    - what the token is, compared against a constant
2$token->name;    // string - what it is called, for messages; may be null
3$token->value;   // string - the exact text it was read from
4$token->offset;  // int    - where that text starts, in bytes
5$token->size;    // int    - how many bytes it covers
6$token->channel; // Channel - the label it was read on

Compare by id, never by name. The identifier is the position of the token's definition in the lexer; names exist for you and for error messages, and the parser does not use them at all - which is why a token is allowed to have none:

1if ($token->id === MyParser::T_DIGIT) {
2    // ...
3}

Generated parsers expose the ids as class constants, so you never have to write the numbers yourself.

offset and size are bytes, not characters, so the two of them address the fragment directly:

1$content = $source->content;
2
3// The exact fragment the token was read from
4$text = \substr($content, $token->offset, $token->size);
5
6// Where the next token starts
7$next = $token->offset + $token->size;

For an ordinary token size is simply strlen($value). It differs only for a token that entered a nested lexer, which is as large as everything that lexer read.

Tokens print themselves in a form meant for error messages, with long values cut off and control characters escaped, so one can go straight into a message without worrying about what is in it:

echo $token;
1"23" (T_DIGIT)      // a named token
2"@" (unknown token) // unrecognized input
3end of input        // the terminal token

Captures

If a token's pattern has capturing groups, whatever they matched is on the token, which saves parsing the value twice:

1// "hi" from  "([^"]*)"      3.14 from  (\d++)\.(\d++)
2$string->captures; // ["hi"]
3$float->captures;  // ["3", "14"]
4
5// Instead of trimming the quotes off $token->value by hand
6$content = $token->captures[0];

Captures are numbered per token, starting at zero - the first group of this token is captures[0], no matter how many groups the tokens above it have. A group that matched nothing still counts, so the numbering stays stable:

1"42"  => captures: ["", "42"]     from  (\+|-)?(\d++)
2"-42" => captures: ["-", "42"]

The Interface

Write against the contract rather than the implementation:

1use Phplrt\Contracts\Lexer\TokenInterface;
2
3function describe(TokenInterface $token): string
4{
5    return \sprintf('%s at %d', $token->name ?? 'anonymous', $token->offset);
6}

Phplrt\Lexer\Token\Token is the standard implementation, and TokenEmbedding extends it for nested lexers. Which tokens reach a reducer at all is a matter of channels.

Building An AST

Here is the whole pattern. Define node classes:

 1abstract class Node
 2{
 3    public function __construct(
 4        public readonly int $offset,
 5    ) {}
 6}
 7
 8final class NumberNode extends Node
 9{
10    public function __construct(int $offset, public readonly float $value)
11    {
12        parent::__construct($offset);
13    }
14}
15
16final class BinaryNode extends Node
17{
18    public function __construct(
19        int $offset,
20        public readonly string $operator,
21        public readonly Node $left,
22        public readonly Node $right,
23    ) {
24        parent::__construct($offset);
25    }
26}

Nothing here knows about the parser: a node is handed exactly the values it needs, which is a little more typing in the grammar and leaves you with plain value objects.

Then build them in the grammar:

 1%skip  T_WHITESPACE  \s++
 2%token T_NUMBER      \d++(?:\.\d++)?
 3%token T_PLUS        \+
 4%token T_MINUS       \-
 5
 6%pragma root Expression
 7
 8Expression -> {
 9    if (!\is_array($children)) {
10        return $children;
11    }
12
13    $node = \array_shift($children);
14
15    // Fold left: 1 + 2 - 3  =>  ((1 + 2) - 3)
16    while ($children !== []) {
17        $operator = \array_shift($children);
18        $right = \array_shift($children);
19
20        $node = new \BinaryNode($node->offset, $operator->value, $node, $right);
21    }
22
23    return $node;
24}
25  : Number() ((<T_PLUS> | <T_MINUS>) Number())*
26  ;
27
28Number -> { return new \NumberNode($offset, (float) $children->value); }
29  : <T_NUMBER>
30  ;

Parsing 1 + 2 - 3 gives you a tree:

1BinaryNode(-)
2├── BinaryNode(+)
3│   ├── NumberNode(1)
4│   └── NumberNode(2)
5└── NumberNode(3)

Note $offset in the Number reducer - that is one of the variables the compiler provides. Keeping an offset on every node is what lets you point at the right place in the source when something goes wrong later, during type-checking or evaluation.

The Context

Every reducer receives a Context as its first argument, describing where the analysis is:

1static function (Context $ctx, mixed $children): mixed {
2    $ctx->rule;   // int - the id of the rule being reduced
3    $ctx->begin;  // int - where this rule starts, in bytes
4    $ctx->length; // int - how many bytes it covers
5    $ctx->source; // the source being parsed
6
7    return null;
8}

In a .pp3 reducer you rarely touch $ctx directly, because the common fields have shorter names - $offset, $length, $end, $source, listed above, and $end is worked out for you.

Returning Nothing

A reducer that returns null leaves the value alone: the children are passed up as if there were no reducer at all. That makes null useful for reducers that only observe:

1Debug -> {
2    \error_log('matched at ' . $offset);
3
4    return null; // do not touch the result
5}
6  : <T_NAME>
7  ;

Reducers Run After Parsing

One thing to be aware of: reducers do not run while the input is being read. The parser first recognizes the whole source, then walks what it recognized and reduces it bottom-up.

This means:

  • a reducer never sees a rule that was tried and rejected - no wasted work, no side effects from a branch that did not win;
  • analyze() in Mode::SyntaxCheck never runs reducers at all;
  • a reducer cannot influence parsing. It cannot look ahead, change what is matched next, or fail the parse to force a different alternative. If a decision depends on the input, express it in the grammar - that is what

In Generated Code

When you generate a parser, reducers become real methods, named after the rule they belong to:

1private static function reduceNumber(\Phplrt\Parser\Context $ctx, mixed $children): mixed
2{
3    return (float) $children->value;
4}

Two practical consequences.

Your code appears verbatim in the generated file. It is debuggable and steppable, and a syntax error in a reducer is a syntax error in that file - so run the generator as part of your build, not at deploy time.

A grammar file has no use statements, so how a short class name resolves depends on where the reducer ends up - the global namespace when the grammar is read on the fly, the generated file's namespace when it is generated. The safe answer is to write class names fully qualified:

1// ✔ works either way
2Number -> { return new \App\Ast\NumberNode($offset, $children->value); }

If the fully qualified names make a big grammar unreadable, you can declare the imports on the generated file instead:

1new Compiler()
2    ->load(FileSource::createFromPathname(__DIR__ . '/grammar.pp3'))
3    ->generate()
4        ->withNamespaceName('App\Parser')
5        ->withClassImport('App\Ast\NumberNode')
6        ->save(__DIR__ . '/Parser.php');
Number -> { return new NumberNode($offset, $children->value); }

The trade-off: that grammar now only works when generated. Pick one approach per project rather than mixing them.