Compilation Process
From source code to on-chain bytecode.
The Matador compiler transforms high-level DSL policies into a compact, efficient bytecode format optimized for the on-chain interpreter. This transformation occurs in four distinct stages.
The Pipeline
Parsing & Validation
The compiler uses Langium to parse the .matador source text into an Abstract Syntax Tree (AST).
Key Operations:
- Grammar Check: Ensures the code follows the valid syntax (brackets, keywords).
- Symbol Resolution: Links references (e.g.,
parameters.token) to their declarations. - ABI Loading: Fetches imported JSON ABIs and parses their definitions.
IR Conversion
The AST is lowered into a Canonical Intermediate Representation (IR). The IR flattens complex expressions (like nested all/any blocks) into a linear sequence of instructions.
Determinism
The IR conversion process is deterministic. Identical logic will always produce the same IR, regardless of whitespace or comment changes in the source.
Semantic Analysis
The compiler performs deep semantic checks on the IR to ensure safety and correctness.
Key Checks:
- Type Compatibility: Verifies that a
uint256parameter isn't passed to an opcode expecting anaddress. - Mutability: Ensures that policies executed in a
viewcontext (like a simulation) do not call state-modifying opcodes. - Module Requirements: Identifies which interpreter modules are needed (e.g.,
Delegation,Storage).
Bytecode Encoding
The final stage encodes the validated IR into the compact binary format used by the on-chain interpreter.
The 2-Pass Encoding Process:
- Instruction Stream: Opcodes and static arguments (fixed-size integers, addresses) are packed into a continuous byte stream. Placeholders are left for dynamic data.
- Dynamic Data Patching: Variable-length data (strings, arrays) is appended to the end of the stream. The placeholders are updated with relative pointers (offsets) to this data.
Bytecode Structure
The final bytecode follows a strict structure to allow for efficient on-chain execution without complex parsing logic.
| Segment | Description |
|---|---|
| Root Group Start | 0x12 (MULTI_CHECK_AND) - Wraps the entire policy. |
| Instruction Stream | Sequence of [Opcode, Args...] bytes. |
| Root Group End | 0x14 (END_MULTI_CHECK) - Terminates execution. |
| Dynamic Heap | Appended data for arrays and bytes, referenced by pointers in the stream. |
Callable state entries
Callable bytecode uses a 48-byte header, 32-byte function entries, and 104-byte state-declaration entries. Each state entry encodes:
declIndex | kind | type | flags/reserved | nameHash | declKey | initialValue
2 bytes 1 1 4 32 32 32initialValue must be zero for transient declarations. Persistent values are
fixed-word literal seeds validated independently by language validation,
semantic analysis, final encoding, and interpreter preflight. The encoder uses
checked integer serialization for every immediate; out-of-range values are
diagnostics rather than silently truncated bytes.
Artifact Generation
The final output is a JSON artifact containing the hexData (bytecode) and metadata required for deployment and verification.
{
"hexData": "0x1201...14...",
"metadata": {
"specVersion": "1.0.0",
"permission": "SwapPolicy",
"checksum": "0x...",
"instructionCount": 12
},
"requiredModules": ["core", "execution"]
}