LEVIATHAN v962456e · 962456eee1

Compiler

The execution engines

The four ways the compiler runs a program, why they always agree, and the few places where one of them has to refuse a program.

since 0.1.0-alpha.1linux

Description

Leviathan has one language and four engines that execute it. You pick the engine with the option you give the compiler:

Option Engine What it is
--run The evaluator Walks the checked syntax tree directly. It defines what every construct means.
--ir The bytecode interpreter Runs a register-machine bytecode that the compiler lowers the program to.
--build <out> The C++ backend Translates the bytecode to a self-contained C++ program and builds it with the system C++ compiler.
--build-native <out> The LLVM backend Translates the bytecode to LLVM IR and links native machine code.

The last three share one lowering step. The program is checked once, lowered to the bytecode once, and the interpreter and the two native backends all start from that same bytecode.

The promise is that the engines agree. A valid program prints the same output and exits with the same status on every engine that accepts it: closures, exceptions, classes, collections and iteration all behave identically. The project's regression corpus is made of real programs with recorded expected output, and every program in it must produce that output on all of the engines that can run it. A difference in output between engines is treated as a bug in the compiler, never as a documented behaviour.

What does differ between engines is speed and the amount of setup. The evaluator and the interpreter start immediately and need no build step. The native backends take a few seconds to build and produce programs that run much faster.

When an engine cannot compile a construct it says so at compile time, with the source position of the construct. It never produces a program that quietly behaves differently. The cases that exist today are listed in lang.native-backends (the C++ backend does not support programs that use files, timers or sockets) and lang.cross-compilation (some targets reject some system features).

The same program on every engine

class Counter {
    int count = 0;
    void bump() {
        count = count + 1;
    }
}

Counter c = Counter();
Array<int> squares = [1, 2, 3, 4].map((n) => n * n);
int total = 0;
for (int s in squares) {
    c.bump();
    total = total + s;
}
console.writeln("sum of squares: ${total}");
console.writeln("bumped ${c.count} times");

try {
    throw RuntimeException("stop");
} catch (RuntimeException e) {
    console.writeln("caught: ${e.message}");
}
sum of squares: 30
bumped 4 times
caught: stop

Rules

  • Every valid program produces the same standard output and exit status on every engine that compiles it.
  • An engine that cannot compile a construct reports a compile-time diagnostic. It does not change the program's meaning.
  • An uncaught exception ends the program with exit status 1 and prints Uncaught <Type>: <message>, and env::exit(n) ends it with status n, on all four engines.
  • The evaluator is the reference. If another engine disagrees with it, the other engine is wrong.

Examples

The program above, saved as engines.lev, gives the same three lines in each of these runs:

leviathan --run engines.lev
leviathan --ir engines.lev
leviathan --build engines_cpp engines.lev && ./engines_cpp
leviathan --build-native engines_native engines.lev && ./engines_native

A program picks its own exit status with env::exit:

Ending a program with an exit status

console.writeln("about to stop");
env::exit(3);

Run with any of the four options, that program prints about to stop and the process exit status is 3.

Notes

The bytecode interpreter has a verification mode, --ir-verify, that checks the compiler's memory analysis while the program runs; see lang.ownership-analysis.

See also