LEVIATHAN v962456e · 962456eee1

Types

Sized floats: float8, float16, float32

The narrow floating-point formats, how every operation re-rounds, and where they overflow or throw.

since 0.1.0-alpha.1linuxwindowswasm

Description

Besides the 64-bit float, there are three narrow floating-point types:

type format largest finite value
float8 OCP E4M3 (1 sign, 4 exponent and 3 mantissa bits) 448
float16 IEEE 754 binary16 65504
float32 IEEE 754 binary32 about 3.4028235 times 10 to the 38th

They default to 0.0. float8 means the same E4M3 format on every machine; it is computed in software and never depends on the host's hardware.

Every arithmetic operator rounds its result into the working format (round to nearest, ties to even). Rounding therefore shows after every single operation, not only when a value is stored.

Per-operation rounding

float16 h = 0.1;
console.writeln(h);
float16 a = 2000;
float16 c = 49;
console.writeln(a + c);
float32 big = 16777216;
float32 one = 1;
console.writeln(big + one);
float8 top = 448;
console.writeln((top + top).isNaN());
console.writeln(top.isInfinite());
float16 near = 60000;
console.writeln((near + near).isInfinite());
console.writeln(h.bits());
console.writeln(float16::fromBits(15360));
0.099976
2048.000000
16777216.000000
true
false
true
11878
1.000000

Rules

Arithmetic.

  • When one operand is a float, the result is a float whose storage width is the wider of the two operands' widths. float with anything is float; int16 + float8 is float16; float16 + int32 is float32; float16 + float32 is float32; any narrow float with int or float is float.
  • Each result is correctly rounded: the exact result is rounded once into the working format. float16 2000 + 49 is exactly 2048, because 2049 is a tie and rounds to even. float32 16777216 + 1 is 16777216.
  • %, the bitwise operators and shifts do not exist for floats at any width; using them is a compile error ("no operator on promoted numeric type").
  • Comparisons convert both sides to the same working type and return bool.

Overflow and special values.

  • float16 and float32 follow IEEE 754: overflow gives infinity, x / 0.0 gives plus or minus infinity and 0.0 / 0.0 gives NaN.
  • float8 has no infinity. Its largest magnitude is 448. An arithmetic result beyond that is NaN, never 448 and never an infinity, and isInfinite() on a float8 is always false. Division by zero at float8 is NaN.
  • Every NaN produced is the single canonical NaN of its format, so every NaN is the same key in a Map.

Literals and conversions.

  • A literal stored in a narrow float is rounded into the format. float16 h = 0.1; holds 0.099976. A literal beyond the largest finite value is a compile error that names it: float8 x = 500.0; and float16 y = 70000; are rejected.
  • An explicit or typed conversion throws a catchable RuntimeException when a finite value is outside the destination's range, and when an infinity is converted to float8. float.toFloat8(), toFloat16() and toFloat32() convert from float. A NaN converts to the NaN of the destination format.
  • toFloat() widens exactly, because every narrow float is exactly representable in a float. toInt() truncates toward zero and throws for NaN, infinity or a value outside the int range.
  • Narrow-to-narrow conversions go through float: f16.toFloat().toFloat8().
  • bits() returns the encoded bit pattern as an int. float8::fromBits(n), float16::fromBits(n) and float32::fromBits(n) build a value from a bit pattern and throw if n has more bits than the format.
  • The sized floats have no radix or hex methods; those belong to the integer types.

Distinct types. Each narrow float is its own type at run time. A float16 and a float32 holding the same number are different union members for is and match.

Examples

Mixed widths and a union:

Mixed float widths

float16 h = 1;
float32 s = 1;
float d = 1;
float8 q = 1;
int16 i = 3;
int32 j = 3;
console.writeln((h + d) is float);
console.writeln((h + s) is float32);
console.writeln((q + h) is float16);
console.writeln((i + q) is float16);
console.writeln((h + j) is float32);
float16 | float32 either = h;
match (either) {
    float16 => console.writeln("half");
    float32 => console.writeln("single");
}
true
true
true
true
true
half

Conversions out of range throw, and rounding inside the range is silent:

Narrowing a float

float big = 500.0;
float tiny = 0.1;
try {
    console.writeln(big.toFloat8());
} catch (RuntimeException e) {
    console.writeln("caught: ${e.message}");
}
console.writeln(tiny.toFloat16().toFloat());
float16 p = 2.9;
console.writeln(p.toInt());
float16 pn = -2.5;
console.writeln(pn.round());
caught: toFloat8: value out of range (max finite 448)
0.099976
2
-3.000000

Notes

Their method set mirrors float: toString, abs, floor, ceil, round (halves round away from zero), trunc, sqrt, pow, isNaN, isInfinite, bits and canonEq; see the library reference. Printing any NaN gives nan.

See also