Types
Sized floats: float8, float16, float32
The narrow floating-point formats, how every operation re-rounds, and where they overflow or throw.
since 0.1.0-alpha.1linuxwindowswasm
Description
Besides the 64-bit float, there are three narrow floating-point types:
| type | format | largest finite value |
|---|---|---|
float8 |
OCP E4M3 (1 sign, 4 exponent and 3 mantissa bits) | 448 |
float16 |
IEEE 754 binary16 | 65504 |
float32 |
IEEE 754 binary32 | about 3.4028235 times 10 to the 38th |
They default to 0.0. float8 means the same E4M3 format on every machine; it is computed in software and never depends on the host's hardware.
Every arithmetic operator rounds its result into the working format (round to nearest, ties to even). Rounding therefore shows after every single operation, not only when a value is stored.
Per-operation rounding
float16 h = 0.1;
console.writeln(h);
float16 a = 2000;
float16 c = 49;
console.writeln(a + c);
float32 big = 16777216;
float32 one = 1;
console.writeln(big + one);
float8 top = 448;
console.writeln((top + top).isNaN());
console.writeln(top.isInfinite());
float16 near = 60000;
console.writeln((near + near).isInfinite());
console.writeln(h.bits());
console.writeln(float16::fromBits(15360));
0.099976
2048.000000
16777216.000000
true
false
true
11878
1.000000
Rules
Arithmetic.
- When one operand is a float, the result is a float whose storage width is the wider of the two operands' widths.
floatwith anything isfloat;int16 + float8isfloat16;float16 + int32isfloat32;float16 + float32isfloat32; any narrow float withintorfloatisfloat. - Each result is correctly rounded: the exact result is rounded once into the working format.
float162000 + 49is exactly2048, because2049is a tie and rounds to even.float3216777216 + 1is16777216. %, the bitwise operators and shifts do not exist for floats at any width; using them is a compile error ("no operator on promoted numeric type").- Comparisons convert both sides to the same working type and return
bool.
Overflow and special values.
float16andfloat32follow IEEE 754: overflow gives infinity,x / 0.0gives plus or minus infinity and0.0 / 0.0gives NaN.float8has no infinity. Its largest magnitude is448. An arithmetic result beyond that is NaN, never448and never an infinity, andisInfinite()on afloat8is alwaysfalse. Division by zero atfloat8is NaN.- Every NaN produced is the single canonical NaN of its format, so every NaN is the same key in a
Map.
Literals and conversions.
- A literal stored in a narrow float is rounded into the format.
float16 h = 0.1;holds0.099976. A literal beyond the largest finite value is a compile error that names it:float8 x = 500.0;andfloat16 y = 70000;are rejected. - An explicit or typed conversion throws a catchable
RuntimeExceptionwhen a finite value is outside the destination's range, and when an infinity is converted tofloat8.float.toFloat8(),toFloat16()andtoFloat32()convert fromfloat. A NaN converts to the NaN of the destination format. toFloat()widens exactly, because every narrow float is exactly representable in afloat.toInt()truncates toward zero and throws for NaN, infinity or a value outside theintrange.- Narrow-to-narrow conversions go through
float:f16.toFloat().toFloat8(). bits()returns the encoded bit pattern as anint.float8::fromBits(n),float16::fromBits(n)andfloat32::fromBits(n)build a value from a bit pattern and throw ifnhas more bits than the format.- The sized floats have no radix or hex methods; those belong to the integer types.
Distinct types. Each narrow float is its own type at run time. A float16 and a float32 holding the same number are different union members for is and match.
Examples
Mixed widths and a union:
Mixed float widths
float16 h = 1;
float32 s = 1;
float d = 1;
float8 q = 1;
int16 i = 3;
int32 j = 3;
console.writeln((h + d) is float);
console.writeln((h + s) is float32);
console.writeln((q + h) is float16);
console.writeln((i + q) is float16);
console.writeln((h + j) is float32);
float16 | float32 either = h;
match (either) {
float16 => console.writeln("half");
float32 => console.writeln("single");
}
true
true
true
true
true
half
Conversions out of range throw, and rounding inside the range is silent:
Narrowing a float
float big = 500.0;
float tiny = 0.1;
try {
console.writeln(big.toFloat8());
} catch (RuntimeException e) {
console.writeln("caught: ${e.message}");
}
console.writeln(tiny.toFloat16().toFloat());
float16 p = 2.9;
console.writeln(p.toInt());
float16 pn = -2.5;
console.writeln(pn.round());
caught: toFloat8: value out of range (max finite 448)
0.099976
2
-3.000000
Notes
Their method set mirrors float: toString, abs, floor, ceil, round (halves round away from zero), trunc, sqrt, pow, isNaN, isInfinite, bits and canonEq; see the library reference. Printing any NaN gives nan.
See also
- Numeric conversions and working types — How operands of different numeric types combine, and when a value is converted on its way into a typed variable, field, parameter or return.
- Sized integers: byte, int8, int16, int32, uint — The fixed-width integer types, how arithmetic wraps, how mixed widths combine, and when conversions throw.
- float — A 64-bit IEEE 754 floating-point number (binary64).