LEVIATHAN v962456e · 962456eee1

Standard Library

class float8

An 8-bit floating-point number in the OCP E4M3 format: 1 sign bit, 4 exponent bits and 3 mantissa bits.

since 0.1.0-alpha.1linuxwindowswasm

Overview

A float8 has only about one significant decimal digit. The largest finite value is 448, the smallest positive normal value is 0.015625 and the smallest positive value is 0.001953125. There is no infinity: an operation whose result is above 448 gives NaN instead, and there is exactly one NaN. Every arithmetic result is rounded to the nearest float8 (ties to even) after each operation. Mixing two float formats computes in the wider one, and mixing with an int or a float computes as float.

A literal assigned to a float8 is rounded to the nearest value and a literal above 448 is a compile error. float.toFloat8() converts a float and throws when the value is out of range. toFloat() converts back exactly.

Examples

Rounding, range and NaN

float8 x = 0.1;
console.writeln(x);
console.writeln(x.bits());
float8 big = 448.0;
float8 two = 2.0;
console.writeln((big * two).isNaN());
console.writeln(big.isInfinite());
float tooBig = 500.0;
try {
    console.writeln(tooBig.toFloat8());
} catch (RuntimeException e) {
    console.writeln("caught: ${e.message}");
}
0.101562
29
true
false
caught: toFloat8: value out of range (max finite 448)

Methods

abs

abs() -> float8

Return the absolute value.

The sign bit is cleared, so -0.0 becomes 0.0.

Returns

this with its sign removed, as a float8.

Examples

float8 neg = -2.5;
console.writeln(neg.abs());
float8 negZero = float8::fromBits(128);
console.writeln(negZero.abs().bits());
2.500000
0

bits

bits() -> int

Return the 8-bit encoding of the value as an int.

The result is the raw interchange bits of the format (1 sign bit, then 4 exponent bits, then 3 mantissa bits), as a non-negative int. Two values that print the same can have different bits, for example 0.0 and -0.0. std.math.float8FromBits turns bits back into a float8 and is also written float8::fromBits.

Returns

The bit pattern of this.

Examples

float8 one = 1.0;
console.writeln(one.bits());
float8 big = 448.0;
console.writeln(big.bits());
float8 x = 0.1;
console.writeln(x.bits());
console.writeln(float8::fromBits(56) == one);
56
126
29
true

See also: float8FromBits

canonEq

canonEq(float8 other) -> bool

Compare two float8 values by their canonical form.

This is the relation used when a float8 is a map key or a field of a struct compared with == on the struct. It differs from the == operator in two cases: every NaN equals every other NaN, and 0.0 equals -0.0. The == operator follows the IEEE rule that NaN is not equal to anything, itself included.

Parameters

other
The value to compare with.

Returns

true when both are NaN or both denote the same number.

Examples

float8 nan = float8::fromBits(127);
console.writeln(nan == nan);
console.writeln(nan.canonEq(nan));
float8 zero = 0.0;
float8 negZero = float8::fromBits(128);
console.writeln(zero == negZero);
console.writeln(zero.canonEq(negZero));
console.writeln(zero.bits());
console.writeln(negZero.bits());
false
true
true
true
0
128

ceil

ceil() -> float8

Round up to a whole number.

Returns

The smallest whole number not less than this, as a float8.

Examples

float8 pos = 2.5;
float8 neg = -2.5;
console.writeln(pos.ceil());
console.writeln(neg.ceil());
3.000000
-2.000000

See also: floor

floor

floor() -> float8

Round down to a whole number.

Returns

The largest whole number not greater than this, as a float8.

Examples

float8 pos = 2.5;
float8 neg = -2.5;
console.writeln(pos.floor());
console.writeln(neg.floor());
2.000000
-3.000000

See also: ceil

isInfinite

isInfinite() -> bool

Test whether the value is infinite.

A float8 has no infinity, so this is always false, even for NaN produced by overflow.

Returns

Always false.

Examples

float8 big = 448.0;
float8 two = 2.0;
console.writeln(big.isInfinite());
console.writeln((big * two).isInfinite());
float8 one = 1.0;
float8 zero = 0.0;
console.writeln((one / zero).isInfinite());
false
false
false

isNaN

isNaN() -> bool

Test whether the value is NaN.

NaN (not a number) is what an undefined operation produces, such as the square root of a negative number. NaN is the only value that is not equal to itself, so x.isNaN() is the reliable test.

Returns

true when this is NaN.

Examples

float8 one = 1.0;
console.writeln(one.isNaN());
float8 big = 448.0;
float8 two = 2.0;
console.writeln((big * two).isNaN());
console.writeln(float8::fromBits(127).isNaN());
false
true
true

pow

pow(float8 e) -> float8

Raise the value to a power.

The result is computed in higher precision and rounded to the nearest float8. Because float8 has no infinity, a result above 448 is NaN.

Parameters

e
The exponent.

Returns

this to the power e, rounded to a float8.

Examples

float8 two = 2.0;
console.writeln(two.pow(two));
float8 half = 0.5;
console.writeln(half.pow(half));
float8 big = 448.0;
console.writeln(big.pow(two));
4.000000
0.687500
nan

round

round() -> float8

Round to the nearest whole number.

A value exactly halfway between two whole numbers rounds away from zero, so 2.5 becomes 3 and -2.5 becomes -3.

Returns

The nearest whole number, as a float8.

Examples

float8 pos = 2.5;
float8 neg = -2.5;
console.writeln(pos.round());
console.writeln(neg.round());
3.000000
-3.000000

See also: trunc

sqrt

sqrt() -> float8

Compute the square root.

The result is rounded to the nearest float8. The square root of a negative number is NaN; it does not throw.

Returns

The square root of this, as a float8.

Examples

float8 nine = 9.0;
console.writeln(nine.sqrt());
float8 two = 2.0;
console.writeln(two.sqrt());
float8 minusOne = -1.0;
console.writeln(minusOne.sqrt());
3.000000
1.375000
nan

toFloat

toFloat() -> float

Convert the float8 to a float.

The conversion is exact: every value of a float8 is representable as a float, so nothing is rounded. Converting a float back with float.toFloat8() rounds to the nearest float8 and throws when the value is out of range.

Returns

The value as a float.

Examples

float8 x = 0.1;
console.writeln(x.toFloat());
console.writeln(x.toFloat() == 0.1);
0.101562
false

See also: toFloat8

toInt

toInt() -> int

Convert the float8 to an int, discarding the fraction.

The value is truncated toward zero, so -2.5 becomes -2.

Returns

The whole-number part of this as an int.

Throws

RuntimeException
when the value is NaN or infinite.

Examples

float8 v = -2.5;
console.writeln(v.toInt());
float8 big = 448.0;
console.writeln(big.toInt());
float8 nan = float8::fromBits(127);
try {
    console.writeln(nan.toInt());
} catch (RuntimeException e) {
    console.writeln("caught: ${e.message}");
}
-2
448
caught: float is not finite or out of int64 range for toInt()

toString

toString() -> string

Format the float8 as decimal text.

The text is the six-decimal form that float prints, for example 3.500000, computed from the value actually stored, so a number that is not exactly representable shows its rounded value. NaN prints as nan and infinity as inf or -inf.

Returns

The decimal text of this.

Examples

float8 x = 0.1;
console.writeln(x.toString());
float8 y = 3.5;
console.writeln("y = " + y.toString());
0.101562
y = 3.500000

trunc

trunc() -> float8

Discard the fraction.

Returns

The whole-number part of this, rounded toward zero, as a float8.

Examples

float8 pos = 2.5;
float8 neg = -2.5;
console.writeln(pos.trunc());
console.writeln(neg.trunc());
2.000000
-2.000000

See also: floor