Standard Library
class float8
An 8-bit floating-point number in the OCP E4M3 format: 1 sign bit, 4 exponent bits and 3 mantissa bits.
since 0.1.0-alpha.1linuxwindowswasm
Overview
A float8 has only about one significant decimal digit. The largest finite value is 448, the smallest positive normal value is 0.015625 and the smallest positive value is 0.001953125. There is no infinity: an operation whose result is above 448 gives NaN instead, and there is exactly one NaN. Every arithmetic result is rounded to the nearest float8 (ties to even) after each operation. Mixing two float formats computes in the wider one, and mixing with an int or a float computes as float.
A literal assigned to a float8 is rounded to the nearest value and a literal above 448 is a compile error. float.toFloat8() converts a float and throws when the value is out of range. toFloat() converts back exactly.
Examples
Rounding, range and NaN
float8 x = 0.1;
console.writeln(x);
console.writeln(x.bits());
float8 big = 448.0;
float8 two = 2.0;
console.writeln((big * two).isNaN());
console.writeln(big.isInfinite());
float tooBig = 500.0;
try {
console.writeln(tooBig.toFloat8());
} catch (RuntimeException e) {
console.writeln("caught: ${e.message}");
}
0.101562
29
true
false
caught: toFloat8: value out of range (max finite 448)
Methods
abs
abs() -> float8Return the absolute value.
The sign bit is cleared, so -0.0 becomes 0.0.
Returns
this with its sign removed, as a float8.
Examples
float8 neg = -2.5;
console.writeln(neg.abs());
float8 negZero = float8::fromBits(128);
console.writeln(negZero.abs().bits());
2.500000
0
bits
bits() -> intReturn the 8-bit encoding of the value as an int.
The result is the raw interchange bits of the format (1 sign bit, then 4 exponent bits, then 3 mantissa bits), as a non-negative int. Two values that print the same can have different bits, for example 0.0 and -0.0. std.math.float8FromBits turns bits back into a float8 and is also written float8::fromBits.
Returns
The bit pattern of this.
Examples
float8 one = 1.0;
console.writeln(one.bits());
float8 big = 448.0;
console.writeln(big.bits());
float8 x = 0.1;
console.writeln(x.bits());
console.writeln(float8::fromBits(56) == one);
56
126
29
true
See also: float8FromBits
canonEq
canonEq(float8 other) -> boolCompare two float8 values by their canonical form.
This is the relation used when a float8 is a map key or a field of a struct compared with == on the struct. It differs from the == operator in two cases: every NaN equals every other NaN, and 0.0 equals -0.0. The == operator follows the IEEE rule that NaN is not equal to anything, itself included.
Parameters
- other
- The value to compare with.
Returns
true when both are NaN or both denote the same number.
Examples
float8 nan = float8::fromBits(127);
console.writeln(nan == nan);
console.writeln(nan.canonEq(nan));
float8 zero = 0.0;
float8 negZero = float8::fromBits(128);
console.writeln(zero == negZero);
console.writeln(zero.canonEq(negZero));
console.writeln(zero.bits());
console.writeln(negZero.bits());
false
true
true
true
0
128
ceil
ceil() -> float8Round up to a whole number.
Returns
The smallest whole number not less than this, as a float8.
Examples
float8 pos = 2.5;
float8 neg = -2.5;
console.writeln(pos.ceil());
console.writeln(neg.ceil());
3.000000
-2.000000
See also: floor
floor
floor() -> float8Round down to a whole number.
Returns
The largest whole number not greater than this, as a float8.
Examples
float8 pos = 2.5;
float8 neg = -2.5;
console.writeln(pos.floor());
console.writeln(neg.floor());
2.000000
-3.000000
See also: ceil
isInfinite
isInfinite() -> boolTest whether the value is infinite.
A float8 has no infinity, so this is always false, even for NaN produced by overflow.
Returns
Always false.
Examples
float8 big = 448.0;
float8 two = 2.0;
console.writeln(big.isInfinite());
console.writeln((big * two).isInfinite());
float8 one = 1.0;
float8 zero = 0.0;
console.writeln((one / zero).isInfinite());
false
false
false
isNaN
isNaN() -> boolTest whether the value is NaN.
NaN (not a number) is what an undefined operation produces, such as the square root of a negative number. NaN is the only value that is not equal to itself, so x.isNaN() is the reliable test.
Returns
true when this is NaN.
Examples
float8 one = 1.0;
console.writeln(one.isNaN());
float8 big = 448.0;
float8 two = 2.0;
console.writeln((big * two).isNaN());
console.writeln(float8::fromBits(127).isNaN());
false
true
true
pow
pow(float8 e) -> float8Raise the value to a power.
The result is computed in higher precision and rounded to the nearest float8. Because float8 has no infinity, a result above 448 is NaN.
Parameters
- e
- The exponent.
Returns
this to the power e, rounded to a float8.
Examples
float8 two = 2.0;
console.writeln(two.pow(two));
float8 half = 0.5;
console.writeln(half.pow(half));
float8 big = 448.0;
console.writeln(big.pow(two));
4.000000
0.687500
nan
round
round() -> float8Round to the nearest whole number.
A value exactly halfway between two whole numbers rounds away from zero, so 2.5 becomes 3 and -2.5 becomes -3.
Returns
The nearest whole number, as a float8.
Examples
float8 pos = 2.5;
float8 neg = -2.5;
console.writeln(pos.round());
console.writeln(neg.round());
3.000000
-3.000000
See also: trunc
sqrt
sqrt() -> float8Compute the square root.
The result is rounded to the nearest float8. The square root of a negative number is NaN; it does not throw.
Returns
The square root of this, as a float8.
Examples
float8 nine = 9.0;
console.writeln(nine.sqrt());
float8 two = 2.0;
console.writeln(two.sqrt());
float8 minusOne = -1.0;
console.writeln(minusOne.sqrt());
3.000000
1.375000
nan
toFloat
toFloat() -> floatConvert the float8 to a float.
The conversion is exact: every value of a float8 is representable as a float, so nothing is rounded. Converting a float back with float.toFloat8() rounds to the nearest float8 and throws when the value is out of range.
Returns
The value as a float.
Examples
float8 x = 0.1;
console.writeln(x.toFloat());
console.writeln(x.toFloat() == 0.1);
0.101562
false
See also: toFloat8
toInt
toInt() -> intConvert the float8 to an int, discarding the fraction.
The value is truncated toward zero, so -2.5 becomes -2.
Returns
The whole-number part of this as an int.
Throws
RuntimeException- when the value is NaN or infinite.
Examples
float8 v = -2.5;
console.writeln(v.toInt());
float8 big = 448.0;
console.writeln(big.toInt());
float8 nan = float8::fromBits(127);
try {
console.writeln(nan.toInt());
} catch (RuntimeException e) {
console.writeln("caught: ${e.message}");
}
-2
448
caught: float is not finite or out of int64 range for toInt()
toString
toString() -> stringFormat the float8 as decimal text.
The text is the six-decimal form that float prints, for example 3.500000, computed from the value actually stored, so a number that is not exactly representable shows its rounded value. NaN prints as nan and infinity as inf or -inf.
Returns
The decimal text of this.
Examples
float8 x = 0.1;
console.writeln(x.toString());
float8 y = 3.5;
console.writeln("y = " + y.toString());
0.101562
y = 3.500000
trunc
trunc() -> float8Discard the fraction.
Returns
The whole-number part of this, rounded toward zero, as a float8.
Examples
float8 pos = 2.5;
float8 neg = -2.5;
console.writeln(pos.trunc());
console.writeln(neg.trunc());
2.000000
-2.000000
See also: floor