LEVIATHAN v962456e · 962456eee1

Standard Library

namespace encoding

Text encodings for byte strings: hexadecimal, base64, URL-safe base64, and percent-encoding.

since 0.1.0-alpha.1linuxwindowswasm

Overview

Every encoder takes a string and works on its bytes, so non-ASCII text is encoded as its UTF-8 bytes. Every decoder is total: input that is not valid for the encoding gives None instead of throwing, so check the result before using it.

Description

The encoding namespace converts between a string of bytes and a printable text form. Each codec comes as an encoder that cannot fail and a decoder that returns string?. Every decoder is total: malformed input produces None, never an exception, because bad data is an expected result, not a programming error.

encoder decoder form
base64Encode(bytes) base64Decode(text) standard base64 with +, / and = padding
base64UrlEncode(bytes) base64UrlDecode(text) URL-safe base64: - and _, no = padding
percentEncode(text) percentDecode(text) RFC 3986 percent encoding
hexEncode(bytes) hexDecode(text) lowercase hexadecimal, two digits per byte

Strings are sequences of bytes here: a non-ASCII character is encoded as its UTF-8 bytes, and a decoder returns the decoded bytes as a string.

The four codecs

console.writeln(encoding::base64Encode("Hello, World!"));
console.writeln(encoding::base64Decode("SGVsbG8sIFdvcmxkIQ==") ?? "malformed");
console.writeln(encoding::base64Encode("???"));
console.writeln(encoding::base64UrlEncode("???"));
console.writeln(encoding::percentEncode("a b&c=d/~"));
console.writeln(encoding::percentDecode("a%20b%26c") ?? "malformed");
console.writeln(encoding::hexEncode("Levi"));
console.writeln(encoding::hexDecode("4c657669") ?? "malformed");
SGVsbG8sIFdvcmxkIQ==
Hello, World!
Pz8/
Pz8_
a%20b%26c%3Dd%2F~
a b&c
4c657669
Levi

Rules

  • base64Decode is strict: the length must be a multiple of four, = may appear only at the end of the final group, and any byte outside the base64 alphabet makes the result None.
  • base64UrlEncode never writes =. base64UrlDecode accepts input with the padding removed and returns None for a length that no base64 text can have (one more than a multiple of four).
  • percentEncode leaves A-Z, a-z, 0-9, -, _, . and ~ as they are and writes every other byte as % followed by two uppercase hexadecimal digits. A space becomes %20, never +.
  • percentDecode turns %XX into a byte, accepting upper or lower case digits, and leaves every other byte alone, including +. A % not followed by two hexadecimal digits makes the result None.
  • hexEncode writes lowercase digits. hexDecode accepts both cases and returns None for an odd length or a non-hexadecimal character.

Examples

Decoders return None on malformed input

console.writeln(encoding::base64Decode("YQ") == None);
console.writeln(encoding::base64Decode("Y!==") == None);
console.writeln(encoding::base64UrlDecode("YQ") ?? "malformed");
console.writeln(encoding::percentDecode("100%") == None);
console.writeln(encoding::percentDecode("a+b") ?? "malformed");
console.writeln(encoding::hexDecode("abc") == None);
console.writeln(encoding::hexDecode("4C657669") ?? "malformed");
true
true
a
true
a+b
true
Levi

Non-ASCII text is encoded byte by byte:

UTF-8 bytes in percent encoding

string s = encoding::percentEncode("café");
console.writeln(s);
console.writeln(encoding::percentDecode(s) ?? "malformed");
console.writeln(encoding::hexEncode("é"));
caf%C3%A9
café
c3a9

Examples

Encode and decode the same text four ways

string text = "hi?>";
console.writeln(encoding::hexEncode(text));
console.writeln(encoding::base64Encode(text));
console.writeln(encoding::base64UrlEncode(text));
console.writeln(encoding::percentEncode(text));
console.writeln(encoding::base64Decode("aGk/Pg==") ?? "invalid");
console.writeln(encoding::hexDecode("xyz") ?? "invalid");
68693f3e
aGk/Pg==
aGk_Pg
hi%3F%3E
hi?>
invalid

Functions

base64Decode

base64Decode(string b64) -> string | None

Decode standard base64 text back into the bytes it describes.

Decoding is strict. The length must be a multiple of four, only the alphabet characters may appear, and = padding is allowed only at the very end of the input. Anything else gives None; in particular, unpadded input such as aGVsbG8 is rejected, and so is URL-safe base64 (use base64UrlDecode for that).

Parameters

b64
The base64 text to decode.

Returns

The decoded string, or None when b64 is not valid padded base64.

Examples

console.writeln(encoding::base64Decode("aGVsbG8=") ?? "invalid");
console.writeln(encoding::base64Decode("aGVsbG8") ?? "invalid");
console.writeln(encoding::base64Decode("aGk_Pg==") ?? "invalid");
hello
invalid
invalid

See also: base64Encode, base64UrlDecode

base64Encode

base64Encode(string bytes) -> string

Encode a string's bytes as standard base64 with = padding.

The alphabet is A-Z, a-z, 0-9, + and /. The output length is always a multiple of four, padded with one or two = characters when the input length is not a multiple of three.

Parameters

bytes
The string whose bytes are encoded.

Returns

The padded base64 text, for example aGVsbG8= for hello.

Examples

console.writeln(encoding::base64Encode("hello"));
console.writeln(encoding::base64Encode("hey"));
console.writeln(encoding::base64Encode("hi"));
aGVsbG8=
aGV5
aGk=

See also: base64Decode, base64UrlEncode

base64UrlDecode

base64UrlDecode(string s) -> string | None

Decode URL-safe base64 text, with or without padding.

The - and _ characters are mapped back to + and /, missing = padding is restored, and the result is decoded like base64Decode. A length that no base64 text can have (one more than a multiple of four) gives None, as does any invalid character.

Parameters

s
The URL-safe base64 text to decode.

Returns

The decoded string, or None when s is not valid URL-safe base64.

Examples

console.writeln(encoding::base64UrlDecode("aGk_Pg") ?? "invalid");
console.writeln(encoding::base64UrlDecode("aGk_Pg==") ?? "invalid");
console.writeln(encoding::base64UrlDecode("a") ?? "invalid");
hi?>
hi?>
invalid

See also: base64UrlEncode, base64Decode

base64UrlEncode

base64UrlEncode(string bytes) -> string

Encode a string's bytes as URL-safe base64 without padding.

This is the variant used in tokens and URLs. It writes - and _ where standard base64 writes + and /, and it leaves off the trailing = padding.

Parameters

bytes
The string whose bytes are encoded.

Returns

The URL-safe base64 text, with no = characters.

Examples

console.writeln(encoding::base64Encode("hi?>"));
console.writeln(encoding::base64UrlEncode("hi?>"));
aGk/Pg==
aGk_Pg

See also: base64UrlDecode, base64Encode

hexDecode

hexDecode(string s) -> string | None

Decode a hexadecimal string back into the bytes it describes.

Both lowercase and uppercase digits are accepted. The input must have an even number of characters and contain only hexadecimal digits; otherwise the result is None. The decoded bytes are returned as a string, so they should form valid UTF-8 if you intend to treat the result as text.

Parameters

s
The hexadecimal text to decode.

Returns

The decoded string, or None when s has an odd length or a non-hexadecimal character.

Examples

console.writeln(encoding::hexDecode("68656c6c6f") ?? "invalid");
console.writeln(encoding::hexDecode("4A4b") ?? "invalid");
console.writeln(encoding::hexDecode("abc") ?? "invalid");
console.writeln(encoding::hexDecode("zz") ?? "invalid");
hello
JK
invalid
invalid

See also: hexEncode

hexEncode

hexEncode(string bytes) -> string

Encode every byte of a string as two lowercase hexadecimal digits.

The result is always twice as long as the input, in bytes. Non-ASCII characters contribute one pair per UTF-8 byte.

Parameters

bytes
The string whose bytes are encoded.

Returns

The lowercase hexadecimal text, for example 68656c6c6f for hello.

Examples

console.writeln(encoding::hexEncode("hello"));
console.writeln(encoding::hexEncode("é"));
68656c6c6f
c3a9

See also: hexDecode

percentDecode

percentDecode(string s) -> string | None

Decode percent-encoded text.

Each %XX group is replaced by the byte it names, and every other byte is copied unchanged. A + stays a +; it is not turned into a space. A % that is not followed by two hexadecimal digits gives None.

Parameters

s
The percent-encoded text.

Returns

The decoded string, or None when s contains a malformed % escape.

Examples

console.writeln(encoding::percentDecode("a%20b%26c") ?? "invalid");
console.writeln(encoding::percentDecode("a+b") ?? "invalid");
console.writeln(encoding::percentDecode("100%") ?? "invalid");
console.writeln(encoding::percentDecode("a%2") ?? "invalid");
console.writeln(encoding::percentDecode("a%zz") ?? "invalid");
a b&c
a+b
invalid
invalid
invalid

See also: percentEncode

percentEncode

percentEncode(string s) -> string

Percent-encode a string for use inside a URL component.

Every byte other than the unreserved characters A-Z, a-z, 0-9, -, _, . and ~ is written as % followed by two uppercase hexadecimal digits. A space becomes %20 (not +), and a non-ASCII character becomes one %XX group per UTF-8 byte. Characters such as /, & and = are encoded too, so the result is safe to use as a single query value or path segment.

Parameters

s
The text to encode.

Returns

The percent-encoded text.

Examples

console.writeln(encoding::percentEncode("a b&c=d/e"));
console.writeln(encoding::percentEncode("café"));
console.writeln(encoding::percentEncode("safe-text_1.0~"));
a%20b%26c%3Dd%2Fe
caf%C3%A9
safe-text_1.0~

See also: percentDecode

See also

  • digest — Message digests and keyed message authentication: MD5, SHA-1, SHA-256 and HMAC-SHA-256.
  • base64Encode — Encode a string's bytes as standard base64 with = padding.
  • hexEncode — Encode every byte of a string as two lowercase hexadecimal digits.
  • percentEncode — Percent-encode a string for use inside a URL component.