~/bend-docscommunity

encoding.bend checks

raw source on the hub · import bend-kit-encoding@0.3.0.0/encoding.bend as Encoding

UTF-8 encoding and decoding between text and Bytes. Source: https://github.com/paymog/bend-kit/tree/main/encoding

2 imports
import Base
import 0x49814d83de8f70993a43e1002be29ecd/bytes.bend as Bytes

Types

type Out source · line 9 · raw

Type

Octets being written: the buffer, the count so far, and the word being filled.

type Utf8 source · line 74 · raw

Data

Decoder state: continuation bytes still needed, the allowed range of the next one, the code point so far, and the output reversed.

Definitions

def put source · line 13 · raw

@o:Out -> @+b:U32 -> Out

Bytes enter at the top of w, as in Bytes.from_string; every fourth byte stores the word.

def done source · line 19 · raw

@o:Out -> 0x49814d83de8f70993a43e1002be29ecd/bytes.Bytes

def utf8.width source · line 24 · raw

@+c:U32 -> U32

def utf8.size source · line 27 · raw

@s:String -> @+n:U32 -> U32

def utf8.cont source · line 34 · raw

@c:U32 -> @n:Nat -> U32

def utf8.push4 source · line 37 · raw

@+c:U32 -> @o:Out -> Out

def utf8.push3 source · line 40 · raw

@+c:U32 -> @o:Out -> @small:Bool -> Out

def utf8.push2 source · line 47 · raw

@+c:U32 -> @o:Out -> @small:Bool -> Out

def utf8.push source · line 54 · raw

@+c:U32 -> @o:Out -> @ascii:Bool -> Out

def utf8.encode.go source · line 61 · raw

@s:String -> @o:Out -> Out

def utf8.encode source · line 69 · raw

@+s:String -> 0x49814d83de8f70993a43e1002be29ecd/bytes.Bytes

Text to octets. One pass sizes the buffer; a second fills it.

def utf8.bad source · line 77 · raw

@out:String -> String

def utf8.lead.multi source · line 80 · raw

@+b:U32 -> @out:String -> Utf8

def utf8.lead.high source · line 86 · raw

@+b:U32 -> @out:String -> @bad:Bool -> Utf8

def utf8.lead source · line 94 · raw

@+b:U32 -> @out:String -> @ascii:Bool -> Utf8

WHATWG: a byte that cannot start a sequence decodes to U+FFFD.

def utf8.more source · line 101 · raw

@+need:U32 -> @cp:U32 -> @out:String -> @done:Bool -> Utf8

def utf8.cont.in source · line 109 · raw

@+need:U32 -> @cp:U32 -> @+b:U32 -> @out:String -> @ok:Bool -> Utf8

WHATWG: a byte that breaks a sequence yields U+FFFD and is read again as a lead.

def utf8.step.if source · line 116 · raw

@+need:U32 -> @+lo:U32 -> @+hi:U32 -> @cp:U32 -> @+b:U32 -> @out:String -> @idle:Bool -> Utf8

def utf8.step.st source · line 123 · raw

@st:Utf8 -> @+b:U32 -> Utf8

def utf8.finish source · line 127 · raw

@st:Utf8 -> String

def utf8.decode.go source · line 131 · raw

@n:Nat -> @r:Pair(Array<U32>, U32) -> @+i:U32 -> @st:Utf8 -> String

def utf8.decode source · line 141 · raw

@b:0x49814d83de8f70993a43e1002be29ecd/bytes.Bytes -> String

Octets to text, as WHATWG decodes UTF-8: malformed input becomes U+FFFD.