~/bend-docscommunity

src/lex.bend checks

raw source on the hub · import emerging-ezjson@1.1.0.0/src/lex.bend as Lex

src/lex: a JSON text as tokens. A string is decoded: \", \\, \/, \b, \f, \n, \r, \t, and \uXXXX. A \u escape of a surrogate pair is one code point.

1 import
import Base

Types

type Tok source · line 7 · raw

Data

the tokens: brackets, :, ,, a string (decoded), or a bare word (a number, true, false, null)

type Class source · line 22 · raw

Data

what a char is: punctuation (with its token), a quote, whitespace, or a word char

type Mode source · line 31 · raw

Data

what the machine is inside of: nothing, a bare word, a string, an escape, a \u escape, a high surrogate waiting for its pair, or a rejected text

type Sur source · line 43 · raw

Data

a \u code point: an ordinary scalar, a high surrogate, or a low surrogate

type Lex source · line 49 · raw

Data

the lexer's state: buf and toks are both reversed

type Span source · line 332 · raw

Data

what the span scanner is inside: nothing, a word, a string, or a rejected text

type Scan source · line 340 · raw

Data

cut is the suffix where the open word or string began, so closing it does not walk the text again. Old means a backslash handed the tail to the copier

Definitions

def classify.go source · line 55 · raw

@+cp:U32 -> Class

a char's class, by code point. Written as comparisons rather than as case '[': arms: a character literal in a pattern is a U32 literal inside a constructor pattern, and bend's C backend pays about 90 MB for each one.

def classify source · line 71 · raw

@ch:Char -> Class

a char's class

def text.go source · line 75 · raw

@buf:List<&2, Char> -> @acc:String -> String

the buffer is reversed, so consing each head onto the front yields the text

def text source · line 83 · raw

@buf:List<&2, Char> -> String

the buffered chars (reversed) as a string

def unescape.go source · line 86 · raw

@+cp:U32 -> Maybe<&2, Char>

def unescape source · line 98 · raw

@+ch:Char -> Maybe<&2, Char>

the char an escape stands for; none when the escape is not one JSON allows

def escape.put source · line 101 · raw

@mb:Maybe<&2, Char> -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def escape.go source · line 109 · raw

@ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> @is_u:Bool -> Lex

the char after a backslash, once it is known whether it is u

def escape source · line 118 · raw

@+ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

the char after a backslash in a string. Tested by code point, not with a case 'u': arm, for the reason classify gives

def hex.ok source · line 121 · raw

@+cp:U32 -> Bool

def hex source · line 127 · raw

@ch:Char -> U32

a hex digit's value

def unicode.sur.lo source · line 132 · raw

@ge:Bool -> @le:Bool -> Sur

def unicode.sur.hi source · line 139 · raw

@+cp:U32 -> @ge:Bool -> @le:Bool -> Sur

def unicode.sur source · line 146 · raw

@+cp:U32 -> Sur

def unicode.pair source · line 149 · raw

@hi:U32 -> @lo:U32 -> U32

def unicode.end.go source · line 152 · raw

@cp:U32 -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> @key:Sur -> Lex

def unicode.end source · line 161 · raw

@+cp:U32 -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def unicode.put source · line 164 · raw

@left:Nat -> @acc:U32 -> @dig:U32 -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def unicode.go source · line 171 · raw

@left:Nat -> @acc:U32 -> @ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> @ok:Bool -> Lex

def unicode source · line 179 · raw

@left:Nat -> @acc:U32 -> @+ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

one hex digit of a \uXXXX escape; the code point once four are in

def unicode.low.end.go source · line 182 · raw

@lo:U32 -> @hi:U32 -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> @key:Sur -> Lex

def unicode.low.end source · line 189 · raw

@+lo:U32 -> @hi:U32 -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def unicode.low.put source · line 192 · raw

@left:Nat -> @acc:U32 -> @hi:U32 -> @dig:U32 -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def unicode.low.go source · line 199 · raw

@left:Nat -> @acc:U32 -> @hi:U32 -> @ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> @ok:Bool -> Lex

def unicode.low source · line 214 · raw

@left:Nat -> @acc:U32 -> @hi:U32 -> @+ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def step.str.go source · line 217 · raw

@ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> @ctrl:Bool -> Lex

def step.str source · line 224 · raw

@+ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def step.hi.go source · line 227 · raw

@hi:U32 -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> @slash:Bool -> Lex

def step.hi source · line 234 · raw

@hi:U32 -> @ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def step.hiesc.go source · line 237 · raw

@hi:U32 -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> @is_u:Bool -> Lex

def step.hiesc source · line 244 · raw

@hi:U32 -> @ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

def step source · line 248 · raw

@mode:Mode -> @cls:Class -> @ch:Char -> @buf:List<&2, Char> -> @toks:List<&2, Tok> -> Lex

one char, by the mode and the char's class

def feed source · line 290 · raw

@+ch:Char -> @st:Lex -> Lex

one char into the machine

def run source · line 295 · raw

@cs:List<&2, Char> -> @st:Lex -> Lex

every char, in order

def finish source · line 304 · raw

@st:Lex -> List<&2, Tok>

the tokens, in order; text that ends inside a string or a word ends in TBad when the string was not closed

def tokens.rev source · line 315 · raw

@src:String -> @+nn:U32 -> @zero:Bool -> @acc:List<&2, Char> -> List<&2, Char>

the first n chars of src, reversed, for the escape fallback's buffer

def tokens.push source · line 324 · raw

@cut:String -> @nn:U32 -> @word:Bool -> @toks:List<&2, Tok> -> List<&2, Tok>

def tokens.hit source · line 344 · raw

@+ch:Char -> @tail:String -> @cut:String -> @+nn:U32 -> @mode:Span -> @toks:List<&2, Tok> -> @cls:Class -> @ctrl:Bool -> Scan

def tokens.done source · line 386 · raw

@cut:String -> @nn:U32 -> @mode:Span -> @toks:List<&2, Tok> -> List<&2, Tok>

def tokens.end source · line 395 · raw

@st:Scan -> List<&2, Tok>

def tokens.feed source · line 402 · raw

@+ch:Char -> @tail:String -> @st:Scan -> Scan

def tokens.go source · line 410 · raw

@txt:String -> @st:Scan -> List<&2, Tok>

match only the text, so a literal unrolls. The tail is where the next span starts

def tokens source · line 418 · raw

@txt:String -> List<&2, Tok>

a text as tokens. Clean words and strings are spans of the text