~/bend-docscommunity

math_fp32.bend checks

raw source on the hub · import stelliferous@0.0.2.0/math_fp32.bend as Math_fp32

Ordered FP32 arithmetic on the shared sequence model. No separate container.

2 imports
import Base
import ./math_sequence.bend as Sequence

Definitions

def multiply_add source · line 5 · raw

@accumulator:F32 -> @input:F32 -> @weight:F32 -> F32

def weighted_add source · line 8 · raw

@weight:F32 -> @accumulator:F32 -> @input:F32 -> F32

def zip_mac source · line 11 · raw

@length:Nat -> @accumulator:0x9be3b13bf249759bc81e1958bcd1a4c0/math_sequence.Values(F32, length) -> @input:0x9be3b13bf249759bc81e1958bcd1a4c0/math_sequence.Values(F32, length) -> @weight:F32 -> 0x9be3b13bf249759bc81e1958bcd1a4c0/math_sequence.Values(F32, length)

def exponential_polynomial source · line 15 · raw

@+r:F32 -> F32

The Cephes expf polynomial after its constant and linear terms.

def from_bits source · line 19 · raw

@value:U32 -> F32

The 32-bit word of a U32, read as an F32.

def exponential_remainder source · line 25 · raw

@+r:F32 -> F32

exp(r)-1-r near |r| <= ln(2)/2. Chebyshev coefficients, rounded to F32; paired evaluation shortens the dependency chain without contracting products.

def logistic_denominator source · line 32 · raw

@q:F32 -> @+base:F32 -> @r:F32 -> @correction:F32 -> F32

b=RN(1+q). Recover the small term lost in that sum before adding the reduced exponential. Writing (1-b)+q keeps the eight lanes together under the default vectorizer; q-(b-1) makes it mix unrelated addition trees.

def selected source · line 35 · raw

@+mask:U32 -> @yes:F32 -> @+no:F32 -> F32

def silu_finish source · line 40 · raw

@+x:F32 -> @value:F32 -> F32

Above 18 SiLU rounds to x. Pass x through, including positive infinity; the negative tail uses the clamped input before the final underflow.

def logistic_scaled source · line 46 · raw

@silu:Bool -> @x:F32 -> @value:F32 -> @power:F32 -> @denominator:F32 -> F32

def sigmoid_reduced source · line 58 · raw

@silu:Bool -> @x:F32 -> @+value:F32 -> @+biased:F32 -> F32

u=-value, n approximates u/ln(2) to the nearest integer, r=u-n*ln(2). The exponent word makes H=2^(64-n), so sigmoid(x)=H/(exp(r)+2^-n)*2^-64. H remains normal; only the last multiplication underflows. The tiny denominator addend q may be floored to the least normal F32 without changing its rounded sum.

def sigmoid_input source · line 65 · raw

@silu:Bool -> @x:F32 -> @+value:F32 -> F32

def sigmoid source · line 71 · raw

@+x:F32 -> F32

At x>=18 sigmoid rounds to 1. Below -110 both sigmoid and SiLU round to zero; multiplying the clamped x before the final scaling preserves -0 for SiLU, including at -infinity. NaNs pass the clamp and propagate.

def activate source · line 74 · raw

@+x:F32 -> @enabled:Bool -> F32

def tanh_finish source · line 83 · raw

@+x:F32 -> @+value:F32 -> F32

Sign restoration and the tiny interval are word selections. For |x|<2^-12, |x-tanh(x)|<|x|^3/3 is below half an F32 ulp, so return x, including -0.

def tanh_magnitude source · line 91 · raw

@+x:F32 -> F32

At |x|>=10, 1-tanh(|x|)<2*exp(-20)<2^-25. Clamping before arithmetic also keeps infinities away from intermediate products; NaNs propagate.

def tanh_small source · line 96 · raw

@+t:F32 -> @+z:F32 -> F32

Chebyshev interpolation of (tanh(sqrt(z))/sqrt(z)-1)/z, |t|<=0.75. The coefficients are F32 values.

def tanh_remainder source · line 100 · raw

@+r:F32 -> F32

One additional term for the expm1 correction in the tanh quotient.

def tanh_exponential source · line 103 · raw

@+a:F32 -> @+biased:F32 -> @x:F32 -> F32

def tanh_input source · line 113 · raw

@+a:F32 -> @x:F32 -> F32

def tanh source · line 116 · raw

@+x:F32 -> F32

def exponential_scale source · line 122 · raw

@biased:F32 -> @x:F32 -> @p:F32 -> F32

Split 2^n at +64 for nonnegative inputs and -64 for negative inputs. The first product is always normal and exact; only the last product can underflow. The exponent words come directly from the biased integer n.

def exponential_precise source · line 128 · raw

@+r:F32 -> F32

FastTwoSum retains the low bits of 1+r before adding the quadratic tail.

def exponential_reduced source · line 133 · raw

@+x:F32 -> @+biased:F32 -> F32

def exponential_input source · line 140 · raw

@+x:F32 -> F32

def exp source · line 143 · raw

@x:F32 -> F32

def logarithm_series source · line 151 · raw

@+r:F32 -> @+n:F32 -> F32

For s near r/(2+r), d = r-(2+r)*s gives the exact real identity log(1+r) = 2*atanh(s) + log(1+d/(1+s)). The odd series ends at s^11. Split r and s into 12-bit heads so their head product is exact; retain the low products in d. (1-s)*(1+s^2+s^4) approximates 1/(1+s). FastTwoSum retains the low bits when adding the exponent's ln(2) part.

def logarithm_normalized source · line 170 · raw

@x:F32 -> @adjustment:F32 -> F32

The exponent carry at mantissa word 0x3504f4 selects the same reduced interval as halving m above sqrt(2), without a floating-point selection.

def logarithm_finish source · line 178 · raw

@+x:F32 -> @value:F32 -> F32

Scaling a subnormal by 2^23 is exact. Zero, negatives, infinities and NaNs are selected explicitly, after normalizing the finite positive input.

def logarithm_input source · line 186 · raw

@+x:F32 -> F32

def log source · line 190 · raw

@x:F32 -> F32