math_fp32.bend checks
raw source on the hub · import stelliferous@0.0.2.0/math_fp32.bend as Math_fp32
Ordered FP32 arithmetic on the shared sequence model. No separate container.
2 imports
import Base import ./math_sequence.bend as Sequence
Definitions
def multiply_add source · line 5 · raw
@accumulator:F32 -> @input:F32 -> @weight:F32 -> F32
def weighted_add source · line 8 · raw
@weight:F32 -> @accumulator:F32 -> @input:F32 -> F32
def zip_mac source · line 11 · raw
@length:Nat -> @accumulator:0x9be3b13bf249759bc81e1958bcd1a4c0/math_sequence.Values(F32, length) -> @input:0x9be3b13bf249759bc81e1958bcd1a4c0/math_sequence.Values(F32, length) -> @weight:F32 -> 0x9be3b13bf249759bc81e1958bcd1a4c0/math_sequence.Values(F32, length)
def exponential_polynomial source · line 15 · raw
@+r:F32 -> F32
The Cephes expf polynomial after its constant and linear terms.
def from_bits source · line 19 · raw
@value:U32 -> F32
The 32-bit word of a U32, read as an F32.
def exponential_remainder source · line 25 · raw
@+r:F32 -> F32
exp(r)-1-r near |r| <= ln(2)/2. Chebyshev coefficients, rounded to F32; paired evaluation shortens the dependency chain without contracting products.
def logistic_denominator source · line 32 · raw
@q:F32 -> @+base:F32 -> @r:F32 -> @correction:F32 -> F32
b=RN(1+q). Recover the small term lost in that sum before adding the reduced exponential. Writing (1-b)+q keeps the eight lanes together under the default vectorizer; q-(b-1) makes it mix unrelated addition trees.
def selected source · line 35 · raw
@+mask:U32 -> @yes:F32 -> @+no:F32 -> F32
def silu_finish source · line 40 · raw
@+x:F32 -> @value:F32 -> F32
Above 18 SiLU rounds to x. Pass x through, including positive infinity; the negative tail uses the clamped input before the final underflow.
def logistic_scaled source · line 46 · raw
@silu:Bool -> @x:F32 -> @value:F32 -> @power:F32 -> @denominator:F32 -> F32
def sigmoid_reduced source · line 58 · raw
@silu:Bool -> @x:F32 -> @+value:F32 -> @+biased:F32 -> F32
u=-value, n approximates u/ln(2) to the nearest integer, r=u-n*ln(2). The exponent word makes H=2^(64-n), so sigmoid(x)=H/(exp(r)+2^-n)*2^-64. H remains normal; only the last multiplication underflows. The tiny denominator addend q may be floored to the least normal F32 without changing its rounded sum.
def sigmoid_input source · line 65 · raw
@silu:Bool -> @x:F32 -> @+value:F32 -> F32
def sigmoid source · line 71 · raw
@+x:F32 -> F32
At x>=18 sigmoid rounds to 1. Below -110 both sigmoid and SiLU round to zero; multiplying the clamped x before the final scaling preserves -0 for SiLU, including at -infinity. NaNs pass the clamp and propagate.
def activate source · line 74 · raw
@+x:F32 -> @enabled:Bool -> F32
def tanh_finish source · line 83 · raw
@+x:F32 -> @+value:F32 -> F32
Sign restoration and the tiny interval are word selections. For |x|<2^-12, |x-tanh(x)|<|x|^3/3 is below half an F32 ulp, so return x, including -0.
def tanh_magnitude source · line 91 · raw
@+x:F32 -> F32
At |x|>=10, 1-tanh(|x|)<2*exp(-20)<2^-25. Clamping before arithmetic also keeps infinities away from intermediate products; NaNs propagate.
def tanh_small source · line 96 · raw
@+t:F32 -> @+z:F32 -> F32
Chebyshev interpolation of (tanh(sqrt(z))/sqrt(z)-1)/z, |t|<=0.75. The coefficients are F32 values.
def tanh_remainder source · line 100 · raw
@+r:F32 -> F32
One additional term for the expm1 correction in the tanh quotient.
def tanh_exponential source · line 103 · raw
@+a:F32 -> @+biased:F32 -> @x:F32 -> F32
def tanh_input source · line 113 · raw
@+a:F32 -> @x:F32 -> F32
def tanh source · line 116 · raw
@+x:F32 -> F32
def exponential_scale source · line 122 · raw
@biased:F32 -> @x:F32 -> @p:F32 -> F32
Split 2^n at +64 for nonnegative inputs and -64 for negative inputs. The first product is always normal and exact; only the last product can underflow. The exponent words come directly from the biased integer n.
def exponential_precise source · line 128 · raw
@+r:F32 -> F32
FastTwoSum retains the low bits of 1+r before adding the quadratic tail.
def exponential_reduced source · line 133 · raw
@+x:F32 -> @+biased:F32 -> F32
def exponential_input source · line 140 · raw
@+x:F32 -> F32
def exp source · line 143 · raw
@x:F32 -> F32
def logarithm_series source · line 151 · raw
@+r:F32 -> @+n:F32 -> F32
For s near r/(2+r), d = r-(2+r)*s gives the exact real identity log(1+r) = 2*atanh(s) + log(1+d/(1+s)). The odd series ends at s^11. Split r and s into 12-bit heads so their head product is exact; retain the low products in d. (1-s)*(1+s^2+s^4) approximates 1/(1+s). FastTwoSum retains the low bits when adding the exponent's ln(2) part.
def logarithm_normalized source · line 170 · raw
@x:F32 -> @adjustment:F32 -> F32
The exponent carry at mantissa word 0x3504f4 selects the same reduced interval as halving m above sqrt(2), without a floating-point selection.
def logarithm_finish source · line 178 · raw
@+x:F32 -> @value:F32 -> F32
Scaling a subnormal by 2^23 is exact. Zero, negatives, infinities and NaNs are selected explicitly, after normalizing the finite positive input.
def logarithm_input source · line 186 · raw
@+x:F32 -> F32
def log source · line 190 · raw
@x:F32 -> F32