Gradient kernels as a production object

The BIF thread ended with a verdict: the loss kernel is the lossy shadow, and sampling it is the expensive way to compute a GEMM. This thread takes the hint and builds the GEMM into an instrument — the preconditioned per-token gradients as they are computed, know the error in advance, and feed the result to things that want an N×N similarity: conductance clustering, influence retrieval, training-dynamics probes. Each piece stands on its own; read top-down for the arc. ← back to Vibe Research