Rust · reverse-mode autograd · no ML frameworks
Synapse is a reverse-mode autograd engine you can read node by node, and a tiny MLP built on top of it, trained on XOR. No ML frameworks, no tensor crates. Here is the real thing, training in your browser.
Not a canned animation. The buttons below drive a scalar autograd engine and a 2 4 4 1 tanh MLP ported line for line from src/value.rs and src/nn.rs, running right here in JavaScript. Every click runs a real forward pass, a real backward pass through the computation graph, and a real gradient descent step.
Loss curve, drawn live from the real MSE at each epoch
Decision boundary, the trained net evaluated over a grid
| x1 | x2 | target | prediction | result |
|---|
The browser demo mirrors the Rust binary. This is the real output of a release build: the loss falling every 20 steps, then the final rounded prediction for each XOR input.
Every value in Synapse remembers how it was made. Add, multiply, subtract, divide, raise to a power, negate, or pass through tanh, and the operation is recorded as a node in a graph with pointers back to its inputs. Call backward() once on the output and gradients for every value in the graph fall out through the chain rule, in reverse topological order, with shared subexpressions correctly accumulating gradient from every place they are used.
A scalar Value wraps its data, its gradient, and the operation and parent nodes that produced it.
One backward() call walks the graph in reverse dependency order and pushes gradients through every op.
A value used twice, like a * a, sums gradient contributions from each use instead of overwriting.
Synapse is written to be read, not to be fast. A handful of deliberate decisions keep the whole engine to a few hundred lines while staying numerically correct.
Every Value is a single f64 node in a graph. There is no tensor abstraction hiding the arithmetic, so you can follow the math one node at a time.
Division is multiplication by pow(-1) and subtraction is addition of a negation, so the only primitive gradients that need to exist are for add, mul, pow, tanh, and neg.
Cloning a Value clones an Rc, not the node. So a * a refers to one node, and its gradient contributions accumulate with +=, matching the multivariable chain rule.
A reverse-topological depth-first walk visits every contributing node exactly once, so a single backward() fills the gradient of every weight and bias in the network.
Weights are seeded from a ChaCha8Rng, so every run reproduces exactly. The browser demo uses the same idea, so Reset always replays the same training.
Four small files and a test suite. That is the whole project.
The Value type and the reverse-mode autograd engine: the operator overloads, the topological build, and the backward pass.
Neuron, Layer, and MLP. Forward pass is plain Value arithmetic, so gradients for every parameter fall out of one backward().
The XOR demo: builds a 2 4 4 1 MLP, trains it with mean squared error, and prints the loss curve and final predictions.
A gradient check against finite differences, a shared-node test that d = a * a gives 2a, and an end-to-end test that XOR trains below a loss threshold.
git clone https://github.com/pavanchow/synapse
cd synapse
cargo run --release # trains XOR, prints the loss curve
cargo test # gradient check, shared-node test, XOR training test