Synapse logo

Rust · reverse-mode autograd · no ML frameworks

A neural network,
built from scratch.

Synapse is a reverse-mode autograd engine you can read node by node, and a tiny MLP built on top of it, trained on XOR. No ML frameworks, no tensor crates. Here is the real thing, training in your browser.

Try the live demo View on GitHub
How to use this playground
Train a from-scratch neural net on XOR and watch the loss fall and the decision boundary form in real time.
Live training run

Not a canned animation. The buttons below drive a scalar autograd engine and a 2 4 4 1 tanh MLP ported line for line from src/value.rs and src/nn.rs, running right here in JavaScript. Every click runs a real forward pass, a real backward pass through the computation graph, and a real gradient descent step.

Loss curve, drawn live from the real MSE at each epoch

Decision boundary, the trained net evaluated over a grid

0
Epoch
1.0000
Loss (MSE)
0/4
XOR points correct
x1x2targetpredictionresult

The same run, in the terminal.

The browser demo mirrors the Rust binary. This is the real output of a release build: the loss falling every 20 steps, then the final rounded prediction for each XOR input.

$ cargo run --release
step 0 loss 1.794003
step 20 loss 0.231178
step 60 loss 0.157512
step 100 loss 0.063884
step 200 loss 0.000137
step 299 loss 0.000000

final predictions:
  [0.0, 0.0] -> 0.0002 (target 0, rounds to 0)
  [0.0, 1.0] -> 0.9998 (target 1, rounds to 1)
  [1.0, 0.0] -> 0.9998 (target 1, rounds to 1)
  [1.0, 1.0] -> 0.0004 (target 0, rounds to 0)
// loss printed every 20 steps; abridged for space

The autograd idea

Every value in Synapse remembers how it was made. Add, multiply, subtract, divide, raise to a power, negate, or pass through tanh, and the operation is recorded as a node in a graph with pointers back to its inputs. Call backward() once on the output and gradients for every value in the graph fall out through the chain rule, in reverse topological order, with shared subexpressions correctly accumulating gradient from every place they are used.

Value graph

A scalar Value wraps its data, its gradient, and the operation and parent nodes that produced it.

Topological backward pass

One backward() call walks the graph in reverse dependency order and pushes gradients through every op.

Gradient accumulation

A value used twice, like a * a, sums gradient contributions from each use instead of overwriting.

Design choices that keep it readable

Synapse is written to be read, not to be fast. A handful of deliberate decisions keep the whole engine to a few hundred lines while staying numerically correct.

Scalar, not tensor

Every Value is a single f64 node in a graph. There is no tensor abstraction hiding the arithmetic, so you can follow the math one node at a time.

Five backward rules

Division is multiplication by pow(-1) and subtraction is addition of a negation, so the only primitive gradients that need to exist are for add, mul, pow, tanh, and neg.

Shared nodes, not copies

Cloning a Value clones an Rc, not the node. So a * a refers to one node, and its gradient contributions accumulate with +=, matching the multivariable chain rule.

One backward call

A reverse-topological depth-first walk visits every contributing node exactly once, so a single backward() fills the gradient of every weight and bias in the network.

Deterministic init

Weights are seeded from a ChaCha8Rng, so every run reproduces exactly. The browser demo uses the same idea, so Reset always replays the same training.

Read it, run it, test it

Four small files and a test suite. That is the whole project.

src/value.rs

The Value type and the reverse-mode autograd engine: the operator overloads, the topological build, and the backward pass.

src/nn.rs

Neuron, Layer, and MLP. Forward pass is plain Value arithmetic, so gradients for every parameter fall out of one backward().

src/main.rs

The XOR demo: builds a 2 4 4 1 MLP, trains it with mean squared error, and prints the loss curve and final predictions.

tests/

A gradient check against finite differences, a shared-node test that d = a * a gives 2a, and an end-to-end test that XOR trains below a loss threshold.

git clone https://github.com/pavanchow/synapse
cd synapse
cargo run --release   # trains XOR, prints the loss curve
cargo test            # gradient check, shared-node test, XOR training test