Chapter 13 · Machines
Case study: the drone
The drone on the front page is flown by 128 spiking neurons that learned with nothing to imitate, by gradients through their own spikes and through the drone's physics. This chapter puts the course to work on it, from the equations of the drone to what the network does in a recovery, and how much of it the flying needs.
The task
A drone hangs in a vertical plane: a bar with a rotor at each end, which can only push along the bar’s up axis. To move sideways it must tilt, to tilt it must push harder on one side, and every tilt costs it lift. Gravity never stops. A network that flies it reads where the target is and how the drone moves, and sets the two rotors’ thrust 100 times a second.
Nobody shows it how. There are no recorded flights to imitate and no reward to estimate. The network learns from one number, how far the drone ends up from its target, and from the derivative of that number with respect to each of its weights, carried back through two seconds of flight.
The drone
The drone has a mass of 1 kg, arms of 25 cm, a moment of inertia of 0.04 kg m², and rotors that push at most 12 N each. Its state is its position, velocity, angle and spin . With thrusts and , and some air drag:
Every 10 ms the physics takes one semi-implicit Euler step: velocity and spin first, then position and angle from the new ones. A drone that hits the edge of its box bounces off at 30% of its speed.
The network
The network reads seven numbers each step: the way to the target, clipped to 1.5 m so that a far target reads as “that way”, the velocity, the sine and cosine of the angle, and the spin. They enter a dense layer as currents. Two layers of 64 LIF neurons follow, with a time constant of 3 steps and a reset to zero, then two leaky integrators. Each integrator’s membrane sets one rotor’s thrust,
so a membrane of zero holds the drone in a hover. Nothing else carries memory: the membranes are the network’s whole state.
import flax.linen as nnimport jaximport jax.numpy as jnp
from sparx.nn import LI, LIF
# 7 readings in: the way to the target, velocity, attitude, spin.# 2 membranes out: the rotors' thrust.pilot = nn.Sequential([ nn.Dense(64), LIF(tau=3.0, reset="zero"), nn.Dense(64), LIF(tau=3.0, reset="zero"), nn.Dense(2), LI(tau=5.0),])
def step(params, carried, readings): """10 ms of the network, its membranes carried in `state`.""" out, mutated = pilot.apply( {"params": params, "state": carried}, readings[None], mutable=["state"]) return out[0], mutated["state"]
params = pilot.init(jax.random.key(0), jnp.zeros((1, 1, 7)))["params"]# 256 drones, every neuron at restmembranes, carried = step(params, {}, jnp.zeros((256, 7)))# Training scans step() and the drone's physics over 2 s of flight# and takes jax.grad of a cost led by the distance to the target:# through the spikes by their surrogate, and through the physics.Learning through the physics
The physics is written in JAX, like the network, so a whole flight is one differentiable function. Each training step flies 256 drones for 2 s from random states, 40% of them at any angle, in boxes from a phone’s shape to an ultrawide’s, chasing targets that jump every 1.2 s on average, through gusts. The loss is
Each is an average, over drones and steps in the first three terms and over neurons in the last. The first term is the distance to the target. The next two discourage tilt and spin. In the last, is the share of steps on which neuron fired: it penalizes a rate outside 2% to 30% of steps, which keeps every neuron in use and none saturated. JAX differentiates through all 200 steps: through the drone’s equations, and through the spikes by their surrogate, chapter 5’s stand-in for the threshold’s missing derivative. This is chapter 6’s backpropagation through time with the world inside the loop. Adam follows the gradient, clipped to a norm of 1, which keeps an occasional exploding one from undoing the rest.
Inside a recovery
Pick a start and watch one flight of 3 s to the green circle. On the right are the spikes of all 128 neurons over the flight, and below them each rotor’s thrust, with the thrust that hovers dashed.
3 s from the start · left and right rotor, dashed at hover · …
From upside down, the drone is upright within 0.81 s and stays within 15 cm of the target from 1.55 s. For about half of its first 0.8 s each rotor is within 0.5 N of a limit, 0 or 12 N, against about a tenth of the last 1.5 s. The network fires most while it turns over. In the first half second it fires 2,010 spikes, 31 a second per neuron, and in a steady hover about 860, 13 a second. Thrown and spinning at 12 rad/s, it arrives in 1.31 s; falling sideways, in 0.95 s.
How well it flies
On 1,000 flights of 6 s to a fixed target from random starts, half of them at any angle, it arrived within 15 cm and stayed there on 100%, in a median of 0.93 s. That includes all 252 starts more than a quarter turn from upright. On 1,000 more, thrown and spun at random, with speeds of 4 m/s and spins of 15 rad/s as standard deviations, it arrived on 100%. After 6 s the median drone was 1.6 cm from its target.
Silencing neurons
On the front page you can silence neurons and watch the flying get worse. Here is the same experiment in sparx: silence neurons at random from both layers, five times for each count, and fly 500 flights each time.
The network is not uniformly redundant. Most sets of 8 or 16 silenced neurons leave it flying, but one set of 8 grounded it entirely: some neurons carry something the rest cannot replace. By 48 the drone mostly falls, and past 64 it never arrives.
The browser runs sparx’s network
The front page loads the trained weights, 4,802 numbers, and steps the network and the physics in TypeScript, in double precision, with the same operations in the same order as sparx. When this page was built, it flew sparx's two recorded float64 flights, 1,400 steps with targets, gusts and spins, and fired the same 38,093 spikes, every one at the same step. The drone's state stayed within 4e-13 of sparx's.
sparx also exports the network through NIR, which chapter 12 showed node by node.
Try this
- In the recovery figure, why do the rotors sit at 0 or 12 N for much of a recovery, instead of in between?
- The loss never asks the drone to stay upright when far from its target. What in it does, and why is its weight small?
- Why does the network need the sine and cosine of the angle, rather than the angle itself?
- Training flew 200 steps at a time. What would go wrong with 2,000?
Answers
- Rolling over or stopping a throw needs the largest torque and force the drone has, and the loss rewards arriving sooner. Gradient descent pushes the readout membranes far past the sigmoid’s middle, where the thrust saturates.
- The tilt term, . To move sideways the drone must tilt, so a large weight would fight the distance term; a small one breaks ties in favour of flying level.
- An angle wraps: and are the same attitude, but as numbers they are far apart, and a turn of changes the number without changing the drone. The sine and cosine are the same for the same attitude and change smoothly.
- The gradient would pass back through ten times as many steps of the drone and the network, multiplying ten times as many Jacobians. Chapter 6 showed what that does: it explodes or vanishes, and the memory to store the flight grows tenfold.
Summary
A spiking network learned to fly a drone from one signal, its distance to the target, by gradients carried back through two seconds of its own spikes and the drone’s physics. It recovers from any attitude, firing hardest while it does, and it still flies with a few neurons silenced, though not with any few. The browser runs the same network, spike for spike.
References
- E. O. Neftci, H. Mostafa and F. Zenke, “Surrogate gradient learning in spiking neural networks”, IEEE Signal Processing Magazine 36, 2019, doi:10.1109/MSP.2019.2931595.
- D. P. Kingma and J. Ba, “Adam: a method for stochastic optimization”, International Conference on Learning Representations, 2015, arXiv:1412.6980.
- R. Pascanu, T. Mikolov and Y. Bengio, “On the difficulty of training recurrent neural networks”, International Conference on Machine Learning, 2013, arXiv:1211.5063.