Chapter 14 · Machines
Case study: the racer
A spiking network drives a car around tracks it has never seen, with nothing but an event camera's output and its own speed to go on. It learned end to end, by gradients through its own spikes, through the car, and through the camera's pixels, whose thresholds are spikes too. This chapter takes it apart, and shows how often such training fails.
The task
A car carries an event camera that looks at the road ahead. The road is dark, its verges are light, and a dashed line runs down its middle. The camera sends nothing about the road itself, only where its brightness changes as the car moves, and a car that stands still sees nothing at all. From those events and its own speed, a network of spiking neurons must steer and set the speed, around a track it has never seen.
Nobody shows it how. Event-camera steering has been learned by imitation: Maqueda et al. trained a conventional network to predict a human driver’s steering from a large recorded dataset. The racer has no driver to imitate, and learns from what happens when it drives.
The car and its camera
The car is a kinematic bicycle: a wheelbase of 0.4 m, steering of at most 0.6 rad, a top speed of 6 m/s. Every 20 ms the network’s two readout membranes and set the steering and a target speed, and the car moves:
The camera has 24 × 12 pixels, each looking at one point of the ground: rows from 0.6 m to 7 m ahead, spaced as a real camera’s rows fall on a flat road, each spanning 0.9 times its distance to either side. Each pixel is chapter 12’s: it sends an ON or OFF event when the logarithm of the brightness it sees has moved 0.15 from the level of its last event. The network reads 577 numbers each step, the 288 pixels’ ON events, their 288 OFF events and the car’s speed, through two layers of LIF neurons, 96 and 64, into the two readout integrators.
events 0 · spikes 0 a step · …
Draw a track with the mouse, a loop around the car, or press “Draw a track” and draw it with a finger. Below the road are the camera’s pixels as they fire, the far rows at the top, orange for ON and blue for OFF, then the 160 neurons, and the steering and speed the readouts set. The track you draw is smoothed and scaled to the size of the tracks the network trained on, which have radii of 6 to 10 m. Draw corners much tighter than those and the car may leave the road; the figure then puts it back at the start.
Gradients through a camera
The car, the camera and the network are written in JAX, so a drive is one differentiable function of the network’s weights. Most of it is smooth: the car’s motion, the brightness of the ground, its logarithm. Two kinds of step are not, the neurons’ spikes and the pixels’ events, and both are sparx.spike, with chapter 5’s surrogate for a derivative. The pixels use a narrower one than the neurons, since their threshold is 0.15 rather than 1.
So the gradient of the loss reaches the weights by every path there is: through the readouts and the spikes, into the events, through each pixel’s threshold into the brightness it saw, through the ground into the car’s position and heading, and back through every earlier step of the drive. Here is one pixel row’s piece of that chain:
import jaximport jax.numpy as jnp
import sparxfrom sparx.surrogate import ATan
threshold, event = 0.15, ATan(alpha=8.0) # the racer's pixels
def brightness(x): """A road's edge: dark road, then a light verge past 0.8 m.""" return 0.12 + 0.7 * jax.nn.sigmoid((x - 0.8) / 0.12)
def on_events(shift): """ON events from 9 pixels across the edge when the car moves `shift` metres toward the verge.""" pixels = jnp.linspace(0.0, 1.6, 9) # metres from the middle before = jnp.log(brightness(pixels)) after = jnp.log(brightness(pixels + shift)) return sparx.spike(after - before - threshold, event).sum()
for shift in (0.05, 0.1, 0.2, 0.4): count, slope = jax.value_and_grad(on_events)(shift) print(f"move {shift:.2f} m: {int(count)} ON events," f" gradient {float(slope):.1f} per m")move 0.05 m: 1 ON events, gradient 29.3 per mmove 0.10 m: 3 ON events, gradient 15.2 per mmove 0.20 m: 3 ON events, gradient 8.0 per mmove 0.40 m: 5 ON events, gradient 5.6 per mThe count of events is a staircase in the car’s position, flat between its steps, so its true derivative is zero almost everywhere. The surrogate’s is never zero: between 0.1 and 0.2 m the count stays at 3, and the gradient still says that moving further brings more events.
The loss
Each training step draws 32 tracks like those you see, puts a car on each at a random point, off the middle, askew and moving, and drives them all for 150 steps, 3 s. With the distance from the middle of the road, m its half width, the speed along the road and the share of steps on which neuron fired:
The first two terms keep the car on the road, the second hard once it is off. The third rewards progress, and the last keeps every neuron in use and none saturated, as in the drone. Adam follows the gradient, clipped to a norm of 1, with a learning rate that warms up and then decays.
How often it fails
Gradients carried back through 150 steps of a car, a camera and a network explode, as chapter 6 warned, and this training does not always survive them. These are the five runs of the final setup on record:
| Run | Seed | Steps | Learning rate | Batch | Unseen tracks finished |
|---|---|---|---|---|---|
| The racer | 1 | 3,000 | 0.001 | 32 | 200 of 200 |
| Earlier | 1 | 1,500 | 0.001 | 32 | 187 of 200 |
| 2 | 3,000 | 0.001 | 32 | 2 of 200 | |
| 3 | 3,000 | 0.001 | 64 | 2 of 200 | |
| 1 | 3,000 | 0.0007 | 32 | 0 of 200 |
The run with seed 2 ended with gradients of norm before clipping, and its cars left the road 45% of the time. The run with the smaller learning rate found a shortcut instead: its cars stay near the middle, 0.4% of the time off the road, but crawl, and none finishes a lap in 30 s. Only seed 1 learned to drive, twice. The car on this page and the front page is that run’s.
How well it drives
Each of 200 tracks it never trained on, drawn by the same rule with another seed, starts the car at rest in the middle of the road, pointing along it, and gives it 30 s. A lap counts when the car has gone the track’s length without straying more than a metre beyond the road’s edge.
It finished all 200, in a median of 20.6 s for tracks of 40 to 75 m, and never strayed that far. Over the 200 drives it spent 0.15% of its time with its centre off the road, on 26 of the tracks. Each step the camera sent 123 events, of 288 possible, since a pixel fires at most once a step, and 16.3 of the 160 neurons fired.
No track failed, so here is the hardest, the one where it spent longest off the road:
Track 164 of the 200, 59.5 m around, with the car’s path over 30 s in orange from the dot. It finished its first lap in 21.8 s and spent 3.1% of its 30 s off the road.
The browser drives the same network from the same weights, in double precision, with the same operations in the same order. Its test holds two drives of 10 s, on two unseen tracks, to sparx’s own float64 run: every event and every spike at the same step, and the car’s path within m.
Try this
- Every drive starts with the car at rest. What does the network see at the first step, and how does the car get moving?
- Draw a track with a hairpin much tighter than any in training. Why might the car fail there, when it drives the 200 unseen tracks without a fault?
- Why does the loss reward speed along the road, rather than speed itself?
- The run with the smaller learning rate stayed on the road and never finished a lap. Which terms of the loss did it satisfy, and why was that a minimum gradient descent could settle in?
Answers
- Nothing: a still camera sends no events, and the speed it reads is 0. But the target speed is , which is never zero, so the car starts to roll, and as it moves the dashes and the road’s edges start to send events.
- The 200 unseen tracks were drawn by the same rule as the training tracks, so their bends are no sharper than those it learned on. A hairpin asks for slowing and steering it never needed in training, and nothing in the loss taught it what to do there.
- Speed itself can be had by driving in circles, or straight off the track. Speed along the road counts only progress round the track.
- The first two: staying near the middle keeps small, and driving slowly makes that easy. A likely reason it stayed there is that each small change toward driving faster first cost more in the distance terms than it earned in progress, so the gradient pointed back to crawling.
Summary
A spiking network learned to drive from an event camera’s output and its own speed, by gradients through its spikes, the car’s motion and the pixels’ own thresholds, and finished every one of 200 tracks it had never seen. The same training failed on two other seeds and a smaller learning rate, so the result is one run that worked, not a recipe that always does.
References
- A. I. Maqueda, A. Loquercio, G. Gallego, N. García and D. Scaramuzza, “Event-based vision meets deep learning on steering prediction for self-driving cars”, Conference on Computer Vision and Pattern Recognition, 2018, arXiv:1804.01310.
- G. Gallego et al., “Event-based vision: a survey”, IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 2022, doi:10.1109/TPAMI.2020.3008413.
- E. O. Neftci, H. Mostafa and F. Zenke, “Surrogate gradient learning in spiking neural networks”, IEEE Signal Processing Magazine 36, 2019, doi:10.1109/MSP.2019.2931595.
- R. Pascanu, T. Mikolov and Y. Bengio, “On the difficulty of training recurrent neural networks”, International Conference on Machine Learning, 2013, arXiv:1211.5063.