All experiments

Experiment · Asteroid Dodger

Can biological neurons learn through feedback?

We gave a biological neural network a more difficult task: control a spacecraft, avoid incoming asteroids, and get better through experience.

Unlike Dino, the correct action wasn't a single jump at the right moment. The neurons had to continuously interpret a changing environment, control movement in real time, and learn from the consequences of their actions.

The task

The network controlled the horizontal movement of a spacecraft while asteroids continuously approached from above.

The neurons received information about where the asteroid was and how close it was getting. From that, the network continuously controlled whether the spacecraft moved left or right.

The Asteroid Dodger game: a green spacecraft near the bottom of a starfield with an asteroid descending, telemetry printed across the top Watch the gameplay
Live telemetry across the top shows the spike counts on each side and the resulting steering bias.
Environment Neurons Action Outcome Feedback
The loop that separates this experiment from Dino: what the neurons do changes what they see next.

Feedback

This time, we added feedback.

Every action had a consequence.

Successful dodge

The network received positive feedback.

Collision

The network received negative feedback and the task restarted.

Through repeated interaction with the environment, the biological network began changing its behaviour.

See. Act. Receive feedback. Adapt. Repeat.

The result

The neurons learned.

  • ~6 min

    Learning begins

    Performance began improving after approximately six minutes of interaction.

  • ~12 min

    Peak performance

    The network reached its highest task performance after approximately twelve minutes.

  • 86.5%

    Peak dodge rate

    At peak performance, the biological network successfully avoided 86.5% of incoming asteroids.

025507510003691215 learning begins 86.5% peak, with feedback without feedback Minutes of interaction Dodge rate %
Drawn from the reported figures: a baseline near 75%, improvement from about six minutes, a peak of 86.5% at about twelve, and a control given the same sensory information without feedback. The shape between those points is indicative rather than a plot of the raw trial data.

The control

Feedback made the difference.

To test whether the improvement was actually associated with feedback, we ran the same task with a control network. It received the same sensory information, but no feedback about whether its actions were successful. The control did not develop the same task proficiency.

With feedback

The neurons progressively improved their ability to control the spacecraft.

Without feedback

Performance did not show the same improvement over time.

Why it matters

Asteroid Dodger demonstrates something fundamentally different from a fixed biological response.

The neurons were placed inside a closed loop where their actions affected the environment, the environment produced feedback, and that feedback shaped future behaviour. Over time, the biological network became better at the task.

The biological substrate learned.

Resource profile

How efficient was the learning?

We also compared the biological network with a conventional AI system trained on the same task. The two systems operate fundamentally differently, so this is not a like for like comparison. Instead, it provides an early look at their different resource profiles.

Time to peak performance

Biological network ~12 min
AI baseline ~17 min

Measured system energy

Biological network ~362 J
AI baseline ~26,168 J

About 72× less energy, on these measurements.

An early benchmark between fundamentally different systems rather than a claim that one beats the other. The AI baseline received a more explicit state representation, including information not given directly to the biological system.

Methodology Stimulation encoding, motor decoding, feedback protocol, performance, and controls

Experimental objective

Continuous sensorimotor control with closed loop feedback, testing whether task performance improves through interaction.

Stimulation encoding

The arena is divided into four vertical strips, each with a corresponding electrode, giving four input channels. Which strip the asteroid occupies carries its position, and the distance between the asteroid and the ship is rate coded.

Motor decoding

Two decoding areas are read, one for left and one for right. The difference in their spike counts within a 20 ms bin becomes movement in the game.

Feedback protocol

Reward is a consistent pattern delivered across all input channels. Punishment is random noise, followed by a period of rest and a reset of the task.

Performance metric

Measured per rally: how many asteroids the network dodged in a single rally, and how long that rally lasted.

Controls and limitations

A control network received the same sensory information without feedback and did not develop comparable proficiency. The AI baseline in the resource comparison received a more explicit state representation than the biological system, so that comparison is indicative rather than like for like.

All experiments Run your own on AxoGrid