Proceedings · Session S-552 · filed October 10, 2026
AI & Emerging Tech in R&DSession paper
Sandia tests reinforcement learning on cooperative drone evasion
Sandia researchers trained four 3.5-inch quadrotor drones to play a strategic version of tag, with cooperating evaders learning to make pursuers collide. The 2026 IEEE ICRA paper targets national-security autonomy applications.
By Priya Raman4 min read770 words
Summary
- Four 3.5-inch quadrotor drones used; each weighs less than a chocolate bar
- Two cooperating evaders trained via reinforcement learning; two pursuers ran proportional navigation
- Conference paper accepted for the 2026 IEEE International Conference on Robotics and Automation
- Intern Christian Llanes (Georgia Tech robotics PhD) began work on the project in fall 2023
- Project funded through Sandia's Laboratory Directed Research and Development program; testing conducted at the CAMINO high bay
Sandia National Laboratories computer scientist Spencer Jensen and his team trained four 3.5-inch quadrotor drones to play a strategic version of tag, with two cooperating evaders learning to manipulate two pursuer drones into colliding mid-flight. The work, published as a conference paper for the 2026 IEEE International Conference on Robotics and Automation, sits inside Sandia's broader AutonomyNM effort to push the limits of machine-learned autonomy.
The hardware choice was deliberate. Each test drone measures roughly 3.5 inches across and weighs less than a chocolate bar, Jensen said. Crashes cost an occasional plastic propeller, not a budget line. The four-drone setup — two evaders and two pursuers — gave the evaders a single objective: reach a protected "base" without being tagged. Rather than flee, the trained evaders learned to maneuver so the pursuers' pursuit paths crossed.
What did reinforcement learning actually solve here?
The evaders' guidance module, which feeds paths to a lower-level control module, came out of simulation-based training. Christian Llanes, a Sandia intern and robotics doctoral student at the Georgia Institute of Technology, started on the project in fall 2023 after a prior summer internship. He designed the reward structure that pushed the algorithm toward cooperative evasion.
"Reinforcement learning is just beginning to be used in commercial robotics," Jensen said. "The really nice thing about reinforcement learning is you can ignore a lot of extremely complex math — that's probably slightly wrong anyway — and just get a best-effort algorithm that is going to be more flexible in rapidly changing scenarios."
The pursuers, by contrast, ran proportional navigation, a non-learning intercept algorithm. The asymmetry was the point: only the cooperative-evader side needed learned behavior. That kept the experiment tractable and let the team isolate what reinforcement learning actually bought them.
How did the team handle the simulation-to-reality gap?
Jensen called the gap between trained-in-simulation performance and physical-hardware behavior the project's central engineering problem. Battery drain changes rotor speed in ways the simulator doesn't fully capture, and one drone's downwash disturbs its neighbor in flight. Communication delays exist in the real world but not the model.
The team worked around that by characterizing the commercial off-the-shelf quadrotors early and folding those measurements back into the simulation. They tested the trained algorithm inside the high bay at Sandia's Center for Advanced Manufacturing and Innovation, known as CAMINO, which uses an infrared motion-capture system to track drone positions precisely.
"The high bay in CAMINO allows us to bridge the gap between early-stage research and high-consequence hardware," Jensen said. "We can do low-cost testing in this facility. A lot of the work we're doing is trying to bridge the simulation-to-reality gap."
What does this buy R&D programs beyond the lab?
Jensen pointed to national-security use cases, including future systems designed to defend critical facilities from hostile drones. Reinforcement learning matters there because the optimal evasion or interception maneuver in a swarm scenario is not something engineers can write down by hand. The team's bet is that policies learned in cheap simulation, then stress-tested on cheap hardware, will eventually transfer to higher-stakes platforms.
Llanes flagged reward shaping — choosing which behaviors to incentivize, and by how much — as the open question. "The challenge with reinforcement learning is you have to do a lot of reward shaping to get the behavior you want," he said.
The project was funded through Sandia's Laboratory Directed Research and Development program, the lab's discretionary in-house research budget. No external funding partners or commercial customers were named.
What the announcement does not yet establish
Sample sizes and validation remain modest. Four drones, an indoor motion-capture volume, and one trained scenario do not show how the policy would scale to outdoor environments, GPS-denied conditions, or larger swarms. The conference paper documents the design and demonstrates one successful evasion pattern; broader benchmarks against alternative learned or hand-engineered policies are not described. The LDRD funding designation also means the work is pre-commercial, and any defense prime interested in fielding the algorithm would carry its own integration cost.
The next milestone for the group, Jensen indicated, is moving from inexpensive quadrotors to "high-cost, high-consequence hardware" — a phrase that hints at larger or more expensive autonomous platforms where the same learned policy would need to perform under tighter tolerances. CAMINO exists in part to make that transition less risky, and the team is now working out which parts of the learned behavior survive the jump.
via newsreleases.sandia.gov (Original)
Filed under
- reinforcement-learning
- drone-technology
- autonomous-systems
- sim2real-transfer
- robotics
More from Priya Raman
References
- Sandia compressed a six-month field test setup into one month
- Sandia completes final Mobile Guardian Transporter crash test
- Sandia's Aires Tide Reaches National Mall After 5-Month AI Build
- Sandia's Agent Bayes Enters Nine-Month Trial Under Genesis Mission
- Canada's NRC funds AirMatrix for AI-driven airspace intelligence