AI models close in on autonomous drone control, Anthropic finds

TL;DR:

  • Anthropic and Andon Labs tested 15 models on autonomously flying a drone to locate and follow a person.
  • The strongest model cleared the human-AI team baseline on four of five sub-tasks, failing on 3D reconstruction.
  • Anthropic warns oversight becomes a cost rather than a safeguard once models pass reliability thresholds.

Anthropic’s Frontier Red Team, working with Andon Labs, has published results from Project Pilot: an evaluation of whether AI models can autonomously fly a quad-rotor drone to find and follow a person in an office. The task was chosen for its policy relevance — it is the core of aerial surveillance, with legitimate uses in search and rescue and obvious potential for abuse.

Not there yet, and the gap is specific

Fifteen models from three developers were tested against Drone-Bench, a benchmark decomposing the mission into five sub-tasks. The trend across model generations is consistent improvement, strongest on detection and following, weakest on reconstruction and localisation. The best performer cleared the baseline on every task except reconstruction — and that single failure cascaded, leaving it unable to navigate between rooms in the physical demonstration.

Consistency is the sharper limitation. Across ten simulation runs, models reached the baseline at least once on four of five tasks, but even the frontier model met it on average on only three of five. The baseline itself is not superhuman: it represents what AI experts, not full-time roboticists, achieved using modern tools.

The governance argument is the substance. Anthropic notes that in early agentic coding, humans approved nearly every tool call; within months, models were trusted with long-horizon work. Once capability and reliability thresholds are passed, “there will be real pressure to treat human oversight as a cost rather than a safeguard” — a pressure this site has seen play out where monitoring systems guarding agents proved subvertible.

Looking forward

The UK timing is pointed. Days earlier the AI minister named drones a central plank of Britain’s reindustrialisation plan, in defence and manufacturing. That policy is being set while frontier labs are still establishing whether models can fly one reliably — and while nobody has settled who is accountable when an autonomous system tracks the wrong person.