- Drone-Bench is a benchmark that tests AI models' ability to write code for drone-based surveillance on low-cost hardware.
- It comprises five tasks: reconstruction, localization, navigation, detection, and following, each scored against a human baseline.
- Tasks are evaluated independently to prevent error compounding, though end-to-end success requires all five tasks.
- Current models can beat the baseline on some tasks, but no model has succeeded at reconstruction, resulting in 0% end-to-end success.
- The benchmark aims to track AI capabilities for public awareness, as models become easier to misuse with increased physical autonomy.
- Future directions include harder tasks, end-to-end runs, and removing feedback scores between submissions.