Skip to content
AI.info

Research

NanoCockpit: Performance-optimized Application Framework for AI-based Autonomous Nanorobotics

Overview Research area: Embedded systems for robotics — specifically hardware/software co-design for sub-100 mW microcontroller-class nano-drones running vision-based TinyML models. Technical level: A

NanoCockpit: Performance-optimized Application Framework for AI-based Autonomous Nanorobotics
arXiv
2601.07476
Published
2026-01-12
Authors
Elia Cereda, Alessandro Giusti, Daniele Palossi

AI summary

Overview

Research area: Embedded systems for robotics — specifically hardware/software co-design for sub-100 mW microcontroller-class nano-drones running vision-based TinyML models.

Technical level: Advanced. The paper assumes familiarity with microcontrollers, DMA, real-time operating systems, SoC architecture, and neural network deployment pipelines (quantization, memory tiling).

Scope in one sentence: NanoCockpit is an open-source application framework for the Bitcraze Crazyflie nano-drone with the AI-deck that pipelines camera acquisition, multi-core inference, inter-MCU data exchange, and Wi-Fi streaming, improving throughput and latency for closed-loop TinyML robotics.

What This Paper Is About

Nano-drones weighing a few tens of grams must run vision-based neural networks on microcontroller units (MCUs) limited to sub-100 mW, and the Crazyflie is the de facto standard platform for this. In practice, roboticists underuse the onboard GAP8, ESP32, and STM32 processors because the existing software layer (the PMSIS runtime on GAP8) is single-threaded and event-callback based, which forces camera capture, inference, and transmission to run one after another instead of in parallel. The paper's goal is to supply the missing software layer — NanoCockpit — so tasks overlap cleanly, cutting latency and increasing throughput without making the programmer's job harder.

Key Contributions

  1. A coroutine-based cooperative multi-tasking layer for GAP8. NanoCockpit extends the PMSIS runtime with stackless co-routines, giving sub-10 µs context switches (on the order of a function call) and 8× lower memory overhead than FreeRTOS on STM32 — 18 B per task versus at least 150 B per task stack in FreeRTOS.
  2. Optimized camera drivers and multi-buffered acquisition. The framework provides fine-grained timing control for the Himax HM01B0 camera, supporting a trigger mode (up to approximately 30 frames, unit as printed) and a streaming mode reaching up to 150 frames with the framework, plus double-buffered acquisition so µDMA fills a new image while the cluster infers and SPI transmits.
  3. A rewritten zero-copy Wi-Fi/CPX stack spanning STM32, GAP8, and ESP32. It overlaps SPI and Wi-Fi transfers, adds hardware-based timestamping for cross-MCU synchronization, and streams 160×160 px images at 72 Hz — 2.4× higher than the original CPX stack.
  4. An open-source reference framework validated in three in-field applications. Code is released on GitHub (https://github.com/idsia-robotics/nanocockpit), and NanoCockpit has been made part of the official Crazyflie software.

Main Findings

  • Abstract-level control improvement: The framework achieves ideal end-to-end latency (zero overhead from serialized tasks), delivering −30% mean position error and a mission success rate increase from 40% to 100% in field experiments.

  • Human pose estimation — throughput beats latency: With onboard PULP-Frontnet (304 k parameters, 14.3 MMAC per inference), measured end-to-end latency stayed at 30.3 ms across 12, 24, and 48 Hz, with mean horizontal error e_xy of 0.96, 1.01, and 0.80 m and angular error e_θ of 0.63, 0.69, and 0.50 rad. The off-board MobileNetV2-based CNN from Cereda et al. (7× more parameters, 90 MMAC) achieved 0.65 m and 0.41 rad at 40 Hz with 168.6 ms latency, and still 0.81 m and 0.46 rad when 500 ms of artificial delay pushed latency to 668.6 ms. The off-board model at 40 Hz reached 19% lower error than PULP-Frontnet despite more than 100 ms higher latency, and scored almost on-par even above 600 ms latency.

  • Drone-to-drone localization — highest throughput never loses the target: The observer tracked a target flying a 10-meter spiral in either 48 s or 24 s (reported average target drone speeds of 0.21 and 0.34 m). At 39 Hz the observer never lost track and reached its best mean position error e_xyz of 0.18 m, versus losing track once at 10 Hz in the slow configuration and in 2 out of 3 runs at 10 Hz in the fast configuration (only one loss at 20 Hz).

  • Nano-drone racing — speed makes throughput decisive: The obstacle-avoidance CNN (331 k parameters, 25 MMAC per frame, up to 30 Hz on GAP8 with the framework) reached a 60% success rate at the highest speed tested. A success was scored when the drone halted between 0.15 and 2 m from the obstacle, having taken off 4 m away; the first 2 m of flight always produced zero collision probability, letting the controller reach its max speed.

  • The Crazyflie dominates nano-drone research: A Scopus survey (update date 9 Jan 2026, years 2020–2024) starting from 710 query results plus 325 Bitcraze-acknowledged publications yielded 586 final entries. The Crazyflie appears in 381 publications over five years (65%), the DJI Tello in 21%, and the Parrot Mambo plus Rolling Spider together in 9%; the remaining nano-quadrotors have only 1 to 5 publications each.

  • Quantization keeps onboard memory small: Onboard CNNs are quantized from float32 to int8 with QuantLib, giving a 4× reduction in memory footprint with marginal performance loss, then compiled to C by DORY using PULP-NN-mixed kernels.

Methodology in Plain English

The authors start from the observation that the Crazyflie's three onboard processors sit idle for large parts of each perception cycle. They rework the software so that work is always being done: the camera driver keeps multiple image buffers in flight, the GAP8 cluster computes on one buffer while the direct memory access unit fills another, and the Wi-Fi stack moves those buffers onward without copying them again. To let programmers write this overlapping behavior naturally, they add coroutines to the GAP8 runtime, so a task can pause and resume without a private stack. They then test the framework in three real tasks — a drone following a walking person, a drone tracking another drone's spiral flight, and a drone approaching an obstacle — measuring throughput, end-to-end latency, tracking errors, and success rates, with position ground truth from a motion capture system. Finally, they survey five years of nano-drone publications to establish how dominant the Crazyflie is, and they repeat each human pose estimation configuration three times and each racing configuration five times.

Why This Matters

Impact on research: The framework removes a software bottleneck that the authors argue has pushed prior state-of-the-art work into suboptimal serialized execution. Because it is open source and has been integrated into the official Crazyflie software, other labs can build on it without reimplementing pipelining, synchronization, and streaming. The paper also contributes a reproducible survey of which nano-drones the field actually uses, plus profiling and debugging tools (GPIO/UART event traces, Wireshark dissectors for CPX traffic).

Real-world applications:

  • Human-following and pose-aware flight, useful for companion or assistant nano-drones that must keep a person centered in view.
  • Drone-to-drone localization, the perception primitive behind peer tracking and coordinated multi-drone flight.
  • Autonomous collision avoidance at speed, the basis of indoor exploration and drone racing scenarios.
  • Onboard image capture streamed to a remote computer or ROS node for dataset collection, where hardware timestamping matters for synchronizing data across MCUs.

Industry relevance: The optimized ESP32 communication stack targets one of the most widely adopted Wi-Fi modules in embedded systems, so it is directly reusable on any platform using that chip. The GAP8 design principles (double-buffered DMA, cluster-level parallelism, memory-aware scheduling) transfer to similar architectures, and the multi-MCU orchestration pattern applies broadly to resource-constrained cyber-physical systems.

Future Directions

  • Porting to GAP9. The paper notes that the open-source GAP9Shield prototype built around the GAP9 Parallel Ultra-Low-Power SoC shares enough hardware and software-stack similarity with GAP8 that NanoCockpit can be moved over with limited effort.
  • Adapting GAP8-specific optimizations. The authors state their GAP8 optimizations are tailored to that SoC and would need adaptation for newer processors, while the underlying design principles should carry over.
  • Extending to newer camera decks. Bitcraze prototype expansion decks rely on the same ESP32 Wi-Fi module as the AI-deck, so the optimized communication stack could be reused directly there.
  • Generalizing beyond nano-drones. The paper raises applicability to other heterogeneous, resource-constrained multi-MCU cyber-physical systems, but does not report experiments outside the Crazyflie platform.

Target Audience

Robotics and embedded-systems researchers working with nano-drones, TinyML, and microcontroller-class hardware; engineers deploying vision models on multi-core low-power SoCs; and developers building applications on the Crazyflie AI-deck who need pipelined perception, wireless streaming, or cross-MCU synchronization. Readers without background in embedded runtimes, DMA, or quantized neural network deployment will find the paper dense, while those already in the field will get directly reusable code and infrastructure rather than only a conceptual contribution.

Authors’ abstract

Autonomous nano-drones, powered by vision-based tiny machine learning (TinyML) models, are a novel technology gaining momentum thanks to their broad applicability and pushing scientific advancement on resource-limited embedded systems. Their small form factor, i.e., a few tens of grams, severely limits their onboard computational resources to sub-100mW microcontroller units (MCUs). The Bitcraze Crazyflie nano-drone is the de facto standard, offering a rich set of programmable MCUs for low-level control, multi-core processing, and radio transmission. However, roboticists very often underutilize these onboard precious resources due to the absence of a simple yet efficient software layer capable of time-optimal pipelining of multi-buffer image acquisition, multi-core computation, intra-MCUs data exchange, and Wi-Fi streaming, leading to sub-optimal control performances. Our NanoCockpit framework aims to fill this gap, increasing the throughput and minimizing the system's latency, while simplifying the developer experience through coroutine-based multi-tasking. In-field experiments on three real-world TinyML nanorobotics applications show our framework achieves ideal end-to-end latency, i.e. zero overhead due to serialized tasks, delivering quantifiable improvements in closed-loop control performance (-30% mean position error, mission success rate increased from 40% to 100%).

Read the original paper