Research
DIJIT: A Robotic Head for an Active Observer
Overview Research area: Robotics — biologically inspired robotic head design, active vision, and saccadic camera control. The paper appears in an IEEE robotics venue (DOI: 10.1109/LRA.2026.3682980) an

- arXiv
- 2512.07998
- Published
- 2025-12-08
- Authors
- Mostafa Kamali Tabrizi, Mingshi Chi, Bir Bikram Dey, Kelly Yuan, Markus D. Solbach, Yiqian Liu, Michael Jenkin, John K. Tsotsos
AI summary
Overview
Research area: Robotics — biologically inspired robotic head design, active vision, and saccadic camera control. The paper appears in an IEEE robotics venue (DOI: 10.1109/LRA.2026.3682980) and was authored by researchers at York University.
Technical level: Intermediate. The mechanical design and hardware descriptions are accessible to a general reader; the saccade-control method (homography estimation, bilinear interpolation) and the experimental comparison table require some background in computer vision and robotics.
Scope: The paper presents DIJIT, an open-source binocular robotic head with nine mechanical degrees of freedom and four optical degrees of freedom per camera, together with a fast data-driven method for executing saccades and an evaluation of its accuracy and speed against human performance and other robotic heads.
What This Paper Is About
Human vision uses an actively controlled system — each eye has three rotational degrees of freedom, the head sits on a controllable neck, and the eyes/head are constantly redirected through rapid movements called saccades — while most robotic vision systems use fewer degrees of freedom and are built for a narrow set of tasks. To study how eye and head movements contribute to visual tasks, and to compare human vision against computer vision methods, researchers need a robot head that mirrors the human system closely enough to be a fair experimental platform. DIJIT is built to fill that gap and to be evaluated on how human-like its motion performance actually is.
Key Contributions
-
A new binocular robotic head design. DIJIT provides nine mechanical degrees of freedom (three per camera: pan, tilt, roll; plus three in the neck: pan, side-bending, and flexion/extension) plus four optical degrees of freedom per camera (focusing distance, exposure time, white balance, sensor gain). The authors state it is the first robotic head with baseline and mechanical degrees of freedom similar to humans — three per camera and three in the neck — in addition to the four optical degrees of freedom.
-
An open-source, lightweight, replicable platform. The 3D part models, parts list, and software code are available online under the MIT license. The body, frames, and joint parts are 3D printed in PA6-CF, a lightweight thermoplastic, in contrast to heads built with heavy metal frames. DIJIT weighs approximately 1.4 kg.
-
A novel data-driven saccade control method. Rather than modeling and calibrating the relationship between motor axes and the camera's optical axis, the method uses homography mappings estimated between images collected in a fast calibration stage to assign motor (pan, tilt) values to each pixel. Calibration takes less than three minutes per camera, compared to the 90 minutes of processing time reported for the most similar prior approach (COG).
-
An experimental evaluation of saccade accuracy and speed against human benchmarks. Across 175 binocular trials, DIJIT achieves a mean saccade error of 1.17° ± 0.86° and 1.14° ± 0.67° for the left and right cameras, and reaches 85% of peak human saccade speed.
Main Findings
-
Saccade speed close to human: Humans perform saccades at speeds up to 700°/s; DIJIT attains 85% of peak human saccade speed, with a reported maximum of up to 600°/s. The paper states this maximum surpasses all other values reported in its comparison table.
-
Saccade accuracy comparable to humans: Mean error was 1.17° ± 0.86° (left) and 1.14° ± 0.67° (right). Reported human saccade errors vary from 0.6° to 2.89° depending on target eccentricity, saccade direction, and task conditions.
-
Median errors: 0.73° and 0.80° horizontally, and 0.29° and 0.37° vertically, for the left and right cameras, respectively.
-
Error after a corrective saccade: 0.93° and 0.95° for the left and right cameras. A corrective saccade is performed only when the primary saccade error exceeds 1°.
-
Error grows with saccade amplitude: Mean error was 1.00° and 0.92° for amplitudes under 6°, 1.33° and 1.50° for amplitudes between 6° and 12°, and 1.56° and 1.28° for amplitudes larger than 12° (left and right cameras). In the right camera, very large saccades had a lower average error than medium-amplitude saccades.
-
Directional error breakdown: Mean horizontal and vertical errors were 1.03° ± 0.88° and 0.40° ± 0.36° for the left camera, and 0.92° ± 0.69° and 0.49° ± 0.40° for the right camera.
-
Hit rates: For 55% (left) and 49% (right) of saccades, error after the primary saccade was under 1°, and for 85% and 87% it was under 2°.
-
Experimental setup and target distribution: 175 trials were run with both cameras moving together. In 91 trials the cameras returned to their home position (looking straight ahead) before the next saccade; in 84 trials the next saccade started from the end direction of the previous one. Target eccentricity was sampled to approximate human distributions: 58% within 6° of the image center, 27% between 6° and 12°, and 15% larger than 12°. For reference, human free-viewing saccade amplitudes are reported as 52% between 0° and 6°, 33% between 6° and 12°, 11% between 12° and 18°, and 4% larger than 18°. The largest saccades in this study were 19.08° and 18.41° (left and right).
-
Timing and pipeline: Computing new motor values for all four motors takes 12 ms on an AMD Ryzen 7 7700X 8-core processor in the master computer. The right camera starts moving 33 ms after the left camera. Motor values are sent sequentially to a Teensy 4.0 microcontroller at 60 Hz.
-
Repeatability: Across 10 ArUco marker positions with 10 repetitions each (100 trials), average saccade amplitude was 11.67° and 11.25° and average error was 1.13° and 1.43° (left and right). Standard deviation of target eccentricity, averaged over positions, was 0.3° horizontal and 0.07° vertical (left), and 0.16° and 0.13° (right). Standard deviation of errors was 0.36° and 0.13° (left) and 0.22° and 0.21° (right), horizontal and vertical respectively.
-
Comparison to other heads: Robotic head saccade errors reported elsewhere include Manfredi at 1.57°, Van Opstal at 1.47°, iCub at under 1° (after a corrective saccade), Tombatossals at more than 5 pixels, Schenck at 1 pixel, and COG at under 1 pixel. The authors note that some of these are evaluated only in simulation, and that pixel-based errors are converted to angular values using camera and lens information supplied by those authors. Several prior methods also assume the camera frame, motors, and optical axes are perfectly aligned — an assumption DIJIT's method does not require.
-
Design ranges: DIJIT's eye pan and tilt ranges are ± 40° and ± 46°; head pan is ± 135°, and head tilt is −40° to +38°. Human ranges cited in the paper are approximately ± 45° in vergence, 28° in elevation, 47° in depression, and ± 10° in torsion for the eye, and ± 60° in yaw, 30° in flexion, 60° in extension, and ± 20° in side-bending for the head. DIJIT's inter-camera baseline is 115 mm, within the reported human range of 45–80 mm... — note: the paper lists the human baseline as 45–80 mm and DIJIT's as 115 mm.
-
Optical and hardware specifications: Each camera is an iDS U3-3881LE-AF with a global shutter, capturing at 59 fps at 3088 × 2076 pixels over USB 3.0 at 5 Gb/s, paired with a Corning Varioptic C-S-25H0-096 liquid 9.6 mm lens, giving a field of view of approximately 40° × 30° (H × V). Two BNO085 IMUs rigidly connected to the camera frames provide rotational states at 60 Hz.
Methodology in Plain English
The authors built a physical robot head whose cameras and neck move along the same axes as human eyes and neck, then designed and tested a way to make it look quickly at a specified point.
The key idea in the saccade method is to avoid modeling exactly how the motors and the camera's optical axis line up — a relationship that is difficult to measure precisely on real hardware. Instead, the researchers place a ChArUco calibration board in front of the head and move the pan and tilt motors in discrete steps (for example, five degrees) across their range, capturing an image at each step along with the motor values. Because a camera rotating about its own center produces images related by a homography, and because the calibration board is planar, they can estimate a homography between any pair of calibration images and use it to predict where the center of one image lands in another.
To execute a saccade, a target point is given in the current image and transferred into the closest matching calibration image. The method then finds which calibration image's center lands closest to that target point, applies bilinear interpolation to handle motor states that were never sampled, and outputs the pan and tilt motor values that move the camera so the target ends up at the image center. This is applied separately to the left and right cameras.
Experiments used an ArUco marker placed at various positions in front of the head; its center was detected with OpenCV and served as the target for both cameras. Target position is not updated during a saccade, mirroring the fact that biological saccades are executed open-loop. After each movement, the marker was re-detected to measure the error, defined as the angle between the back-projected ray of the landing point and that of the image center.
Why This Matters
Impact on research. The paper argues that robotic and human visual systems differ significantly — robotic systems typically use fewer degrees of freedom and target a small number of tasks — and that a human-like platform is needed to compare how eye and head movements solve visual tasks versus how computer vision methods do. Active vision has been shown to benefit tasks such as SLAM and object recognition. DIJIT is offered as a platform for active vision research, for studying eye and head-neck motion and their interrelationships, and for evaluating theories in cognitive science and neuroscience. The authors also claim a design difference from most prior methods: many saccade approaches assume the camera coordinate system, camera frame axes, and motor axes are perfectly aligned, while DIJIT's data-driven method does not.
Real-world applications
- Mobile robotic platforms that need human-like gaze control, including deployment in an autonomous system to support wheelchair-bound individuals (named as ongoing work).
- Active binocular 3D reconstruction, listed as ongoing work.
- Laboratory studies of eye movements beyond saccades — the authors state work on other eye movements is ongoing.
- Affordable replication for other research groups: the design and software are open source under the MIT license, and the paper describes the hardware and 3D-printed parts as inexpensive.
Industry relevance. The paper highlights a compact, lightweight, 3D-printed design using PA6-CF thermoplastic rather than heavy metal frames, and reports that low-resolution motor encoders still yield better or comparable saccade accuracy than other binocular heads. That combination — low-cost components, open-source availability, fast calibration (under three minutes per camera versus 90 minutes reported for a comparable method), and near-human speed and accuracy — is directly relevant to robotics groups building mobile platforms where weight, size, and cost matter.
Future Directions
- Other eye movements. The authors list ongoing work on eye movements beyond saccades.
- Active binocular 3D reconstruction. Named explicitly as ongoing work.
- Autonomous deployment. Deploying DIJIT within an autonomous system to support wheelchair-bound individuals.
- Generalizing to cyclotorsion. The current method uses only pan and tilt for a simple saccade; the paper states it can easily generalize to cyclotorsion, which may be needed for off-horopter binocular fusion, and identifies this as ongoing work.
- Target motion and decision-making. The paper states that visual processing, decision-making, and control for moving targets are outside the scope of the current work but are topics of active research; the method for deciding "where to look next" is also explicitly out of scope.
Target Audience
Robotics researchers and engineers working on biologically inspired robot design, active vision, and gaze control; computer vision researchers interested in comparing human eye/head motion against machine vision methods; cognitive scientists and neuroscientists who need a hardware platform to evaluate theories of human vision; and labs that want an affordable, open-source, replicable binocular head for mobile robots.
Authors’ abstract
We present DIJIT, a novel binocular robotic head expressly designed for mobile agents that behave as active observers. DIJIT's unique breadth of functionality enables active vision research and the study of human-like eye and head-neck motions, their interrelationships, and how each contributes to visual ability. DIJIT is also being used to explore the differences between how human vision employs eye/head movements to solve visual tasks and current computer vision methods. DIJIT's design features nine mechanical degrees of freedom, while the cameras and lenses provide an additional four optical degrees of freedom. The ranges and speeds of the mechanical design are comparable to human performance. DIJIT attains 85\% of the peak human saccade speed. Our design includes the ranges of motion required for convergent stereo, namely, vergence, version, and cyclotorsion. Here, we present DIJIT and some aspects of its performance. We also present a novel method for saccadic camera movements, using a direct relationship between camera orientation and motor values. The resulting saccadic camera movements are close to human movements in terms of their accuracy, with 1.17$^\circ$ and 1.14$^\circ$ mean error for the left and right cameras, respectively.