Skip to content
AI.info

Research

EvtSlowTV -- A Large and Diverse Dataset for Event-Based Depth Estimation

EvtSlowTV — A Large and Diverse Dataset for Event-Based Depth Estimation Overview Research area: Computer vision / neuromorphic (event-based) sensing, specifically large-scale dataset construction and

EvtSlowTV -- A Large and Diverse Dataset for Event-Based Depth Estimation
arXiv
2511.02953
Published
2025-11-04
Authors
Sadiq Layi Macaulay, Nimet Kaygusuz, Simon Hadfield

AI summary

EvtSlowTV — A Large and Diverse Dataset for Event-Based Depth Estimation

Overview

Research area: Computer vision / neuromorphic (event-based) sensing, specifically large-scale dataset construction and self-supervised monocular depth estimation.

Technical level: Intermediate to Advanced. The paper assumes familiarity with event cameras, contrast maximization, photometric/geometric self-supervision, and encoder–decoder depth–pose networks.

Scope in one sentence: The paper introduces EvtSlowTV, a large event-stream dataset generated from 45 curated YouTube videos, and demonstrates that training a self-supervised, purely event-driven depth–pose network on it improves generalization to complex scenes and motions.

What This Paper Is About

Event cameras record per-pixel brightness changes asynchronously with high dynamic range and low latency, which makes them attractive for depth estimation in low-light, high-speed, or high-dynamic-range conditions where ordinary cameras fail. Progress is blocked by the fact that existing event-based depth datasets are small and narrow in scene diversity, and that most learning pipelines rely on synchronized RGB, LiDAR, or stereo ground truth, which discards the event camera's asynchronous advantage. The goal of this work is to provide a much larger and more varied event dataset, drawn from real-world footage, and to show it can be used to train a depth estimator without any external sensor annotations or frame-based labels.

Key Contributions

  1. A large-scale event dataset for depth estimation. EvtSlowTV is described as an order of magnitude larger than existing event-based datasets, offering unconstrained depth variation and high scene diversity. It is reported as containing more than 13B events (the abstract) or over 10B events (the introduction), aggregated from 45 YouTube videos totalling approximately 2M frames, and organized as 40 sequences with a total duration of 9000 minutes.

  2. A self-supervised depth learning framework that needs no external sensors. The method removes dependence on RGB, LiDAR, or stereo supervision and preserves the asynchronous nature of event data, training instead with a contrast maximization loss plus a student–teacher distillation scheme.

  3. Validation that the generated event dataset improves generalization and accuracy. The authors show that depth maps can be learned directly from a spatiotemporal event representation (event volumes) without auxiliary sensor annotation, and that exposure to the dataset's scale and variety improves results on outdoor day, outdoor night, and indoor flying sequences.

  4. An automated, adaptive event-generation toolchain. The pipeline converts video into event streams with ESIM, using adaptive frame sampling driven by log-irradiance gradients rather than uniform sampling, and the authors release tools to add more videos or re-encode events into other formats such as quantized event tensors.

Main Findings

  • Dataset scale. Table 1 reports EvtSlowTV at 40 sequences and 9000 minutes of total duration, compared with MVSEC (5 sequences, 160 min), EventScape (758, 800 min), DESEC (53, 500 min), DDD17 (39, 35 min), and DDD20 (12, 600 min). The paper describes this as an order of magnitude more data.

  • Category coverage. Table 1 marks EvtSlowTV as covering indoor, outdoor, hiking, driving, and flying, and it is the only listed dataset flagged for hiking. Note that the same table marks the "Natural" category as No for EvtSlowTV, and marks the "Depth" column as No, even though the abstract and introduction describe the setting as unconstrained and naturalistic — the text and table are not fully aligned.

  • Depth accuracy versus baselines. On the MVSEC indoor_flying test sequences (Table 2), the proposed method achieves the lowest absolute mean error: 0.1887, 0.1659, and 0.1882 on flying1, flying2, and flying3. Baselines are EMVS (0.3937, 0.3142, 0.3054), ESVO (0.2339, 0.2042, 0.2429), and Ghosh et al. (0.2253, 0.1820, 0.1949).

  • A weakness in scale consistency. The same table shows the proposed method has much worse rms_log than all three baselines: 0.5675, 0.4708, and 0.4735 versus 0.1411, 0.1359, and 0.1201 for the best baseline (Ghosh et al.). The authors attribute this to proportional depth estimation being unreliable when the teacher model has only been exposed to limited data variability.

  • Self-supervision beats supervised training. In the ablation (Table 3, RMS error), moving from supervised EvtSL to self-supervised variants improves results across sequences. For outdoor day1 at 0–10 m, error drops from 0.4184 to 0.2571 (finetuning) and 0.2570 (teacher–student); for outdoor night1 at 0–10 m, from 0.3734 to 0.2962 for both; for indoor flying1 at 0–30 m, from 0.0493 to 0.0646 (finetuning) and 0.0396 (teacher–student).

  • Finetuning and teacher–student are close. Table 3 shows the two self-supervised strategies are nearly identical on the outdoor sequences, while teacher–student is better on the more varied indoor flying1 sequence at the 0–20 m cut-off (0.0595 versus 0.0803) and 0–30 m cut-off (0.0396 versus 0.0646).

  • Shorter range, better predictions. The paper reports reduced error within a 10-metre cut-off, which it interprets as the models making better predictions when reliable motion and edge information is available, since closer objects undergo more pronounced apparent transformations than distant ones.

  • Evaluation setup. The test samples used in the comparison form roughly a 200 second-long sequence with 4000 frames of ground-truth depth map and a maximum depth distance of 8.40 meters.

Methodology in Plain English

Building the dataset. The authors start from SlowTV, a public collection of long-duration real-world videos, and pick 45 of them covering hiking, flying, driving, and underwater exploration. Rather than sampling video frames at a fixed rate, they sample adaptively: they compute how fast the logarithm of brightness changes across space and time, and when brightness changes quickly they sample frames more often, and when it is static they sample less often. Consecutive frame times are set using the inverse of the maximum brightness change rate over the image (Equation 1). They feed those frames into ESIM, an open event camera simulator, which triggers an event whenever the change in log-intensity at a pixel since the last event exceeds a contrast threshold C (Equations 2 and 3). Each event is stored as a pixel location, a timestamp, and a polarity of +1 or −1 for a brightness increase or decrease. The authors describe the result as a large-scale synthetic event dataset generated from real-world video sequences. This keeps events sparse and temporally precise while avoiding the cost of physically capturing such large volumes of event data.

Turning events into something a network can consume. Raw events are sparse and asynchronous, so a slice of the stream is aggregated into a spatiotemporal tensor. The chosen window is 0.1665 seconds, corresponding to five consecutive frames at 30 FPS, and events are accumulated into 5 bins using a linearly weighted kernel (Equation 4). This event volume is the network input.

Depth and pose prediction. A skip-connection encoder–decoder network takes the event volume and outputs both a depth map and a camera pose (Equation 5). Depth is used to back-project events into 3D world points using the camera intrinsics (Equation 6), and the predicted pose supplies the rotation and translation that transform points between timestamps (Equation 7).

The self-supervision signal. The core idea is that events arise mainly at the boundaries of moving objects, so if depth and camera motion are correct, all back-projected and warped events should line up into a sharp, high-contrast edge map. The network is therefore trained by maximizing the variance of the Image of Warped Events (Equations 8 and 9), which is minimized as a loss (L_contrast). No labels are needed for this term.

Teacher–student training. Because contrast maximization alone is unstable, the authors also train a supervised teacher on MVSEC using a scale-invariant loss (L_si, Equation 10) and a scale-invariant gradient matching loss built from Sobel operators (L_grad, Equation 11), combined as L_teacher = L_si + λ·L_grad (Equation 12). The student is then trained on EvtSlowTV with L_student = L_contrast + (1−λ)·L1(D1, D2) (Equation 13), blending the self-supervised contrast loss with an L1 match to the frozen teacher's depth predictions, with λ = 0.6. A separate ablation compares plain finetuning (EvtSL → EvtSSL) against this teacher–student scheme (EvtSL → EvtSSL̄).

Training details. Models are implemented in PyTorch, trained to convergence with learning rate 10⁻⁴, weight decay 10⁻³, the AdamW optimizer, and batch size 16 on a single NVIDIA GeForce RTX 3060 (12GB). The teacher is trained on the first 80% of the duration of most MVSEC sequences (excluding indoor_flying), with the last 20% used for testing.

Why This Matters

Impact on research. Event-based depth estimation has been bottlenecked by data scarcity and by a dependence on synchronized frame-based ground truth, which contradicts the asynchronous nature of event sensors. This work attacks both at once: it supplies far more data than prior benchmarks (40 sequences, 9000 minutes, spanning hiking, driving, flying, and underwater) and shows that a purely event-driven self-supervised pipeline can be trained without auxiliary RGB, LiDAR, or stereo supervision. It also honestly reports the trade-off it exposes, namely strong absolute error but poor proportional (scale-consistent) error, which gives the field a concrete failure mode to work on.

Real-world applications:

  • Robotics and autonomous navigation in low-light or high-speed conditions where frame cameras blur or saturate.
  • Autonomous vehicles, particularly the driving day and night scenarios evaluated in the ablation.
  • Drone and aerial perception, matching the flying footage in the dataset and the MVSEC indoor_flying evaluation.
  • Augmented reality and medical imaging, both cited in the paper as domains where depth estimation and 3D scene understanding are needed.

Industry relevance. Neuromorphic sensor and event-camera vendors benefit from larger training corpora that make their hardware viable for learned perception. Automotive and drone perception teams can use a dataset harvested from existing video rather than requiring costly new multi-sensor capture rigs. The released tooling for adding videos to the dataset lowers the barrier for companies to extend it with proprietary footage, since the pipeline runs from ordinary video rather than from a physical event camera.

Future Directions

  • Fixing scale consistency. The rms_log results (0.5675, 0.4708, 0.4735) are far worse than the baselines the paper beats on absolute mean error, and the authors attribute this to a teacher trained on limited data variability. A teacher trained on more diverse data, or a different distillation objective, is the obvious next step.

  • Closing the simulation gap. EvtSlowTV events are generated by ESIM from ordinary video, so a sim-to-real gap remains relative to events captured by physical sensors. Validating the learned models on real event camera streams is an open question.

  • Addressing the missing ground truth. Table 1 flags EvtSlowTV as not providing depth ground truth, so depth supervision must come from elsewhere (here, MVSEC). Building an EvtSlowTV-derived benchmark with some form of depth validation would allow more direct measurement on the dataset itself.

  • Expanding and diversifying further. The released tools let users add videos and convert events to other formats such as quantized event tensors; extending coverage beyond the current 40 sequences and reconciling the table's category flags with the paper's descriptions (for example the "Natural" and "Depth" columns) would strengthen the benchmark.

Target Audience

Researchers and engineers working on event-based vision, neuromorphic computing, and self-supervised depth estimation will get the most from this paper, as will those building datasets for 3D scene understanding. Practitioners in autonomous driving, drone perception, robotics, and AR who need depth in low-light or high-dynamic-range conditions will find the benchmarking comparison and the ablation on training strategies directly useful. Readers with only a general machine learning background should expect to need some background in event camera representation and geometric self-supervision to follow Sections 4 and 5.

Authors’ abstract

Event cameras, with their high dynamic range (HDR) and low latency, offer a promising alternative for robust depth estimation in challenging environments. However, many event-based depth estimation approaches are constrained by small-scale annotated datasets, limiting their generalizability to real-world scenarios. To bridge this gap, we introduce EvtSlowTV, a large-scale event camera dataset curated from publicly available YouTube footage, which contains more than 13B events across various environmental conditions and motions, including seasonal hiking, flying, scenic driving, and underwater exploration. EvtSlowTV is an order of magnitude larger than existing event datasets, providing an unconstrained, naturalistic setting for event-based depth learning. This work shows the suitability of EvtSlowTV for a self-supervised learning framework to capitalise on the HDR potential of raw event streams. We further demonstrate that training with EvtSlowTV enhances the model's ability to generalise to complex scenes and motions. Our approach removes the need for frame-based annotations and preserves the asynchronous nature of event data.

Read the original paper