Research
SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
Overview Research area: Computer vision — multi-object tracking (MOT), specifically end-to-end transformer trackers that unify object detection and identity association in a single model. Technical le

- arXiv
- 2511.20279
- Published
- 2025-11-25
- Authors
- Fabian Gülhan, Emil Mededovic, Yuli Wu, Johannes Stegmaier
AI summary
Overview
Research area: Computer vision — multi-object tracking (MOT), specifically end-to-end transformer trackers that unify object detection and identity association in a single model.
Technical level: Advanced. The paper assumes familiarity with DETR-style query-based detection, Hungarian matching, decoder self-attention, and standard tracking metrics (HOTA, DetA, AssA, IDF1, MOTA).
Scope: The paper analyzes why joint detection–association decoding in transformer trackers suppresses detection quality, and proposes SelfMOTR, a detector-free method that generates its own detection priors internally to decouple proposal discovery from association.
What This Paper Is About
End-to-end transformer trackers such as MOTR unify detection and association into one learnable process, but they consistently detect worse than dedicated detectors because track queries and detect queries interfere with each other during joint decoding. Prior fixes either add generous label assignment, denoising, or — most effectively — inject detection priors from an external pretrained detector, which reintroduces a second model and partially breaks the self-contained design. SelfMOTR instead asks the tracker to generate its own detection priors: a detection-only forward pass produces 4D anchor proposals that are then fused with propagated track queries in the tracking pass, removing the interference without any external detector.
Key Contributions
-
A diagnostic for query interference. The paper introduces Track Attention Mass, a cardinality-corrected decoder self-attention statistic (paired with normalized Shannon entropy), inspired by attention-sink analyses in large language models, to measure how detect queries distribute attention between track keys and detection keys.
-
Evidence of latent detection capacity. The authors reproduce MOTR training on DanceTrack and evaluate every checkpoint twice — with and without track queries — showing that removing track queries at inference produces higher and more stable mAP throughout training, i.e. the detector exists but is suppressed by joint decoding.
-
A systematic comparison of four prior-injection strategies. Detection pretraining, query pretraining, distillation, and anchor proposal are each evaluated on the DanceTrack validation set, with anchor proposal giving the largest single gain (HOTA 51.2 → 58.0).
-
SelfMOTR, a detector-free end-to-end tracker. The model generates self-proposals from its own features and a confidence threshold, fuses them with track queries in a shared decoder, and reaches 69.2 HOTA on DanceTrack, 71.1 HOTA on BFT, and best-among-end-to-end results on AnimalTrack.
Main Findings
-
Detection collapses when tracking is switched on in MOTR. With track queries disabled, MOTR reaches 66.3 AP, 85.5 AP50 and 70.7 AP75; enabling track queries drops these to 60.0, 78.1 and 63.7 (a −6.3 AP loss). SelfMOTR shows only a −0.3 AP drop (71.2 → 70.9), effectively removing the detection–association conflict.
-
Anchor proposal is the strongest single injection strategy on DanceTrack validation. Starting from MOTR (HOTA 51.2, DetA 68.8, AssA 38.4, IDF1 49.1, MOTA 74.4), detection pretraining gives 51.5 HOTA, query pretraining 52.7 HOTA with AssA 41.1, distillation 52.1 HOTA with DetA 72.0 and MOTA 80.1, and anchor proposal 58.0 HOTA (+6.8), AssA 47.5 (+9.1) and IDF1 59.5 (+10.4).
-
DanceTrack test performance. SelfMOTR attains 69.2 HOTA, 80.9 DetA, 59.3 AssA, 72.5 IDF1 and 89.9 MOTA, described as on par with CO-MOT (69.4 HOTA, 82.1 DetA, 58.9 AssA, 71.9 IDF1, 91.2 MOTA) while achieving the highest association score (59.3 AssA) among the end-to-end methods compared.
-
Best result on Bird Flock Tracking. SelfMOTR leads all evaluated non-end-to-end and end-to-end methods on BFT with 71.1 HOTA, 82.7 IDF1 and 77.6 MOTA.
-
Strong result on the small AnimalTrack benchmark. SelfMOTR reaches 45.5 HOTA, 53.7 IDF1 and 49.5 MOTA, a reported gain of +14.5 HOTA over TrackFormer (31.0 HOTA) among tracking-by-propagation approaches.
-
Attention polarization is avoided. In MOTR, detect-query attention mass is broadly dispersed and skewed toward high track attention, with a drop in normalized entropy at both extremes. SelfMOTR concentrates query mass into a balanced regime centered near 0.4 with entropy approximately 0.85 in the final decoder layer.
-
Data scaling favors the decoupled design. Adding HSV augmentation and joint CrowdHuman training raises the learnable-anchor variant from 55.8 to 60.5 HOTA and the self-proposal variant from 58.2 to 64.3 HOTA.
-
A low proposal threshold is better for association. A confidence threshold of 0.05 yields 58.2 HOTA and 47.4 AssA, versus 57.7 HOTA and 46.4 AssA at 0.50, while the higher threshold slightly improves DetA (72.1 vs 71.9).
-
Shared decoder weights matter for association, not detection. A dual-decoder variant with separate parameters per pass reaches comparable detection but only 54.5 AssA / 69.1 IDF1 / 66.2 HOTA, versus 59.3 AssA / 72.5 IDF1 / 69.2 HOTA for the shared version.
-
Efficiency. SelfMOTR runs at 20.7 FPS on a Tesla L40S versus 24.8 FPS for MOTR with learnable 4D anchors, using 42M parameters versus 141M for MOTRv2 and 99M for the YOLOX detector MOTRv2 relies on. SelfMOTR's proposal AP is 71.2 versus YOLOX's 79.7.
-
MOT17 remains difficult for end-to-end tracking. SelfMOTR scores 57.5 HOTA, 57.3 DetA, 58.1 AssA, 70.6 IDF1 and 70.0 MOTA; an MOT17-optimized regime reducing clip length from 5 to 3 improves DetA by +2.7, HOTA by +0.9 and MOTA by +3.5 to 58.4 HOTA, but AssA decreases slightly and IDF1 stays nearly constant.
Methodology in Plain English
The authors start by proving the problem exists. They train MOTR normally on DanceTrack, then re-evaluate each saved checkpoint twice: once as usual, and once with the track queries switched off. The detection-only runs score clearly higher, which shows the model already knows how to detect and that tracking is what degrades it.
They then test four ways of putting detection strength back in. Detection pretraining trains MOTR as a plain detector on the target data first, then fine-tunes it for tracking. Query pretraining transfers and freezes only the detect query embeddings. Distillation freezes a detection-trained MOTR as a teacher and adds a Hungarian-matched distillation loss to the student's tracking loss. Anchor proposal converts the teacher's boxes into 4D anchor boxes used during both training and inference. Anchor proposal wins by a wide margin but still depends on a separately trained frozen detector.
SelfMOTR removes that dependency. At each frame, the model first runs a detection-only decoder pass using detect queries, producing box predictions with confidence scores. Predictions above a proposal threshold are converted into proposal queries: the 4D box becomes the positional part, and the content part is a learned shared proposal embedding plus a sine–cosine encoding of the confidence. Ten learned proposal anchors are appended to guard against missed detections. These proposals are concatenated with the track queries propagated from the previous frame, and a single shared decoder refines and associates them jointly. The detection-only pass is supervised with the standard Deformable DETR set-based matching loss (focal classification, L1 box regression, generalized IoU), weighted by λ_prop = 0.5.
To explain why this works, the authors borrow a tool from language model analysis. For each detect query they compute Track Attention Mass — the ratio of average attention per track key to the average over track and detection keys combined, which corrects for the fact that the number of keys varies per frame. They also compute the normalized Shannon entropy of each query's attention row to measure whether attention is concentrated or spread out. Plotting these across decoder layers shows MOTR polarizing into track-dominated and track-ignoring queries, while SelfMOTR stays in a balanced, high-entropy range.
Why This Matters
Impact on research. The paper reframes the detection–association conflict as an attention-allocation problem rather than a capacity problem, and shows the hidden detector inside a joint transformer tracker can be harvested at inference time without any external model. This argues for internal prior generation as a default design principle for unified tracking architectures, and it imports a diagnostic tool (attention sinks) from language modeling into vision tracking, which may be reusable for other query-based multi-task decoders.
Real-world applications.
- Sports and dance analytics: tracking performers with highly uniform appearance and unpredictable, non-linear motion, exactly the DanceTrack setting.
- Wildlife monitoring and ecology: the BFT setting of dense, fast-moving, continuously deforming bird flocks where individual appearance cues are weak.
- Animal behavior research and livestock monitoring: AnimalTrack-style scenes with heavy occlusion and dense target concentrations (averaging 33 targets per sequence).
- Video surveillance and crowd monitoring: identity-preserving tracking in crowded scenes, where the method's high AssA (59.3) directly addresses identity-continuity failures.
Industry relevance. The framework is self-contained and compact (42M parameters versus 141M for MOTRv2), avoids maintaining a separate detector, and tolerates shallow proposal decoders, so practitioners can trade one or two decoder layers for speed. That matters for deployment on constrained hardware, though the reported 20.7 FPS on a Tesla L40S shows it is not yet a real-time solution on that hardware.
Future Directions
-
Adaptive reuse of latent capacity. The authors state that joint-decoding transformers have sufficient capacity for robust detection and tracking and call for more adaptive ways to leverage this capacity beyond a fixed detection-only pass.
-
Closing the proposal-quality gap. SelfMOTR's internal proposal AP (71.2) trails YOLOX's (79.7), which carries over to tracking AP (70.9 vs 75.3). Improving internal proposals without external detectors is the obvious lever for DetA and MOTA.
-
Improving association under low-data or short-clip regimes. On MOT17, reducing clip length from 5 to 3 improved detection metrics but did not improve association (AssA decreased slightly, IDF1 nearly constant), leaving open how to convert detection gains into identity gains there.
-
Understanding why end-to-end trackers still lose to tracking-by-detection on MOT17. The paper notes that even the strongest end-to-end method, CO-MOT, falls short of non-end-to-end HOTA on that benchmark, and that MeMOTR and SambaMOTR get competitive results only by offloading association to extra modules.
Target Audience
Researchers and graduate students working on multi-object tracking, query-based transformer detection, or end-to-end multi-task vision models, particularly those interested in removing external detectors from tracking pipelines. It is also relevant to practitioners who need identity-stable tracking in crowded or visually homogeneous scenes (sports, wildlife, surveillance) and to anyone looking for a template for diagnosing attention imbalance in decoder-based architectures. Readers should already be comfortable with DETR-style set prediction, Hungarian matching, and standard MOT metric definitions.
Authors’ abstract
End-to-end transformer architectures have driven significant progress in multi-object tracking by unifying detection and association into a single, heuristic-free framework. Despite these benefits, poor detection performance and the inherent conflict between detection and association in a joint architecture remain critical concerns. Recent approaches aim to mitigate these issues by employing advanced denoising or label assignment strategies, or by incorporating detection priors from external object detectors. In this paper, we propose SelfMOTR, a simple yet highly effective detector-free alternative that decouples proposal discovery from association using self-generated internal detection priors. Through extensive analysis and ablation studies, we show that end-to-end transformer trackers with joint detection-association decoding retain substantial hidden detection capacity, and we provide a practical detector-free mechanism for leveraging it. To shed light on these joint decoding dynamics, we draw inspiration from attention sink analyses in large language models, leveraging Track Attention Mass to show that standard generic queries exhibit unbalanced attention, frequently struggling to weigh track context against novel object discovery. SelfMOTR achieves highly competitive performance in complex, dynamic environments, yielding 69.2 HOTA on DanceTrack and leading with 71.1 HOTA on the Bird Flock Tracking (BFT) dataset. Project page: https://medem23.github.io/SM