Research
Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Overview Research area: Multimodal representation learning, specifically omni-modal retrieval/embedding models that place text, speech, audio, images, video, and visually-rich documents into one share

- arXiv
- 2610.02148
- Published
- 2026-10-01
- Authors
- Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
AI summary
Overview
- Research area: Multimodal representation learning, specifically omni-modal retrieval/embedding models that place text, speech, audio, images, video, and visually-rich documents into one shared embedding space.
- Technical level: Advanced. The paper assumes familiarity with contrastive learning (SigLIP, InfoNCE), LoRA adapters, knowledge distillation, Matryoshka representation learning, and hard-negative mining, and it reports results on MTEB-v2, MAEB, MMEB-V2 and ViDoRe-V3.
- Scope: A single paper presenting Omni-Embed-Mini, a 0.9B-parameter omni-modal embedder built by freezing a text backbone and distilling each new modality toward that same frozen backbone's embedding of a dense generated caption, plus a 2.3B variant built by swapping in a native vision-language backbone.
What This Paper Is About
Adding new modalities to a text embedding model normally damages text retrieval (catastrophic forgetting), and existing omni-modal embedders avoid this by scaling to billions of parameters (the paper calls this the "multimodal expansion dilemma"). Omni-Embed-Mini instead treats text as an immovable anchor: no text-side parameter is ever updated, and every media modality is aligned to the frozen backbone's own embedding of a dense cascaded caption for that sample. The goal is a sub-billion-parameter model that covers six modalities in one cosine space while leaving text retrieval exactly untouched.
Key Contributions
- Zero-regression omni embedding. A 0.9B model scores 49.57 nDCG@10 on MTEB-v2 BEIR-8 while adding five modalities; the text branch's weights are bit-identical to the stock 0.6B backbone, so training cannot regress it.
- Cascaded-caption self-distillation through a shared backbone. Teacher and student share the same frozen weights, so cosine alignment in the native pooled hidden space suffices, with no projection head, no separate embedding teacher, and fully cacheable targets.
- A parameter-efficient training recipe combining Matryoshka SigLIP contrastive loss, an online hybrid hard-negative miner whose negatives sharpen as the encoder improves, and phased LoRA, which carries over to a 2.3B variant by swapping the backbone for a native vision-language model.
- The smallest open omni embedder reported at the time of writing, covering text, speech, audio, image, video and visually-rich documents in one cosine space.
Main Findings
- Text retrieval is preserved by construction. Omni-Embed-Mini-0.9B scores 49.57 nDCG@10 on MTEB-v2 BEIR-8, and the text-only inference path touches only unchanged backbone weights. Independently trained checkpoints agree on all eight tasks to four decimal places; across three random seeds (42, 71, 1234) text is identical at 49.57.
- Size advantage. At 0.9B the model is roughly 2.7x smaller than the next entry (BidirLM-Omni-2.5B) and roughly 9.5x smaller than LCO-Embedding-Omni-7B. Inference parameters are 935.35M, with 869.98M frozen and 68.32M trained (7.3%); the 68.32M splits into 65.37M of projectors and pooler plus 2.95M of r=16 LoRA adapters, which are merged at export.
- Text ranking among omni embedders. At 49.57, the 0.9B model is second among the compared open omni embedders only to omni-embed-nemotron-3B (50.48), which is five times larger.
- Speech and audio trail at 0.9B, which the paper describes as the expected cost of adding those modalities through small projectors rather than retraining the backbone.
- The 2.3B variant leads the omni block on video. Built on a frozen Qwen3-VL-Embedding-2B, it reaches an overall-modality average of 51.39, with Video 55.18, Image 64.80 and Visual-Doc 58.10. It edges ahead of the closed gemini-embedding-2 on the overall average (51.39 against 49.51), leading on image, video and visual documents while trailing on text, speech and audio.
- Comparison with e5-omni. e5-omni-7B leads on the overall average (52.93 against 51.39), while the 2.3B edges ahead of e5-omni-3B (51.39 against 49.28). The paper reports leading on video (55.18 against 32.92 and 44.73) at a half to a quarter of their parameters, and notes e5-omni's own text suite falls from 47.80 to 45.34 going from 3B to 7B.
- The 0.9B vision scores are limited by encoder capacity. Image 26.29 and Video 18.48, because its vision encoder is extracted from a 0.8B backbone with a comparatively small ViT, trained on roughly 86K image and roughly 36K video rows.
- Unfreezing the backbone behaves differently at the two scales. At 0.9B, backbone LoRA improves every media modality (+1.10 to +5.66) while text drops 3.13 points; at 2.3B the same intervention degrades every modality, most sharply the two the adapters were trained for (speech -10.53, audio -8.40), with text, image and video shifting by under a point. Text improved at neither scale, so the backbone is kept frozen.
- Dense captions are the single most important data-side ingredient. On the 0.9B mining-ablation checkpoint, dense captions over original captions give Speech +2.68, Audio +2.70, Image +8.60, Video +14.06, and Overall +7.01 (23.78 to 30.79).
- Combined mining is best. Only combined (text plus media) mining separates clearly from no mining (+2.6 on the mean over the five media modalities); text-only mining slightly hurts several modalities, and visual documents are near-insensitive to mining (all four configurations within 0.6 nDCG@10, 46.42 to 46.98). Replacing dense captions with original captions costs 4.8 nDCG@10 on that modality (46.98 to 42.21).
- Within-modality retrieval is retained without dedicated training. On NIGHTS and FashionIQ (Table 7) the 2.3B scores 68.00 and 34.50 against its frozen Qwen3-VL-Embedding-2B backbone's 67.80 and 37.50; the 0.9B, never trained on an image-to-image objective, keeps 53.90 against 59.30 on NIGHTS and 7.40 against 3.80 on FashionIQ.
- Matryoshka behaviour is inherited from the backbone. Averaged-probe AVG falls only slightly from 1024 to 128 dimensions for both variants, whereas the omni-specific baselines fall away far more steeply. For the 0.9B, a 128-dim truncation reduces on-disk footprint 8x: a 10M-item omni index in float32 fits in roughly 5 GB rather than roughly 41 GB.
- Seed stability on the media modalities. Across three seeds of the 10-epoch no-mining configuration: image 27.48 ± 0.43, general audio 33.96 ± 0.59, video 19.82 ± 0.95, speech 43.10 ± 1.03, visual-document 45.48 ± 1.76.
- Parameter budget relative to baselines. Table 1 lists 2.3B variant training at 141.68M of 2.44B inference (5.8%), versus BidirLM-Omni-2.5B at 1.72B trained of 2.45B (70.4%), and e5-omni/LCO/nemotron variants at 4.70B or 8.93B inference with far smaller trained fractions. gemini-embedding-2 is omitted because its parameter count is not disclosed.
Methodology in Plain English
The model keeps one thing completely fixed: an off-the-shelf text embedding backbone. The 0.9B version uses Qwen3-Embedding-0.6B (hidden size 1024) with a ViT extracted from Qwen3.5-0.8B; the 2.3B version uses Qwen3-VL-Embedding-2B (hidden size 2048) with its native vision module. Because that backbone is frozen, its text behaviour cannot change during training.
Each training media sample is paired with a dense caption generated by Qwen3-Omni-30B-A3B-Instruct at temperature T=0.7, with the source dataset's ground-truth caption supplied as a "Reference caption" inside the system prompt (visual-document rows use an expert document analyst prompt with no reference caption, since none exists). The teacher target is simply the frozen backbone's L2-normalised EOS-pooled embedding of that caption, computed once and cached; the student is the same pooling applied to a token sequence with media spliced in at modality-tagged placeholder positions. Since teacher and student share weights, they live in identical geometry and no projection head is needed.
Media enters through encoders. Speech and audio both run in parallel: Whisper (mel) and Dasheng (raw waveform) each emit 128 projected tokens per 30-second chunk, which are temporally interleaved (w1, d1, w2, d2, ...) into 256 tokens, described as a zero-parameter structural inductive bias. Vision for the 0.9B passes through a ViT, a 2x2 spatial merger, a dimension projector, and a projection to the backbone hidden size, with a cross-attention spatial pooler against a learnable 196-query budget for video. The 2.3B uses the backbone's own visual module and injects audio through a forward hook on the embedding layer.
Training happens in two phases: Phase 1 trains projectors only, and Phase 2 (after 20% of steps) injects LoRA (r=16, alpha=32) on the media encoders. The loss aggregates a SigLIP pair-wise sigmoid contrastive loss across Matryoshka dimensions {128, 256, 512, 1024} (plus 2048 for 2.3B), with per-dimension weights proportional to d^(-1/2), plus a cosine self-distillation term with λ=1. Hard negatives come from a background subsystem maintaining FAISS-CPU inner-product indices: a text cycle re-embeds captions roughly 10x per epoch and a media cycle re-embeds samples roughly 5x per epoch, using a 0.92 cosine cutoff. Because the text branch never receives gradient, the text-side index is perfectly stable while only the media-side index drifts.
Data comes from the caption slice of an omni-modal corpus: 590,858 total rows filter to 242,080 single-turn samples, and a duration filter reduces this to 227,704 effective training rows. Training used 8x AMD Instinct MI210 (64 GB HBM2e each), bf16, DDP, with AdamW (wd 0.01, clip 1.0), 2% linear warm-up to cosine, LR 2x10^-4 for audio projectors and 1x10^-4 for vision projectors and encoder LoRA. Evaluation covers MTEB-v2 BEIR-8 (text, nDCG@10), MAEB (22 English audio tasks, mixed metrics), MMEB-V2 (16 tasks, 10 image and 6 video, hit@1) and ViDoRe-V3 (7 English visually-rich domains, nDCG@10).
Why This Matters
Impact on research. The paper reframes cross-modal alignment as self-distillation against a frozen anchor rather than joint contrastive retraining, arguing that text should be an immovable anchor rather than a co-learner. It reports evidence that jointly updating a backbone puts native retrieval behaviour at risk (for example, e5-omni's text suite falling from 47.80 to 45.34 between its 3B and 7B versions), and it shows that only combined text and media mining gives a clear gain, with visual-document retrieval responding to caption quality rather than negative sampling. It also leaves a methodological door open: because the backbone is frozen, a better text embedder could in principle be adopted by re-running only the projector and LoRA training, as the swap from the 0.6B text encoder to Qwen3-VL-Embedding-2B illustrates.
Real-world applications
- Text-first enterprise or web search that needs to accept image, video, audio or document queries without risking degradation of the existing text index.
- Edge or on-premise retrieval at scale, where the 0.9B inference footprint (935.35M parameters) and 128-dim truncation to roughly 5 GB for a 10M-item index matter.
- Visually-rich document retrieval (slide decks, reports, scanned pages) using the ViDoRe-V3 setting, where the 2.3B scores 58.10 nDCG@10.
- Media archive search over speech, environmental audio and video, where the 2.3B leads its omni comparison block on video at 55.18.
Industry relevance. The paper positions size-versus-performance as its main selling point: a 7.3% (0.9B) or 5.8% (2.3B) trained-parameter footprint, with LoRA adapters merged at export so inference parameter counts do not grow. It also argues for deployment predictability, since the deployed text inference path is unchanged and training cannot regress it, which matters when a text index is a production dependency.
Future Directions
- Better video handling. Replace per-frame mean-pooling for video with a native spatial pooler that better preserves temporal and spatial structure; the paper notes that using MMEB-V2's own 8 pre-extracted frames keeps comparisons uniform but caps achievable video performance.
- Learned mixing of the dual audio-encoder tokens instead of the fixed temporal interleaving, which the authors note may be suboptimal for inputs dominated by only one audio type.
- **
Authors’ abstract
Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: https://omniembed.cvmbzuai.com