Research
RenderRank: Learning to Rerank Text with Compressed Visual Tokens
Overview Research area: Information retrieval (cs.IR), specifically document reranking for retrieval-augmented generation, combined with vision-language model text compression. Technical level: Advanc

- arXiv
- 2609.35069
- Published
- 2026-09-28
- Authors
- Seongtae Hong, Youngjoon Jang, Jungseob Lee, Hyeonseok Moon, Heuiseok Lim
AI summary
Overview
Research area: Information retrieval (cs.IR), specifically document reranking for retrieval-augmented generation, combined with vision-language model text compression.
Technical level: Advanced. The paper assumes familiarity with cross-encoder reranking, knowledge distillation, contrastive (InfoNCE) objectives, and vision-language model architecture.
One-sentence scope: RenderRank renders document text as images, encodes them into compressed visual tokens with a frozen vision encoder, and trains only the language decoder to score how relevant each rendered document is to a text query.
What This Paper Is About
Conventional rerankers feed a query and each candidate document into a language model as text tokens; because every candidate is scored separately, document length multiplies the cost of evaluating a whole candidate set. RenderRank instead renders each document as page images, runs them through a vision encoder to produce a shorter sequence of visual tokens, and learns to predict relevance scores from those compressed representations. The goal is to keep ranking quality high while using fewer input tokens, which in turn lowers computation and increases throughput.
Key Contributions
-
A reranker that scores relevance from compressed visual document representations. RenderRank renders candidate documents as images at a fixed configuration (Roboto Regular, 12pt, line spacing 1.0, 96 DPI, width 896 pixels, maximum height 896 pixels) and encodes them into visual tokens instead of using text token sequences.
-
A two-stage training recipe for cross-modal relevance learning. Cross-Modal Relevance Distillation minimizes mean squared error between the student's visual-input score and a text-based teacher's score, followed by Query-Local Relevance Discrimination, an InfoNCE contrastive loss that orders the positive document above negatives sharing the same query.
-
Query-independent, precomputable document encoding. Document images are encoded in advance, independently of the query, and the resulting visual representations are used to predict scores conditioned on the textual query — so document representations can be cached rather than recomputed per query.
-
A joint evaluation of ranking quality, token counts, estimated FLOPs, and measured GPU throughput. The paper reports results on 11 BEIR datasets and four long-document datasets (MLDR, 2WikiMQA, QMSum, SummScreenFD), plus ablations over training stages, rendering font size, and sequence-length budgets.
Main Findings
-
BEIR ranking quality and token savings: Across 11 BEIR datasets, RenderRank (2.1B parameters) reaches an average NDCG@10 of 55.96 using an average of 290.07 input tokens per query–document pair, which is 16.5–35.5% fewer tokens than the evaluated text-based rerankers. It outperforms all evaluated text-based baselines below 4B parameters and some larger models; the paper states gains of 7.66% and 7.26% in average NDCG@10 over zerank-2-reranker and LightOn-rerank-PW-4B respectively.
-
Long-document performance: Across the four long-document datasets, RenderRank achieves an average NDCG@10 of 88.27 with approximately half the average input token count of the evaluated text-based rerankers. On MLDR it reaches NDCG@10 of 99.74 using an average of 4.20K input tokens; on the three LongEmbed datasets it uses 53–57% fewer input tokens than the most token-efficient text baseline.
-
Throughput: On BEIR, RenderRank averages 76.83 PPS (pairs per second) measured on an NVIDIA A6000 48GB, approximately twice the throughput of bge-reranker-v2-gemma at comparable reranking performance, and higher than the 67.55 PPS measured for the smaller LAMAR-600m despite RenderRank's higher estimated TFLOPs per pair. On the four long-document datasets it delivers the highest throughput of the compared models, averaging 4.51 PPS versus 2.66 PPS for the fastest baseline, a 1.70× improvement, with per-dataset speedups from 1.55× to 1.91×.
-
Estimated computational cost: RenderRank averages 0.63 TFLOPs per query–document pair and a QPP (queries per petaFLOP) of 20.1, compared with 1.10 TFLOPs per pair and QPP of 12.3 for bge-reranker-v2-gemma.
-
Fixed sequence-length budgets: Under identical maximum sequence lengths on MLDR, RenderRank scores 97.9 at 2K, 99.5 at 4K, and 99.7 at 8K — the best at each length. Its 2K result exceeds gte-reranker-modernbert (95.8) and LightOn-rerank-PW-4B (97.3) at 4K, and its 4K result is comparable to text-based rerankers at 8K.
-
Training-stage ablation: Removing Query-Local Relevance Discrimination lowers average NDCG@10 from 55.96 to 54.86; removing both stages lowers it to 50.20, showing Cross-Modal Relevance Distillation raises performance with image inputs from 50.20 to 54.86.
-
Caching effect: With cached image inputs, throughput is 76.83 PPS; when visual encoding is included in the measurement, it drops to 37.31 PPS.
-
Rendering density trade-off: Varying font size from 10pt to 14pt raises average input length from approximately 236 to 370 tokens and lowers throughput from approximately 94 to 63 PPS, while improving NDCG@10 from 55.04 to 56.43. The default 12pt setting gives 55.96 NDCG@10 at approximately 290 tokens and 77 PPS, a relative difference of approximately 0.83% versus 14pt.
-
Rendering configuration selection: On QASPER and GovReport generation tasks used to choose the rendering setup, with line spacing fixed at 1.0, 12pt outperformed 10pt on both tasks; for Qwen3.6, QASPER F1 rose from 41.32 at 10pt to 48.72 at 12pt. Increasing line spacing at 12pt added visual tokens without consistent performance gains.
Methodology in Plain English
The authors start from the observation that a reranker must score many candidates per query, so any per-document token reduction is multiplied across candidates. They test how document rendering choices — font size and line spacing, evaluated on QASPER question answering (F1) and GovReport summarization (ROUGE-L) — affect both how well a model understands the rendered text and how many document tokens the rendering produces. They settle on Roboto Regular at 12pt with line spacing 1.0, rendered at 96 DPI at 896 pixels wide with a height in 32-pixel increments, chosen partly to match the 16×16-pixel patches and 2×2 spatial merging of the Qwen3-VL vision encoder. Long documents become multiple ordered images processed as one candidate.
RenderRank is initialized from Qwen3-VL-Reranker-2B. Training has two stages. First, Cross-Modal Relevance Distillation trains the model to match the relevance scores of a text-based teacher (Qwen3-Reranker-4B) on original text query–document pairs, using mean squared error — aligning at the score level rather than requiring token-level correspondence between text and images. Second, Query-Local Relevance Discrimination uses an InfoNCE contrastive loss over each query's candidate set (one positive plus K negatives) so that positives score above negatives for the same query. The vision encoder and visual feature mergers stay frozen across both stages; LoRA is applied only to the text decoder, and following ReLoRA's merge-and-reinitialize principle the first-stage adapters are merged into the backbone before fresh adapters are initialized for the second stage.
Training data: the English fine-tuning data released by Sourty et al. (2026), comprising 1.57M training records with one positive and ten negatives per query; the authors retain all records and select one positive plus three negatives per query, yielding 6.28M query–document pairs. The second stage uses RLHN-100K.
Evaluation: on BEIR, the top 100 documents per query retrieved by BM25 are reranked, using NDCG@10. For long documents, the English subset of MLDR (following MMTEB) plus LongEmbed's 2WikiMQA, QMSum, and SummScreenFD, with the top eight documents per query retrieved by Qwen3-Embedding-0.6B. All baselines use text queries and documents; RenderRank keeps queries as text and renders only documents as images. The authors assume document visual embeddings are precomputed. Efficiency is measured two ways: estimated cost via the E2R-FLOPs approximation (reported as TFLOPs per pair and QPP), and measured Pairs per Second on an NVIDIA A6000 48GB, taking the maximum over successfully measured batch sizes starting at 8 and doubling until out-of-memory.
Why This Matters
Impact on research. The paper shows that relevance scoring is possible from compressed visual document representations without aligning text and image features in a shared space, and that score-level distillation transfers a text reranker's judgments to a visual-input student. It also frames reranking as a token-budget problem, reporting performance at matched maximum sequence lengths (2K/4K/8K) rather than only at each model's native limit, which makes representations comparable under equal constraints.
Real-world applications (illustrative):
- Retrieval-augmented generation pipelines where reranking many long candidates dominates latency and cost.
- Search over long or dense documents (reports, filings, transcripts) where the evidence needed to judge relevance sits far beyond a truncated text window.
- Serving scenarios with fixed GPU memory or strict compute budgets, where precomputed document embeddings can be cached and reused across queries.
- Systems that already hold documents as page images (for example, scanned or paginated archives) and can reuse the same rendering pipeline for downstream generation.
Industry relevance. Because the vision encoder is frozen and only the decoder is adapted with LoRA, the approach reuses an existing vision-language backbone rather than requiring a jointly trained text compressor. The reported throughput gains (76.83 PPS on BEIR; 4.51 PPS average on long documents) and the caching effect (76.83 PPS cached versus 37.31 PPS with encoding included) are the numbers most directly relevant to production deployment decisions.
Future Directions
- Extending to longer documents. The authors frame this work as opening a path toward representing and reranking longer text documents; the current rendering uses a maximum page height of 896 pixels with continuation onto additional images.
- Tuning rendering density per deployment. The font-size analysis shows compression directly trades ranking quality for tokens and throughput, leaving open how to select or adapt the rendering configuration for a given quality or latency target.
- Reducing or removing reliance on precomputation. The reported throughput assumes document visual embeddings are computed in advance; the paper does not report an end-to-end pipeline cost that includes offline encoding.
- Comparing against other compression strategies. The paper's related work describes jointly trained text compressors with fixed memory tokens, query-dependent key–value selection, and multi-ratio compression; the paper does not report head-to-head comparisons against those approaches.
Target Audience
Information retrieval and RAG researchers studying reranking efficiency; engineers building retrieval or reranking services with latency or memory constraints; and vision-language researchers interested in whether rendered text can substitute for text tokens in ranking tasks. Readers should be comfortable with NDCG@10, knowledge distillation, and contrastive losses, though the paper's main argument — that fewer tokens per document can preserve ranking quality — is stated plainly enough for a broader applied audience.
Authors’ abstract
Rendering document text as images allows vision-language models to encode documents as visual tokens, which can reduce input sequence length compared with text input. This reduction in input length is particularly useful for reranking, where each query involves scoring multiple candidate documents and token savings apply to each candidate evaluation. We introduce RenderRank, a reranker that learns query-dependent relevance scoring from compressed visual document representations instead of the text token sequences used by conventional text-based rerankers. Training first aligns relevance scores from visual inputs with those of a text-based teacher, then refines the relative scores of positive and negative documents for the same query. Across 11 datasets from BEIR, RenderRank uses 16.5-35.5% fewer input tokens while achieving an average NDCG@10 of 55.96, outperforming all evaluated text-based baselines below 4B parameters and some larger models. Across four long-document datasets, it achieves an average NDCG@10 of 88.27 with approximately half the average input token count of the evaluated text-based rerankers. In this setting, RenderRank delivers 1.70x the highest average throughput of the evaluated baselines. These results demonstrate that compressed visual representations can support accurate document relevance scoring, providing an alternative to text token representations for reranking.