Research
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
Overview Research area: High-resolution multimodal (vision-language) modeling, specifically image-tiling strategies for large multimodal models, framed as a reproducibility study of the Monkey VLM (Li

- arXiv
- 2512.11167
- Published
- 2025-12-11
- Authors
- Anatole Jacquin de Margerie, Alexis Roger, Irina Rish
AI summary
Overview
- Research area: High-resolution multimodal (vision-language) modeling, specifically image-tiling strategies for large multimodal models, framed as a reproducibility study of the Monkey VLM (Li et al. 2023b, CVPR24).
- Technical level: Intermediate. The paper assumes familiarity with vision-language architectures, training phases, LoRA, and standard VQA benchmarks, but the core idea (splitting images into tiles) is easy to grasp.
- Scope: The authors reimplement the Monkey VLM tiling pipeline inside LLaVA, evaluate 2×2 and 3×3 tilings with and without a resized global view against a 1×1 single-image baseline across six benchmarks, and report where the original qualitative claims hold and where results diverge.
What This Paper Is About
Most vision encoders accept only a fixed input resolution, so high-resolution images must be resized, which destroys fine details. The Monkey VLM framework proposed splitting large images into tiles, encoding each tile separately, and fusing their representations, optionally adding a downsampled full-image "global context." This paper asks whether that approach is robust across downstream tasks, how much the global view restores lost coherence, and whether independent researchers can reproduce the reported performance under realistic compute budgets.
Key Contributions
- A full reproduction of the Monkey VLM training pipeline, including tiling configurations and combined global-context inputs, reimplemented within the LLaVA architecture using open checkpoints and a reimplemented training pipeline.
- A systematic evaluation across six benchmarks (ScienceQA Image, GQA, TextVQA, MMVET, LLaVA-Bench, POPE), analyzing trade-offs between local detail and global coherence.
- A reflection on the reproduction process, documenting that tiling effectiveness is task-dependent and that several under-specified implementation details hindered reproducibility.
- Reporting of training energy costs in kWh for each configuration (Table 1), showing that adding global context incurred modest cost increases (5–12% during finetuning).
Main Findings
- Detail-oriented tasks benefit from tiling plus global context. On ScienceQA-Image, the 2×2 split without global context (73.53%) outperformed the 1×1 baseline (71.34%), and the 3×3 split with global view reached 80.17%, surpassing all other configurations.
- Global context fully mitigates coherence loss on GQA. The 2×2 configuration with global context scored 54.20%, essentially matching the baseline (54.08%).
- Multi-capability benchmarks can improve. On MMVET, the 3×3 split with global context (29.50%) outperformed the baseline (28.30%).
- Text-centric tasks did not improve. On TextVQA, every tiling configuration lagged the baseline (52.99%); the best tiling result was 2×2 with global context at 49.96%. The authors suggest simple concatenation may be suboptimal for OCR and reasoning about text scattered across image regions.
- Open-ended generation degraded sharply. On LLaVA-Bench, all tiling configurations performed much worse than the baseline (58.40%); the best tiling result was 2×2 with global context at 30.60%, and 3×3 without global context fell to 16.70%.
- Hallucination results were mixed. On POPE (average F1 over Random, Popular, and Adversarial), both 2×2 configurations exceeded the baseline (78.57%): 2×2 without global context reached 81.86% and 2×2 with global context 81.53%. The 3×3 split without global context performed poorly at 33.01%.
- Extreme fragmentation without context is harmful. The 3×3 split without global context also dropped to 66.09% on ScienceQA-Image and 38.54% on GQA.
- Computational overhead was modest. Training energy ranged from 22.28 kWh (pretrain) and 50.84 kWh (finetune) for the 1×1 baseline up to 23.04 kWh (pretrain) and 63.40 kWh (finetune) for 3×3 with global context. Pretraining cost differences across splits were minimal, and the authors contrast this with Qwen-VL, whose training required thousands of GPU days.
- Reproduction confirmed qualitative trends but not uniform superiority. Tiling improved recognition of fine-grained details, the global view mitigated coherence loss, and compute overhead stayed low — but the original paper's uniformly positive framing did not hold, as optimal performance on one benchmark did not translate to others.
- Implementation fragility. Under-specified details — the precise ordering of patch embeddings, the data preprocessing pipeline, and finetuning hyperparameters — produced measurable performance differences.
Methodology in Plain English
The authors rebuilt Monkey's image-splitting idea inside the LLaVA framework. They used a ViT-S/16 vision encoder pre-trained with SigLIP and OpenHermes-2.5-Mistral-7B as the language backbone — a deliberate deviation from the original paper's Qwen-VL, which they justify by noting both models are of the same size and similar performance. Images are divided into non-overlapping tiles; each tile, and optionally a resized copy of the whole image, is encoded independently by separate encoder instances; all visual tokens are then concatenated before the projection layer and the language model. No cross-tile attention was used, mirroring Monkey's design.
They compared 2×2 and 3×3 tilings, each with and without global context, against a 1×1 single resized image baseline. Training followed two phases: pretraining the MLP projection layer on the BLIP-LAION-CC-SBU-558k dataset with a learning rate of 1e-3 while the vision encoder and language model stayed frozen; then finetuning on the Monkey training data (1.44m samples drawn from public datasets such as COCO Caption, TextCaps, VQAv2, OKVQA, GQA, ScienceQA, VizWiz, TextVQA, OCRVQA, AI2D, DocVQA, ChartQA, InfoVQA, DeepForm, KLC, WTQ, TabFact, and VisualMRC) with a learning rate of 2e-5 and LoRA for the language model, updating the vision encoder, the MLP projection layer, and the LoRA adapters. Evaluation used six benchmarks: ScienceQA Image, GQA, TextVQA, MMVET, LLaVA-Bench, and POPE.
Why This Matters
- Impact on research: The paper makes the case that complex multimodal models often lack transparent implementation details and accessible training infrastructure, and that even minor documentation gaps (patch normalization and ordering, dataset filters, finetuning hyperparameters) can block replication. It calls for complete training scripts and preprocessing pipelines rather than configuration summaries.
- Real-world applications:
- Medical imaging, where both local detail and global anatomical context matter.
- Document understanding, including charts and tables.
- Remote sensing and aerial imagery analysis.
- Text-in-image tasks such as scene-text reading, OCR-based VQA, and document VQA.
- Industry relevance: The results show a practical, low-cost route to high-resolution capability built on open off-the-shelf components, with finetuning energy in the tens of kWh rather than the thousands of GPU days the authors attribute to full-model retraining approaches like Qwen-VL. The trade-off is task dependence: the approach pays off on detail-oriented benchmarks but not on open-ended narrative generation.
Future Directions
- Adaptive splitting: move beyond fixed 2×2 or 3×3 grids toward models that dynamically adjust their processing strategy based on image content and task requirements.
- Content-aware tiling: apply finer splits to detailed regions while using coarser processing for homogeneous areas, and in specialized domains optimize different vision towers for specific regions of interest (the authors give the example of a dedicated heart split in a full chest scan).
- Better fusion: replace the simple concatenation of tile and global representations with attention-based fusion mechanisms, which the authors expect could improve coherence in generative tasks.
- Reproducibility infrastructure: publish full training scripts and preprocessing pipelines, and explore learning-based attention that identifies regions needing high-resolution analysis, mirroring human foveal vision.
Target Audience
Researchers and engineers working on multimodal and vision-language systems, particularly those interested in high-resolution image understanding, efficient training, and model reproducibility. It is also useful for practitioners in medical imaging, document understanding, and remote sensing who need to weigh detail preservation against global scene coherence, and for anyone attempting to reproduce or extend the Monkey VLM approach. Note that the paper does not report specific GPU hardware, GPU-hour counts for its own runs, or wall-clock training times — only energy in kWh.
Authors’ abstract
Reproducibility remains a cornerstone of scientific progress, yet complex multimodal models often lack transparent implementation details and accessible training infrastructure. In this work, we present a detailed reproduction and critical analysis of the Monkey Vision-Language Model (VLM) (Li et al. 2023b) published in CVPR24, a recent approach to high-resolution image understanding via image tiling. The original paper proposed splitting large images into tiles to recover fine-grained visual details while maintaining computational efficiency. Our study replicates this strategy using open checkpoints and reimplements the training pipeline. We confirm the key finding of the original Monkey VLM work, namely that tiling effectively recovers local details. We then extend this work further, by investigating the effect of the inclusion of the global context, which provide practical insights for future high-resolution multimodal modeling. However, we also report deviations in the results, with the magnitude of these effects depending heavily on task type and tile granularity.