Skip to content
AI.info

Research

A Matter of Time: Revealing the Structure of Time in Vision-Language Models

Overview Research area: Computer vision and multimodal representation learning, specifically the temporal awareness of vision-language models (VLMs). Technical level: Intermediate. The paper assumes s

A Matter of Time: Revealing the Structure of Time in Vision-Language Models
arXiv
2510.19559
Published
2025-10-22
Authors
Nidham Tekaya, Manuela Waldner, Matthias Zeppelzauer

AI summary

Overview

Research area: Computer vision and multimodal representation learning, specifically the temporal awareness of vision-language models (VLMs).

Technical level: Intermediate. The paper assumes some familiarity with contrastive VLMs and embedding spaces, but its core ideas (a timeline derived from embeddings) are described in largely accessible terms.

Scope: The paper introduces a benchmark (TIME10k) and a methodology for measuring whether 37 open-vocabulary VLMs implicitly encode when human-made objects first appeared, whether that temporal information has geometric structure in the embedding space, and whether an explicit "timeline" can be derived from it.

What This Paper Is About

Vision-language models such as CLIP are trained on web-scale image-text pairs whose alt-text often contains temporal hints ("a vintage car from the 1960s"). The authors ask whether these models have absorbed any genuine sense of time, meaning the ability to place a depicted human-made object at the point in history when it first appeared. Their goal is to measure that awareness systematically, to understand how time is arranged inside the models' shared image-text embedding space, and to turn whatever structure exists into an efficient and explicit timeline that can date new images without fine-tuning.

Key Contributions

  1. TIME10k, a temporally annotated benchmark of over 10,000 images (10,091 total) spanning six object classes (Aircraft, Cars, Instruments, Mobile phones, Ships, Weapons & Ammunition) with year-level ground truth covering 1715 to 2024, curated from Wikipedia and Wikimedia Commons.
  2. A comprehensive evaluation framework that probes the time-awareness of 37 state-of-the-art VLMs across different architectures, backbones, training datasets, and prompt formulations, without fine-tuning.
  3. An embedding space analysis showing that temporal information is organized along a low-dimensional, non-linear manifold rather than a linear subspace, a finding the authors contrast with prior work reporting linear temporal subspaces in language models.
  4. Two timeline representations derived from the embedding space: an implicit UMAP-based projection optimized with the Tree-structured Parzen Estimator (TPE), and an explicit Bézier curve approximation. Both aim to model chronological progression and are reported as competitive to superior in accuracy versus the prompt-based baseline at lower computational cost.

Main Findings

  • Time awareness varies widely by architecture and training data. Using prompt P7, the strongest results come from EVA-CLIP (EVA02-CLIP-L-14-336: MAE 6.20, TAI 0.86) and OpenCLIP (ViT-bigG-14-quickgelu: MAE 6.31, TAI 0.85), followed by ImageBind (MAE 7.16, TAI 0.83) and ViT-Lens (MAE 7.78, TAI 0.85). Weak performers include OpenCLIP ViT-B-32 trained on CommonPool (MAE 144.54, TAI 0.08), SigLIP nillb-clip-large-siglip trained on V1 (MAE 96.46, TAI 0.32), and ViTamin-S (MAE 75.80, TAI 0.48). Backbone scale matters: ViTamin-XL-384 reaches MAE 6.48 and TAI 0.86 while ViTamin-S reaches MAE 75.80 and TAI 0.48.

  • Prompt wording strongly affects temporal prediction. Minimal prompts perform poorly (P1 "[year]": MAE 48.28, TAI 0.70 for CLIP and MAE 48.14, TAI 0.73 for EVA-CLIP; P2 "Year [year]": MAE 19.89 / 16.34), while descriptive formulations do much better. The best prompt is P7 "Was built in the year [year]" (CLIP MAE 8.79, TAI 0.86; EVA-CLIP MAE 7.44, TAI 0.89).

  • Accuracy is highly class-dependent. For EVA02-CLIP-L-14-336, Cars (MAE 2.64, TAI 0.95) and Mobile Phones (MAE 2.65, TAI 0.95) are predicted accurately, Aircraft (MAE 14.25, TAI 0.76) and Ships (MAE 18.12, TAI 0.76) moderately, while Music Instruments (MAE 32.94, TAI 0.24) and Weapons & Ammunition (MAE 33.50, TAI 0.49) are poorly positioned. The authors attribute this to the era of photographic documentation and to differing annual ranges per class (Aircraft [1893,2017], Cars [1888,2024], Mobile Phones [1984,2024], Music Instruments [1715,2009], Ships [1744,1999], Weapons & Ammunition [1939,2003]). Class-specific scores do not average directly to the overall table because of differing class cardinalities.

  • Time probing's similarity scores are reliable. Analyzing EVA02-CLIP-L-14-336, the highest-scoring year is most often the correct one, followed by the second-highest, and top-scoring years align well with ground truth, with some outliers for recent years.

  • Temporal information forms a non-linear, low-dimensional manifold. 3D KPCA projections of time embeddings from ViT-B/32 (CLIP) for years 1700 to 2024 arrange into a manifold-like structure with a chronological color progression (violet for 1700 to yellow for 2024), and the same behavior appears for EVA-CLIP. This contrasts with prior language-model findings of linear temporal subspaces.

  • Chronological order largely survives 1D projection. With KPCA, Spearman's rho is 0.96 (CLIP) and 0.92 (EVA-CLIP), Kendall's tau 0.84 and 0.77, and the modified normalized Damerau-Levenshtein distance 0.84 and 0.77. With default UMAP parameters, results are weaker and less stable (CLIP rho 0.80, tau 0.55, delta 0.55; EVA-CLIP rho -0.70, tau -0.48, delta 0.74), motivating hyperparameter optimization.

  • Timelines outperform or match prompting efficiently. The abstract reports that the timeline approaches achieve competitive to superior accuracy compared to the prompt-based baseline while being computationally efficient. Detailed section 6.3 result tables are not included in the provided content.

Methodology in Plain English

The authors build a test set by scraping Wikipedia's category system, which organizes objects by introduction year (for example, "Cars introduced in 2022"). They collect images for six object classes, verify them manually, and keep year-level labels, giving 10,091 images (Cars 4,393; Mobile Phones 4,337; Ships 841; Instruments 436; Aircraft 69; Weapons 15).

They then run three experiments. First, time probing: they write sentences like "Was built in the year 1957" for every year from 1700 to 2024 (325 candidate years), encode each sentence with the model's text encoder, encode a query image with the image encoder, and pick the year whose text embedding has the highest dot-product similarity to the image embedding. This is the baseline, and it treats each year independently.

Second, embedding space analysis: they take the 325 time embeddings and use Kernel PCA (with a cosine kernel matching the contrastive training metric) and UMAP to project them into 1D, 2D, and 3D spaces, then check whether the resulting order matches the true chronological order using ranking metrics. They also project image embeddings into the same low-dimensional space to see whether images land near their correct years.

Third, timeline modeling: instead of comparing against 325 separate prompts, they build a single sequential representation of time. One version learns a UMAP projection whose hyperparameters (neighborhood size, minimum distance, metric) are tuned by the Tree-structured Parzen Estimator to maximize Spearman correlation with the true year order; image embeddings are then projected to the same 1D line and matched to the nearest year. The other version fits a Bézier curve through the time embeddings using the de Casteljau algorithm with 200 control points sampled uniformly and 1,000 curve samples, then maps images onto the curve either in the full embedding space or in a KPCA-reduced subspace. Predictions use either nearest-neighbor matching or interpolation between the two nearest time points, which yields fractional years.

Evaluation uses ranking metrics (Spearman's rho, Kendall's tau, and a modified normalized Damerau-Levenshtein distance based on adjacent swaps, all ranging from -1 to 1) plus accuracy metrics: Mean Absolute Error and a new Time-Adaptive accuracy measure (TAI) that grants more tolerance for older objects and less for modern ones, with thresholds set to 20 years tolerance and 50 years intolerance at the earliest year, and 5 and 15 years at the latest year, interpolated linearly. All models are used zero-shot with frozen backbones.

Why This Matters

The work shows that temporal structure is genuinely present in contrastive VLM embeddings and can be extracted without any fine-tuning, which reframes time estimation from a supervised dating task into a representation-analysis task. It also establishes a measurable benchmark and a standardized set of metrics where none existed before, since time annotations are scarce and inconsistent across existing datasets.

Real-world applications:

  • Digital archives and libraries: assigning approximate dates to undated historical photographs and cataloguing large collections where metadata is missing or incomplete.
  • Cultural heritage and historical research: constraining when an artifact or object type could have appeared in a given image, supporting provenance and dating arguments.
  • E-commerce and secondhand marketplaces: estimating the era of a product from a photo, useful for vintage goods, used vehicles, or collectibles.
  • Content moderation and misinformation detection: flagging images whose claimed date is implausible given the objects they depict.

Industry relevance centers on the efficiency argument: maintaining and querying 325 prompt embeddings per image is expensive, whereas a compact timeline representation could be embedded directly into retrieval pipelines or multimodal search systems built on top of CLIP-style backbones, which the paper notes serve as the foundation for many recent generative and multimodal models.

Future Directions

  • Improve and stabilize the timeline representations. The default-parameter UMAP 1D projection produced a negative Spearman correlation for EVA-CLIP (-0.70), showing that the derivation method is sensitive and needs more robust optimization or better geometric assumptions.
  • Extend beyond the "time of first appearance." The current setup targets when an object was introduced, not when a specific image was taken, so dating photographs of long-established objects remains open.
  • Broaden the temporal and class coverage. TIME10k is skewed toward recent decades with pre-1830s images mostly being drawings, and the class distribution is highly uneven (Cars 4,393 versus Weapons 15), so a more balanced dataset could test whether the manifold finding generalizes.
  • Reconcile modality differences. Prior work found linear temporal subspaces in language models while this work finds non-linear manifolds in VLMs; understanding why text-only and multimodal models organize time differently is an open theoretical question.

Target Audience

Researchers in computer vision and multimodal representation learning interested in interpretability and in the geometry of VLM embedding spaces; practitioners who need to date or temporally categorize image collections without labeled training data; and digital humanities, archival, and cultural heritage specialists who work with historical image sets and want to know how far off-the-shelf vision-language models can be trusted on temporal questions.

Authors’ abstract

Large-scale vision-language models (VLMs) such as CLIP have gained popularity for their generalizable and expressive multimodal representations. By leveraging large-scale training data with diverse textual metadata, VLMs acquire open-vocabulary capabilities, solving tasks beyond their training scope. This paper investigates the temporal awareness of VLMs, assessing their ability to position visual content in time. We introduce TIME10k, a benchmark dataset of over 10,000 images with temporal ground truth, and evaluate the time-awareness of 37 VLMs by a novel methodology. Our investigation reveals that temporal information is structured along a low-dimensional, non-linear manifold in the VLM embedding space. Based on this insight, we propose methods to derive an explicit ``timeline'' representation from the embedding space. These representations model time and its chronological progression and thereby facilitate temporal reasoning tasks. Our timeline approaches achieve competitive to superior accuracy compared to a prompt-based baseline while being computationally efficient. All code and data are available at https://tekayanidham.github.io/timeline-page/.

Read the original paper