Skip to content
AI.info

The Pulse

Tempus Targets 100,000 Whole Genomes for AI Research

Tempus AI is building a de-identified dataset of 100,000 disease-specific whole genomes linked to longitudinal clinical outcomes. The company plans to make the platform broadly available by mid-2027 and eventually expand it to one million g

Tempus Targets 100,000 Whole Genomes for AI Research

AI.info Team ·

Tempus AI is building a whole-genome dataset that will initially connect 100,000 disease-specific genomes with longitudinal clinical outcomes, then expand toward one million genomes. The Chicago-based company announced the initiative on September 11, describing it as a research platform designed specifically for artificial intelligence rather than a genome archive assembled mainly for population studies.

The distinction matters because a genome rarely explains a patient’s disease or treatment response on its own. Tempus says researchers will be able to examine whole-genome sequencing alongside clinical histories, imaging, pathology and outcomes inside its existing de-identified data environment. The initial dataset is already available to selected participants through an Early Adopter Program, while general availability is planned for mid-2027.

Tempus’ announcement places the company in a race to build biomedical datasets that combine molecular measurements with what happens to patients over time. Its stated objective is not simply to increase the number of genomes, but to create a machine-learning resource that can support pre-training, post-training, model validation and research workflows without requiring customers to move the underlying data between systems.

Tempus said the platform will focus on disease populations and patient outcomes, a structure intended to help researchers investigate disease progression, treatment response and the biological factors associated with both.

Tempus starts with 100,000 genomes and a one-million target

The first phase calls for 100,000 whole-genome sequences linked to longitudinal clinical information over the next several years. Tempus did not provide a detailed timetable for completing that first phase, disclose the number of genomes already assembled, or publish a disease-by-disease breakdown of the planned cohort.

The company did say development is already underway. Researchers and model developers can access the initial dataset through an Early Adopter Program, with additional members expected to join in waves as the resource grows. Tempus plans general availability for mid-2027, a date that gives potential users a clearer point at which the project is expected to move beyond its initial group of participants.

After reaching 100,000 genomes, Tempus plans to pursue a much larger goal of one million. That target is a company objective rather than an existing dataset size, and the announcement does not establish that the full one-million-genome resource will be completed on a fixed schedule.

Tempus also does not describe the project as an open public database. The release frames it as a research platform available through the company’s environment, with access beginning through the Early Adopter Program. Commercial terms, eligibility requirements and the precise conditions for use were not included in the announcement.

Why whole genomes change the data problem

Many clinical genomics programs rely on targeted panels or other tests that examine selected genes and regions. Those approaches can answer defined clinical questions efficiently, but they do not capture every variant across the genome. Tempus says its new effort expands beyond targeted gene panels to give researchers a deeper view of how genomic variation may relate to disease progression and treatment response.

Whole-genome sequencing produces a broad molecular record, but breadth alone does not guarantee useful research. A genome becomes more informative when researchers can connect it to diagnoses, therapies, laboratory results, imaging, pathology and the patient’s subsequent course. Tempus is presenting that linkage as the central feature of its project.

Existing population-scale genome programs, according to Tempus, are generally drawn from broad populations and are not designed around longitudinal disease and treatment outcomes. A disease-centered dataset can support different questions: whether certain variants appear more often in patients with a particular condition, whether genetic patterns correlate with treatment benefit or resistance, and whether genomic signals improve predictions made from clinical data alone.

Those questions require careful study design. Disease populations can introduce selection effects, and outcomes recorded in routine care may vary in completeness across health systems and patient groups. Tempus’ announcement describes the intended structure of the resource but does not present validation results showing how representative the cohort will be or how consistently its clinical outcomes are captured.

The platform joins DNA with notes, scans and outcomes

The planned whole-genome resource will sit inside Tempus’ existing de-identified multimodal data environment. The company says researchers will be able to analyze genomic information alongside clinical histories, imaging, pathology and patient outcomes through Tempus Lens, its research and analytics platform.

That arrangement is designed to reduce a common technical burden in biomedical machine learning: assembling data from separate repositories, aligning identifiers, standardizing formats and preserving the timing of events. Tempus says its system will allow researchers to build and validate AI models without moving datasets between systems. The release also describes a computational setup intended for both pre-training and post-training workflows.

Eric Lefkofsky, founder and chief executive officer of Tempus, said the value of the project depends on what researchers can do with the data rather than on the genome count alone.

“A large dataset is only valuable if you can turn it into insight,” said Eric Lefkofsky, founder and CEO of Tempus.

Lefkofsky added that the company has spent the past decade connecting diagnostics, multimodal clinical data and AI infrastructure, and said that adding whole-genome information linked to longitudinal outcomes will give researchers a richer foundation for building models and generating research findings.

Tempus brings an oncology data record to the project

Tempus is not starting from a blank slate. The company says it has built one of the largest multimodal real-world oncology databases and has used that data in hundreds of drug-development decisions. Its existing resources include molecular information, clinical records, pathology and imaging, although the whole-genome initiative represents a move beyond the company’s earlier emphasis on targeted sequencing and other genomic tests.

Tempus has previously described a data operation spanning more than 500 petabytes and more than 45 million de-identified patient journeys. In a May 2026 announcement about its multimodal foundation-model work, the company said it had more than 1.5 million patients with sequenced data and more than 400,000 cancer records containing full genomic, transcriptomic, imaging and clinical information.

The same announcement described a model trained on 2.5 million longitudinal records, more than 250 million pages of clinical notes, 450,000 digitized medical images and 500,000 genomic and transcriptomic sequences. Those figures refer to Tempus’ broader multimodal AI efforts, not to the new whole-genome dataset. The distinction is important: the newly announced resource has a stated target of 100,000 whole genomes, while the company’s larger data library contains many types of records and sequencing assays.

Tempus has also been developing disease-specific foundation models. In August, the company announced results for PRISM2, a pathology model developed with Microsoft researchers. Tempus said PRISM2 was trained on 2.3 million whole-slide images and 14 million diagnostic question-and-answer pairs derived from nearly 700,000 pathology reports. That work shows the company’s broader strategy: combine large datasets with specialized models and software that can turn them into research tools.

Patient privacy and access will shape the project

A dataset linking genomes to clinical histories carries a higher privacy burden than a collection of isolated genomic sequences. Tempus describes the planned resource as de-identified, but the announcement does not specify the de-identification methods, consent framework, governance process or security controls that will apply to the whole-genome records.

Those details will matter to hospitals, biopharmaceutical companies and academic researchers deciding whether to participate. Whole genomes can contain information relevant to a patient’s relatives as well as to the individual whose sample was sequenced. Linking them to treatment histories and clinical outcomes adds research value, but it also increases the sensitivity of the combined record.

The announcement also leaves unanswered how Tempus will measure representation across ancestry groups, disease subtypes, health systems and treatment settings. A model trained on a large but uneven cohort may perform well for the populations most represented in the data and less reliably for others. Tempus has not yet released a cohort profile for the initial 100,000-genome target.

Commercial access raises a separate question. Tempus operates both a clinical diagnostics business and a data and analytics business serving life-sciences companies. The release says researchers and model developers will be able to work through Tempus Lens, but it does not say whether access will be subscription-based, project-based, partnership-based or structured under another model.

Mid-2027 is the first public checkpoint

For now, the most concrete milestone is not the one-million-genome ambition but the planned mid-2027 general availability date. Before then, Tempus must grow the dataset, complete the clinical and genomic linkage, establish access procedures and demonstrate that the resource can support reproducible research across disease populations.

The company’s announcement makes a specific claim about design: it calls the planned resource the first de-identified multimodal whole-genome sequencing dataset built around disease populations and patient outcomes and optimized for AI-driven research. That claim describes Tempus’ characterization of the project; the release does not provide an independent comparison of every existing genome database.

The next evidence will come from the Early Adopter Program. Researchers will want to see the number of genomes available at each stage, the diseases represented, the proportion with complete longitudinal outcomes, the quality of the sequencing data and the performance of models trained on the combined records. They will also need clear information about patient permissions, data governance and the limits placed on downstream model development.

Tempus has announced a 100,000-genome starting target, a long-term goal of one million genomes and a mid-2027 date for general availability. The unresolved questions are whether the company can assemble those records with consistent outcomes, how broadly the cohort represents patients outside its existing networks, and what researchers will be permitted to build with the data once access expands.

Source

Tempus AI

Explore

More articles