Skip to content
AI.info

The Pulse

Tether Releases 191 Billion-Token STEM Dataset for Smaller Models

Tether AI Research has released QVAC Genesis III, a 191.43-billion-token synthetic dataset for training smaller models on STEM reasoning. The company says models trained on the corpus outperformed Cosmopedia-v2 and other open baselines acro

Tether Releases 191 Billion-Token STEM Dataset for Smaller Models

AI.info Team ·

Tether Puts Its Bet on Better Data, Not Bigger Models

Tether AI Research released QVAC Genesis III on September 23, presenting a 191.43-billion-token synthetic dataset as an alternative to the industry’s continued push toward larger models and heavier computing systems.

The company says Genesis III helps smaller models answer STEM questions while also explaining the reasoning behind those answers. Its stated target is local AI: tutors, technical assistants and research tools that can run on laptops, phones and local servers instead of sending every request to a cloud model.

“Most of the AI industry has focused on making models bigger and giving them more computing power,” Paolo Ardoino, CEO of Tether, said in the announcement. “Genesis III shows what can happen when you focus instead on making the data smaller models learn from better.”

The release is accompanied by a research paper submitted to arXiv on September 17. Tether says the paper has also been accepted for presentation at the 2026 Conference on Language Modeling.

What Genesis III Adds to Earlier QVAC Releases

Genesis III expands on QVAC Genesis I and Genesis II, released in October and December 2025. The new corpus contains 159.6 million documents covering 19 STEM areas, including biology, chemistry, physics, mathematics, computer science, medicine, astronomy, electrical engineering, statistics and machine learning.

The dataset is available through QVAC’s open Genesis platform and is offered under a CC-BY-NC 4.0 license for research and education. QVAC’s page provides a one-line loading example using the Hugging Face datasets library, allowing researchers to stream or download the material for training runs.

The project is designed around two forms of synthetic reasoning data. Failure Analysis takes mistakes made by a smaller student model and turns them into corrective explanations. A stronger teacher model identifies the error, explains the misconception and works through a solution.

Option-Level Reasoning uses questions the student model answered correctly. It generates explanations for why the selected answer is right and why the other answer choices are wrong. The approach is intended to give models more information about errors and alternatives than a conventional question-and-answer pair would provide.

The Reported Gains Against Cosmopedia-v2

The paper evaluates 1.7-billion-parameter models trained from scratch and compares Genesis III with Cosmopedia-v2, an open synthetic training corpus. Tether says the Option-Level dataset produced gains across several model architectures, including Qwen, Llama, SmolLM and Gemma.

Against a token-matched Cosmopedia-v2 model, the Genesis III results improved scores by 28.57 percentage points on ARC-Easy, 21.35 points on ARC-Challenge and 15.03 points on the MMLU STEM benchmark. The paper also reports that a Genesis III model outperformed the publicly released Cosmo-1B model, which is larger and trained with additional mathematics and code data.

One 1.7-billion-parameter model trained with the Option-Level data achieved a Valid Answer Rate of 99.45% in the study’s evaluation. The paper describes that measure as an answer-parsing protocol that tracks whether a model produces a valid final answer, rather than relying only on exact benchmark accuracy.

The reported results come from Tether’s paper and announcement. They show benchmark performance under the study’s stated training and evaluation setup, not a general guarantee that every small model trained on Genesis III will match those results.

Why Tether Wants Reasoning on the Device

Tether’s QVAC initiative is built around running AI locally, with inference taking place on hardware controlled by the user. Genesis III fits that strategy by focusing on the training data used by smaller models rather than only increasing parameter counts.

Local execution could reduce the need to send sensitive questions to an external service and may help tools operate where connectivity is limited. Tether points to education as one application: a local tutor could work through a physics problem, explain a wrong mathematical step and identify where a student’s reasoning failed.

The same design could support specialized assistants for researchers and professionals. Tether describes those systems as tools that could run locally without continuously transmitting technical questions or other sensitive information to outside servers.

Open Access, With a Research-Only License

The dataset’s availability gives independent researchers access to a large synthetic STEM corpus that is explicitly organized around difficulty levels, educational styles and reasoning traces. QVAC describes Genesis III as its largest release to date and lists Genesis I at 41 billion tokens and Genesis II at 107 billion tokens.

The noncommercial license limits how companies can use the release, even as it allows research and education. That distinction matters for developers who may want to build commercial tutors or technical assistants from the corpus: the dataset is open to inspect and use for research, but the license does not provide unrestricted commercial rights.

For now, Genesis III’s strongest claim is specific rather than universal. Under the experiments reported by its authors, carefully generated explanations and contrastive reasoning help 1.7-billion-parameter models compete with larger or differently trained systems on selected STEM tests. The dataset is available through QVAC’s Genesis page and its linked Hugging Face release.

Source

Tether

Explore

More articles