Research
WarpRec: Unifying Academic Rigor and Industrial Scale for Responsible, Reproducible, and Efficient Recommendation
Overview Research area: Recommender systems infrastructure — evaluation frameworks, scalability, reproducibility, green AI, and agentic AI integration. Technical level: Intermediate. No heavy mathemat

- arXiv
- 2602.17442
- Published
- 2026-02-19
- Authors
- Marco Avolio, Potito Aghilar, Sabino Roccotelli, Vito Walter Anelli, Chiara Mallamaci, Vincenzo Paparella, Marco Valentini, Alejandro Bellogín, Michelantonio Trizio, Joseph Trotta, Antonio Ferrara, Tommaso Di Noia
AI summary
Overview
Research area: Recommender systems infrastructure — evaluation frameworks, scalability, reproducibility, green AI, and agentic AI integration.
Technical level: Intermediate. No heavy mathematics is required, but familiarity with recommender-system evaluation (splitting, metrics, statistical testing) and with MLOps concepts (distributed training, hyperparameter optimization) helps.
Scope: A systems paper introducing WarpRec, an open-source recommendation framework intended to let the same codebase run from a laptop to a distributed cluster while enforcing reproducibility, energy tracking, and agent interoperability.
What This Paper Is About
Recommender systems research is split between academic libraries that are easy to use but cannot scale past a single machine, and industrial engines that scale but lack the flexible evaluation and statistical rigor that science requires — a divide the authors call the "Deployment Chasm." WarpRec is presented as a single framework that removes this trade-off through a backend-agnostic architecture, so a model defined once can move from local prototyping to distributed training on Ray clusters without rewriting. The paper's goal is to demonstrate that scalability need not be bought at the cost of scientific integrity or sustainability.
Key Contributions
-
A backend-agnostic, modular architecture. WarpRec is built on Narwhals as a compatibility layer and composed of five decoupled engines (Reader Module, Data Engine, Recommendation Engine, Evaluation, Writer Module), plus an Application Layer, enabling a "write-once, run-anywhere" paradigm from local execution to Ray-based distributed training.
-
A large built-in catalog. The framework ships with 55 algorithms spanning 6 model classes (the abstract states "50+"), 40 metrics organized into accuracy, rating, coverage, novelty, diversity, bias, and fairness families, and 19 filtering and splitting strategies (13 filtering, 6 splitting).
-
Green AI as a first-class feature. WarpRec integrates CodeCarbon for real-time energy and carbon tracking, which the authors claim makes it the first recommendation framework to enforce ecological accountability, and it is described as the only framework in their comparison table to natively support multi-objective metrics.
-
Rigor and agent readiness by default. The framework automates significance testing (paired and independent tests) with Bonferroni and FDR corrections for multiple comparisons, and natively implements a Model Context Protocol server alongside a REST API, turning the recommender into a callable tool for LLM agents.
Main Findings
-
Small scale is unproblematic for everyone. On MovieLens-1M, every evaluated framework completed the full pipeline. WarpRec's serial end-to-end totals were 30.85s for EASER, 2m 2s for NeuMF, 2m 20s for LightGCN, and 11m 37s for SASRec, against competitors such as RecBole (1m 39s EASER, 15m 25s NeuMF, 5m 13s LightGCN, 44m 53s SASRec) and Microsoft Recommenders (35m 51s NeuMF, 7m 54s LightGCN).
-
Scale separates the frameworks. As data grew to MovieLens-32M and NetflixPrize-100M, several established frameworks failed to complete the pipeline, either from RAM or VRAM exhaustion or from exceeding time limits, with the notation (x/6) recording how many hyperparameter configurations finished before timeout. WarpRec consistently executed the end-to-end workflow across all three benchmarks.
-
LightGCN is the most computationally expensive model. Across frameworks, it was the heaviest workload; on NetflixPrize-100M only RecBole and WarpRec managed to process it. The authors attribute RecBole's low nDCG@10 there (0.0247) to failing to converge within the time limit, versus 0.1779 for WarpRec.
-
EASER is where WarpRec claims the widest advantage. Because EASER requires inverting a dense Gram matrix, the authors report that WarpRec outperforms all competitors across all dataset scales for this model, and that for memory reasons EASER was only considered under serial execution. On NetflixPrize-100M, WarpRec's total time was 2h 26m versus 5h 49s for RecBole and 3h 22m for Cornac.
-
Energy is driven by training duration more than peak power. On NetflixPrize-100M, SASRec recorded the highest peak GPU power (278.6 W) but a moderate total footprint (0.32 kWh); LightGCN drew less average power (157.4 W) yet consumed 2.42 kWh and produced the highest emissions (0.0095 kg CO2eq). EASER consumed roughly 95% less energy (0.115 kWh) than the deep graph-based baselines, which the authors present as evidence that strong accuracy can come at low environmental cost.
-
Full Green AI profile reported. Table 5 lists, per model, emissions in kg CO2eq (ItemKNN 0.0002, EASER 0.0005, NeuMF 0.0004, LightGCN 0.0095, SASRec 0.0012), emissions rate in kg CO2eq/h, CPU and GPU power in watts, CPU/GPU/RAM energy in kWh, total energy consumed in kWh, and peak RAM usage in GB.
-
Agentic inference works end to end. A SASRec model trained on MovieLens-32M, served through the MCP interface, produced recommendations for a three-movie viewing history and achieved an nDCG@10 of 0.9257 with 100 negative samples; because the model is sequential it needs no pre-trained user embedding, and the agent added a natural-language explanation on top of the raw list.
-
Accuracy and speed are not perfectly aligned in the reported tables. For EASER on NetflixPrize-100M, the table lists WarpRec's nDCG@10 as 0.3035, below RecBole's 0.3283 and Cornac's 0.3143, even though WarpRec claims the efficiency advantage for that model; on MovieLens-32M LightGCN, RecBole's faster run yielded 0.0237 against WarpRec's 0.2061.
Methodology in Plain English
The authors built a framework and then benchmarked it head-to-head against existing tools.
- Comparison set. Five frameworks that provide ready-to-use experiment pipelines: Cornac, DaisyRec, Elliot, and RecBole for academic benchmarking, plus Microsoft Recommenders for enterprise-scale deployment. NVIDIA Merlin was excluded from the runtime comparison because it requires error-prone custom implementations rather than an out-of-the-box pipeline.
- Models tested. Five algorithms spanning different learning paradigms: EASER, NeuMF, LightGCN, SASRec, and ItemKNN, the last included as a reference for environmental impact comparisons.
- Data. Three datasets of increasing size (MovieLens-1M, MovieLens-32M, NetflixPrize-100M), with statistics reported in Table 2. The authors state they deliberately limited the dataset scale so that competitor frameworks could run at all, noting WarpRec's capacity extends further.
- Protocol. A 90-10 random holdout for EASER, LightGCN, NeuMF, and ItemKNN, and a temporal holdout for SASRec. Models were tuned with exhaustive grid search over 6 configurations per model drawn from established literature, trained for 10 epochs on MovieLens and 2 on NetflixPrize-100M, selected by validation nDCG@10, with a batch size of 8,192 and a 24-hour timeout per trial.
- Measurement. Time is decomposed into preprocessing (ingestion and splitting), training, evaluation (average per-trial optimization and validation), HPO wall-clock, and total. Serial versus parallel (concurrent via Ray) execution is compared, with resource allocations by dataset and mode detailed in Table 3, using 16-core CPUs and 64GB NVIDIA A100 GPUs. Final quality is reported as nDCG@10 under full ranking evaluation so that speed
Authors’ abstract
Innovation in Recommender Systems is currently impeded by a fractured ecosystem, where researchers must choose between the ease of in-memory experimentation and the costly, complex rewriting required for distributed industrial engines. To bridge this gap, we present WarpRec, a high-performance framework that eliminates this trade-off through a novel, backend-agnostic architecture. It includes 50+ state-of-the-art algorithms, 40 metrics, and 19 filtering and splitting strategies that seamlessly transition from local execution to distributed training and optimization. The framework enforces ecological responsibility by integrating CodeCarbon for real-time energy tracking, showing that scalability need not come at the cost of scientific integrity or sustainability. Furthermore, WarpRec anticipates the shift toward Agentic AI, leading Recommender Systems to evolve from static ranking engines into interactive tools within the Generative AI ecosystem. In summary, WarpRec not only bridges the gap between academia and industry but also can serve as the architectural backbone for the next generation of sustainable, agent-ready Recommender Systems. Code is available at https://github.com/sisinflab/warprec/