Skip to content
AI.info

Research

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth Overview Research area: Language model evaluation and post-training, using an executable simulation environment (cu

arXiv
2608.20574
Published
2026-08-20
Authors
Josef Chen, Erim Hayretci

AI summary

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

Overview

Research area: Language model evaluation and post-training, using an executable simulation environment (culinary reasoning) as a deterministic reward source instead of a model judge or human preference panel.

Technical level: Advanced. The paper combines benchmark construction, frontier-model API orchestration, cluster bootstrap inference, Holm-corrected paired testing, and LoRA supervised fine-tuning.

Scope: The paper introduces FlavourBench, a benchmark that compiles dense 0–100 reward maps for three-of-eight ingredient-selection tasks from a versioned culinary runtime called Epicure, uses those maps to rank 27 frontier endpoints on 534 shared tasks, and then uses the same maps as supervision for a controlled fine-tuning study.

What This Paper Is About

Open-ended evaluation of language models usually substitutes a model judge, a small human preference panel, or exact-match factual questions for a missing answer key. That entangles the evaluator with the systems being tested and discards useful differences between plausible answers. FlavourBench instead compiles a dense answer map from a versioned culinary environment: each task asks for three ingredients out of eight, Epicure scores all 56 legal portfolios before any model runs, and a response is graded by table lookup with no post-hoc interpretation.

Key Contributions

  1. An executable culinary benchmark with 29,904 model-independent portfolio scores and a single interpretable 0–100 metric.
  2. A 27-model, 534-task complete-core evaluation with simultaneous uncertainty and paired, multiplicity-controlled tests (351 pairwise hypotheses, 101 resolved after Holm correction).
  3. Post-collection tests of task-filter, metric, family-weight, and reward-map dependence, including a rescore against three public Epicure checkpoints (Cooc, Core, Chem) and a held-out, human-observed substitution check against Recipe1MSubs.
  4. A preregistered three-seed study showing that Epicure-optimal SFT transfers to unseen reward maps beyond format learning, backed by a public-map replication and a content-addressed release.

Main Findings

  • No unique best endpoint is identified. Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% interval [61.0, 69.2]), but 101 of 351 prespecified paired contrasts remain significant after Holm correction, and corrected evidence does not identify a unique best model.

  • The leaderboard spans a wide range. Grok 4.6 (65.1), Gemini 3.1 Pro (65.0), GPT-5.6 Sol Pro (64.2), Muse Spark 1.2 (63.8), and GPT-5.6 Terra Pro (63.7) sit in the top group, while Command R+ ranks last at 47.9 (rank 95% CI [27, 27], group 3).

  • Rankings replicate across independently compiled panels. Panel-1 versus panel-2 model-level Pearson correlation is 0.89 and rank correlation is 0.80, computed on disjoint task IDs with the same frozen metric and 89 tasks per family.

  • 534 tasks are more than enough for a stable global score, but not for naming a winner. The crossed-design generalizability coefficient is 0.936 at 534 tasks, and the same variance model estimates 329 balanced tasks for 0.90. In 5,000 stratified half-size subsamples (270 tasks), the median rank correlation with the complete point order is 0.952 (empirical 95% range 0.897–0.980), median top-five overlap is 80%, and the complete point leader is retained in only 46.1% of subsets.

  • The task filter does not determine the order. The endpoint whose candidate-panel validity most constrained the common-core filter was Claude Fable 5; omitting it retains 71.0% of the official tasks, scores for the remaining 26 endpoints differ by 0.65 points on average and at most 1.37 points, rank correlation is 0.955, 92.9% of pair directions agree, and the point leader is unchanged. Other omissions alter at most three tasks per stratum, and no comparison among 42 numerical or category diagnostics resolves after Holm correction.

  • Alternative score summaries recover the same broad order. Rank correlations range from 0.938 to 0.985 with pair-order agreement of at least 89.2% across chance-adjusted gain, action percentile, and exact-optimum rate. Across 696 moderate family-weight combinations, rank correlation with equal weighting never falls below 0.980.

  • The order survives replacement of the reward map. Scoring the fixed 534 tasks and fixed 14,418 model selections with Epicure-Cooc, Epicure-Core, and Epicure-Chem yields median task-level rank correlations of 0.660 to 0.752 and exact-optimum agreement on 27.7% to 38.4% of tasks, yet model-rank correlation ranges from 0.903 to 0.957 with 86.9% to 91.7% of model-pair directions agreeing. Grok 4.6 has the largest point estimate under all three.

  • Public checkpoints recover human-observed substitutions above chance. Exact matching maps 3,282 of 10,747 Recipe1MSubs test events to the public 1,790-ingredient vocabulary, yielding 1,469 unique directed pairs over 357 source ingredients. Cooc, Core, and Chem place the observed target at equal-source within-food-group rank percentiles 0.806, 0.800, and 0.780, all rejecting the 0.5 chance null after Holm correction (p_Holm < 10^-4). On 594 directed pairs absent from the Recipe1MSubs training split, the percentiles remain 0.754, 0.735, and 0.718. Full-vocabulary Hit@10 is 0.133–0.172 against an analytic random baseline of 0.0056.

  • Reward supervision transfers to unseen maps. LoRA SFT of a pinned Qwen3-0.6B checkpoint on 270 Epicure-optimal answers improves its score on 84 anchor-disjoint maps by 13.30 points over a format- and label-matched control (95% CI 6.52–20.29, p = 1.70 × 10^-4), and replicates on all 534 public maps at +11.73 points (95% CI 8.98–14.54, p = 1.00 × 10^-5).

  • Format training alone is not the explanation. Both trained arms parse 100.00% of responses while the base parses 89.29%. Format training changes the base score by only 0.90 points (95% CI −4.91–6.85, p = 0.790). On the public maps the control still improves on the base by 3.05 points (95% CI 0.46–5.68, p = 0.021), consistent with removing parse failures, and treatment-minus-control differences are positive in all six evaluation-by-family cells.

  • Aggregate scores conceal differing profiles. Family-level results show similar aggregate scores arising from different strengths; constraint tasks penalize infeasible portfolios directly while pairing and substitution reward graded semantic structure.

  • Continuous scoring separates plausible answers. In the released case studies, two non-optimal portfolios for the same task differ sharply: for a boursin cheese substitution, "caciocavallo, fromage blanc, grana padano" scores 100 while "fromage blanc, quail egg, goose egg" scores 15.

Methodology in Plain English

The researchers built a benchmark where the answer key consists of a table rather than an opinion. A culinary runtime called Epicure represents 1,790 ingredients in a 300-dimensional space and exposes deterministic operations for substitution, pairing, dietary feasibility, and regional composition. A task compiler turns those operations into selection problems: eight candidate ingredients, choose three. Because there are only 56 possible three-item subsets, the compiler can score every legal portfolio in advance and freeze a continuous 0–100 map, normalized within each task so that the best portfolio gets 100 and every task has a unique optimum.

The 534-task core is split evenly across three decision families (substitution, pairing, and constraints), with 178 tasks each, drawn as 89 tasks per family from each of two independently compiled panels and spanning 534 unique anchor ingredients. Candidate sets are selected across validation strata and labeled by a task-derived hash; no evaluated model writes an item, distractor, or score. Prompts end with a FINAL_SELECTION marker followed by three distinct A–H labels; the parser takes the final marker line, sorts the labels as an unordered set, and performs one lookup.

Twenty-seven frontier endpoints from sixteen organizations faced the same tasks, with routes chosen before each scored block and automatic fallback disabled. Complete-core eligibility was decided score-blind: a task entered the shared core only if all 27 models completed it and produced a parseable three-item set, with selection under a fixed SHA-256 ordering that does not load ingredients or scores until task IDs are fixed. Inference used 50,000 bootstrap samples over ingredient anchors (tasks sharing an anchor move together), simultaneous max-t intervals, and two-sided anchor-cluster sign-flip tests with 100,000 Monte Carlo draws and Holm correction at familywise alpha = .05.

For the training study, the team pinned Qwen3-0.6B at a single revision, used 270 training and 72 validation tasks and 84 primary-transfer tasks with disjoint anchors, candidate sets, and prompt hashes, and ran three seeds of three epochs of all-linear LoRA SFT (rank 16, effective batch size 16, learning rate 10^-4, completion-only loss, final checkpoint without validation-based selection). To isolate reward learning from syntax practice, the control rotated the same A–H portfolios onto different prompts within each family and source panel, so treatment and control have identical prompts, row counts, completion lengths, and portfolio-label histograms, but the control contains no accidental task optimum.

Why This Matters

Impact on research. The paper offers a template for evaluation where the answer key is an executable, content-addressed program rather than another model. It pairs that with inference that resolves only 101 of 351 paired model contrasts after correction, showing how many leaderboard orderings would not survive proper multiplicity control. It also demonstrates that a benchmark's dense reward surface can double as a supervision signal, with a preregistered three-seed transfer test rather than a post-hoc ablation.

Real-world applications:

  • Ingredient substitution and pairing assistance for recipe platforms, meal-planning apps, and grocery recommendation systems, where graded alternatives matter more than one correct answer.
  • Dietary-constraint filtering, since the constraint family scores portfolios against diet and maximum NOVA processing level and assigns zero to infeasible sets.
  • Reward-model and post-training research, where a programmatic environment can supply exact 0–100 targets at scale instead of crowdsourced pairwise preferences.
  • Evaluation infrastructure for any domain with a deterministic simulator, such as database state, software tests, or scored game outcomes.

Industry relevance. The release contains prompts, reward maps, raw responses, routes, training and evaluation manifests, code, and offline verifiers, all content-addressed and reconstructable without provider access. That matters to teams that need auditable model comparisons under frozen routes and pinned checkpoints, and to anyone considering whether a domain-specific reward function can be distilled into a small model.

Future Directions

  • Independent training replication. The public-map result reuses the same adapters, so it is an independent task replication rather than an independent training replication; a fresh training run is not reported.
  • Larger and more varied training. The transfer study covers only LoRA SFT of one 0.6B checkpoint over three seeds, so the paper does not compare training algorithms or model scales; the authors note the same maps expose SFT, preference, and scalar-reward interfaces for larger follow-up studies, but only SFT is tested here.
  • Recovering or replacing the primary runtime. The original training run, seed, and source revision for the primary Epicure runtime were not recovered, and the artifact is not identified as the published Cooc, Core, or Chem checkpoint. A newly compiled public-checkpoint benchmark that also selects its own candidate sets is not reported.
  • External validation beyond Recipe1MSubs. Recipe1MSubs supplies held-out human-observed substitution labels but shares Recipe1M ancestry with part of Epicure's corpus, and it does not test the pairing or constraint components or validate the unrecovered primary runtime. No human sensory or cooking-outcome validation is reported.

Target Audience

Researchers and engineers working on language model evaluation, benchmark design, and reward-based post-training, particularly those interested in executable environments, partial-credit scoring, cluster-based uncertainty estimation, and multiplicity control. The paper is also relevant to applied teams in food technology, recipe recommendation, or any domain where a simulator can serve as a scoring oracle. Readers without a statistics or fine-tuning background will find the leaderboard and case studies accessible, but the inference and training sections assume familiarity with bootstrap resampling, Holm correction, and LoRA adaptation.

Authors’ abstract

Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.

Read the original paper