Skip to content
AI.info

The Pulse

Redwood Finds GPT-6 Astra Gains From Filler Tokens

Redwood Research reports that GPT-6 Astra answers some reasoning tasks more accurately when its prompts contain meaningless filler tokens. The researchers warn that the gains raise questions about evaluating reasoning that does not appear i

Redwood Finds GPT-6 Astra Gains From Filler Tokens

AI.info Team ·

“I think your results are likely partially tracking Astra being underelicited.”

Neel Nanda, in a comment on Redwood Research’s post

Redwood Research’s September 23 tests found that GPT-6 Astra often solves harder problems more accurately when its prompt includes tokens with no relevant information, such as dots, counting sequences or repetitions of the question. The model was told to answer immediately without reasoning. On four-hop natural-facts questions, accuracy rose from roughly 10–20% without filler to about 50% with it.

The result points to a gap between what a model can work out and what it shows in its written reasoning. Redwood’s researchers—Dylan Xu, Sebastian Prasanna and Alek Westover—say the gains complicate efforts to assess models by monitoring their chain of thought. They do not establish that Astra is concealing a harmful plan; the tests measure answer accuracy under prompts designed to suppress visible reasoning.

Dots and repeated questions improve scores

The team tested GPT-6 Astra on multi-hop fact questions, generated arithmetic and math competition problems, as well as Humanity’s Last Exam and LiveBench. In the main experiments, filler usually took the form of dots appended after the question. Other tests used counting sequences or repeated the question, and the authors say the methods produced broadly similar results.

For the natural-facts task, each question requires chaining information across several steps. Astra’s accuracy climbed as the researchers increased the amount of filler, while other tested models generally gained less. Redwood also found that filler improved Astra on some broader tests, including older AIME problems, where accuracy rose from roughly 60–70% to about 90%.

The test blocks visible reasoning, not computation

Redwood instructed Astra not to reason and required a one-line answer. Because the API did not offer a “no reasoning” setting for Astra, the researchers used low reasoning effort and checked the API’s reported reasoning-token counts; they say all included outputs reported zero. That protocol supports a narrower conclusion than the phrase “hidden reasoning” might suggest: filler improved performance without a visible explanation, but the experiment does not show exactly how the model used those tokens.

The authors connect the result to chain-of-thought monitoring, a safety practice that examines a model’s written reasoning for signs of unwanted behavior. If a model can make meaningful progress without putting that process into its explanation, a clean-looking answer may reveal less about how it reached its conclusion. Redwood argues that evaluations intended to measure performance without reasoning should test prompts with filler, rather than treating a no-filler score as the model’s full capability.

Prompt format remains a live caveat

Nanda’s comment on the post raises a competing explanation: Astra may be “underelicited,” meaning that a prompt does not draw out its existing ability. He said a clearer format and ten examples erased most of the filler gains in his own multi-hop test. In the numbers he shared, a particular prompt format with ten examples rose from 32.5 to 37 with dots, a gain he described as not statistically significant.

Redwood’s own appendix reports that ten-shot prompting made Astra’s filler improvement milder, though the authors say Astra still gained more than the comparison models. The results therefore show a strong sensitivity to filler in the tested setups, not a settled account of why it happens or how broadly it applies. On Redwood’s four-hop test, the reported change is concrete: accuracy moved from roughly 10–20% without filler to about 50% with it.

Source

Explore

More articles