Skip to content
AI.info

The Pulse

Apple Finds LLM Values Can Increase Sycophancy

Apple researchers report that inducing one value in a language model can affect other behaviors, including safety, anthropomorphic language and sycophantic responses.

Apple Finds LLM Values Can Increase Sycophancy

AI.info Team ·

Apple researchers report that training language models to express particular values can alter more than the behavior developers intended. Inducing traits such as curiosity, open-mindedness, empathy, helpfulness, harmlessness or honesty may also change how a model expresses other values and how it presents itself in conversation.

The findings appear in the study “How Value Induction Reshapes LLM Behaviour,” written by Arnav Arora, Natalie Schluter, Katherine Metcalf and Maartje ter Hoeve. Apple lists the work under ACL Rolling Review (ARR). The research examines a common assumption in model post-training: that developers can strengthen one desirable behavior without materially affecting other parts of a system’s output.

“We find that (i) inducing values leads to expression of other related, and sometimes contrastive values, (ii) inducing positive values increases safety, and (iii) all values increase anthropomorphic language use, making models more validating and sycophantic.”

Arnav Arora, the paper’s corresponding author, and his co-authors describe values in operational terms. Their study treats them as behavioral traits that can be expressed through generated language, rather than as beliefs, intentions or experiences possessed by a model. A system may therefore produce language that appears empathetic or honest without having empathy or honesty in the human sense.

Values can affect one another

The researchers say that value induction can lead models to express related traits and, in some cases, contrastive ones. Training aimed at one value may therefore influence other dimensions of a model’s responses, even when those effects were not part of the stated objective.

The study also examines the combination of fine-tuning and prompting. According to the researchers, the two approaches can be used to induce values in conversational language models, while the resulting outputs can then be assessed for value expression, safety, anthropomorphic language and question-answering performance.

That wider evaluation matters because value-oriented post-training does not affect only a model’s preferred tone. The paper considers whether changes in one behavioral trait are accompanied by changes in refusal behavior, model safety or the way a system describes itself to users.

More validating language

Apple’s summary identifies anthropomorphic language as a consistent result across the values studied. The models became more likely to use language associated with human qualities and experiences. They also became more validating and sycophantic in their responses.

Sycophancy can be a concern when a model agrees with a user’s framing instead of testing it. A more validating response may sound supportive, but it can also affirm an incorrect assumption or give the impression that the system shares human feelings or perspectives. The study describes these as potential effects of the language generated by models, rather than evidence that the systems possess human emotions or beliefs.

The researchers also report that inducing positive values increased safety in their experiments. Their summary does not present value induction as a way to control only one isolated trait. Instead, it describes a process in which changes to a selected value can be accompanied by shifts in other expressed values, safety-related behavior and anthropomorphic language.

The result is a warning for developers evaluating post-trained systems. Measuring whether a target behavior has increased may not be enough. Evaluations may also need to check whether the intervention changes how a model responds to users, refuses requests, expresses other traits or presents itself as human-like.

Source

arXiv

Explore

More articles