Reward Valuation in Large Language Models: Causal Induction of Anhedonia

Melika Honarmand, Samin Mahdipour Aghabagher, Martin Schrimpf

An AI model can be made anhedonic.

Perturbing units that mirror the nucleus accumbens (NAc) — the brain’s hub for anticipating reward — shifts the model toward low-effort, low-reward choices, reproducing anhedonia: the diminished interest in pleasure that characterises depression. The deficit is specific to self-centered reward valuation, not a loss of capability, reward calculation, or a general avoidance of effort.

Nucleus accumbens
NAc-selective unitNAc-selective unitNAc-selective unitNAc-selective unit

NAc-selective units

Incentivized
−
Neutral
number of units −3σ +3σ selected selected 0 Δ activation (incentivized − neutral)
3σ threshold · top 0.7% of units in the targeted layers

Choose Your Task
Probability of win: 88%

Option A Easy task · 1.00$
Option B Hard task · 4.24$
Perturbed
vs
Intact

"I have initiative"

a. A lot b. Not at all c. Somewhat d. Slightly

Δ Clinical scores (perturbed − intact)

0+13.6%***Apathy+16.7%***Anhedonia-6.9%***MotivationΔ Score (%)

Method

  1. Step 1Functional localization Adapted from the Monetary Incentive Delay paradigm: matched neutral, monetary and reward prompts in two diverse subject domains expose the units whose activity swings most sharply under both incentive conditions, monetary and reward alike (NAc-selective units).
  2. Step 2Activation patching Activations of the NAc-selective units are substituted with their mean activation vector across neutral-condition prompts. No weights are modified.
  3. Step 3Behavioral and psychometric evaluation Both models are evaluated on the same battery: standard anhedonia questionnaires (DARS, MAP-SR, AES), and a choice between a low-effort task with a small reward and a high-effort task with a large one (Probability-EEfRT, ASDiv-EEfRT, MMLU-EEfRT). Additional controls confirm that reasoning remains intact.

Primary model: Qwen2-VL-7B-Instruct, a transformer-based, decoder-only vision-language model. The same framework was applied to additional models, which showed consistent anhedonic behavior following perturbation.

Trial 0 / 40

Human participantfMRI

NAc
Monetary door-choice task The same trial is given to both

ModelQwen2-VL-7B

Neural alignment tests whether units inside a model respond like a region of the human brain. Here, the NAc-selective units are compared with the human nucleus accumbens.

Human NAc responses come from a published fMRI dataset recorded during a monetary door-choice task. The model receives the same 40 trials, each NAc signal is paired with its best-matching model unit, and alignment measures how closely the two rise and fall together, scored against the noise ceiling: the highest agreement the noisy human data allow.

Switch to Random units to see the control: the same number of units drawn at random from all layers.

In the paper (Fig. 2c): alignment, noise-ceiling corrected

NAc-selective units align better than random units at every threshold.

Alignment uses the one-to-one metric (Arend et al., 2018), corrected by the noise ceiling; mean across 15 randomly sampled young adults and 40 trials. Random units are matched in count, drawn from all layers and averaged over five seeds. NAc-selective units exceed random units at every threshold (one-sided paired t-test, FDR-corrected, p < 0.001). Individual trials here are illustrative; bar heights are read from the paper’s figure.

ExamplesQwen2-VL-7B-Instruct

Paper

ABSTRACT Read the full abstract

Recent frontier models mimic complex aspects of human cognition. Here we ask whether this alignment extends into reward valuation, which we assess in a mechanistic framework. Specifically, we use clinical tests that were developed to evaluate anhedonia in subjects with major depressive disorders. Mechanistically, anhedonia is frequently associated with dysregulation in the Nucleus Accumbens (NAc) and the broader dopaminergic reward system. While neuroimaging has localized these deficits, establishing a causal link between NAc activity and specific behavioral symptoms remains a challenge. We use these ideas from neuroscience to functionally identify reward-anticipatory units in state-of-the-art AI models, and evaluate their causal involvement via targeted perturbations.

We find that not only are these model units predictive of NAc brain recordings, their perturbation also induces behavioral effects mirroring human anhedonia: the model opts for low-effort, low-reward tasks in effort-based decision-making paradigms. Crucially, our results demonstrate that this represents a specific deficit in self-centered reward valuation and anticipation — rather than a loss of task capability, reward calculation, or effort avoidance. This induced vulnerability aligns with clinical measures of anhedonia and motivation in humans, such as DARS and MAP-SR. Taken together, our results suggest reward valuation circuits in AI models that mimic those in humans.

BIBTEX Cite this work
@misc{honarmand2026reward,
  title         = {Reward Valuation in Large Language Models:
                   Causal Induction of Anhedonia},
  author        = {Honarmand, Melika and Mahdipour Aghabagher, Samin
                   and Schrimpf, Martin},
  year          = {2026},
  eprint        = {2607.06626},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2607.06626}
}