Steering vectors as a control surface for LLM agents: what reinforcement learning adds

Two literatures that grew up separately converged over 2024 and 2025. Activation steering adds a small vector to a transformer’s residual stream to push behaviour in a chosen direction; reinforcement-learning post-training rewrites millions of weights to teach a model to reason or act. The bridge is to treat the steering vector, or a router that composes several of them, as the only trainable object and to train it with a reward. On mathematical reasoning that recovers most of what full RL fine-tuning delivers with roughly 0.0016 percent of the parameters. For agents the same move turns exploration, procedural memory and persona into settings that can be inspected at inference time.

The object: a direction in activation space

The premise is the linear representation hypothesis: many behaviours a language model can exhibit correspond to directions in its residual stream, so adding a scaled vector at one or more layers shifts the behaviour without touching the weights. Subramani, Suresh and Peters 2022 showed that frozen decoders already contain latent steering vectors that regenerate a target sentence almost perfectly, and Turner et al. 2023 made the idea practical with activation addition, taking the difference between the activations of two contrasting prompts and adding it during generation. Zou et al. 2023 generalised the idea into representation engineering, a top-down programme for reading and controlling concepts such as honesty and power-seeking from population-level activations. Li et al. 2023 applied a similar intervention to attention heads to raise truthfulness, and Panickssery et al. 2024 systematised the contrastive recipe into contrastive activation addition, with vectors computed from paired multiple-choice completions. The most striking single result in this family is Arditi et al. 2024: across thirteen open chat models, refusal is mediated by a single direction, and ablating it disables refusal while adding it induces refusal on harmless prompts.

The survey by Wehner et al. 2025 organises all of this as a pipeline of representation identification, operationalisation and control, and is the right entry point for a reader who wants the full taxonomy. Two things matter for what follows. First, the classical vectors are extracted, not learned: they come from activation differences on curated prompt pairs. Second, the intervention is static: the same vector, at the same strength, on every token of every input.

Learning the vector instead of extracting it

Fitting the vector with a training objective, rather than reading it off activation differences, takes several forms. Ackerman 2024 tuned honesty directions into a model with a loss that combines cosine similarity to the vector with the usual token loss, finding a stronger effect than online steering and better generalisation than token-loss fine-tuning alone. Wu et al. 2024 made this a general fine-tuning paradigm, representation finetuning, in which low-rank interventions on hidden representations replace weight updates and match or beat parameter-efficient adapters with far fewer parameters. Cao et al. 2024 fit steering vectors by preference optimisation on contrastive pairs, and Wu et al. 2025 extended that to a bidirectional objective in which one learned intervention both promotes and suppresses a concept, narrowing the gap with prompting on a large steering benchmark while resisting the prompt-based jailbreaks that defeat prompting.

The reinforcement-learning version arrived with Sinii et al. 2025. They freeze the base model, add one trainable vector per layer to the residual stream, and optimise only those vectors with an online policy-gradient recipe of the kind used to train reasoning models, with a binary reward for a correct boxed answer. Trained on the DeepScaleR problems and evaluated on six mathematical benchmarks (AIME 2024 and 2025, AMC 2023, MATH-500, MinervaMath, OlympiadBench) across seven base models from the Qwen 2.5 and Llama 3.1 families, the steered models match the fully RL-tuned ones, with one exception: the base Llama 3.1 8B recovers about seventy percent of the fine-tuning gain. The cost side is the memorable number. On the 14-billion-parameter Qwen model the trainable set is 245 thousand parameters against 14.7 billion, optimiser memory falls from 13.8 gigabytes to 240 kilobytes, and an epoch takes 34 seconds instead of 52 minutes. The authors are careful about the interpretation. Under the working assumption that an additive vector can only amplify features the network already has, the result says that wherever the base model already knows how to solve the task, elicitation is enough; where it does not, a low-rank adapter closes the remaining gap, which they take to mean that a single global vector, not the elicitation idea, is what limits the method. A companion mechanistic study, Sinii et al. 2025b, reads the trained vectors with a logit lens. The last-layer vector behaves like a token-substitution bias concentrated on the first generated token, favouring openers such as “To” and “Step”. The penultimate-layer vector leaves attention largely alone and works through the MLP and the unembedding. And the vectors transfer to other models of the same family. For anyone who has wondered what RL teaches a model, this is the cleanest evidence so far that a large part of the answer is a low-dimensional additive nudge.

From one vector to a policy over vectors

A static vector cannot adapt to the input. The 2026 work makes the intervention state-dependent, which is the point at which steering stops being a knob and becomes a policy. Ye et al. 2026 build a library of reusable reasoning vectors and train a lightweight router, by reinforcement learning under task-level rewards, to compose them per input. Across seven benchmarks the router improves zero-shot accuracy by 3.4 to 6.5 percent on average over the base model and beats chain-of-thought prompting at two to three times the token efficiency; the paper reports that the learned compositions are interpretable. A smaller parallel line replaces the learned router with a classical controller: Bharadwaj 2025 modulates the steering strength during generation with a PID loop driven by a chunk-level redundancy classifier, because a constant coefficient strong enough to suppress the worst reasoning chunks over-suppresses the productive ones. The author presents it as a proof of concept on a subset of GSM8K with a 1.5-billion-parameter model, and it should be read as one. Oozeer et al. 2025 drop the linearity assumption altogether, training a single non-linear multi-label classifier on activations and steering along its gradient, which composes several behaviours without storing a vector per attribute.

Read together, these papers describe a small control stack: a dictionary of directions, a state estimate, and a policy that maps the state to a mixture and a gain. That is a reinforcement-learning problem in miniature, with the residual stream as the actuator.

Agents: exploration, memory and persona as steering

Three agent-specific results show why this matters beyond benchmarks.

Exploration. Rahn, D’Oro and Bellemare 2024 study in-context LLM agents in sequential decision tasks and find them overconfident: the entropy of the implied action distribution collapses and the agent stops exploring. Token-level sampling tricks do not fix this. Their entropic activation steering computes an entropy-weighted steering vector and raises the agent’s action entropy at the activation level, restoring exploration and, in passing, changing the uncertainty the agent expresses.

Procedural memory. Zhao et al. 2026 argue that retrieval-augmented text guidelines suffer from a text-to-action disconnect: the instruction is in the context but the internal mechanism it should trigger stays dormant. Their neural procedural memory distils contrastive past experiences into steering vectors and applies them during task execution. On four agent benchmarks the implicit memory matches explicit textual instructions, and combining the two is better than either, with representational analysis showing that the vectors encode consistent task logic.

Persona. Chen et al. 2025 extract persona vectors for traits such as sycophancy and the propensity to hallucinate from a natural-language description alone. The vectors monitor personality drift at deployment time. Both intended and unintended personality changes after fine-tuning project onto the same directions, which lets the authors flag training data that will shift a trait before training begins, and steer preventatively during it. For an agent that will be fine-tuned on interaction data, this is a monitoring instrument first and a control second.

Where it breaks

Two results temper the enthusiasm. Tan et al. 2024 had already found that steering vectors generalise unevenly and that in-distribution reliability varies widely across inputs; Braun et al. 2025 measure steering reliability across seven prompt types and find a net positive effect with very high sample-level variance, frequently the opposite of the intended effect on individual inputs. Steering works when the target behaviour is a coherent direction, as measured by the cosine similarity of the training-set activation differences, and fails when it is not. A steering vector is therefore a population-level intervention with an individual-level failure rate that must be measured, not assumed. And Korznikov et al. 2025 show that steering systematically degrades alignment safeguards: even a random direction raises harmful compliance from zero to between one and thirteen percent across model families, benign features from a sparse autoencoder do comparable damage, and twenty random vectors that each jailbreak one prompt combine into a universal attack. Precise control over internals is not precise control over behaviour. The same actuator that adds procedural memory removes refusal.

The reinforcement-learning variants inherit a further risk that the extraction variants did not have: whatever the reward rewards, the vector will find. A router trained on task reward has no reason to preserve the behaviours the reward does not measure. I have not found such an evaluation in the papers above, and none of their abstracts mentions one; it is the gap I would most want filled.

What this means for agents that act on estimates

The agents I care about act on estimated quantities: a treatment effect, an elasticity, a forecast. Steering offers a cheap and inspectable way to change how such an agent explores or which procedure it follows, without retraining. The vector itself can be filed as part of the audit record. But an intervention on activations is a treatment applied to a policy, and its effect on downstream decisions is a causal quantity like any other. The evaluation should be run the way any treatment is evaluated: steered and unsteered conditions assigned at random over held-out prompts, with the primary metrics fixed in advance and chosen to include the behaviours the reward never measured. Check the coherence diagnostic of Braun et al. before trusting a vector at all. Two open questions are worth an experiment each, on open-weight models and synthetic environments with known ground truth. Does the individual-level failure rate that Braun et al. measure shrink when the vector is learned rather than extracted? And does preventative steering during agent fine-tuning preserve calibration on ground-truth effects, rather than only on the stated trait? A third, whether a learned router generalises out of distribution better than the prompt it replaces, is the obvious one and the least likely to surprise.

Limits of this reading

Most results sit on models of eight billion parameters or fewer and on mathematical benchmarks; the agent results use small benchmark suites. The 2026 papers are preprints or findings-track publications and have not yet been replicated across base models. Comparisons across papers are confounded by different base models, layers and evaluation protocols. Of the papers above, Sinii et al. 2025 was read in full; the others are summarised from their abstracts and the passages quoted here, and every claim about them should be treated as abstract-level until the paper is read. The four I would read next, because the note leans on them, are Ye et al. 2026, Zhao et al. 2026, Braun et al. 2025 and Korznikov et al. 2025.

References

  • Ackerman, C. M. (2024). Representation tuning. arXiv:2409.06927.
  • Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., Nanda, N. (2024). Refusal in language models is mediated by a single direction. arXiv:2406.11717.
  • Bharadwaj, A. R. (2025). Adaptive activation steering for efficient LLM reasoning via closed-loop PID control. arXiv:2506.18831.
  • Braun, J., Eickhoff, C., Krueger, D., Bahrainian, S. A., Krasheninnikov, D. (2025). Understanding (un)reliability of steering vectors in language models. arXiv:2505.22637.
  • Cao, Y., Zhang, T., Cao, B., Yin, Z., Lin, L., Ma, F., Chen, J. (2024). Personalized steering of large language models: versatile steering vectors through bi-directional preference optimization. NeurIPS 2024. arXiv:2406.00045.
  • Chen, R., Arditi, A., Sleight, H., Evans, O., Lindsey, J. (2025). Persona vectors: monitoring and controlling character traits in language models. arXiv:2507.21509.
  • Korznikov, A., Galichin, A., Dontsov, A., Rogov, O. Y., Oseledets, I., Tutubalina, E. (2025). The Rogue Scalpel: activation steering compromises LLM safety. arXiv:2509.22067.
  • Li, K., Patel, O., Viégas, F., Pfister, H., Wattenberg, M. (2023). Inference-time intervention: eliciting truthful answers from a language model. arXiv:2306.03341.
  • Oozeer, N., Marks, L., Jain, S., Barez, F., Abdullah, A. (2025). Beyond linear steering: unified multi-attribute control for language models. Findings of EMNLP 2025. arXiv:2505.24535.
  • Panickssery, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., Turner, A. M. (2024). Steering Llama 2 via contrastive activation addition. arXiv:2312.06681.
  • Rahn, N., D’Oro, P., Bellemare, M. G. (2024). Controlling large language model agents with entropic activation steering. arXiv:2406.00244.
  • Sinii, V., Balagansky, N., Gerasimov, G., Laptev, D., Aksenov, Y., Kurochkin, V., Gorbatovski, A., Shaposhnikov, B., Gavrilov, D. (2025b). Small vectors, big effects: a mechanistic study of RL-induced reasoning via steering vectors. arXiv:2509.06608.
  • Sinii, V., Gorbatovski, A., Cherepanov, A., Shaposhnikov, B., Balagansky, N., Gavrilov, D. (2025). Steering LLM reasoning through bias-only adaptation. EMNLP 2025 (main). arXiv:2505.18706.
  • Subramani, N., Suresh, N., Peters, M. E. (2022). Extracting latent steering vectors from pretrained language models. Findings of ACL 2022. arXiv:2205.05124.
  • Tan, D., Chanin, D., Lynch, A., Kanoulas, D., Paige, B., Garriga-Alonso, A., Kirk, R. (2024). Analyzing the generalization and reliability of steering vectors. NeurIPS 2024. arXiv:2407.12404.
  • Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., MacDiarmid, M. (2023). Steering language models with activation engineering. arXiv:2308.10248.
  • Wehner, J., Abdelnabi, S., Tan, D., Krueger, D., Fritz, M. (2025). Taxonomy, opportunities, and challenges of representation engineering for large language models. TMLR. arXiv:2502.19649.
  • Wu, Z., Arora, A., Wang, Z., Geiger, A., Jurafsky, D., Manning, C. D., Potts, C. (2024). ReFT: representation finetuning for language models. arXiv:2404.03592.
  • Wu, Z., Yu, Q., Arora, A., Manning, C. D., Potts, C. (2025). Improved representation steering for language models. NeurIPS 2025. arXiv:2505.20809.
  • Ye, W., Yuan, X., Bin, Y., Zeng, P., Jin, H., Peng, L., Shen, H. T. (2026). RISER: orchestrating latent reasoning skills for adaptive activation steering. Findings of ACL 2026 (aclanthology.org/2026.findings-acl.226). arXiv:2601.09269.
  • Zhao, C., Tan, Y., He, S., Wang, Y., Zhao, J., Liu, K. (2026). Neural procedural memory: empowering LLM agents with implicit activation steering. arXiv:2606.29824.
  • Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, J. Z., Hendrycks, D. (2023). Representation engineering: a top-down approach to AI transparency. arXiv:2310.01405.