Universal Activation VerbalizerA Unified Framework for
Cross-Model Activation Explanation

Haiyan Zhao1Zirui He1Guanchu Wang2Ali Payani3Yingcong Li1Mengnan Du4,†
1 New Jersey Institute of Technology2 University of North Carolina at Charlotte3 Cisco Research4 The Chinese University of Hong Kong, Shenzhen

† Corresponding author

A SHARED LANGUAGE FOR MODEL ACTIVATIONS

Different models.
One reusable verbalizer.

UAV turns hidden activations into natural-language answers. Donor- and layer-specific adapters translate representations into soft tokens that a shared decoder can read.

UAV framework: a donor model's hidden activation is mapped through an MLP or Q-Former adapter to soft tokens for a verbalizer. Stage 1 trains the adapter to reconstruct text; Stage 2 trains the adapter and decoder-side LoRA for question answering.
The UAV framework. Activation-to-text alignment teaches the adapter to translate hidden states. Instruction-following alignment trains the verbalizer to answer questions about the information they encode.

Abstract

Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each model explains only its own activations. We introduce Universal Activation Verbalizer (UAV), a shared-decoder framework that uses donor- and layer-specific adapters to explain activations from heterogeneous donor models. UAV learns a lightweight adapter that converts donor activations into soft tokens in decoder’s embedding space, and further supports adapter-only transfer by reusing a frozen decoder-side LoRA while training only a new adapter for another donor. Across classification, fact retrieval, and gist summarization, UAV remains competitive with strong self-explanation baselines while enabling cross-model verbalization across model families and scales. Ablations show that decoder-side tuning mainly improves task behavior, whereas the adapter provides the activation-grounded factual and semantic evidence. Code and data are available at our GitHub repository.

THE METHOD

Align activations. Learn to answer. Transfer.

The verbalizer receives activation-derived soft tokens and a question. The source text is withheld from the verbalizer at evaluation time.

01

Activation-to-text alignment

Train an MLP or Q-Former adapter to reconstruct the input prefix from a donor activation. The decoder stays frozen while the adapter learns to produce decoder-readable soft tokens.

Stage 1 · Adapter only
02

Instruction-following alignment

Train the adapter together with decoder-side LoRA on explanation-oriented questions. The decoder backbone remains frozen as the verbalizer learns classification, factual retrieval, and gist tasks.

Stage 2 · Adapter + LoRA
03

Adapter-only transfer

For another donor, reuse the adapted decoder and freeze its LoRA. Train a new donor-specific adapter to connect the new activation space to the existing verbalizer.

Reuse the adapted decoder

“Universal” refers to decoder reuse across donors. New donors require their own trained adapters; the framework reads out recoverable information and does not establish causal mechanisms.

CROSS-MODEL VERBALIZATION

A shared decoder across model families.

Qwen3-4B-Instruct-2507 serves as the decoder for Llama, Gemma, and Yi donor models.

4BShared decoderQwen3-4B-Instruct-2507
4B–34BCross-model donor sizesConfigurations evaluated in Table 3
3Activation readout tasksClassification, fact retrieval, and gist summarization

Cross-donor results

Table 3
Cross-donor results from Table 3 using Qwen3-4B-Instruct-2507 as the shared decoder.
Donor modelValidation loss ↓ROUGE-L ↑BERTScore ↑
Llama-3.1-8B-Instruct1.63880.2740.369
Gemma-3-4B-IT1.67900.2660.354
Gemma-3-12B-IT1.65890.2690.355
Yi-1.5-34B-Chat1.58370.2990.377

Mean generation scores are shown; per-example standard deviations are reported in the paper. Donor size alone does not determine how readily its activations can be verbalized.

Comparison with trained self-explanation baselines

Table 2 · Overall scores
Selected training-based methods from Table 2, with overall ROUGE-L and BERTScore.
Method / adaptationDecoderROUGE-L ↑BERTScore ↑
Donor: Qwen3-4B-Instruct-2507
Activation OracleQwen3-4B0.1980.234
LatentQAQwen3-4B0.2350.327
UAV · FullQwen3-4B0.2540.344
UAV · AOT from LlamaQwen3-4B0.2600.347
Donor: Llama-3.1-8B-Instruct
Activation OracleLlama-8B0.2730.362
LatentQALlama-8B0.2850.370
UAV · Full, self-decodingLlama-8B0.2860.379
UAV · Full, cross-decodingQwen3-4B0.2740.369
UAV · AOT from QwenQwen3-4B0.2570.346

Full: two-stage adaptation on the current donor–decoder pair. AOT: adapter-only transfer with a frozen decoder-side LoRA learned from the named source donor. The paper includes task-level results, training-free baselines, standard deviations, and paired-bootstrap comparisons.

WHAT EACH COMPONENT CONTRIBUTES

The adapter supplies activation-grounded evidence.

Decoder-side LoRA improves task behavior. Recovering input-specific facts and meaning also depends on the trained activation adapter.

Figure 3 compares Token F1, ROUGE-L, and BERTScore across six ablation variants. Full UAV improves fact and gist readout over decoder-side LoRA alone.
Component ablation (Figure 3). Fine-tuned LoRA alone achieves a fact-retrieval BERTScore of 0.196; full UAV reaches 0.402. These are ablation-specific values, separate from the main comparison settings.

TRAINING WITH CACHED ACTIVATIONS

A smaller decoder lowers training cost.

Caching donor activations allows verbalizer training without keeping the donor model in GPU memory.

Training cost per epoch from Appendix D.1 and Table 12.
ConfigurationStage 1 GPU-hours ↓Stage 2 GPU-hours ↓Validation loss ↓
12B self-decoding7.126.51.6004
4B UAV cross-decoding2.711.61.6594

Appendix D.1, measured per epoch under the same setup. Lower training cost comes with a validation-loss increase of 0.059 in this comparison.

ACTIVATIONS IN NATURAL LANGUAGE

What can the verbalizer recover?

Selected fact-retrieval examples from Table 19. Source excerpts are shown for context; the verbalizer receives the question and activation-derived tokens.

Successful factual readout

ROUGE-L 1.000

SOURCE EXCERPT

Westfield Corporation was founded with the spin-off of the Westfield Group in 2014…

QUESTION

What was the name of the group this entity split from?
ReferenceWestfield Group
UAVWestfield Group

A factual failure

ROUGE-L 0.333

SOURCE EXCERPT

…an old place name for a part of Nishinari-ku in Osaka, Japan.

QUESTION

What part of Osaka is this place located in?
ReferenceNishinari-ku
UAVthe Kita-ku district

A plausible answer can still contain the wrong fact. These examples illustrate both recoverable information and the limits of activation verbalization.

REFERENCE

Citation

@inproceedings{zhao2026universal,
  title     = {Universal Activation Verbalizer: A Unified Framework
               for Cross-Model Activation Explanation},
  author    = {Zhao, Haiyan and He, Zirui and Wang, Guanchu and
               Payani, Ali and Li, Yingcong and Du, Mengnan},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods
               in Natural Language Processing},
  year      = {2026},
  url       = {https://arxiv.org/abs/2605.25903}
}