All Layers Splitter¶
AllLayersSplitter captures the residual stream before the first transformer block and after every block. It also
applies the wrapped model's native normalization, pooling, and prediction head to one or more residual states.
interpreto.AllLayersSplitter
¶
AllLayersSplitter(model_or_repo_id, *, automodel=AutoModelForCausalLM, tokenizer=None, config=None, device_map=None, layer_path=None, **kwargs)
Bases: LanguageModel
Extract the residual stream before and after every transformer block.
The transformer blocks are inferred from the model configuration or selected
explicitly with layer_path. Activations are returned in model order: the
input to the first block followed by the output of every block. A model with
L transformer blocks therefore returns L + 1 tensors of shape
(1, sequence_length, model_width).
This splitter is intended for methods that compare representations across
model depths, such as Logit Lens and Tuned Lens. It does not implement the
concept-specific BaseSplitter interface because there is no single split
point or latent representation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
str | PreTrainedModel
|
Hugging Face repository ID, local checkpoint path, or preloaded model. |
required |
|
type[AutoModel]
|
Hugging Face AutoClass used when loading a model from a repository ID or local path. |
AutoModelForCausalLM
|
|
PreTrainedTokenizerBase | None
|
Tokenizer associated with a preloaded model. |
None
|
|
PretrainedConfig | None
|
Optional model configuration passed to the model loader. |
None
|
|
device | str | None
|
Device map passed to the model loader. |
None
|
|
str | None
|
Path to the transformer block |
None
|
|
Any
|
Additional arguments passed to NNsight's |
{}
|
Raises:
| Type | Description |
|---|---|
InitializationError
|
If a preloaded model is provided without a tokenizer. |
ValueError
|
If a tokenizer cannot be inferred for a repository ID. |
Example
from transformers import AutoModelForCausalLM from interpreto import AllLayersSplitter splitter = AllLayersSplitter("gpt2") activations = splitter.get_activations("Interpreto is useful.") len(activations) == len(splitter.split_points) + 1 True
Source code in interpreto/concepts/splitters/all_layers_splitter.py
activation_names
property
¶
Names of the residual states returned by :meth:get_activations.
apply_head
¶
apply_head(activations)
Apply the wrapped model's prediction head to residual activations.
The transformer blocks are skipped and activations are used as
their output. The wrapped model then executes its own downstream
normalization, pooling, and prediction head. This avoids
architecture-specific head names and preserves functional operations
implemented in model forward methods.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
Tensor
|
Residual activations with shape
|
required |
Returns:
| Type | Description |
|---|---|
Tensor
|
torch.Tensor: Logits returned by the wrapped model for every activation in the leading dimension. |
Source code in interpreto/concepts/splitters/all_layers_splitter.py
get_activations
¶
get_activations(inputs)
Extract the residual stream for one text input.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
str
|
Text passed to the wrapped model. |
required |
Returns:
| Type | Description |
|---|---|
list[Tensor]
|
list[torch.Tensor]: Input to the first transformer block followed by
every transformer block output in |