intermediate-outputs

intermediate-outputs is a skill for Claude Code from zjunlp/Mechanist. It costs 45 tokens per session (5,709 once invoked), scanned A, original, MIT.

A set of methods for finding the internal parts of a language model that produce a behavior. It includes activation patching, attribution patching, and Layer-wise Relevance Propagation, which traces how much each layer contributes to an output.

In plain words
What is it for?
Use it to discover model circuits, patch internal activations during experiments, measure which components matter, and apply relevance analysis to transformer models.
Why use it?
It helps investigate why a model gives an answer by tracing information through its internal calculations.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin.

Part of the mechanist plugin — 54 skills, 4 agents shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/zjunlp/mechanist/intermediate-outputs
Any agent
npx skills add zjunlp/Mechanist --skill intermediate-outputs
Clone the repo
git clone --depth 1 https://github.com/zjunlp/Mechanist

Made for: Claude Code.

Or install mechanist, the plugin that ships this one along with the rest of its 54 skills, 4 agents.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for intermediate-outputs

README.md
[![agentmods](https://agentmods.dev/badge/skills/zjunlp/mechanist/intermediate-outputs.svg)](https://agentmods.dev/skills/zjunlp/mechanist/intermediate-outputs)
Your own site
<a href="https://agentmods.dev/skills/zjunlp/mechanist/intermediate-outputs"><img src="https://agentmods.dev/badge/skills/zjunlp/mechanist/intermediate-outputs.svg" alt="Measured on agentmods" height="20"></a>
Per session 45 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 5,709 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00045 $0.05709
Opus 5 $0.00023 $0.02854
Sonnet 5 $0.00009 $0.01142
Haiku 4.5 $0.00005 $0.00571

Measured 6d ago against content hash 225c23ad9e28, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

intermediate-outputs scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (scripts/basic_relp_analysis.py, scripts/ioi_task_analysis.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/mechanism-skills/gradient-detection/intermediate-outputs/SKILL.md · 752 lines

How it starts

The opening of the file, as written. The whole thing — 752 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Demo Scripts

scripts/basic_relp_analysis.py

#!/usr/bin/env python3
"""
Basic RelP Analysis Script

This script demonstrates how to use RelP (Relevance Patching) for circuit discovery
in transformer language models using the enhanced TransformerLens library.

Requirements:
    - Install RelP: git clone https://github.com/FarnoushRJ/RelP.git && cd RelP/TransformerLens && pip install -e .
"""

import torch
import numpy as np
from typing import Dict, List, Optional, Tuple
import transformer_lens
from transformer_lens import HookedTransformer, ActivationCache


def setup_model_with_lrp(
    model_name: str = "gpt2-small",
    lrp_rules: Optional[List[str]] = None,
    device: str = "cuda" if torch.cuda.is_available() else "cpu"
) -> HookedTransformer:
    """
    Load a transformer model with LRP (Layer-wise Relevance Propagation) enabled.
    
    Args:
        model_name: Name of the pretrained model to load
        lrp_rules: List of LRP rules to apply. Defaults to standard rules.
        device: Device to load the model on
    
    Returns:
        HookedTransformer model with LRP configuration
    """
    print(f"Loading model: {model_name}")
    model = HookedTransformer.from_pretrained(model_name, device=device)
    
    # Enable LRP
    model.cfg.use_lrp = True
    
    # Set LRP rules (use defaults if not specified)
    if lrp_rules is None:
        lrp_rules = ['LN-rule', 'Identity-rule', 'Half-rule']
    model.cfg.LRP_rules = lrp_rules
    
    print(f"Model loaded with LRP rules: {lrp_rules}")
    return model


def analyze_text_with_relp(
    model: HookedTransformer,
    text: str,
    return_logits: bool = True
) -> Tuple[torch.Tensor, ActivationCache]:
    """
    Analyze a text input using RelP to get relevance scores and activations.
    
    Args:
        model: The transformer model with LRP enabled
        text: Input text to analyze
        return_logits: Whether to return logits along with activations
    
    Returns:
        Tuple of (logits, activation_cache) containing model outputs and internal states
    """
    print(f"Analyzing text: '{text}'")
    
    # Run the model with cache to capture all activations
    logits, cache = model.run_with_cache(text)
    
    # Display basic information about the outputs
    print(f"Logits shape: {logits.shape}")
    print(f"Number of cached activations: {len(cache)}")
    
    return logits, cache


def extract_attention_patterns(
    cache: ActivationCache,
    layer: int = 0,
    head: int = 0
) -> np.ndarray:
    """
    Extract attention patterns from a specific layer and head.
    
    Args:
        cache: ActivationCache containing model activations
        layer: Layer index to extract from
        head: Attention head index
    
    Returns:
        Attention pattern as numpy array
    """
    # Get attention pattern for specified layer and head
    attn_pattern_key = f"blocks.{layer}.attn.hook_pattern"
    
    if attn_pattern_key in cache:
        attn_patterns = cache[attn_pattern_key]
        # Shape: [batch, head, seq_len, seq_len]
        pattern = attn_patterns[0, head].cpu().numpy()
        return pattern
    else:
        print(f"Warning: Attention pattern not found for layer {layer}")
        return np.array([])


def analyze_mlp_contributions(
    cache: ActivationCache,
    layer: int = 0
) -> Dict[str, torch.Tensor]:
    """
    Analyze MLP (feedforward) layer contributions using cached activations.
    
    Args:
        cache: ActivationCache containing model activations
        layer: Layer index to analyze
    
    Returns:
        Dictionary containing MLP-related activations
    """
    mlp_info = {}
    
    # Extract MLP pre-activation
    mlp_pre_key = f"blocks.{layer}.mlp.hook_pre"
    if mlp_pre_key in cache:
        mlp_info['pre_activation'] = cache[mlp_pre_key]
    
    # Extract MLP post-activation
    mlp_post_key = f"blocks.{layer}.mlp.hook_post"
    if mlp_post_key in cache:
        mlp_info['post_activation'] = cache[mlp_post_key]
    
    # Extract MLP output
    mlp_out_key = f"blocks.{layer}.hook_mlp_out"
    if mlp_out_key in cache:
        mlp_info['output'] = cache[mlp_out_key]
    
    return mlp_info


def compute_relevance_scores(
    model: HookedTransformer,
    text: str,
    target_token_idx: int = -1
) -> Dict[str, torch.Tensor]:
    """
    Compute relevance scores for different model components using RelP.
    
    Args:
        model: Transformer model with LRP enabled
        text: Input text
        target_token_idx: Index of target token to compute relevance for
    
    Returns:
        Dictionary of relevance scores for different components
    """
    # Tokenize input
    tokens = model.to_tokens(text)
    
    # Run forward pass
    logits, cache = model.run_with_cache(tokens)
    
    # Get output for target token
    if target_token_idx == -1:
        target_token_idx = tokens.shape[1] - 1
    
    target_logits = logits[0, target_token_idx]
    
    # Compute relevance scores (simplified example)
    relevance_scores = {}
    
    # Store some key activation magnitudes as proxy for relevance
    for layer in range(model.cfg.n_layers):
        # Residual stream relevance
        resid_key = f"blocks.{layer}.hook_resid_post"
        if resid_key in cache:
            resid_relevance = cache[resid_key][0, target_token_idx].abs().mean()
            relevance_scores[f"layer_{layer}_residual"] = resid_relevance
        
        # MLP relevance
        mlp_key = f"blocks.{layer}.hook_mlp_out"
        if mlp_key in cache:
            mlp_relevance = cache[mlp_key][0, target_token_idx].abs().mean()
            relevance_scores[f"layer_{layer}_mlp"] = mlp_relevance
        
        # Attention relevance
        attn_key = f"blocks.{layer}.hook_attn_out"
        if attn_key in cache:
            attn_relevance = cache[attn_key][0, target_token_idx].abs().mean()
            relevance_scores[f"layer_{layer}_attention"] = attn_relevance
    
    return relevance_scores


def compare_lrp_rules(
    model_name: str = "gpt2-small",
    text: str = "The cat sat on the mat",
    rules_sets: Optional[List[List[str]]] = None
) -> None:
    """
    Compare different LRP rule configurations on the same input.
    
    Args:
        model_name: Model to use for comparison
        text: Input text for analysis
        rules_sets: List of LRP rule sets to compare
    """
    if rules_sets is None:
        rules_sets = [
            ['LN-rule', 'Identity-rule', 'Half-rule'],
            ['LN-rule', '0-rule', 'AH-rule'],
            ['Identity-rule', '0-rule', 'Half-rule']
        ]
    
    print(f"\nComparing LRP rules on: '{text}'\n")
    
    for rules in rules_sets:
        print(f"Testing rules: {rules}")
        model = setup_model_with_lrp(model_name, lrp_rules=rules)
        
        # Compute relevance scores
        relevance = compute_relevance_scores(model, text)
        
        # Display top relevance scores
        sorted_relevance = sorted(relevance.items(), key=lambda x: x[1], reverse=True)
        print("Top 5 components by relevance:")
        for component, score in sorted_relevance[:5]:
            print(f"  {component}: {score:.4f}")
        print()
        
        # Clean up
        del model
        torch.cuda.empty_cache()


def main():
    """
    Main demonstration of RelP functionality.
    """
    print("=" * 60)
    print("RelP (Relevance Patching) Demonstration")
    print("=" * 60)
    
    # Example 1: Basic model setup and analysis
    print("\n1. Basic Setup and Analysis")
    print("-" * 40)
    
    model = setup_model_with_lrp("gpt2-small")
    text = "The capital of France is Paris"
    logits, cache = analyze_text_with_relp(model, text)
    
    # Example 2: Extract attention patterns
    print("\n2. Attention Pattern Analysis")
    print("-" * 40)
    
    for layer in [0, 5, 11]:  # Sample from early, middle, late layers
        attn_pattern = extract_attention_patterns(cache, layer=layer, head=0)
        if attn_pattern.size > 0:
            print(f"Layer {layer}, Head 0 - Attention pattern shape: {attn_pattern.shape}")
            print(f"  Max attention: {attn_pattern.max():.4f}")
            print(f"  Mean attention: {attn_pattern.mean():.4f}")
    
    # Example 3: Analyze MLP contributions
    print("\n3. MLP Contribution Analysis")
    print("-" * 40)
    
    for layer in [0, 5, 11]:
        mlp_info = analyze_mlp_contributions(cache, layer=layer)
        print(f"Layer {layer} MLP:")
        for key, tensor in mlp_info.items():
            if tensor is not None:
                print(f"  {key}: shape={tensor.shape}, mean={tensor.mean().item():.4f}")
    
    # Example 4: Compute relevance scores
    print("\n4. Component Relevance Scores")
    print("-" * 40)
    
    relevance_scores = compute_relevance_scores(model, text)
    
    # Find most relevant components
    sorted_scores = sorted(relevance_scores.items(), key=lambda x: x[1], reverse=True)
    print("Top 10 most relevant components:")
    for component, score in sorted_scores[:10]:
        print(f"  {component}: {score:.4f}")
    
    # Example 5: Compare different LRP rules
    print("\n5. LRP Rule Comparison")
    print("-" * 40)
    
    compare_lrp_rules(
        model_name="gpt2-small",
        text="Machine learning is transforming technology",
        rules_sets=[
            ['LN-rule', 'Identity-rule', 'Half-rule'],
            ['LN-rule', '0-rule', 'AH-rule']
        ]
    )
    
    print("\n" + "=" * 60)
    print("RelP demonstration complete!")
    print("=" * 60)


if __name__ == "__main__":
    main()

Read the full file on GitHub · 752 lines

Files

What ships with it

3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 752 lines · 45 tokens per session scan A 225c23ad9e28

Subscribe to this mod's changes

intermediate-outputs is a skill published in the GitHub repository zjunlp/Mechanist (55 stars, last pushed 10d ago), licensed MIT. It adds 45 tokens to every session and 5,709 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.