Checkpoint-aware Jacobian-lens, typed-site causal tracing, and J-space toolkit for Gemma models
  • Python 97.8%
  • Shell 1.2%
  • HTML 0.9%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Turner Monroe 71fdc653b8
Some checks failed
ci / test (3.10) (push) Failing after 9s
ci / test (3.12) (push) Failing after 2s
Update local qualification runner
2026-08-30 14:37:08 +02:00
.github/workflows Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
containers Run qualification containers with Podman 2026-07-30 01:55:55 +02:00
docs Allow reviewed release after Flower withdrawal 2026-08-30 14:36:38 +02:00
examples Preregister cross-shard spider signal confirmation 2026-07-29 21:32:42 +02:00
locks Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
scripts Update local qualification runner 2026-08-30 14:37:08 +02:00
src/gemma_jlens Queue stock Gemma 4 12B lens qualification 2026-07-29 23:24:36 +02:00
tests Allow reviewed release after Flower withdrawal 2026-08-30 14:36:38 +02:00
.dockerignore Add operational control panel and validation workflows 2026-07-29 21:28:29 +02:00
.gitignore Update local qualification runner 2026-08-30 14:37:08 +02:00
CAUSAL_STUDIES.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
CHANGELOG.md Allow reviewed release after Flower withdrawal 2026-08-30 14:36:38 +02:00
gemma-jacobian-lens-0.4.0-source.zip Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
gemma_jlens-0.4.0-py3-none-any.whl Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
gemma_jlens-0.4.0.tar.gz Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
JSPACE_DECOMPOSITION.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
LICENSE Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
MANIFEST.in Add operational control panel and validation workflows 2026-07-29 21:28:29 +02:00
OFFLINE_DOCTOR.json Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
PLE_CHANNELS.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
pyproject.toml Add operational control panel and validation workflows 2026-07-29 21:28:29 +02:00
README.md Queue stock Gemma 4 12B lens qualification 2026-07-29 23:24:36 +02:00
REAL_MODEL_VALIDATION.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
RELEASE_MANIFEST.json Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
RELEASE_NOTES.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
RELEASE_VALIDATION.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
REPRODUCIBILITY.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
SHA256SUMS.txt Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
STRIX_HALO_RUNBOOK.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
STUDY_DAGS.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
THIRD_PARTY_NOTICES.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
TILED_STORAGE.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
TYPED_HOOK_SITES.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00
VALIDATION_LADDER.md Initial import of Gemma J-Lens v0.4.0 2026-07-26 14:17:31 +02:00

Gemma J-Lens v0.4

A checkpoint-aware Jacobian-lens, typed-site causal tracing, and J-space research toolkit for Gemma-family Hugging Face models.

v0.4 turns the project from a collection of strong research primitives into a coherent causal-tracing system. It adds explicit hook-site addresses, same-layer Per-Layer Embedding analysis, residual-aware sparse decomposition, delayed readouts, multiple null distributions, joint residual/PLE factorial studies, structured prompt receipts, lazy rectangular lens storage, and resumable study DAGs.

The exact estimator remains Anthropic's average input-output Jacobian:

lens_s_to_t(x_s) = decode_t(J_s_to_t @ x_s)
J_s_to_t         = E[d x_t / d x_s]

For a conventional residual-to-final lens this reduces to:

lens_l(h_l) = unembed(J_l @ h_l)
J_l         = E[d h_target / d h_l]

The exact numerical core is a frozen, Apache-2.0-compatible copy of Anthropic's reference implementation at commit:

581d398613e5602a5af361e1c34d3a92ea82ba8e

It ships inside the wheel, so installation and offline validation do not depend on a live Git repository. The vendored implementation is the deterministic default. Advanced parity testing can explicitly select an external compatible jlens package with GEMMA_JLENS_CORE=external, or an intentionally injected module with GEMMA_JLENS_CORE=auto. Merely installing or importing another package does not change the estimator in the default mode.

What is new in v0.4

Typed causal locations

A layer number alone is ambiguous. v0.4 addresses activations by both layer and site:

resid_pre@12
attn_out@12
resid_mid@12
mlp_out@12
ple_input@12
resid_post@12
final_norm_in
final_norm_out

This permits exact questions such as:

ple_input@12 -> resid_post@12
attn_out@8   -> final_norm_in
resid_pre@4  -> resid_post@20

The causal validator rejects reverse, identical, and unsupported paths before an expensive fit begins.

Same-layer PLE analysis

Gemma 4 E2B/E4B can inject an auxiliary Per-Layer Embedding input into a decoder block. v0.4 can measure the direct same-layer derivative:

J = d resid_post@L / d ple_input@L

as well as longer PLE-to-final paths. The result is a conditional derivative with other inputs held fixed. It is not automatically an additive attribution of the whole model computation.

Residual-aware sparse J-space decomposition

The default solver now rescans the full eligible token dictionary against the current reconstruction error after every selected concept. It is no longer a one-shot candidate screen wearing an OMP nametag.

Available solvers:

  • stable-nn-omp: nonnegative residual-aware pursuit
  • stable-signed-omp: signed residual-aware pursuit
  • screened-nn-omp: retained legacy screened approximation

The report uses explained_energy_fraction, reconstruction cosine, relative error, effective sparsity, selected coefficients, and dictionary-pass history.

Probe taxonomy

v0.4 separates two artifacts that answer different questions:

  • ResidualConceptProbe: a discriminative direction learned directly in the residual space
  • JSpaceConcept: a residual probe projected into a sparse span of transported token directions

This makes it possible to measure how much of an abstract concept is captured by verbalizable J-space rather than quietly calling every linear probe a J-space concept.

Delayed causal readouts

Interventions now distinguish:

patch_positions
readout_positions

You can alter an earlier token state and measure a later answer position instead of being forced to patch and score the same token.

Explicit null distributions

A target intervention can be evaluated against multiple separately reported null families:

  • geometry-gaussian: isotropic directions matched to within-layer pair geometry
  • activation-covariance: directions sampled from the empirical activation covariance subspace
  • jspace-token: actual transported vocabulary directions, with cosine-matched pairs for swaps

Reports preserve each null family's assumptions. A small empirical p-value is only evidence relative to the stated null, not a universal badge of causality.

Joint residual/PLE factorial studies

For one concept, v0.4 can measure:

Residual channel PLE channel Condition
intact intact baseline
changed intact residual-only intervention
intact changed PLE-only intervention
changed changed joint intervention

It reports the interaction contrast:

joint - residual_only - ple_only + baseline

This can reveal redundancy, repair, or synergy between the two channels.

Typed rectangular storage

Exact residual, PLE, and typed-site Jacobians can be stored as independently hashed row tiles in safetensors. The format supports square and rectangular operators, verifies index and tile geometry, rejects traversal and duplicate-file metadata, and streams:

x @ J.T
J.T @ w

without loading a complete large matrix into RAM.

Declarative studies

JSON or YAML study files define a DAG of gemma-jlens commands. The runner:

  • validates dependencies and cycles
  • invokes commands with an argument vector, never a shell
  • constrains declared outputs beneath the study output directory
  • hashes resolved commands, dependencies, outputs, and environment identity
  • reuses only valid byte-identical completed nodes
  • persists failure and resume state atomically

Declared study outputs may be regular files or complete directory artifacts. Directory trees are hashed recursively, symbolic links are rejected, and cached nodes are reused only when every declared output remains byte-identical.

Supported model classes

Strongest expected support:

  • Gemma 3 text causal-LM checkpoints
  • Gemma 4 causal-LM and multimodal wrapper checkpoints in Hugging Face PyTorch
  • Gemma 4 E2B/E4B PLE-enabled text paths
  • Gemma 4 12B and 31B text paths
  • full fine-tunes retaining the architecture
  • active LoRA/PEFT adapters

Requires an explicit conformance and backward smoke test:

  • quantized or QAT kernels
  • custom model wrappers
  • automatic CPU/disk device maps
  • nonstandard attention implementations
  • Gemma 4 MoE variants

Not directly supported:

  • GGUF, llama.cpp, Ollama, MLX, or remote inference APIs
  • runtimes without a differentiable PyTorch graph and hookable decoder modules
  • image, audio, or video Jacobian tracing in v0.4

v0.4 is text- and PLE-channel focused. Multimodal causal tracing is deliberately left for a later release rather than shipped as an unvalidated tentacle.

Installation

Create an environment containing the correct PyTorch build for the host first. On ROCm, PyTorch still exposes AMD accelerators through the torch.cuda APIs.

python -m venv .venv
source .venv/bin/activate

# Install the correct CPU, CUDA, or ROCm PyTorch wheel first.
pip install -e '.[peft]'
hf auth login

Reference constraints and the exact CPU validation environment are under locks/. The ROCm and CUDA constraint files intentionally do not pretend one universal accelerator wheel exists.

First run

1. Exercise the offline instrument

gemma-jlens doctor --output out/00-doctor.json
gemma-jlens receipt --output out/00-environment.json

The v0.4 doctor exercises:

  • adapter discovery and final-logit conformance
  • a finite, nonzero residual VJP
  • typed-site fitting and VJP
  • delayed patch/readout intervention
  • residual-aware sparse recovery
  • structured prompt receipts
  • two explicit null families
  • lazy typed-site tiled storage
  • the frozen exact estimator bundled inside the package
  • optional parity selection of an explicitly injected external compatible core

2. Inspect supported hook sites

gemma-jlens site-list \
  google/gemma-4-E2B-it \
  --prompt-mode chat \
  --output out/01-sites.json

3. Prove a same-layer PLE path

gemma-jlens site-smoke \
  google/gemma-4-E2B-it \
  --prompt-mode chat \
  --source ple_input@12 \
  --target resid_post@12 \
  --output out/02-ple-same-layer-vjp.json

When SDPA forward succeeds but backward does not, repeat the same gate with:

--attn-implementation eager

Change one runtime variable at a time.

Local operational control panel

The control panel queues GPU work one job at a time and provides forms for:

  • hypothesis-first token-rank trajectories from a fitted residual lens
  • new prompts and interactive HTML readouts
  • steer, ablate, and swap interventions with matched random controls
  • resumable exact fits against one fixed corpus
  • retained logs, reports, lenses, and rendered artifacts

Bind it only to a trusted local or Tailscale address. The server invokes the CLI with argument vectors rather than a shell, confines generated paths beneath its workspace, and rejects cross-origin mutations.

gemma-jlens control-panel \
  google/gemma-4-E2B-it \
  out/e2b-residual.pt \
  out/corpora/wikitext103-validation-fit-v1.jsonl \
  --workspace out/control-panel \
  --host 100.118.137.30 --port 8765 \
  --device cuda --dtype bfloat16 --prompt-mode raw

Open http://100.118.137.30:8765/. The top-token grid is exploratory. Prefer the tracked-token trajectory: enter a concrete single-token hypothesis such as euro and compare its full-vocabulary rank across layers and positions.

A deterministic WikiText fitting corpus can be prepared without making PyArrow a runtime dependency:

uv run --no-project --with pyarrow --with huggingface-hub \
  scripts/prepare_wikitext_corpus.py \
  out/corpora/wikitext103-validation-fit-v1.jsonl \
  --count 120 --seed 20260726

The generated manifest pins the dataset revision, source digest, sampling rule, ordered prompt hash, and fit-only role. Keep evaluation and covariance-control rows disjoint.

Fit a typed-site lens

gemma-jlens site-fit \
  google/gemma-4-E2B-it \
  examples/prompts.txt \
  out/e2b-sites.pt \
  --prompt-mode raw \
  --sources resid_post@8,ple_input@12,resid_post@20 \
  --target final_norm_in \
  --dim-batch 2 \
  --checkpoint out/e2b-sites.fit.pt \
  --checkpoint-every 1 \
  --min-prompts 100 \
  --stop-at-delta 0.002 \
  --stop-window 10 \
  --report out/e2b-sites-fit.json

The fitter provides:

  • exact average Jacobians
  • fail-closed checkpoint identity
  • ordered corpus binding
  • atomic checkpoints
  • convergence history and stopping
  • automatic dim_batch reduction after an accelerator OOM
  • optional one-VJP preflight

A one-prompt fit proves plumbing, not scientific convergence.

Apply and intervene at typed sites

gemma-jlens site-apply \
  google/gemma-4-E2B-it \
  out/e2b-sites.pt \
  "The country shaped like a boot uses the currency" \
  --sources resid_post@20 \
  --positions -1 \
  --top-k 10 \
  --output out/site-readout.json

Patch one token and score a later token:

gemma-jlens site-intervene \
  google/gemma-4-E2B-it \
  out/e2b-sites.pt \
  "Italy is shaped like a boot. Its currency is" \
  --sources resid_post@20 \
  --kind steer \
  --token " euro" \
  --patch-positions -3 \
  --readout-positions -1 \
  --strength 1 \
  --output out/delayed-readout.json

Generate paired continuations while recomputing the full prefix at every step:

gemma-jlens site-generate \
  google/gemma-4-E2B-it \
  out/e2b-sites.pt \
  "Italy is shaped like a boot. Its currency is" \
  --sources resid_post@20 \
  --kind steer \
  --token " euro" \
  --patch-positions -1 \
  --strength 0.75 \
  --max-new-tokens 32 \
  --do-sample \
  --num-samples 16 \
  --temperature 0.8 \
  --seed 7 \
  --output out/paired-generation.json

Sampled baseline and intervention trajectories use common Gumbel noise. Each pair's causal horizon ends when either trajectory emits EOS, preventing padded tokens from becoming artificial continuation context. The full-prefix execution is slower than KV-cached generation, but preserves exact Python position semantics for arbitrary typed hook sites.

Sparse decomposition

The existing decompose command defaults to the residual-aware solver in v0.4:

gemma-jlens decompose \
  google/gemma-4-E2B-it \
  out/e2b-residual.pt \
  "Italy is shaped like a boot. Its currency is" \
  --layer 20 \
  --position -1 \
  --solver stable-nn-omp \
  --max-concepts 24 \
  --output-artifact out/currency.decomposition.pt \
  --report out/currency.decomposition.json

A sparse token explanation is an approximation. Inspect the unexplained remainder, fit-shard stability, and prompt-perturbation stability before treating a selected token as unique.

Build a J-space-constrained concept

First build a residual probe with concept-build, then project it through a fitted lens dictionary:

gemma-jlens concept-build-jspace \
  google/gemma-4-E2B-it \
  out/e2b-residual.pt \
  out/clarification.probe.pt \
  examples/concepts/clarification_positive_train.jsonl \
  examples/concepts/clarification_background_train.jsonl \
  out/clarification.jspace.pt \
  --solver stable-signed-omp \
  --max-concepts 24 \
  --report out/clarification-jspace-fit.json

Evaluate it on held-out prompts:

gemma-jlens concept-evaluate-jspace \
  google/gemma-4-E2B-it \
  out/clarification.jspace.pt \
  examples/concepts/clarification_positive_test.jsonl \
  examples/concepts/clarification_background_test.jsonl \
  --output out/clarification-jspace-eval.json

Multi-null causal study

gemma-jlens multi-null-study \
  google/gemma-4-E2B-it \
  out/e2b-residual.pt \
  "Italy is shaped like a boot. Its currency is" \
  --kind steer \
  --token " euro" \
  --layers 20 \
  --patch-positions -3 \
  --readout-positions -1 \
  --effect-token " euro" \
  --strengths=-2,-1,-0.5,0,0.5,1,2 \
  --null-families geometry-gaussian,jspace-token \
  --controls 100 \
  --output out/euro-multi-null.json

activation-covariance additionally requires --covariance-prompts. Control families are reported separately because they encode different exchangeability assumptions.

Joint residual/PLE factorial study

gemma-jlens joint-channel-study \
  google/gemma-4-E2B-it \
  out/e2b-residual.pt \
  out/e2b-ple.pt \
  "Italy is shaped like a boot. Its currency is" \
  --kind ablate \
  --token " euro" \
  --layers 20 \
  --patch-positions -3 \
  --readout-positions -1 \
  --effect-token " euro" \
  --controls 100 \
  --output out/euro-residual-ple-factorial.json

The residual and PLE control banks are matched within each channel. v0.4 does not claim they form a fully matched cross-channel trajectory null.

Pack a large lens into row tiles

gemma-jlens pack-tiled \
  out/31b-sites.pt \
  out/31b-sites.tiled \
  --tile-rows 512 \
  --dtype float16 \
  --verify \
  --report out/31b-sites-pack.json

gemma-jlens verify-tiled out/31b-sites.tiled \
  --output out/31b-sites-verify.json

The source lens manifest is copied and rebound to the tiled artifact. The tiled index checksum detects accidental metadata corruption; the manifest artifact hash is the external trust anchor.

Structured prompt receipts

gemma-jlens prompt-record \
  google/gemma-4-E2B-it \
  --messages '[{"role":"user","content":"What currency does Italy use?"}]' \
  --prompt-mode chat \
  --output out/italy.prompt.json

Chat truncation fails closed. Raw truncation also fails unless --allow-truncation is explicitly supplied, in which case the receipt records that decision and the final exact token IDs.

Declarative study DAG

gemma-jlens study-validate examples/study.yaml
gemma-jlens study-plan examples/study.yaml out/study \
  --output out/study-plan.json
gemma-jlens study-run examples/study.yaml out/study \
  --output out/study-run.json

See docs/STUDY_DAGS.md for the schema and placeholder rules.

Fine-tunes and adapters

A lens is checkpoint-specific. Fit separate lenses for:

  • base and instruction-tuned checkpoints
  • each full fine-tune
  • each active LoRA/PEFT adapter
  • materially different prompt/template modes
  • final and penultimate targets when comparing them

v0.4 manifests bind local model and adapter files, Hub revisions when available, tokenizer and processor state, ordered corpus hash, sites, runtime settings, and artifact identity. Deliberate stale use requires an explicit override and is labeled in output reports.

Command groups

The CLI currently exposes 56 commands across these groups:

  • model inspection, calibration, conformance, and smoke tests
  • residual, PLE, sketched, tuned, and typed-site fitting
  • lens application, rendering, comparison, and storage
  • token directions, sparse decomposition, and contextual concepts
  • interventions, multi-null studies, and residual/PLE factorial studies
  • prompt sharding, prompt receipts, provenance, and study DAG execution

Use:

gemma-jlens --help
gemma-jlens COMMAND --help

Validation status

The source release passes:

  • 140 automated tests with warnings promoted to errors
  • 76% total source statement coverage
  • dependency-free static checks across 93 Python files
  • bytecode compilation of source, tests, and scripts
  • all 55 CLI help surfaces
  • the v0.4 offline doctor
  • source archive, source distribution, and isolated wheel checks documented in RELEASE_VALIDATION.md

The build environment was CPU-only. The release does not claim successful execution of gated Gemma 4 weights, ROCm/CUDA backward kernels, a real LoRA, or a converged scientific fit. Those are explicit gates in docs/REAL_MODEL_VALIDATION.md.

Documentation

License

Gemma J-Lens is released under Apache License 2.0. The frozen Anthropic-compatible exact core is derived from an Apache-2.0 reference implementation; see THIRD_PARTY_NOTICES.md.