- Python 97.8%
- Shell 1.2%
- HTML 0.9%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .github/workflows | ||
| containers | ||
| docs | ||
| examples | ||
| locks | ||
| scripts | ||
| src/gemma_jlens | ||
| tests | ||
| .dockerignore | ||
| .gitignore | ||
| CAUSAL_STUDIES.md | ||
| CHANGELOG.md | ||
| gemma-jacobian-lens-0.4.0-source.zip | ||
| gemma_jlens-0.4.0-py3-none-any.whl | ||
| gemma_jlens-0.4.0.tar.gz | ||
| JSPACE_DECOMPOSITION.md | ||
| LICENSE | ||
| MANIFEST.in | ||
| OFFLINE_DOCTOR.json | ||
| PLE_CHANNELS.md | ||
| pyproject.toml | ||
| README.md | ||
| REAL_MODEL_VALIDATION.md | ||
| RELEASE_MANIFEST.json | ||
| RELEASE_NOTES.md | ||
| RELEASE_VALIDATION.md | ||
| REPRODUCIBILITY.md | ||
| SHA256SUMS.txt | ||
| STRIX_HALO_RUNBOOK.md | ||
| STUDY_DAGS.md | ||
| THIRD_PARTY_NOTICES.md | ||
| TILED_STORAGE.md | ||
| TYPED_HOOK_SITES.md | ||
| VALIDATION_LADDER.md | ||
Gemma J-Lens v0.4
A checkpoint-aware Jacobian-lens, typed-site causal tracing, and J-space research toolkit for Gemma-family Hugging Face models.
v0.4 turns the project from a collection of strong research primitives into a coherent causal-tracing system. It adds explicit hook-site addresses, same-layer Per-Layer Embedding analysis, residual-aware sparse decomposition, delayed readouts, multiple null distributions, joint residual/PLE factorial studies, structured prompt receipts, lazy rectangular lens storage, and resumable study DAGs.
The exact estimator remains Anthropic's average input-output Jacobian:
lens_s_to_t(x_s) = decode_t(J_s_to_t @ x_s)
J_s_to_t = E[d x_t / d x_s]
For a conventional residual-to-final lens this reduces to:
lens_l(h_l) = unembed(J_l @ h_l)
J_l = E[d h_target / d h_l]
The exact numerical core is a frozen, Apache-2.0-compatible copy of Anthropic's reference implementation at commit:
581d398613e5602a5af361e1c34d3a92ea82ba8e
It ships inside the wheel, so installation and offline validation do not depend
on a live Git repository. The vendored implementation is the deterministic
default. Advanced parity testing can explicitly select an external compatible
jlens package with GEMMA_JLENS_CORE=external, or an intentionally injected
module with GEMMA_JLENS_CORE=auto. Merely installing or importing another
package does not change the estimator in the default mode.
What is new in v0.4
Typed causal locations
A layer number alone is ambiguous. v0.4 addresses activations by both layer and site:
resid_pre@12
attn_out@12
resid_mid@12
mlp_out@12
ple_input@12
resid_post@12
final_norm_in
final_norm_out
This permits exact questions such as:
ple_input@12 -> resid_post@12
attn_out@8 -> final_norm_in
resid_pre@4 -> resid_post@20
The causal validator rejects reverse, identical, and unsupported paths before an expensive fit begins.
Same-layer PLE analysis
Gemma 4 E2B/E4B can inject an auxiliary Per-Layer Embedding input into a decoder block. v0.4 can measure the direct same-layer derivative:
J = d resid_post@L / d ple_input@L
as well as longer PLE-to-final paths. The result is a conditional derivative with other inputs held fixed. It is not automatically an additive attribution of the whole model computation.
Residual-aware sparse J-space decomposition
The default solver now rescans the full eligible token dictionary against the current reconstruction error after every selected concept. It is no longer a one-shot candidate screen wearing an OMP nametag.
Available solvers:
stable-nn-omp: nonnegative residual-aware pursuitstable-signed-omp: signed residual-aware pursuitscreened-nn-omp: retained legacy screened approximation
The report uses explained_energy_fraction, reconstruction cosine, relative
error, effective sparsity, selected coefficients, and dictionary-pass history.
Probe taxonomy
v0.4 separates two artifacts that answer different questions:
ResidualConceptProbe: a discriminative direction learned directly in the residual spaceJSpaceConcept: a residual probe projected into a sparse span of transported token directions
This makes it possible to measure how much of an abstract concept is captured by verbalizable J-space rather than quietly calling every linear probe a J-space concept.
Delayed causal readouts
Interventions now distinguish:
patch_positions
readout_positions
You can alter an earlier token state and measure a later answer position instead of being forced to patch and score the same token.
Explicit null distributions
A target intervention can be evaluated against multiple separately reported null families:
geometry-gaussian: isotropic directions matched to within-layer pair geometryactivation-covariance: directions sampled from the empirical activation covariance subspacejspace-token: actual transported vocabulary directions, with cosine-matched pairs for swaps
Reports preserve each null family's assumptions. A small empirical p-value is only evidence relative to the stated null, not a universal badge of causality.
Joint residual/PLE factorial studies
For one concept, v0.4 can measure:
| Residual channel | PLE channel | Condition |
|---|---|---|
| intact | intact | baseline |
| changed | intact | residual-only intervention |
| intact | changed | PLE-only intervention |
| changed | changed | joint intervention |
It reports the interaction contrast:
joint - residual_only - ple_only + baseline
This can reveal redundancy, repair, or synergy between the two channels.
Typed rectangular storage
Exact residual, PLE, and typed-site Jacobians can be stored as independently
hashed row tiles in safetensors. The format supports square and rectangular
operators, verifies index and tile geometry, rejects traversal and duplicate-file
metadata, and streams:
x @ J.T
J.T @ w
without loading a complete large matrix into RAM.
Declarative studies
JSON or YAML study files define a DAG of gemma-jlens commands. The runner:
- validates dependencies and cycles
- invokes commands with an argument vector, never a shell
- constrains declared outputs beneath the study output directory
- hashes resolved commands, dependencies, outputs, and environment identity
- reuses only valid byte-identical completed nodes
- persists failure and resume state atomically
Declared study outputs may be regular files or complete directory artifacts. Directory trees are hashed recursively, symbolic links are rejected, and cached nodes are reused only when every declared output remains byte-identical.
Supported model classes
Strongest expected support:
- Gemma 3 text causal-LM checkpoints
- Gemma 4 causal-LM and multimodal wrapper checkpoints in Hugging Face PyTorch
- Gemma 4 E2B/E4B PLE-enabled text paths
- Gemma 4 12B and 31B text paths
- full fine-tunes retaining the architecture
- active LoRA/PEFT adapters
Requires an explicit conformance and backward smoke test:
- quantized or QAT kernels
- custom model wrappers
- automatic CPU/disk device maps
- nonstandard attention implementations
- Gemma 4 MoE variants
Not directly supported:
- GGUF, llama.cpp, Ollama, MLX, or remote inference APIs
- runtimes without a differentiable PyTorch graph and hookable decoder modules
- image, audio, or video Jacobian tracing in v0.4
v0.4 is text- and PLE-channel focused. Multimodal causal tracing is deliberately left for a later release rather than shipped as an unvalidated tentacle.
Installation
Create an environment containing the correct PyTorch build for the host first.
On ROCm, PyTorch still exposes AMD accelerators through the torch.cuda APIs.
python -m venv .venv
source .venv/bin/activate
# Install the correct CPU, CUDA, or ROCm PyTorch wheel first.
pip install -e '.[peft]'
hf auth login
Reference constraints and the exact CPU validation environment are under
locks/. The ROCm and CUDA constraint files intentionally do not
pretend one universal accelerator wheel exists.
First run
1. Exercise the offline instrument
gemma-jlens doctor --output out/00-doctor.json
gemma-jlens receipt --output out/00-environment.json
The v0.4 doctor exercises:
- adapter discovery and final-logit conformance
- a finite, nonzero residual VJP
- typed-site fitting and VJP
- delayed patch/readout intervention
- residual-aware sparse recovery
- structured prompt receipts
- two explicit null families
- lazy typed-site tiled storage
- the frozen exact estimator bundled inside the package
- optional parity selection of an explicitly injected external compatible core
2. Inspect supported hook sites
gemma-jlens site-list \
google/gemma-4-E2B-it \
--prompt-mode chat \
--output out/01-sites.json
3. Prove a same-layer PLE path
gemma-jlens site-smoke \
google/gemma-4-E2B-it \
--prompt-mode chat \
--source ple_input@12 \
--target resid_post@12 \
--output out/02-ple-same-layer-vjp.json
When SDPA forward succeeds but backward does not, repeat the same gate with:
--attn-implementation eager
Change one runtime variable at a time.
Local operational control panel
The control panel queues GPU work one job at a time and provides forms for:
- hypothesis-first token-rank trajectories from a fitted residual lens
- new prompts and interactive HTML readouts
- steer, ablate, and swap interventions with matched random controls
- resumable exact fits against one fixed corpus
- retained logs, reports, lenses, and rendered artifacts
Bind it only to a trusted local or Tailscale address. The server invokes the CLI with argument vectors rather than a shell, confines generated paths beneath its workspace, and rejects cross-origin mutations.
gemma-jlens control-panel \
google/gemma-4-E2B-it \
out/e2b-residual.pt \
out/corpora/wikitext103-validation-fit-v1.jsonl \
--workspace out/control-panel \
--host 100.118.137.30 --port 8765 \
--device cuda --dtype bfloat16 --prompt-mode raw
Open http://100.118.137.30:8765/. The top-token grid is exploratory. Prefer the
tracked-token trajectory: enter a concrete single-token hypothesis such as
euro and compare its full-vocabulary rank across layers and positions.
A deterministic WikiText fitting corpus can be prepared without making PyArrow a runtime dependency:
uv run --no-project --with pyarrow --with huggingface-hub \
scripts/prepare_wikitext_corpus.py \
out/corpora/wikitext103-validation-fit-v1.jsonl \
--count 120 --seed 20260726
The generated manifest pins the dataset revision, source digest, sampling rule, ordered prompt hash, and fit-only role. Keep evaluation and covariance-control rows disjoint.
Fit a typed-site lens
gemma-jlens site-fit \
google/gemma-4-E2B-it \
examples/prompts.txt \
out/e2b-sites.pt \
--prompt-mode raw \
--sources resid_post@8,ple_input@12,resid_post@20 \
--target final_norm_in \
--dim-batch 2 \
--checkpoint out/e2b-sites.fit.pt \
--checkpoint-every 1 \
--min-prompts 100 \
--stop-at-delta 0.002 \
--stop-window 10 \
--report out/e2b-sites-fit.json
The fitter provides:
- exact average Jacobians
- fail-closed checkpoint identity
- ordered corpus binding
- atomic checkpoints
- convergence history and stopping
- automatic
dim_batchreduction after an accelerator OOM - optional one-VJP preflight
A one-prompt fit proves plumbing, not scientific convergence.
Apply and intervene at typed sites
gemma-jlens site-apply \
google/gemma-4-E2B-it \
out/e2b-sites.pt \
"The country shaped like a boot uses the currency" \
--sources resid_post@20 \
--positions -1 \
--top-k 10 \
--output out/site-readout.json
Patch one token and score a later token:
gemma-jlens site-intervene \
google/gemma-4-E2B-it \
out/e2b-sites.pt \
"Italy is shaped like a boot. Its currency is" \
--sources resid_post@20 \
--kind steer \
--token " euro" \
--patch-positions -3 \
--readout-positions -1 \
--strength 1 \
--output out/delayed-readout.json
Generate paired continuations while recomputing the full prefix at every step:
gemma-jlens site-generate \
google/gemma-4-E2B-it \
out/e2b-sites.pt \
"Italy is shaped like a boot. Its currency is" \
--sources resid_post@20 \
--kind steer \
--token " euro" \
--patch-positions -1 \
--strength 0.75 \
--max-new-tokens 32 \
--do-sample \
--num-samples 16 \
--temperature 0.8 \
--seed 7 \
--output out/paired-generation.json
Sampled baseline and intervention trajectories use common Gumbel noise. Each pair's causal horizon ends when either trajectory emits EOS, preventing padded tokens from becoming artificial continuation context. The full-prefix execution is slower than KV-cached generation, but preserves exact Python position semantics for arbitrary typed hook sites.
Sparse decomposition
The existing decompose command defaults to the residual-aware solver in v0.4:
gemma-jlens decompose \
google/gemma-4-E2B-it \
out/e2b-residual.pt \
"Italy is shaped like a boot. Its currency is" \
--layer 20 \
--position -1 \
--solver stable-nn-omp \
--max-concepts 24 \
--output-artifact out/currency.decomposition.pt \
--report out/currency.decomposition.json
A sparse token explanation is an approximation. Inspect the unexplained remainder, fit-shard stability, and prompt-perturbation stability before treating a selected token as unique.
Build a J-space-constrained concept
First build a residual probe with concept-build, then project it through a fitted
lens dictionary:
gemma-jlens concept-build-jspace \
google/gemma-4-E2B-it \
out/e2b-residual.pt \
out/clarification.probe.pt \
examples/concepts/clarification_positive_train.jsonl \
examples/concepts/clarification_background_train.jsonl \
out/clarification.jspace.pt \
--solver stable-signed-omp \
--max-concepts 24 \
--report out/clarification-jspace-fit.json
Evaluate it on held-out prompts:
gemma-jlens concept-evaluate-jspace \
google/gemma-4-E2B-it \
out/clarification.jspace.pt \
examples/concepts/clarification_positive_test.jsonl \
examples/concepts/clarification_background_test.jsonl \
--output out/clarification-jspace-eval.json
Multi-null causal study
gemma-jlens multi-null-study \
google/gemma-4-E2B-it \
out/e2b-residual.pt \
"Italy is shaped like a boot. Its currency is" \
--kind steer \
--token " euro" \
--layers 20 \
--patch-positions -3 \
--readout-positions -1 \
--effect-token " euro" \
--strengths=-2,-1,-0.5,0,0.5,1,2 \
--null-families geometry-gaussian,jspace-token \
--controls 100 \
--output out/euro-multi-null.json
activation-covariance additionally requires --covariance-prompts. Control
families are reported separately because they encode different exchangeability
assumptions.
Joint residual/PLE factorial study
gemma-jlens joint-channel-study \
google/gemma-4-E2B-it \
out/e2b-residual.pt \
out/e2b-ple.pt \
"Italy is shaped like a boot. Its currency is" \
--kind ablate \
--token " euro" \
--layers 20 \
--patch-positions -3 \
--readout-positions -1 \
--effect-token " euro" \
--controls 100 \
--output out/euro-residual-ple-factorial.json
The residual and PLE control banks are matched within each channel. v0.4 does not claim they form a fully matched cross-channel trajectory null.
Pack a large lens into row tiles
gemma-jlens pack-tiled \
out/31b-sites.pt \
out/31b-sites.tiled \
--tile-rows 512 \
--dtype float16 \
--verify \
--report out/31b-sites-pack.json
gemma-jlens verify-tiled out/31b-sites.tiled \
--output out/31b-sites-verify.json
The source lens manifest is copied and rebound to the tiled artifact. The tiled index checksum detects accidental metadata corruption; the manifest artifact hash is the external trust anchor.
Structured prompt receipts
gemma-jlens prompt-record \
google/gemma-4-E2B-it \
--messages '[{"role":"user","content":"What currency does Italy use?"}]' \
--prompt-mode chat \
--output out/italy.prompt.json
Chat truncation fails closed. Raw truncation also fails unless
--allow-truncation is explicitly supplied, in which case the receipt records
that decision and the final exact token IDs.
Declarative study DAG
gemma-jlens study-validate examples/study.yaml
gemma-jlens study-plan examples/study.yaml out/study \
--output out/study-plan.json
gemma-jlens study-run examples/study.yaml out/study \
--output out/study-run.json
See docs/STUDY_DAGS.md for the schema and placeholder
rules.
Fine-tunes and adapters
A lens is checkpoint-specific. Fit separate lenses for:
- base and instruction-tuned checkpoints
- each full fine-tune
- each active LoRA/PEFT adapter
- materially different prompt/template modes
- final and penultimate targets when comparing them
v0.4 manifests bind local model and adapter files, Hub revisions when available, tokenizer and processor state, ordered corpus hash, sites, runtime settings, and artifact identity. Deliberate stale use requires an explicit override and is labeled in output reports.
Command groups
The CLI currently exposes 56 commands across these groups:
- model inspection, calibration, conformance, and smoke tests
- residual, PLE, sketched, tuned, and typed-site fitting
- lens application, rendering, comparison, and storage
- token directions, sparse decomposition, and contextual concepts
- interventions, multi-null studies, and residual/PLE factorial studies
- prompt sharding, prompt receipts, provenance, and study DAG execution
Use:
gemma-jlens --help
gemma-jlens COMMAND --help
Validation status
The source release passes:
- 140 automated tests with warnings promoted to errors
- 76% total source statement coverage
- dependency-free static checks across 93 Python files
- bytecode compilation of source, tests, and scripts
- all 55 CLI help surfaces
- the v0.4 offline doctor
- source archive, source distribution, and isolated wheel checks documented in
RELEASE_VALIDATION.md
The build environment was CPU-only. The release does not claim successful
execution of gated Gemma 4 weights, ROCm/CUDA backward kernels, a real LoRA, or a
converged scientific fit. Those are explicit gates in
docs/REAL_MODEL_VALIDATION.md.
Documentation
- Design and scientific boundaries
- Control panel and interpretation guide
- Global-workspace replication gates
- Stock Gemma 4 12B SolarKyle reproduction
- Typed hook sites
- Causal studies and nulls
- J-space decomposition
- PLE channels
- Structured prompts and study DAGs
- Tiled storage
- Experiment protocol
- Reproducibility
- Real-model validation gates
- Strix Halo runbook
- Validation ladder
License
Gemma J-Lens is released under Apache License 2.0. The frozen Anthropic-compatible
exact core is derived from an Apache-2.0 reference implementation; see
THIRD_PARTY_NOTICES.md.