Roadmap¶
This page is the backlog for pulsatrix after v1.0. It says what we plan to build, in what order,
and why. Every item has an ID (for example FND-1) so issues, pull requests and chronicles can
point at it. The evidence behind each item, including its sources, how we would know it was a
mistake, and how it is known to fail, lives in Research Notes.
Nothing here is a promise. Priorities move as items land and as measurements come in.
Where v1.0 stands¶
v1.0 is strong on explainability and thin on the foundations that larger models need.
- Strong. LRP covers every layer and is checked against Zennit and LXT. Saliency, Integrated Gradients and Grad-CAM are checked against Captum. LIME, KernelSHAP, PDP, logit lens, activation patching, circuit graphs, probes and a basic sparse autoencoder are all in.
- Missing. There is no way to save or load model weights, no import from PyTorch or Hugging
Face, and everything is
float32. Tensors have no views or broadcasting. There are no learning-rate schedulers, no AdamW and no LoRA. Attention has no mask. - Done since v1.0. The whole FND epic: named parameters, freezing, top-k, eigensolver/QR/SVD, BatchNorm eval mode and folding, Conv2D stride and padding, seeding and deterministic mode, and device checks.
- HIP backend. It works on gfx1151 (Strix Halo), but every op synchronizes the stream,
reductions run one thread per row or channel, and every tensor is a raw
hipMalloc.
The goal that orders this list¶
Explain real pretrained models. That means you can load a Hugging Face model (a small LLM such as SmolLM2-135M, or a vision model such as ResNet18), fine-tune it with LoRA on a Strix Halo machine, and explain it with the tools pulsatrix already has. Items on that path come first.
How to read this page¶
- ID. The epic prefix plus a number, for example
TRN-7. - Priority. P0 means the next release depends on it. P1 means it's important and planned. P2 means valuable but it can wait. P3 means maybe, or only if someone needs it.
- Effort. S is under a week, M is one to three weeks, L is one to two months, and XL is longer than that. These are rough and assume one person.
- Depends on. The IDs that have to land first.
Rules that apply to every item¶
- Every new interpretability tool ships with baselines. Run it on a randomly initialized copy of the model and compare it with a plain linear probe. Attribution maps, probes and sparse autoencoders all produce plausible-looking output on random networks, so a result that looks the same on a random network means nothing.
- Train in bf16, explain in fp32. Once mixed precision exists, LRP, Integrated Gradients and the conservation checks still run in fp32. LRP's stabilized divisions and conservation tests don't survive bf16's 8-bit mantissa.
- Never unpickle in C++.
.pt,.pthand.pklfiles are code, not data. They go through an optional converter that uses PyTorch's restricted loader. The C++ side reads safetensors only. - Every importer passes a golden-logit test. "It loaded" is not a test. Compare logits against Hugging Face Transformers on real weights and real text. A wrong RoPE layout loads fine and produces fluent-looking nonsense.
- Every new module implements
propagate_relevanceand has a conservation test. This is what makes pulsatrix pulsatrix. - Python is optional. Agents, the MCP server, notebook output and model loading for safetensors checkpoints are native C++. Python appears only in the legacy-pickle converter and in optional notebook conveniences on top of the existing bindings.
Where each requested area went¶
| Area | Epics |
|---|---|
| 1. New tools: embedding analysis, Goodfire BSF, surrogates, student-teacher | INT, TDA, FEAT, KD |
| 2. Fine-tuning methods | TRN |
| 3. Translating model files | IO |
| 4. Visualization pack | VIZ, NB |
| 5. LLMs and agent orchestration | LLM, AGT |
| 6. Built-in XAI framework (4 reasons, 9 question categories) | XAI |
| 7. New patterns and architectures | ARCH |
| 8. Kitchen sink | KS, FND |
| 9. HIP efficiency for training and tuning | HIP |
Milestones¶
Each milestone lists the items it contains. The order follows the dependencies.
v1.1 "Foundations and trust"¶
The plumbing everything else needs, plus the checks that keep explanations honest.
- FND-1 to FND-8 (done): named parameters, freezing, top-k, eigensolver, BatchNorm eval mode, Conv2D stride and padding, seeding, device checks
- IO-1, IO-2: safetensors and the native checkpoint format
- TRN-1 to TRN-6: parameter groups, AdamW, clipping, schedulers, gradient accumulation, full fine-tuning
- XAI-5, XAI-6: explanation-quality metrics and the random-model baseline harness
- HIP-1, HIP-2, HIP-4, HIP-5: profiling, parallel reductions, fewer syncs, multi-block
dot/sum - VIZ-1, VIZ-2: the JSON export format and the SVG renderer
- KS-1, KS-2:
install()and a benchmark suite - HIP-9: a ROCm 10.0 evaluation image and a pinned host kernel
v1.2 "Real pretrained models"¶
Load SmolLM2-135M and ResNet18, run them, and match the reference implementations.
- IO-3 to IO-6: the pickle converter, name mapping, Hugging Face configs, bf16 upcast
- LLM-1 to LLM-7: attention upgrade, tied LM head, native tokenizer, generation, KV cache, the golden-logit harness, AttnLRP parity with LXT
- HIP-3, HIP-6, HIP-7: caching allocator, fused kernels, bounded-memory Conv2D
- VIZ-3, VIZ-6a: Vega-Lite HTML and the token relevance view
- NB-1, NB-2: native rich display and the
.ipynb/HTML report writer - AGT-1 to AGT-4: the native orchestrator core
- KS-8, KS-9: GPU CI on gfx1151 and a model zoo (ResNet18, SmolLM2)
v1.3 "Tune and explain"¶
Fine-tune the models from v1.2 and explain what the fine-tuning changed.
- TRN-7 to TRN-11: LoRA, LRP through LoRA, LoReFT, IA³, model diffing
- KD-1, KD-2: distillation and teacher–student explanation agreement
- XAI-1 to XAI-4: the question-driven planner, explanation reports, counterfactuals, surrogates, Anchors
- INT-1 to INT-6: neighbors, CKA, probe controls, intrinsic dimension, PCA, TCAV and CRP
- TDA-1 to TDA-8, TDA-11: persistent homology and RTD
- KS-3, KS-4: uncertainty and adversarial robustness
- Llama-3.2-1B and Qwen2.5-0.5B pass the golden-logit test
v1.4 "Features and agents"¶
Feature discovery on real models, and agents that can drive the library.
- FEAT-1 to FEAT-7: the featurizer interface, TopK and its family, metrics, transcoders, BSF, steering
- TDA-9, TDA-10, TDA-12: topology views, the manifold verifier for featurizers, layer-wise topology
- VIZ-4, VIZ-6b, VIZ-6c: Neuronpedia export, feature dashboards, the embedding projector
- AGT-5, AGT-8: the native MCP server and automated feature descriptions
- NB-3: Python notebook display
- LLM-8, LLM-9: tuned lens, AtP*, Gemma 3
- KS-5: TracIn data attribution
- IO-7:
.npy/.npz, for published SAE dictionaries - ARCH-1 to ARCH-3: B-cos layers, concept bottleneck models, mixture of experts
- HIP-8: hipBLASLt and bf16 GEMM
Backlog, not yet scheduled¶
These P2 and P3 items wait until a milestone needs them or someone asks: TRN-12, TRN-13, IO-10, IO-11, INT-7, INT-8, TDA-13 to TDA-19, KD-3, KD-4, FEAT-8, ARCH-4 to ARCH-10, VIZ-7, HIP-10, AGT-6, AGT-7, AGT-9, NB-4, NB-5, KS-6 and KS-7.
v2.0¶
Larger changes that touch every module or backend.
- A dtype system and full bf16 training (HIP-11)
- Strided views and broadcasting (KS-10)
- Op-level autograd (KS-12)
- WMMA, HIP graphs and FlashAttention for training only (HIP-12)
- Attribution graphs and parameter decomposition (FEAT-9, FEAT-10)
- ARCH items at P1 and P2
- ONNX, GGUF and QLoRA (IO-8, IO-9, TRN-14)
- Python wheels and a vcpkg port (KS-11)
FND: Foundations¶
These unblock most of the other epics. All eight are done (2026-10-04, PRs #33 to #41).
| ID | Item | Why | Depends on | P | Effort | Status |
|---|---|---|---|---|---|---|
| FND-1 | named_parameters() with hierarchical names (blocks.3.attn.q_proj.weight) |
Saving, loading, freezing by name, LoRA targeting and Hugging Face name mapping all need names | — | P0 | M | Done, #33 |
| FND-2 | Parameter freezing: a per-parameter requires_grad flag, and a "skip the weight gradient, still compute the input gradient" branch in every backward |
Fine-tuning and LoRA. Skipping dW is where the memory and compute savings come from |
FND-1 | P0 | M | Done, #34 |
| FND-3 | Top-k selection, on the host first and then on the device | TopK and BSF featurizers, MoE routers, top-k sampling, nearest neighbors, landmarks | — | P0 | M | Done, #35 |
| FND-4 | Symmetric eigensolver, power iteration, and a small SVD/QR (Jacobi) | PCA, BSF stable rank, LoReFT's orthonormal projection, PiSSA, intruder-dimension checks, spectral distances | — | P0 | M | Done, #36 (see below) |
| FND-5 | BatchNorm running statistics and an eval mode, plus a canonizer that folds BatchNorm into the previous layer before LRP (plans/lrp_issues.md #8) |
Today a sample's explanation depends on the rest of its batch | — | P0 | S | Done, #37 |
| FND-6 | Conv2D stride and padding | Needed to load VGG and ResNet, the standard LRP benchmark models | — | P0 | M | Done, #38 |
| FND-7 | A seeding and determinism API (set_seed, a flag that forbids nondeterministic paths) |
Reproducible explanations. Most of the pieces already exist | — | P0 | S | Done, #39 (see below) |
| FND-8 | Device-consistency checks at module and loss boundaries, and fix Tensor::Stack tagging (plans/gpu_review.md #1, #2) |
A CPU tensor fed to a GPU module aborts the process on gfx1151 | — | P0 | S | Done, #40 |
How the FND work departed from the plan¶
- FND-4. The eigensolver is Householder tridiagonalization plus implicit-shift QL, not Jacobi. Cyclic Jacobi took 12.7 s at n = 512, and PCA on SmolLM2's 576-wide activations would hit that on every call; QL takes 0.77 s. The SVD is one-sided Jacobi, as planned.
- FND-7. Deterministic mode is on by default, unlike PyTorch. It forbids hipBLAS and cuBLAS atomics.
Follow-ups the FND work surfaced¶
| Follow-up | Found in | Belongs with |
|---|---|---|
Overlapping and padded MaxPool2DModule, and an adaptive average pool: the ResNet stem uses MaxPool2d(3, stride 2, padding 1) |
FND-6 | KS-9 (model zoo) |
A named_buffers() so checkpoints can save BatchNorm running statistics, which aren't parameters |
FND-5 | IO-2 |
BatchNormModule initializes gamma to 0; PyTorch uses 1 |
FND-5 | IO-5 (configs) or a separate fix |
std::normal_distribution and friends differ between libstdc++ and MSVC, so seeded runs match on one platform only |
FND-7 | KS-1, or its own item |
named_parameters() and set_requires_grad() in the Python bindings |
FND-1, FND-2 | NB-3 |
| A device top-k for large k (radix or bitonic select); the current kernel is O(cols·k) per row | FND-3 | FEAT-2 if k grows |
GaussianMutationTest.ZeroSigma… aborts in Debug: std::normal_distribution rejects σ = 0 under libstdc++ assertions |
found while testing | a bug fix |
IO: Serialization and model import¶
One format does everything: safetensors. It is a small JSON header followed by raw little-endian data, it cannot run code, and nearly every Hugging Face model ships in it. It covers pulsatrix's own checkpoints, optimizer state, LoRA adapters, imported models, published sparse autoencoder dictionaries and a model zoo.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| IO-1 | Native safetensors reader and writer. Check every offset against the file size, reject overlaps and holes, and use overflow-safe size arithmetic | The single file format | — | P0 | S–M |
| IO-2 | Native checkpoint format: safetensors with format_version metadata, optimizer state in a sibling file, and a migration table between versions |
Save and resume training; the passing test is a bit-identical forward pass and loss curve after reload | IO-1, FND-1 | P0 | M |
| IO-3 | An optional converter for legacy pickle files (.pt, .pth, .pkl). It uses torch.load(weights_only=True) with torch 2.6 or newer and writes safetensors |
Pickle is code, so it stays out of C++. Not needed for models that already ship safetensors | IO-1 | P1 | M |
| IO-4 | A name-mapping manifest with transforms: transpose (pulsatrix Linear stores (in, out), PyTorch stores (out, in)), RoPE layout permutation, splitting fused QKV, weight tying. Strict mode fails on unmapped or extra keys |
Turns Hugging Face names and layouts into pulsatrix modules | FND-1, IO-1 | P1 | M |
| IO-5 | Read Hugging Face config.json and sharded model.safetensors.index.json |
Every small LLM on the Hub uses these | IO-1 | P1 | S |
| IO-6 | Upcast bf16 and fp16 weights to fp32 on load (exact) | Most published weights are bf16 | IO-1 | P1 | S |
| IO-7 | .npy and .npz reading, with object and big-endian dtypes rejected |
Many SAE dictionaries and Python users' arrays | — | P1 | S–M |
| IO-8 | ONNX import, weights only | Graph import would mean pattern-matching ONNX ops back into modules, which is research | IO-1 | P3 | M |
| IO-9 | GGUF import for F32, F16 and Q8_0, with every count and size checked for overflow | Some small models only exist as GGUF. GGUF parsers have a long CVE history | IO-1 | P3 | M |
| IO-10 | Keras and TensorFlow weights, through the Python converter only. Never call load_model on an untrusted file |
Small user base, repeated safe_mode bypasses |
IO-3 | P3 | M |
| IO-11 | joblib/sklearn pickles and TorchScript: not supported. Document how to export from them instead | TorchScript is deprecated upstream. Pickle is code | — | P3 | — |
TRN: Training and fine-tuning¶
The infrastructure comes first. Most LoRA failures in the literature are configuration problems, not method problems: a study of LoRA variants from January 2026 found that, once the learning rate is tuned, plain LoRA matches or beats most of them.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| TRN-1 | Optimizer parameter groups, each with its own learning rate and weight decay | No weight decay on norms and biases; LoRA+ | FND-1 | P0 | S |
| TRN-2 | AdamW (decoupled weight decay) and SGD with momentum and Nesterov | The default for transformer fine-tuning | TRN-1 | P0 | S |
| TRN-3 | Global gradient-norm clipping | Stability at LoRA's higher learning rates | — | P0 | S |
| TRN-4 | Learning-rate schedulers: warmup, cosine, linear, constant | Every published recipe assumes them | — | P0 | S |
| TRN-5 | Gradient accumulation that divides by the total token count, not the mean of micro-batch means | Averaging micro-batch means is wrong when sequence lengths vary | — | P0 | S |
| TRN-6 | A full fine-tuning recipe on a small pretrained model | The baseline every other method is compared against | FND-2, IO-2 | P0 | S |
| TRN-7 | LoRALinear: a frozen base weight plus a low-rank update, with merge and unmerge, adapter save and load using PEFT key names, and PEFT's defaults. Documentation follows "LoRA Without Regret": apply it to all layers, use about 10× the full fine-tuning learning rate |
The main fine-tuning method | FND-2, IO-1 | P1 | M |
| TRN-8 | LRP through LoRA that splits relevance between the base path and the adapter path. Document that the gamma and z+ rules give different results for merged and unmerged adapters | Shows what the fine-tuning changed, per input. No other library does this | TRN-7 | P1 | M |
| TRN-9 | LoReFT (representation fine-tuning) and LRP through it. The intervention is affine, so the epsilon rule is exact | Fine-tuning that is itself an interpretable subspace | FND-4 | P1 | M |
| TRN-10 | IA³ (learned rescaling vectors) | Tiny and trivially explainable | FND-2 | P2 | S |
| TRN-11 | Model diffing: relevance of the fine-tuned model minus relevance of the base model, plus an intruder-dimension check on the weights | Shows what a fine-tune changed across a dataset | TRN-8, FND-4 | P2 | M |
| TRN-12 | Activation checkpointing per block. Off in explain mode, because LRP needs the cached activations | Memory for 1B-parameter models | — | P2 | M |
| TRN-13 | rsLoRA, LoRA+ and DoRA | Small gains that a tuned learning rate often matches | TRN-7 | P2 | S–M |
| TRN-14 | PiSSA, VeRA, AdaLoRA, GaLore, QLoRA, prefix and prompt tuning, BitFit, adapter layers, and 2025–26 LoRA variants | Low value per the 2026 variant study, or blocked by missing dtypes | TRN-7 | P3 | — |
LLM: Running real language models¶
Target order: SmolLM2-135M, then Qwen2.5-0.5B, Llama-3.2-1B, Qwen3-0.6B and Gemma 3 270M. SmolLM2 needs the fewest changes. Gemma 3 has the most quirks, but Google's Gemma Scope 2 publishes sparse autoencoders and transcoders for every layer, so it pays off for the FEAT epic.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| LLM-1 | Attention upgrade: grouped-query attention (num_kv_heads), head_dim separate from d_model / num_heads, optional QKV bias, causal and padding masks, a RoPE layout flag (Hugging Face "rotate half" vs. adjacent pairs), and a position offset |
Without these, no current small LLM loads | — | P0 | M |
| LLM-2 | A tied LM head that shares the embedding matrix | SmolLM2, Qwen3 and Gemma 3 all tie their embeddings | — | P0 | S |
| LLM-3 | A native byte-level BPE tokenizer that reads Hugging Face tokenizer.json, with a hand-written pre-tokenizer and a generated Unicode category table (std::regex can't match Unicode categories). It passes when its ids match Hugging Face on a 10,000-line multilingual corpus |
No Python needed to tokenize | — | P0 | M |
| LLM-4 | Generation: greedy, temperature, top-k and top-p sampling, seeded, with EOS handling | Run the model, not just score it | LLM-1, FND-3 | P1 | S |
| LLM-5 | A preallocated KV cache | Generation without recomputing the whole prefix | LLM-1 | P1 | M |
| LLM-6 | Golden-logit harness: compare against Hugging Face on real text, fp32, maximum absolute difference under 1e-3 | Catches layout bugs that still "load" | IO-4 | P0 | S |
| LLM-7 | AttnLRP parity with LXT on SmolLM2 (per-token relevance correlation above 0.99), including the relevance split across shared key and value heads | The headline result: pulsatrix explains a real LLM and matches the reference | LLM-1 to LLM-3, LLM-6 | P1 | M |
| LLM-8 | Tuned lens and AtP* (corrected attribution patching) | Better versions of the logit lens and patching that already exist | LLM-1 | P2 | S–M |
| LLM-9 | Gemma 3 support: sliding-window attention, (1 + w) RMSNorm, embedding scaling |
Unlocks Gemma Scope 2 dictionaries | LLM-1 | P3 | M |
XAI: Question-driven explainability framework¶
This epic builds two frameworks into the library:
- Four reasons to explain (Adadi and Berrada, 2018): to justify a decision, to control the model (find its flaws), to improve it, and to discover something new from it.
- Nine question categories (Liao, Gruen and Miller, CHI 2020): Input (data), Output, Performance, How (global), Why, Why not, What if, How to be that (a different prediction), and How to still be this (the same prediction).
You state the question and the reason. Pulsatrix picks the explainers that answer that question, runs the checks that reason requires, and returns a report that says what the result does and doesn't support.
| Question | Primary method | Status in pulsatrix |
|---|---|---|
| Why | LRP composites, Integrated Gradients | Have |
| Why not | Contrastive LRP (relevance of logit P minus logit Q) | Have |
| What if | PDP and ICE, re-running with an edited input | Have |
| How (global) | Global surrogate tree, PDP, TCAV, CRP | Partial; XAI-4, INT-6 |
| How to be that | Counterfactual search | Gap; XAI-4 |
| How to still be this | Anchors | Gap; XAI-4 |
| Performance | Metrics, calibration, per-slice error, uncertainty | Partial; KS-3 |
| Input (data) | Dataset statistics, nearest training examples, data attribution | Partial; INT-1, KS-5 |
| Output | Output schema and label documentation | Template only |
The reason changes the defaults:
- Justify: faithfulness and stability metrics are required, and the report carries uncertainty.
- Control: the model-randomization test and per-slice error are required.
- Improve: attributions are aggregated over the dataset to surface spurious features.
- Discover: concept and featurizer tools run, each with a random-model baseline.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| XAI-1 | XaiQuestion (9 values) and XaiReason (4 values), and a capabilities() declaration on every explainer |
Machine-readable "what this explainer answers" | — | P0–P1 | S |
| XAI-2 | ExplanationPlanner::plan(question, reason, model_traits), returning ranked explainers and the checks that must run. Optional, and never hides the direct explainer APIs |
The framework users actually call | XAI-1 | P1 | M |
| XAI-3 | An Explanation report: the attribution plus question, reason, method, settings, metric results and caveats. Caveats are filled in automatically from known failure modes |
Explanations that say what they don't support | XAI-2 | P1 | M |
| XAI-4 | The missing explainers: counterfactuals (gradient-based first, then a diverse variant), a global decision-tree surrogate with a fidelity score, local fidelity output for LIME, Anchors, prototypes from nearest neighbors, and pertinent negatives | Covers How to be that, How to still be this and How (global) | XAI-1, INT-1 | P1 | M each |
| XAI-5 | Explanation-quality metrics in the Quantus families: deletion and insertion curves (with the ROAD correction), the model-parameter randomization test, sparseness and complexity | Today pulsatrix has conservation and stability checks only | — | P0 | M |
| XAI-6 | NullModelBaseline: re-run any explainer, probe or featurizer on a re-initialized copy of the model and report the difference |
The baseline rule above, as one call | FND-1 | P0 | S–M |
INT: Embedding and representation analysis¶
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| INT-1 | Nearest neighbors and mutual nearest neighbors over embeddings or activations, with optional mean-centering | The cheapest useful embedding tool; also a cross-model alignment score | FND-3 | P0 | S |
| INT-2 | Linear CKA (debiased) | Compare layers across models, training runs and backends | — | P0 | S |
| INT-3 | Control tasks for LinearProbe (selectivity = probe accuracy minus control accuracy) |
High probe accuracy can come from the probe, not the representation | — | P0 | S |
| INT-4 | Anisotropy, effective rank and TwoNN intrinsic dimension | Explains why raw cosine similarity misleads; one number per layer | INT-1 | P1 | S |
| INT-5 | PCA projections | 2D and 3D views, steering directions, preprocessing for TDA | FND-4 | P1 | M |
| INT-6 | TCAV (with significance against random concepts) and Concept Relevance Propagation | Concept-level explanations, and CRP reuses pulsatrix's LRP rules | INT-3 | P1 | M |
| INT-7 | RSA and SVCCA | Mostly covered by CKA | FND-4 | P2 | — |
| INT-8 | UMAP and t-SNE: export the data only, don't implement | Expensive, and they distort distances and cluster sizes | VIZ-1 | P3 | — |
TDA: Topological data analysis¶
Persistent homology measures the shape of a point cloud: its clusters (H0), loops (H1) and voids (H2), and how long each one lasts as the scale grows. The plan treats it as a typed pipeline:
PointCloud → DistanceMatrix → Filtration → Diagram → Vectorization
The category theory pays off in one place: stability. Each stage has a bound on how much its
output can move when its input moves, and those bounds compose. So every result carries a
certified error bar (Certified<T>{value, bound, provenance}) that adds up the fp32 distance
error, the subsampling error and the vectorization error. The API doesn't expose Functor
classes that nothing checks.
Pulsatrix writes H0 itself (a minimum spanning tree, on CPU and GPU), vendors Ripser (MIT) for H1 and H2, and writes the distances, vectorizations and significance tests itself. GUDHI's GPL-licensed parts and giotto-ph (AGPL) are avoided.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| TDA-1 | A PointCloud built from ActivationSnapshots, with flatten and pooling options, concatenated across batches |
The input to everything below | — | P0 | S |
| TDA-2 | Pairwise distances via GEMM on every backend, with an fp32 error bound | Distances are the only part that touches the model's dimension | TDA-1 | P0 | S |
| TDA-3 | H0 as a minimum spanning tree | Exact, fast and GPU-friendly | TDA-2 | P0 | S |
| TDA-4 | Vendored Ripser for H1 and H2, returning critical edges and cocycles | Loops and voids | TDA-2 | P0 | M |
| TDA-5 | Diagram, bottleneck and Wasserstein distances, Certified<T>, and property tests for stability |
Compare layers, models and seeds with certified error | TDA-3, TDA-4 | P0 | M |
| TDA-6 | Betti curves, persistence landscapes, persistence images, persistent entropy | Turn diagrams into numbers for probes and plots | TDA-5 | P0 | S |
| TDA-7 | Farthest-point landmarks, which give the subsampling error bound | Persistent homology can't run on 100,000 activations | FND-3 | P0 | S |
| TDA-8 | Significance tests: a universal null distribution, bootstrap confidence bands, Gaussian and random-network baselines | Every reported loop needs a p-value | TDA-5 | P0 | M |
| TDA-9 | Barcode, persistence diagram and Betti curve views, with JSON and SVG export | See the result | TDA-5, VIZ-1 | P1 | M |
| TDA-10 | A manifold verifier for featurizer blocks: decide whether a BSF or SAE block forms a circle or a torus, and recover a circular coordinate to steer along. It passes by reproducing the GPT-2 day-of-week circle (Engels et al.) against a random-network control | Goodfire measures block dimension with a linear method. Nobody has published a topological check, so this is new | TDA-4, TDA-7, TDA-8, FEAT-6 | P1 | M |
| TDA-11 | Representation Topology Divergence next to CKA in one compare_representations() report |
Catches clusters and loops that CKA misses | TDA-4, INT-2 | P1 | S |
| TDA-12 | Betti numbers layer by layer through a network, with a random-network baseline | Reproduces "topology simplifies through the layers" (Naitzat et al., 2020) | TDA-7, TDA-8 | P1 | M |
| TDA-13 | Spectral and distance-to-measure distances | Raw distances concentrate in high dimensions and hide loops | INT-1, FND-4 | P2 | M |
| TDA-14 | Mapper, with an instability score and a parameter sweep, never a single graph | Mapper output depends heavily on its parameters | FND-4, TDA-3 | P2 | M |
| TDA-15 | A differentiable persistence loss | Topology as a training signal | TDA-4 | P2 | M |
| TDA-16 | Probes on diagram features for out-of-distribution and trojan detection, compared against non-topological features | Only worth it if it beats simpler statistics | TDA-6 | P2 | M |
| TDA-17 to TDA-19 | Image/kernel persistence across layers, zigzag persistence, multiparameter persistence | Research-grade | TDA-4 | P3 | L–XL |
Not planned: "neural persistence" on weight graphs. A 2023 follow-up showed it mostly measures weight variance.
KD: Surrogates and student-teacher¶
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| KD-1 | A knowledge-distillation loss (softened teacher outputs, with the T² scaling) | A wrapper around the existing KLDivergenceLoss |
— | P1 | S |
| KD-2 | A teacher–student explanation-agreement test: compare LRP or IG maps with CKA or rank correlation | Turns distillation into a test of whether the student uses the same evidence | KD-1, INT-2 | P1 | S |
| KD-3 | Feature and attention distillation, soft decision trees, and rule extraction that feeds the Datalog engine | Interpretable students; rules connect to pulsatrix's neuro-symbolic side | KD-1 | P2 | — |
| KD-4 | Born-again networks and a DistilBERT-style recipe | Small and inconsistent gains | KD-1, IO-2 | P3 | — |
FEAT: Featurizers, sparse autoencoders and Goodfire BSF¶
Goodfire's Block-Sparse Featurizers (BSF, June 2026) are like sparse autoencoders, except the unit of sparsity is a small block of 2–4 dimensions rather than one direction. That lets one feature be a curve or a circle instead of a line. BSF is the newest member of a family pulsatrix can't represent yet, so the plan builds the family first, behind one interface.
Sparse autoencoders have real limits: they lose to linear probes out of distribution, steering with their features loses to simple difference-of-means vectors, and BSF still splits features. They ship here as discovery tools with metrics and baselines, not as detectors.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| FEAT-1 | A Featurizer interface (encode, decode, loss, decoder normalization), with the existing SAE ported to it and given unit-norm decoders, an L0 metric and dead-latent tracking |
One interface for the whole family | — | P0 | M |
| FEAT-2 | TopK SAE with the auxiliary loss for dead latents | Removes L1 shrinkage; sets sparsity directly | FEAT-1, FND-3 | P0 | M |
| FEAT-3 | Core SAEBench metrics: explained variance, loss recovered when the reconstruction is spliced back into the model, dead and dense latents, feature absorption. Random-model and probe baselines on by default | Without these, no featurizer result can be trusted | FEAT-1, XAI-6 | P0 | M |
| FEAT-4 | BatchTopK, JumpReLU and Matryoshka SAEs | Fix specific failures of TopK | FEAT-2 | P1 | — |
| FEAT-5 | Transcoders and skip transcoders | Reported to be more interpretable than SAEs | FEAT-1 | P1 | M |
| FEAT-6 | BSF: the vanilla, Grassmannian and group-lasso variants, then tournament top-k, with MDL and stable-rank metrics | The newest member of the family | FEAT-2, FND-4 | P1 | L |
| FEAT-7 | Steering with difference-of-means by default and featurizer directions as an option, with a reliability report | The simple baseline usually wins | FEAT-1 | P1 | S |
| FEAT-8 | Crosscoders, including the Delta-Crosscoder for fine-tuning diffs | Model diffing across layers and models | FEAT-1, IO-2 | P2 | — |
| FEAT-9 | Parameter decomposition (SPD and VPD) | Interpretability in weight space; few libraries have it | FND-4 | P2 | L–XL |
| FEAT-10 | Attribution graphs with cross-layer transcoders | The existing circuit graph plus LRP may cover most of the value first | FEAT-5 | P2 | XL |
Open: no LRP rule exists yet for propagating relevance through a block top-k. FEAT-6 needs one.
ARCH: New architectures¶
These are ranked by how well they fit an explainability-first library.
| ID | Item | Why | Depends on | P |
|---|---|---|---|---|
| ARCH-1 | B-cos layers | Self-explaining networks whose explanation can be compared directly with LRP | — | P1 |
| ARCH-2 | Concept bottleneck models | Supports "what if this concept were true" interventions. Watch for concept leakage | — | P1 |
| ARCH-3 | Mixture of experts with router attribution | The router's choice is a free, discrete explanation | FND-3 | P1 |
| ARCH-4 | Prototype networks (ProtoPNet) | "This looks like that" explanations | INT-1 | P2 |
| ARCH-5 | Graph neural networks with GNN-LRP | A new data type; gather_rows and scatter_add_rows already exist |
— | P2 |
| ARCH-6 | Mamba-2 / SSD | Extends the existing Mamba and MambaLRP | — | P2 |
| ARCH-7 | xLSTM | Extends the recurrent family | — | P2–P3 |
| ARCH-8 | Kolmogorov–Arnold networks | Interpretable by construction, but they lose to MLPs outside symbolic tasks | — | P2 |
| ARCH-9 | Diffusion transformers (adaLN) | The timestep embedding and noise schedules already exist | — | P2 |
| ARCH-10 | Gated DeltaNet and Titans, Hopfield layers, JEPA, self-explaining neural networks, hypernetworks, neural ODEs and liquid networks | Lower fit or poor match with hand-written backward passes | — | P3 |
VIZ: Visualization pack¶
Interactive visualization libraries for interpretability tend to go stale, while stable data formats last. The plan keeps ImGui for live views and puts every view on top of a versioned JSON format, with static and web renderers next to it.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| VIZ-1 | A versioned JSON document model (pulsatrix.<kind>.v1) for attributions, heatmaps, token relevance, circuit graphs, training logs and feature dashboards. The ImGui widgets read it too. NaN and infinity encode as null plus a flag |
One source of truth for every renderer | — | P0 | M |
| VIZ-2 | A dependency-free SVG renderer: bar, waterfall, heatmap, token strip, beeswarm | Publication figures with no GPU or display, including in CI | VIZ-1 | P0 | M |
| VIZ-3 | Self-contained Vega-Lite HTML, with the JavaScript inlined or loaded from a CDN | Hover, zoom and export for free, in a browser or notebook | VIZ-1 | P1 | S |
| VIZ-4 | Export circuit graphs in the attribution-graph schema that Neuronpedia and circuit-tracer read | Large graphs get a mature viewer for free | VIZ-1 | P1 | M |
| VIZ-6 | New views: (a) token relevance for text, (b) feature dashboards, (c) an embedding projector, (d) an attention head grid, (e) SHAP force, decision and dependence plots | Fill the gaps between the current widgets and the reference tools | VIZ-1 | P2 | M each |
| VIZ-7 | A node editor for circuit graphs | Only if graphs outgrow the current view; VIZ-4 covers large ones | — | P3 | M |
(VIZ-5, the notebook path, moved to the NB epic.)
NB: Notebook layer¶
Notebook output needs neither Python nor a running kernel. The layers build on each other. ImGui can't draw inline in a notebook, so it stays the desktop tool and shares the VIZ-1 data.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| NB-1 | pulsatrix::mime_bundle_repr(const T&) overloads that return MIME bundles: tensor summaries, attribution heatmaps as SVG, token relevance, circuit graphs and persistence diagrams as Vega-Lite. The C++ Jupyter kernel xeus-cpp finds them automatically, so pulsatrix doesn't depend on it |
Rich display in a C++ notebook | VIZ-1, VIZ-2 | P1 | S |
| NB-2 | A C++ Report builder that writes .ipynb files (nbformat 4.5) and self-contained HTML: markdown, code shown as text, and rich outputs |
Notebook-format results with no Python and no kernel; they render on GitHub, in VS Code and in Quarto | NB-1 | P1 | S–M |
| NB-3 | Python _repr_mimebundle_ on the bound types, delegating to NB-1 so both languages render the same (golden test), then interactive views with anywidget |
Notebook users on the Python side | NB-1 | P2 | S → M |
| NB-4 | xeus-cpp support: a pulsatrix_notebook.hpp header, example notebooks, and a Linux CI smoke test. Windows and GPU code inside notebook cells are marked experimental. Every precondition reachable from a cell throws instead of aborting, because a crash kills the kernel |
Run pulsatrix interactively in C++ | NB-1 | P3 | M |
| NB-5 | A WebAssembly build of the CPU backend for JupyterLite | Notebooks in the browser with nothing installed | NB-4 | Deferred | XL |
HIP: Training efficiency on AMD GPUs¶
The bottlenecks are structural, so measure first and fix those before anything exotic. On the gfx1151 APU almost every non-GEMM kernel in pulsatrix is limited by memory bandwidth at best. In practice they are limited by launch and sync overhead and by too little parallelism.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| HIP-1 | scripts/profile_hip.sh: a rocprofv3 wrapper that writes kernel time per op type to CSV |
The baseline, and the falsifier for every item below | — | P0 | S |
| HIP-2 | Reductions with one workgroup or one wave per row or channel (softmax, norms, column sums): wave shuffles using the runtime warpSize and 64-bit masks, then a shared-memory stage. Small channel counts get a deterministic two-stage grid reduction, no atomics |
BatchNorm at N=64, C=3, 224×224 currently runs 3 threads over 3.2 million elements | HIP-1 | P0 | M |
| HIP-3 | A caching allocator: size classes, a pool per stream, and a configurable budget instead of hipMemGetInfo, which overstates what the APU can allocate. Don't use hipMallocAsync (open corruption bugs on RDNA) |
Raw hipMalloc and hipFree on every tensor |
— | P0 | M |
| HIP-4 | Remove the 56 per-op hipStreamSynchronize calls. Sync only when the host reads a result, and add PULSATRIX_HIP_SYNC_DEBUG=1 to bring them back for debugging |
Launch and sync cost more than the kernels on small models | HIP-3, FND-8 | P0 | S |
| HIP-5 | dot and sum across many blocks (partials, then a second pass), deterministic |
They run on a single block today | — | P0 | S |
| HIP-6 | Fused kernels: AdamW across all parameters in one launch, bias plus activation, softmax plus cross-entropy. In explain mode they still write the values LRP needs | Fewer launches and less memory traffic | TRN-2 | P1 | — |
| HIP-7 | Conv2D that runs im2col and GEMM in batch chunks | The first layer's im2col buffer at N=64 is about 350 MB, which competes with system RAM on an APU | FND-6 | P1 | S |
| HIP-8 | A hipBLASLt probe on the pinned container, then bf16 GEMM through hipblasGemmEx |
Reports conflict on whether hipBLASLt works on gfx1151 in ROCm 7.2.4; measure it | HIP-1 | P1 | — |
| HIP-9 | A ROCm 10.0 evaluation image, the first release that officially lists gfx1151. Also pin the host kernel version (6.18.4 or newer, or the Ubuntu OEM kernel with the VGPR fix) | Known gfx1151 crashes depend on both | — | P1 | S |
| HIP-10 | Zero-copy staging buffers on APUs, enabled only when the device reports itself as integrated | Saves a copy on Strix Halo without slowing discrete GPUs | HIP-3 | P2 | — |
| HIP-11 | Full bf16 training | Halves memory traffic and reaches the matrix cores. Needs a dtype in Tensor |
IO-6 | P2 | XL |
| HIP-12 | HIP graphs, WMMA or rocWMMA kernels, FlashAttention for training only, MIOpen | Last: graphs have measured slowdowns on gfx11, and FlashAttention never builds the attention matrix that AttnLRP needs | HIP-4 | P3 | — |
AGT: Agents, native C++¶
Pulsatrix can run its own agent orchestrator without Python. There is no official C++ SDK for the Anthropic API or for MCP, but neither is hard to write against. The plain HTTP and JSON API is stable, and the 2026-07-28 MCP specification dropped the handshake and sessions, so a stdio server that only offers tools is small. The genuinely hard parts are narrow:
- TLS certificate trust on Windows and macOS;
- validating JSON Schema 2020-12;
- getting a 135M-parameter model to pick tools reliably.
LLM agents live in pulsatrix::orch (target pulsatrix_orchestrator), because
pulsatrix::Agent already names the reinforcement-learning policy contract.
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| AGT-1 | An optional PULSATRIX_ENABLE_ORCHESTRATOR target, off by default, with nlohmann/json (MIT). The same JSON library serves VIZ-1 and NB-1 |
Keeps the core free of networking dependencies | — | P1 | S |
| AGT-2 | An HttpTransport interface. cpp-httplib by default, with an optional libcurl backend that uses the operating system's TLS on Windows. A shared server-sent-events parser that ignores unknown event types |
Streaming LLM responses | AGT-1 | P1 | M |
| AGT-3 | Providers and the tool loop: an Anthropic provider and an OpenAI-compatible provider (llama-server, vLLM, Ollama), with one internal message format | The orchestrator itself | AGT-2 | P1 | M |
| AGT-4 | A tool registry: name, description, input schema and a std::function. Schemas use only the keywords that mean the same in JSON Schema draft-7 and 2020-12, checked by a small built-in validator. The MCP server and the in-process loop share one registry |
Define a tool once, use it everywhere | AGT-1 | P1 | S–M |
| AGT-5 | A native MCP server over stdio (2026-07-28 specification, with a fallback for older clients) that exposes explain, attribute, logit lens, patching and featurizer tools | Lets Claude and other agents drive pulsatrix | AGT-4, IO-1 | P2 | M |
| AGT-6 | A local provider that runs pulsatrix's own LLM, with tool calls constrained by a token-level mask | Agents with no network, and a model pulsatrix can explain | LLM-1 to LLM-5, AGT-4 | P3 | M |
| AGT-7 | Explain an agent's tool choice: run AttnLRP from the logit of the first token that distinguishes the chosen tool, back over the prompt, the tool descriptions and the history. It passes if deleting the top-attributed tokens flips the choice more often than deleting random tokens | No other library can explain why an agent picked a tool | AGT-6, LLM-7 | P3 | L |
| AGT-8 | Automated descriptions of featurizer latents: an LLM describes the top activations, then the description is scored by how well it predicts activations | Labels thousands of features | AGT-3, FEAT-1 | P2 | M |
| AGT-9 | An interpretability agent that runs experiments on the model (MAIA-style) | Research-grade | AGT-3, AGT-4 | P3 | L |
AGT-3 has to follow the current Anthropic API rules:
- Forcing a specific tool returns HTTP 400 on the current Claude models; use automatic tool choice with strict tool schemas instead.
- Tool results come first in the next user message.
- A 429 from a spend limit has no
retry-afterheader and must not be retried. - Back off with jitter on 5xx, 529 and mid-stream overload errors.
- Cache the system prompt and tool definitions.
- Read the API key from the environment and never log it.
AGT-5 has to follow these security rules:
- Write nothing to stdout except protocol messages.
- Resolve every path and confine it to allowed directories.
- Load safetensors, GGUF and pulsatrix-native formats only, never pickle.
- Cap tensor and request sizes.
- Use opaque handles that expire.
- Treat tool output as untrusted.
KS: Kitchen sink¶
| ID | Item | Why | Depends on | P | Effort |
|---|---|---|---|---|---|
| KS-1 | CMake install() and a package config, so other projects can find_package(pulsatrix) |
There's no install step today | — | P0 | S |
| KS-2 | A benchmark suite: step time, explanation time, conservation error | Measures every HIP item and catches regressions | — | P1 | S |
| KS-3 | Uncertainty: Monte Carlo dropout, deep ensembles, split conformal prediction | Answers the Performance question | — | P1 | S |
| KS-4 | FGSM and PGD attacks, also used as a robustness test for explanations | Robustness and the Control reason | — | P1 | S |
| KS-5 | TracIn data attribution | Answers "which training examples caused this" | FND-1 | P2 | M |
| KS-6 | Fairness metrics, drift detection and a model-card generator | The Input question, and documentation | — | P2 | S each |
| KS-7 | Wire the existing thread pool into DataLoader's num_workers |
Built but not connected | — | P2 | M |
| KS-8 | GPU CI on a self-hosted gfx1151 runner | The HIP backend is only tested by hand today | HIP-9 | P1 | M |
| KS-9 | A model zoo: ResNet18, VGG16 and SmolLM2 with reference heatmaps | Reproducible examples on real models | IO-4, FND-6 | P1 | M |
| KS-10 | Strided views and broadcasting | Removes copies everywhere; touches every kernel | — | P2 | L |
| KS-11 | Python wheels and a vcpkg port | Easier installation | KS-1 | P2 | — |
| KS-12 | EK-FAC influence functions, quantization, op-level autograd, distributed training | Large or low priority for an explainability library | — | P3 / v2 | — |
Open questions¶
These came up in the research and aren't settled yet.
- BSF details. The exact MDL formula and the block-size defaults beyond the published example need a full read of the papers.
- Missing LRP rules. No published rule exists for relevance through a block top-k (BSF), through a LoReFT intervention, or through LoRA under the gamma rule. The epsilon-rule split for LoRA is our own derivation and needs a numerical check.
- Qwen3 in LXT. LXT marks Qwen3 as experimental, with relevance skewed toward the first token. The cause isn't documented.
- hipBLASLt on gfx1151. Sources disagree on whether it works on ROCm 7.2.4. HIP-8 settles it.
- rsLoRA scaling. It conflicts with the claim that 1/r scaling makes the learning rate independent of rank. Only an experiment settles it.
- Licenses. Llama 3.2's license terms for redistributing converted weights: ship converters, not weights. The licenses of PHAT, Dionysus, Hera and Eirene weren't checked; none are needed for the plan.
- Mamba-2 and xLSTM. No published LRP rules found.
- MCP security statistics. The figures cited come from secondary reports, not primary scans.
- xeus-cpp. How version 0.10 loads prebuilt libraries, and whether prebuilt HIP kernels can be called from notebook cells, needs a hands-on test.
- Windows TLS. Whether the open cpp-httplib Windows certificate issue affects
api.anthropic.comneeds a CI probe. - TDA numerics. The fp32 distance error between near-duplicate activations hasn't been measured. Don't trust loops smaller than that error.