Roadmap Research Notes¶
This page holds the evidence behind the Roadmap. The research was done on 2026-10-03. For each item it gives:
- Source: the primary source and its date.
- Falsifier: the result that would tell us not to build the item, or that it doesn't work.
- Fails when: known failure modes from papers, issue trackers and post-mortems.
Some claims about AMD's gfx1151 rest on community reports rather than AMD documentation, so they carry tags:
- [V]: verified in a primary source (vendor docs, a spec, a paper).
- [C]: a community report (a GitHub issue or benchmark). Real, but not reproduced by us.
- [I]: our own inference. Treat every speedup estimate as [I] until HIP-1 measures it.
Sources and tools age quickly in this field. Check the date before relying on one.
Cross-cutting¶
- Random networks give plausible explanations.
- Méloux et al., "The Dead Salmons of AI Interpretability", arXiv 2512.18792 (2025-12): attribution, probing, sparse autoencoders and causal analyses all produce plausible output on randomly initialized networks.
- Adebayo et al., "Sanity Checks for Saliency Maps", NeurIPS 2018: some saliency methods barely change when the model's weights are randomized.
- Heap et al., arXiv 2501.17727 (2025-01): SAEs on randomly initialized transformers produce features that look interpretable.
This is the basis for the baseline rule and for XAI-6. - Explainers disagree with each other. Krishna et al., "The Disagreement Problem in Explainable Machine Learning", arXiv 2202.01602 (2022-02). - Surrogates can be fooled. Slack et al., AIES 2020: a model can be built to look innocent to LIME and SHAP while staying biased.
FND: Foundations¶
All FND items are done; see the roadmap's FND section for what changed from the plan. The notes below record the reasoning they were built on.
- FND-1, FND-2.
Module::parameters()returnedstd::vector<ParamRef>with no names. It is now derived fromnamed_parameters(). - Falsifier for freezing: a frozen parameter's bytes change after N steps, or the input gradient no longer matches finite differences.
- FND-3, FND-4. The demand for these comes from items in other epics. FEAT-2 (Gao et al., TopK SAE, arXiv 2406.04093, 2024-06) and ARCH-3 need top-k. INT-5, FEAT-6, TRN-9 (Wu et al., ReFT, arXiv 2404.03592, 2024-04) and TRN-11 (Shuttleworth et al., arXiv 2410.21228, 2024-10) need an eigensolver or SVD.
linear_algebra.hpphad only a dense solve;matrix_decompositions.hppnow has the rest. - Fails when: eigenvector sign and order are arbitrary, so tests must align them before comparing.
- FND-5.
- Sources:
plans/lrp_issues.md#8. Zennit's canonizers fold BatchNorm into the preceding layer before LRP. - Falsifier: none. Explanations that depend on the rest of the batch are a correctness bug.
- FND-6. VGG and ResNet are the standard LRP benchmark models, and both need stride and padding.
- FND-7.
plans/gpu_review.mdlists "no atomics, so deterministic" as a current strength. Keep that. - FND-8.
plans/gpu_review.md#1 (no device check inModule::forward) and #2 (Tensor::Stacktags the wrong device).
IO: Serialization and model import¶
- IO-1.
- Format: huggingface/safetensors (Apache-2.0) is an 8-byte length, a JSON header, then a packed data section. The header is at most 100 MB, and
__metadata__maps strings to strings. - Audit: Trail of Bits audited the format (published 2023-05).
- Option: the header-only C++ reader safetensors-cpp (MIT) exists, but its README says shape validation and mmap are incomplete. Fuzz both it and a hand-written reader before choosing.
- Fails when: an offset runs past the end of the file, ranges overlap, or the element-count product overflows.
- IO-2. Falsifier: save, then load, doesn't give a bit-identical forward pass, or resumed training doesn't reproduce the loss curve.
- IO-3: pickle is code.
- CVE-2025-32434 (GHSA-53q9-r3pm-6pq6, 2025-04, CVSS 9.8):
torch.load(weights_only=True)was remote code execution before torch 2.6.0. - Scanners get bypassed. "nullifAI" (ReversingLabs, 2025-02) used broken pickles and 7z-compressed
.ptfiles; CVE-2025-46417 is another bypass; ShadowPickle (arXiv 2607.17503, 2026-07) shows evasion continues. - Format:
torch.savewrites a zip ofdata.pklplus storage blobs, and tensors are rebuilt with explicit strides. Non-contiguous tensors must be made contiguous before writing. - IO-4: layout bugs load fine and give garbage.
- A missing Q/K RoPE permutation gives fluent nonsense (huggingface/transformers #25199). Random-weight tests can't catch it.
- Llama 3.2's
rope_scalingfactor is 32, not 8; the wrong value only hurts past about 80k tokens (meta-llama/llama-models #241). - GPT-2's
c_attnis a Conv1D stored[in, out], the opposite ofnn.Linear. - Gemma's RMSNorm multiplies by
(1 + w). rms_norm_epsis 1e-5 for Llama and SmolLM2, 1e-6 for Qwen3 and Gemma 3.- IO-5. These are Hugging Face
config.jsonfields, fetched from the Hub on 2026-10-03:
| Model | Notable config |
|---|---|
| SmolLM2-135M | d=576, 9 heads / 3 KV heads, tied, theta 1e5 |
| Qwen2.5-0.5B | QKV bias, theta 1e6 |
| Llama-3.2-1B | rope_scaling llama3, factor 32 |
| Qwen3-0.6B | head_dim 128 ≠ d/h, QK-norm |
| Gemma 3 270M | head_dim 256, sliding window 512 |
- IO-6. bf16 to fp32 is a 16-bit shift, so it's exact. fp16 to fp32 is also exact.
- IO-7. NumPy's format spec covers the
.npyheader.allow_picklecode execution has existed since CVE-2019-6446, so never accept the object dtype. - IO-8. onnx/onnx (Apache-2.0). Protobuf caps a message at 2 GB, so larger models use external data (onnx #3273). External-data paths had a path-traversal CVE (CVE-2024-27318, fixed in onnx 1.16).
- IO-9. The format spec is
docs/gguf.mdin ggml-org/ggml (MIT). Its parsers have a long CVE history: - TALOS-2024-1913 and TALOS-2024-1914 (2024-03);
- CVE-2025-53630 (integer overflow in
gguf_init_from_file_impl); - CVE-2026-27940 (a heap overflow);
- six more posted to oss-security on 2026-05-15.
- IO-10. Keras's
safe_modekeeps being bypassed: CVE-2025-1550, then CVE-2025-8747 (which bypassed that fix), then CVE-2025-9905 (.h5loading ignoressafe_mode). - IO-11. TorchScript is deprecated as of PyTorch 2.10 (release blog); 2.11 recommends
torch.export.
TRN: Training and fine-tuning¶
- Overall falsifier: Shuttleworth et al. (arXiv 2601.22708, 2026-01) found that with a tuned learning rate, plain LoRA matches or beats most of its variants. If a variant doesn't beat tuned vanilla LoRA on our benchmark, don't ship it.
- "LoRA Without Regret" (Thinking Machines, 2025-09-29):
- Attention-only LoRA underperforms; apply it to all layers.
- The best learning rate is about 10× full fine-tuning's.
- LoRA tolerates large batches worse.
- Rank 1 is enough for RL.
- "LoRA learns less and forgets less" (Biderman et al., arXiv 2405.09673, 2024-05): LoRA is clearly worse than full fine-tuning on code and math, and it forgets less.
- TRN-2. Loshchilov and Hutter, "Decoupled Weight Decay Regularization", arXiv 1711.05101. Test: decay is applied as
w -= lr * wd * w, not added to the gradient. - TRN-3, TRN-4. Global-norm clipping and warmup-plus-cosine schedules are the defaults in every recipe cited above (for example "LoRA Without Regret"). Clip after accumulation and before the optimizer step.
- TRN-6. Full fine-tuning is the baseline in Biderman et al. and in "LoRA Without Regret". Adam in fp32 needs about 16 bytes per parameter plus activations: about 2.2 GB for 135M parameters, about 16 GB for 1B.
- TRN-5. Hugging Face Trainer fixed this gradient-accumulation normalization bug in v4.46 (2024-10). Falsifier: 4×8 accumulated doesn't equal 1×32 on ragged batches.
- TRN-7. Hu et al., "LoRA", arXiv 2106.09685 (2021-06).
- Merging into a quantized base changes the outputs (huggingface/peft #2321), so merge in fp32.
- Falsifier: merged and unmerged outputs differ by more than 1e-5.
- TRN-8. Under the epsilon rule, LRP is linear in the contributions, so relevance splits exactly between W₀x and the adapter term. This is our own derivation and needs a numerical check. Gamma and z+ act on the positive part of the summed weight, so merged and unmerged models give different maps.
- Related: Soligo et al., arXiv 2506.11618 (2025-06), read rank-1 LoRA adapters as interpretable directions.
- TRN-9. Wu et al., "ReFT", arXiv 2404.03592 (2024-04).
- It fails on arithmetic (GSM8K).
- It fails badly at concept suppression compared with rank-1 steering vectors (arXiv 2505.20809, 2025-05).
- TRN-10. Liu et al., "IA³", arXiv 2205.05638 (2022-05).
- TRN-11.
- Shuttleworth et al., "LoRA vs Full Fine-tuning: An Illusion of Equivalence", arXiv 2410.21228 (2024-10): LoRA introduces "intruder" singular vectors.
- Minder et al., arXiv 2510.13900 (ICLR 2026): narrow fine-tunes leave readable traces in activation differences. Mixing pretraining data into the fine-tune removes the trace, so a null diff doesn't prove nothing changed.
- TRN-12. Activation checkpointing conflicts with LRP's cached activations.
- TRN-13.
- rsLoRA: Kalajdzievski, arXiv 2312.03732 (2023-12).
- LoRA+: Hayou et al., arXiv 2402.12354 (2024-02).
- DoRA: Liu et al., arXiv 2402.09353 (2024-02).
- TRN-14.
- PiSSA: arXiv 2404.02948.
- VeRA: arXiv 2310.11454.
- AdaLoRA: arXiv 2303.10512.
- GaLore: arXiv 2403.03507.
- QLoRA: arXiv 2305.14314.
- Prefix tuning: arXiv 2101.00190.
- Prompt tuning: arXiv 2104.08691.
- BitFit: arXiv 2106.10199. Most modern LLMs have no biases, so BitFit is nearly useless for them.
- Houlsby adapters: arXiv 1902.00751.
- bf16 updates round away small steps. Zamirai et al., arXiv 2010.06192. Keep fp32 master weights.
LLM: Running real language models¶
- LLM-1. SmolLM2 needs grouped-query attention. Qwen3 and Gemma 3 need
head_dimseparate fromd_model / num_heads. - pulsatrix's
RoPEModulerotates adjacent pairs, while Hugging Face Llama-family checkpoints use "rotate half". - Fails when: the AttnLRP rules stop conserving relevance under a mask. Re-validate against LXT.
- LLM-2. SmolLM2, Qwen3 and Gemma 3 set
tie_word_embeddings. Without views, an untied copy of Gemma 3 270M's embedding would cost about 680 MB in fp32. - LLM-3.
- mlc-ai/tokenizers-cpp (Apache-2.0) exists but wraps the Rust tokenizers library, which pulls in a Rust toolchain.
- A C++ port of the Hugging Face tokenizer found double-space mismatches in 4 of 15 million lines (wangkuiyi/huggingface-tokenizer-in-cxx).
- Falsifier: id mismatches on the 10,000-line corpus.
- LLM-4. Sampling must use the seeded RNG from FND-7 so generations are reproducible.
- LLM-5. Falsifier: cached and uncached logits differ.
- LLM-6. The threshold (fp32, max abs diff under 1e-3) follows from IO-4's layout pitfalls: random-weight tests miss them, real weights on real text don't.
- LLM-7.
- Achtibat et al., "AttnLRP", arXiv 2402.05602 (ICML 2024).
- LXT 2.1 (2025-07-10) supports Gemma 3 and Qwen3, and marks Qwen3 experimental with attribution skewed toward the first token.
- Falsifier: per-token relevance correlation with LXT is below 0.99 on Llama-family models.
- LLM-8.
- Tuned lens: Belrose et al., arXiv 2303.08112 (2023-03).
- AtP*: Kramár et al., arXiv 2403.00745 (2024-03). Plain attribution patching misses effects through attention saturation.
- LLM-9. Gemma 3 technical report, arXiv 2503.19786 (2025-03). Gemma Scope 2 (Google DeepMind, 2025-12) publishes SAEs and transcoders for every layer of Gemma 3 270M and 1B.
XAI: Question-driven framework¶
- Four reasons to explain. Adadi and Berrada, "Peeking Inside the Black-Box", IEEE Access 6 (2018): justify, control, improve, discover. The exact wording comes from secondary summaries; the PDF section still needs reading.
- Nine question categories. Liao, Gruen and Miller, "Questioning the AI", CHI 2020, arXiv 2001.02478. The categories were checked against the paper's Figure 1.
- Designers report that "why not" questions are common and poorly served.
- Question-to-method mapping. Liao et al., "Question-Driven Design Process for Explainable AI User Experiences", arXiv 2104.03483 (2021-04). It notes an "availability bias": practitioners default to LIME and SHAP.
- Evaluation levels. Doshi-Velez and Kim, arXiv 1702.08608 (2017-02): application-grounded, human-grounded and functionally grounded evaluation. A library can only own the functionally grounded level.
- XAI-1. The enums mirror Liao et al. (2020) and Adadi and Berrada (2018) exactly, so reports can be compared against the papers' question bank.
- XAI-2, XAI-3. Falsifier: users ignore the planner and call LRP directly. That's why the planner is optional and never hides the direct APIs.
- XAI-4.
- Counterfactuals: Wachter et al., arXiv 1711.00399 (2017-11). DiCE: Mothilal et al., arXiv 1905.07697. On image models counterfactuals turn into adversarial examples (Freiesleben, arXiv 2009.05487).
- Anchors: Ribeiro et al., AAAI 2018. Coverage can be tiny, and precision is estimated, not guaranteed.
- TREPAN surrogates: Craven and Shavlik, NeurIPS 1996. Falsifier: fidelity under about 85% at a depth a person can read.
- Pertinent negatives (CEM): Dhurandhar et al., NeurIPS 2018.
- ProtoDash: Gurumoorthy et al., ICDM 2019.
- XAI-5.
- Quantus: Hedström et al., JMLR 2023, arXiv 2202.06861. It groups metrics into six families: faithfulness, robustness, localisation, complexity, randomisation and axiomatic.
- ROAD: Rong et al., ICML 2022.
- The model-parameter randomization test comes from Adebayo et al., NeurIPS 2018.
- XAI-6. See the cross-cutting section above.
INT: Embedding and representation analysis¶
- INT-1. Huh et al., "Platonic Representation Hypothesis", arXiv 2405.07987 (2024-05), defines mutual-kNN alignment.
- Fails when: a few points become everyone's neighbor in high dimensions (hubness).
- INT-2. Kornblith et al., ICML 2019, arXiv 1905.00414.
- Falsifier: CKA between a trained and a random-init network stays high. That would mean it measures the architecture, not the learning.
- Fails when: outliers sway it, and it can be manipulated (Davari et al., arXiv 2210.16156).
- INT-3. Hewitt and Liang, EMNLP 2019. MDL probing: Voita and Titov, EMNLP 2020.
- INT-4.
- Anisotropy: Ethayarajh, EMNLP 2019.
- TwoNN: Facco et al., Scientific Reports 2017; applied to deep networks by Ansuini et al., NeurIPS 2019.
- INT-6.
- TCAV: Kim et al., ICML 2018, arXiv 1711.11279. Concept vectors are unstable across random negative sets.
- CRP: Achtibat et al., Nature Machine Intelligence 2023, arXiv 2206.03208.
- INT-7.
- RSA: Kriegeskorte et al., 2008.
- SVCCA: Raghu et al., NeurIPS 2017.
- PWCCA: Morcos et al., NeurIPS 2018.
- INT-8. t-SNE and UMAP distort global distances and cluster sizes (Wattenberg et al., Distill 2016; Chari and Pachter, PLoS Comp Bio 2023).
TDA: Topological data analysis¶
- Core theory.
- Persistent homology: Edelsbrunner, Letscher and Zomorodian, DCG 2002; Zomorodian and Carlsson, DCG 2005.
- Stability: Cohen-Steiner, Edelsbrunner and Harer, DCG 2007.
- Persistence modules as functors, and the interleaving distance: Chazal, de Silva, Glisse and Oudot, The Structure and Stability of Persistence Modules (2016).
- Single linkage is the unique stable functorial clustering: Carlsson and Mémoli, JMLR 2010.
- Mapper as a cosheaf: de Silva, Munch and Patel, DCG 2016.
- TDA-2, TDA-3. Distances via ‖x‖² + ‖y‖² − 2xyᵀ suffer fp32 cancellation for near-duplicate points; clamp at zero and carry the error as a diagram bound (stability gives d_B ≤ ‖D − D′‖∞). H0 equals single-linkage clustering, which is the minimum spanning tree (Carlsson and Mémoli 2010).
- TDA-4. Ripser: Bauer, JACT 2021, arXiv 1908.02518. MIT, about 1000 lines, no dependencies.
- giotto-ph is AGPLv3, so it can't be vendored.
- GUDHI's CGAL-dependent modules are GPLv3.
- Practical limits on CPU [I]: H1 up to about 2–5k points, H2 up to about 1k.
- TDA-5.
- Bottleneck distance: Kerber, Morozov and Nigmetov 2017.
- Wasserstein stability: Skraba and Turner, arXiv 2006.16824.
- TDA-6.
- Persistence landscapes: Bubenik, JMLR 2015.
- Persistence images: Adams et al., JMLR 2017.
- Betti curves are not stable under the bottleneck distance.
- TDA-7. The landmark subsampling error is at most twice the covering radius. Witness complexes: de Silva and Carlsson 2004.
- TDA-8.
- Bootstrap confidence sets: Fasy et al., Annals of Statistics 2014.
- Permutation tests: Robinson and Turner 2017.
- The universal null: Bobrowski and Skraba 2023.
- TDA-9. Barcodes and persistence diagrams follow VIZ-1's JSON format; the significance band from TDA-8 is drawn on the diagram.
- TDA-10.
- Engels et al., "Not All Language Model Features Are Linear", ICLR 2025, arXiv 2405.14860: days of the week and months form circles in GPT-2 and Mistral.
- Goodfire's BSF vision post (2026-07-07) measures block dimension by stable rank, a linear measure.
- Circular coordinates: de Silva, Morozov and Vejdemo-Johansson, DCG 2011.
- Falsifier: the dominant H1 bar isn't significant against the null, or it also appears on random-init activations.
- TDA-11. RTD: Barannikov, Trofimov et al., ICML 2022.
- TDA-12. Naitzat, Zhitnikov and Lim, "Topology of Deep Neural Networks", JMLR 2020, arXiv 2004.06093.
- TDA-13.
- Distance concentration hides loops in high dimensions: Damrich, Berens and Kobak, NeurIPS 2024, arXiv 2311.03087.
- Distance-to-measure: Chazal, Cohen-Steiner and Mérigot 2011.
- TDA-14.
- Mapper: Singh, Mémoli and Carlsson 2007.
- Instability and parameter dependence: Belchí et al., JMLR 2020; Carrière, Michel and Oudot, JMLR 2018.
- TDA-15. Carrière et al., ICML 2021. Topological autoencoders: Moor et al., ICML 2020.
- TDA-16. TopoTroj: Zheng et al., NeurIPS 2021, arXiv 2106.06469.
- Falsifier: a probe on norms or kNN distances matches it.
- TDA-17 to TDA-19. Zigzag persistence on LLM layers: Gardinazzi et al., ICML 2025, arXiv 2410.11042. Multiparameter persistence has no complete invariant (Carlsson and Zomorodian 2009).
- Neural persistence. Rieck et al., ICLR 2019. It mostly reduces to weight variance (Girrbach et al., arXiv 2307.10865, 2023).
- Caution on recent LLM-topology papers. Most from 2025–26 use H0 only and are unreviewed preprints. H0 is a function of the minimum spanning tree, so check those results against plain kNN statistics.
KD: Surrogates and student-teacher¶
- KD-1. Hinton, Vinyals and Dean, arXiv 1503.02531 (2015-03).
- Students often don't match teachers even with enough capacity: Stanton et al., "Does Knowledge Distillation Really Work?", NeurIPS 2021.
- KD-2. If prediction agreement is high but explanation agreement is low, that is the interesting signal, not a bug.
- KD-3.
- FitNets: arXiv 1412.6550.
- Attention transfer: arXiv 1612.03928.
- Soft decision trees: Frosst and Hinton, arXiv 1711.09784.
- Boolean rule column generation: Dash, Günlük and Wei, NeurIPS 2018.
- KD-4.
- Born-again networks: Furlanello et al., arXiv 1805.04770.
- DistilBERT: Sanh et al., arXiv 1910.01108.
FEAT: Featurizers¶
- FEAT-6, BSF.
- Fel et al., "Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds", arXiv 2606.25234 (2026-06-23).
- The code is at github.com/goodfire-ai/block-sparse-featurizer. The variants are:
- vanilla: keep the top-k blocks by norm;
- Grassmannian: tied, with orthonormal block frames;
- group lasso: a block soft-threshold.
- The paper evaluates with minimum description length and stable rank.
- The follow-up, Jerpelea and Ananthram, "A Deeper Analysis of Block-Sparse Featurizers", arXiv 2608.27515 (2026-08-27), reports that BSF still suffers from feature splitting and composition, and proposes tournament top-k.
- Falsifier: on pulsatrix-scale models, BSF fails to beat a TopK SAE at matched L0.
- FEAT-2. Gao et al., arXiv 2406.04093 (2024-06), introduces the auxiliary loss for dead latents.
- FEAT-3.
- SAEBench: Karvonen et al., arXiv 2503.09532 (2025-03). Its authors note that proxy metrics don't reliably predict usefulness.
- Feature absorption: Chanin et al., arXiv 2409.14507.
- FEAT-4.
- BatchTopK: arXiv 2412.06410.
- JumpReLU: arXiv 2407.14435.
- Matryoshka: arXiv 2503.17547.
- FEAT-5.
- Transcoders: Dunefsky et al., arXiv 2406.11944.
- "Transcoders Beat Sparse Autoencoders for Interpretability": Paulo et al., arXiv 2501.18823.
- FEAT-7. AxBench (Wu et al., arXiv 2501.17148, 2025-01): SAE steering underperforms simple baselines. Steering vectors are unreliable in distribution and misgeneralize out of it (Tan et al., NeurIPS 2024).
- FEAT-8. Crosscoders: Lindsey et al., Transformer Circuits (2024-10). Delta-Crosscoder: arXiv 2603.04426 (2026-03).
- FEAT-9.
- APD: arXiv 2501.14926.
- SPD: Bushnaq, Braun and Sharkey, arXiv 2506.20790.
- VPD: Goodfire (2026-05). Its technical report URL returned 404, so the details are unverified.
- FEAT-10. Ameisen, Lindsey et al., "Circuit Tracing", Transformer Circuits (2025-03-27). The authors list limits: QK circuits aren't explained, there is reconstruction error, and inactive features are missed.
- SAE limits.
- Linear probes beat SAE probes out of distribution: DeepMind, "Negative Results for SAEs on Downstream Tasks" (2025-03-26). After this the team deprioritized SAE research.
- SAEs trained on different seeds don't agree: Paulo and Belrose, arXiv 2501.16615.
- Position SAEs as discovery tools: Peng et al., arXiv 2506.23845.
ARCH: New architectures¶
- ARCH-1. B-cos: Böhle, Fritz and Schiele, CVPR 2022, arXiv 2205.10268. B-cos language models: arXiv 2502.12992.
- ARCH-2. Koh et al., ICML 2020, arXiv 2007.04612. Concept leakage: Mahinpei et al., arXiv 2106.13314.
- ARCH-3. Switch Transformer: arXiv 2101.03961. OLMoE router analysis: arXiv 2409.02060. Experts tend to specialize by token and syntax rather than domain.
- ARCH-4. ProtoPNet: arXiv 1806.10574. Human and model similarity don't always match: Hoffmann et al., arXiv 2105.02968.
- ARCH-5.
- GNNExplainer: arXiv 1903.03894.
- GNN-LRP: Schnake et al., TPAMI 2021.
- Explainer benchmarks are unreliable: Faber et al. 2021.
- ARCH-6. Dao and Gu, arXiv 2405.21060. No SSD LRP reference implementation was found.
- ARCH-7. Beck et al., NeurIPS 2024, arXiv 2405.04517.
- ARCH-8. KAN: arXiv 2404.19756. In "KAN or MLP: A Fairer Comparison", arXiv 2407.16674, KANs lose outside symbolic formulas.
- ARCH-9. Peebles and Xie, arXiv 2212.09748.
- ARCH-10.
- Gated DeltaNet: arXiv 2412.06464.
- Titans: arXiv 2501.00663.
- Hopfield layers: arXiv 2008.02217.
- V-JEPA 2: arXiv 2506.09985.
- SENN: arXiv 1806.07538.
- Hypernetworks: arXiv 1609.09106.
- Neural ODEs: arXiv 1806.07366.
- Liquid time-constant networks: arXiv 2006.04439.
VIZ: Visualization pack¶
- Tools go stale. CircuitsVis's last release was about a year ago, and token-hover regressions are still open (#98) [C]. Captum Insights was deprecated in Captum 0.8 (2025-03) [C].
- VIZ-1.
- Fails when: NaN and infinity appear in JSON (they're invalid), or the schema churns. Version it from day one and add golden-file tests.
- VIZ-2. ImPlot has no vector export, so publication figures need a separate renderer [I]. Share
colormap.hppso SVG and ImGui colors match. - VIZ-3. Vega-Lite embedding docs. Inlining the JavaScript adds roughly 800 KB per file [I].
- VIZ-4. circuit-tracer (Anthropic open-sourced it, 2025). Neuronpedia publishes an attribution-graph JSON schema and a validator at neuronpedia.org/graph/validator.
- Falsifier: the schema needs transcoder-specific fields pulsatrix can't fill. Check against the validator first.
- VIZ-6. Reference designs: SAEDashboard (jbloomAus), Neuronpedia feature pages, the TensorBoard projector, and SHAP's plots.
- VIZ-7.
- imnodes: maintained, but has no zoom.
- imgui-node-editor: upstream is stale; the imgui_bundle fork is maintained (discussion #428, 2026-01) [C].
NB: Notebook layer¶
- NB-1. xeus-cpp's
xcpp::displaycallsmime_bundle_reprwithout qualification, so argument-dependent lookup finds an overload in thepulsatrixnamespace. This was read frominclude/xcpp/xdisplay.hpp. - Falsifier:
xcpp::display(tensor)in xeus-cpp 0.10 doesn't pick up the overload. - NB-2. The nbformat v4 format description: since 4.5, every cell needs an
id. MIME types ending in+jsonare stored as objects, and PNG data is base64. - Falsifier: the output fails
nbformat.validate, or doesn't render on GitHub. - NB-3. anywidget (MIT, 0.11.0, 2026-04-27) works in Jupyter, JupyterLab, Colab, VS Code and marimo.
- NB-4. xeus-cpp 0.10.0 (BSD-3-Clause, released 2026-04-02).
- Kernel crashes: a segfault or failed assert restarts the kernel (#434), and lambdas with xwidgets crash it (#476).
- GPU: CUDA math functions return 0 inside JIT-compiled kernels (#467).
- Windows: "Does not work on Windows" is still open (#144).
- xeus-cling hasn't been updated since 2025-10, so don't target it.
- NB-5. xeus-cpp-lite (Jupyter blog, 2025-06-19) loads libraries as emscripten-forge side modules. Threads need SharedArrayBuffer headers [I].
HIP: Training efficiency on AMD GPUs¶
- Hardware ceiling.
- Memory bandwidth: 256 GB/s theoretical (LPDDR5X-8000, 256-bit), about 212 GB/s measured on the GPU [C] (llm-tracker Strix Halo page).
- fp16/bf16 peak: about 59 TFLOPS from WMMA [V] (GPUOpen, "WMMA on RDNA3"). About 37 TFLOPS bf16 measured with an optimized stack [C].
- fp32 peak: about 15 TFLOPS without dual issue [I].
- So nearly every non-GEMM kernel is bandwidth-bound or worse [I].
- warpSize and shuffles. In ROCm 7.0,
warpSizeis no longerconstexpr, and*_syncshuffles take 64-bit masks [V] (ROCm 7.0 release notes; ROCm/HIP #3667). - gfx1151 ISA. The LLVM AMDGPU target table lists gfx1151 as an APU with no XNACK [V] (LLVM AMDGPUUsage).
- HIP-1.
- rocprofv3 (ROCprofiler-SDK) is the standard profiler in ROCm 10.0 [V].
- rocprof-compute support on gfx115x landed in TheRock PR #8299 [C]. ROCm 10.0 fixed its roofline precision list [V]. It's unverified on 7.2.4.
- HIP-2.
- Falsifier: rocprofv3 shows these kernels take under 5% of step time on real workloads.
- Fails when:
- a 32-bit mask truncates the shuffle on 64-wide waves;
- results differ from the CPU in the last bit (use tolerances);
- a wave64 build changes the reduction order.
- HIP-3.
- Design: PyTorch's caching allocator, described by Zach DeVito (2022) and in the PyTorch devlog on fragmentation (2026-06-01).
- Avoid
hipMallocAsyncon RDNA [C]:- silent corruption on gfx1030 (CTranslate2 #2090);
- OOM reported while memory is still consumed (rocm-systems #11642);
- double free on pool failure (rocm-systems PR #11643).
- VRAM on the APU [C]:
- ROCm adds VRAM and GTT together when reporting capacity, which overstates what can be allocated (ROCm #6004).
- AMD recommends a 0.5 GB carve-out on Strix Halo [V] (AMD Strix Halo system optimization guide).
- HIP-4. Falsifier: the GPU is already more than 90% busy with the syncs in place.
- Fails when: the host reads stale data, or an error surfaces in a later op.
- HIP-5.
dotandsumlaunch a single 256-thread block today (src/gpu_kernels.cuh). Use per-block partials and a second pass; no atomics, to keep results deterministic. - HIP-6. Fails when: the fused training path and the unfused explain path drift apart numerically. Gate them with a conservation test.
- HIP-7. At N=64, a 3×3 kernel on 224×224 input needs about 350 MB of im2col buffer [I].
- MIOpen on gfx1151 has a history of trouble [C]: missing Composable Kernel libraries (TheRock #5105, closed), and a Winograd lockup (TheRock #5581).
- HIP-8. Sources disagree [C].
- Against: ROCm #5643 (ROCm 7.1.0) reports "unsupported architecture".
- For: the changelog and community guides say ROCm 7.2.4 ships gfx1151 kernels.
- An older run measured hipBLASLt slower than rocBLAS.
- A 30-line probe settles it.
- HIP-9.
- ROCm 10.0.0 (2026-08-26) is the first release whose notes list the Ryzen AI Max parts as gfx1151 [V] (release notes, compatibility matrix).
- ROCm 7.2 page faults on host-to-device copies (ROCm #5890, open) [C].
- A wrong VGPR count crashed gfx1151 through ROCm 7.1.x. The user-mode fix landed in 7.2; the kernel-mode fix needs Ubuntu's OEM kernel 1018 or newer (TheRock #2991) [C].
- AMD asks for host kernel 6.18.4 or newer on other distributions [V].
- HIP-10. HIP's unified-memory documentation [V]: without XNACK,
hipMallocManagedandhipHostMallocgive pinned, zero-copy host memory. - HIP-11. LRP's stabilized divisions and conservation checks degrade in bf16, which has an 8-bit mantissa [I]. That's why the rule is to explain in fp32.
- HIP-12. Reports that this work may not pay off [C]:
- hipGraph is 2–5% slower than eager execution on gfx11 (legacy-rocm-build #6685);
- HIP graph branches get serialized on gfx1201 (rocm-systems #12330);
- Composable Kernel's flash attention on RDNA3 is forward-only (CK #1434);
- rocWMMA flash attention regresses prefill by 41% on gfx1151 (llama.cpp #24437).
AGT: Agents¶
- No official C++ SDKs.
- Anthropic's official SDKs are Python, TypeScript, C#, Go, Java, PHP and Ruby [V] (client SDK docs, checked 2026-10-03).
- MCP's official SDK tiers [V] (modelcontextprotocol.io/docs/sdk):
- Tier 1: TypeScript, Python, C#, Go, Rust and Ruby.
- Tier 2: Java.
- Tier 3: Swift, PHP and Kotlin.
- Community C++ MCP libraries.
- gopher-mcp (Apache-2.0, v0.1.18): its own issue #294 says 2026-07-28 support had only been tested against itself.
- cpp-mcp (MIT): targets older spec revisions.
- Neumann-Labs/mcp-cpp: GPL-3.0, so it can't go into an MIT library.
- AGT-1. nlohmann/json (MIT, v3.12.0). xeus-cpp's display API also uses it, so one dependency serves AGT, VIZ-1 and NB-1. Throughput doesn't matter at these payload sizes.
- AGT-2.
- cpp-httplib (MIT, v0.59.0): requires OpenSSL 3.0 or newer. Issue #2596 (opened 2026-09-29) reports Windows certificate verification failures.
- cpr (MIT-style) wraps libcurl, which can use Windows' own TLS stack.
- Falsifier: Windows CI fails to verify
api.anthropic.com. Then switch the default to libcurl. - AGT-3: Anthropic API behavior [V] (define-tools, handle-tool-calls, rate-limits and streaming docs, 2026-10-03):
tool_choice: anyandtool_choice: toolreturn HTTP 400 on the current 5.x models.tool_resultblocks must come first in the next user message.- A 429 from a spend limit has error code
enforced_spend_limit_reachedand noretry-afterheader. - New stream event types can appear at any time.
- Streamed tool input arrives as
input_json_deltafragments, parsed atcontent_block_stop. - AGT-4.
- MCP schemas default to JSON Schema 2020-12, but the main C++ validator (pboettch/json-schema-validator 2.4.0) supports draft-7 only.
- Falsifier: more than two real tools need
$reforoneOf. - AGT-5.
- MCP 2026-07-28 changelog [V]: no
initializehandshake, no sessions, andserver/discoveris required. - stdio transport [V]: nothing but protocol messages on stdout, and no embedded newlines.
- tools spec [V]: servers must validate inputs, apply access controls, rate-limit and sanitize output.
- Security statistics about popular MCP servers (for example, 43% with command injection) come from secondary vendor reports (Checkmarx), not primary scans.
- AGT-6.
- Grammar-constrained decoding is fragile at the edges of the schema language (llama.cpp #25923, #22314, #29006).
- XGrammar (Apache-2.0, v0.2.8) has a C++ core.
- SmolLM2-135M reaches about 22% exact match on xLAM function calling even after full fine-tuning (chloe-mp/smollm2-xlam-function-calling) [C]. Grammar constraints guarantee well-formed calls, not correct ones.
- AGT-7. No prior art or benchmark was found for attributing an agent's tool choice. The falsifier is the deletion test in the roadmap.
- AGT-8.
- "Language Models Can Explain Neurons": Bills et al., OpenAI (2023-05).
- Delphi: Paulo et al., arXiv 2410.13928 (2024-10). Its explanations fail about 38% of the time (arXiv 2605.01555, 2026-05).
- Output-centric descriptions: arXiv 2501.08319.
- AGT-9. MAIA: Shaham et al., arXiv 2404.14394 (ICML 2024). Transluce's Docent (Apache-2.0) analyzes agent transcripts.
- Prior art. mozilla-ai/agent.cpp (MIT, 2025-12) is a C++17 agent loop tied to llama.cpp, with no MCP. It's a design reference, not a dependency.
KS: Kitchen sink¶
- KS-1. The README says "There's no
install()step yet." - KS-2. The benchmark suite provides the numbers HIP-1 through HIP-8 are judged by, plus conservation error so performance work can't silently break LRP.
- KS-3.
- Monte Carlo dropout: Gal and Ghahramani 2016.
- Deep ensembles: Lakshminarayanan et al. 2017.
- Split conformal prediction: Vovk et al.; Angelopoulos and Bates, arXiv 2107.07511.
- KS-4.
- FGSM: Goodfellow et al., arXiv 1412.6572.
- PGD: Madry et al., arXiv 1706.06083.
- KS-5. TracIn: Pruthi et al., NeurIPS 2020. Captum 0.8 added influence functions [C]. Kronfluence's EK-FAC covers only linear and conv layers [C].
- KS-6.
docs/data-pipeline/index.mdcurrently calls drift detection and fairness metrics out of scope. Model cards: Mitchell et al., FAT* 2019. - KS-7.
DataThreadPoolandBoundedQueueexist but aren't wired intoDataLoader(data_loader.hpp). - KS-8. The HIP backend is tested on real gfx1151 hardware but not in CI (README Status).
- Risk: runner flakiness from the gfx1151 driver issues listed under HIP-9.
- KS-9. Llama 3.2 is under the Llama 3.2 Community License. Ship converters, not converted weights.
- KS-11. Python wheels with GPU backends are the hard part. A CPU-only wheel and a vcpkg port (mlpack and dlib both ship one) come first.
- KS-10, KS-12. These close gaps against the other libraries we compared:
| Library | What it has that pulsatrix lacks |
|---|---|
| PyTorch, JAX | Views, broadcasting, op-level autograd |
| tinygrad | Automatic fusion |
| ggml and llama.cpp | GGUF and quantization |
| Captum | Influence functions |
| Quantus | Evaluation metrics |
| TransformerLens and nnsight | Hook points |
| SAELens | SAE training at scale |