A Mechanistic Interpretability Toolkit for Cross-Layer Transcoder Training and Attribution-Graph Visualization
-
Updated
Jul 30, 2026 - Python
A Mechanistic Interpretability Toolkit for Cross-Layer Transcoder Training and Attribution-Graph Visualization
Dyno Lab — an AI safety, alignment and interpretability research workbench for Apple Silicon. Explore activations, probes, SAEs and interventions with local inference, a Python SDK, APIs and MCP.
Fast, standardized, and easy-to-use interpretability engine.
Universal probing and interpretability tool for MLX language models on Apple Silicon
AI Safety research platform for studying personality drift in AI systems using mechanistic interpretability and clinical assessment tools. Complete simulation framework with neural circuit analysis, statistical drift detection, and intervention protocols.
Framework for evaluating and steering generative image systems using geometry-first metrics, structural stress testing, and constraint-based analysis. Designed to expose compositional collapse, spatial priors, and model failure modes without accessing training data or model internals.
OKI TRACE: Local LLM observability. See step-by-step, layer-by-layer what your AI thinks. Logit Lens & Attention for HuggingFace models.
NOMOS v0.3.0 is an auditable decision orchestration layer built on deterministic v0.2 foundation, scoring 95/100 in IMDA AI Verify. Equipped with native hash-chained audit engine, it records every decision step for independent verification, generates rational candidates only based on input facts and reserves final judgment to humans.
Reproduce Emotion Concepts and their Function in a Large Language Model on Qwen 3.6 27b.
A J-space-inspired AI visual art and interpretability playground for watching hidden-state word candidates swarm and collapse into language.
Human Retention Layer for AI Work — hand-solvable math shadow models + session reasoning distillation. Cross-platform agent skills for Claude Code, Copilot, Cursor, Windsurf, Cline, Codex CLI, Gemini CLI, and 10+ more.
👽 An Alien Mind — The Epistemic Operating System for the AI Age. Open-source epistemic layer between humans and AI: Cognitive Firewall, multidimensional Trust Profile, five-agent Alien Council, Thought DNA provenance, behavioral fingerprints, and Monte Carlo simulation. Not a chatbot. Not a guardrail. Not a lie detector.
Conformal Geometric Algebra (CGA) with efficient sequence modeling by introducing a recurrent rotor mechanism and a novel bit-masked hardware kernel that solves the computational bottleneck of Clifford products.
Official code of the project "Query Circuits: Explaining How Language Models Answer User Prompts" accepted to ICML 2026
An interactive, serverless WebAssembly dashboard demonstrating the statistical fragility of AI interpretability tools. Built for the alphaXiv Hackathon to simulate the multiple comparisons problem.
Toy 6. An interactive phase-space instrument mapping Ψ = S/D — the ratio of capability to modeling depth that determines whether a system is in the viable, transitional, or failure-mode-dominant regime. Includes the Inner Crossing animation. Companion simulation for The Inner Crossing — Series 2, Part 3.
Do LLMs think like brains? We test GPT-2, BERT, Mistral, DeepSeek & Qwen+SAE against EEG data. Sparse features yield a 4.3× alignment jump. Working paper included.
A NeuroAI project using Bernoulli-inspired fluid-flow analogy to explore how information moves through neural networks. The signal strength in the NN is defined as the "pressure" from Bernoulli's equation, the speed of information propagation as the "flow speed of fluid" and, the activation level as the "opening and closing of valves".
I Asked It to Forget, but It Didn't — A Case of Miscommunication Between AI and Humans
Open protocol for reusable AI artifacts, knowledge continuity, and check-before-create workflows across AI tools.
To associate your repository with the ai-interpretability topic, visit your repo's landing page and select "manage topics."