Open interactive terminal ↗ · Try research or skills, then trace a decision.
I'm Wenhao XU, a Machine Learning PhD student at Université de Montréal & Mila, working on rule alignment for language models.
Do models actually use the rules we give them? I study how explicit rules shape model decisions, with an eye toward AI we can understand and trust.
AI Safety · Alignment · Trustworthy AI · Mechanistic Interpretability
I build open-source agent skills that turn "looks good" into something you can inspect, replay, or challenge.
| Skill | What it does |
|---|---|
| RageClick | Break a web app like an impatient user. Replay the failure. |
| SkillClash | Find conflicting agent instructions. Trace them to the source. |
| RepoQuest | Turn a real bug into a playable debugging mystery. |
| PaperCourt | Put empirical ML claims on trial against their evidence. |
| ProveIt | Check that a regression test actually catches the bug. |
Play RepoQuest → · Watch RageClick →
TaxAgent — local tax preparation / ReviewBudget — pull request verification planning.
Montreal / LinkedIn / All repositories ↗


