Skip to content
View zanwenfu's full-sized avatar
🎯
Hustling
🎯
Hustling

Highlights

  • Pro

Block or report zanwenfu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
zanwenfu/README.md

Zanwen Fu, founder and engineer. Make something people want. VYNN AI, Robinhood, AutoCodeRover (acquired by Sonar), Binance.

zanwenfu.com  ·  LinkedIn  ·  X  ·  zanwen.fu@duke.edu

Agents are easy to demo and hard to depend on. I build the ones people depend on.

I founded VYNN AI and built every part of it myself. Before that I was an early employee at AutoCodeRover, one of the first coding agents, which Sonar acquired. I've shipped agentic AI at Robinhood and reliability infrastructure for Binance's Web3 Wallet. What I took from all of it: the model is rarely what decides whether an agent holds up. Everything around it is. That's what I build now.

Building

VYNN AI: a personal financial analyst for every investor
Founder and sole engineer · 2025 to now · how it's built · agent code

Ask about any company, fund, coin or prediction market and get a sourced report and a live Excel model in about two minutes, for about three cents. More than 5,000 investors have signed up, and the first 500 came without a dollar of marketing. The model never writes a number. Code computes every figure, and when VYNN disagrees with Wall Street, it stands by its number and shows you both.

Agent OS: an operating system for AI agents
Creator · 2026 · how it works · the thesis

Every team building agents rebuilds the same plumbing: memory, rollback, checks, budgets. Agent OS is the layer underneath them. A central brain plans the work, each piece runs in its own process, and a monitor that can't touch anything signs off. Every step is a commit in git, so nothing unverified ships and nothing is ever lost.

Errata-Bench: a self-improving benchmark of whether coding agents tell the truth
Creator · 2026 · leaderboard · code · dataset

Coding agents end their work with a report, and developers act on it. Errata-Bench checks every claim in that report against what the agent actually did, and it grows from real moments where a developer caught one misreporting. No model was reliably honest: 44–73% of each model's answers claimed something it hadn't established. When an agent left a bug unfixed, 2.5% of its reports said so.

Shipped

Robinhood: Agentic AI team
Machine Learning Engineer Intern · 2026 · Robinhood Cortex

I solo-designed and shipped a proactive agent for Robinhood Cortex that decides when market news deserves a customer's attention, instead of waiting to be asked. It cut false positives five-fold in backtesting with no material event missed and scaled coverage 30× at flat latency. I also caught and fixed a production delivery failure that monitoring had missed.

AutoCodeRover: one of the first coding agents, acquired by Sonar
Early employee · 2024 to 2025 · IDE plugin · the story

In 2024, before Claude Code or Codex, it fixed real GitHub issues on its own. I worked on lifting it to 51.6% on SWE-bench Verified and built Self-Fix, which traces a rejected patch back to the step that went wrong. I also built the JetBrains plugin end to end, which merges the agent's fix into the developer's latest code. After Sonar acquired it in 2025, the former AutoCodeRover team's Foundation Agent reached #1 on SWE-bench's unfiltered leaderboard.

Research

  • LUMINA · first author. Four agents that screen studies for medical systematic reviews: 98.2% sensitivity across 15 published reviews, with 35× fewer missed studies than a published baseline, at under a cent per citation.
  • architectural-damping · Prompt injections fooled VYNN's language model every time. In a 12-case pilot, the calculator behind it stopped 10 of them from reaching what users see, and reading its source predicted which ones would get through.
  • speculative-decoding-t4 · Sequoia's cost model predicts a 1.68× speedup on a T4. I measured 0.56×, and one measured cost explains the gap to within 1.1%.
  • football-llm · My fine-tuned Llama seemed to beat XGBoost at World Cup predictions. With team names hidden, its exact-score accuracy fell from 43.8% to 10.9%. It had memorized the 2022 tournament.

Writing


Off the keyboard: fifteen years of clarinet.
Making something people want? I'd like to hear about it: zanwen.fu@duke.edu

Pinned Loading

  1. Agentic-Analyst/stock-analyst Agentic-Analyst/stock-analyst Public

    The agent behind VYNN AI, a personal AI financial analyst with 5K+ users. Ask any market question; code computes every number and every claim is sourced.

    Python 129 8

  2. taste-is-all-you-need taste-is-all-you-need Public

    Agent OS: an operating system for AI agents. A central brain plans, worker processes act, observe-only monitors certify, and git is the memory.

    Python

  3. jetbrains-ide-plugin jetbrains-ide-plugin Public

    AutoCodeRover's JetBrains plugin: autonomous code repair inside the IDE, with live agent streaming, developer feedback, and three-way AST merges that keep your edits.

    Kotlin

  4. agentic-reviewers-for-SRMA agentic-reviewers-for-SRMA Public

    LUMINA: four agents that screen citations for medical systematic reviews. 98.2% sensitivity across 15 published reviews. First-author research.

    Python 2 1