Heterogeneous Multi-Model AI Agent Swarm · Autonomous Security Automation
無敵 · Project Muteki is an open-source, multi-agent security framework built with Remix Engineering. A Coordinator schedules different Agent engines toward one goal, manages their context, and lets them collaborate through shared evidence. The current Worker engines are Claude, Codex, Cursor, Pi, OMP, Kimi, Grok, and OpenCode. The project will keep expanding beyond CTF toward a broader Agent workbench.
Use Muteki only against challenges and targets you own or are authorized to test. Muteki drives CLI Agents that can run commands, invoke security tools, and access target services. It does not isolate malicious challenges.
A dedicated disposable VPS, VM, or machine without sensitive data is recommended. Avoid shared hosts and production systems. See SECURITY.md. I often run it directly on my own computer because that is easier to set up, but that does not remove the risk.
At RIFFHACK 2026, Muteki ran for three hours without human takeover, solved every challenge, and placed eighth.
On the iChunQiu Yunjing blackmaze penetration-testing range, which had seen no solves for three months, Muteki took first blood in about two hours of actual solving. The platform shows 39 hours because debugging and multi-Flag development were interleaved with that run.
It also cleared the Yunjing badge scenarios and Hack The Box challenges across Insane and Hard difficulty levels. The NYU CTF Bench evaluation reached 200/200; those figures reflect the tested setup and model versions, not a promise about every new challenge.
The TSecBench hosted-mode run with deepseek-flash ranked 13th: result page.
After a month of engineering work, tuning, and fixes, Muteki is open source without a paid tier or private subgroup. Please report problems in Issues and help improve it. If the discussion group is full, the QR code below is another way to get in touch.
The current version is 0.4.0. See CHANGELOG.md for new features, behavior changes, removals, and upgrade notes.
Start with the single-challenge CTF workspace. The home page shows this workspace by default. Conversation, competitions, and custom extensions are still being tested and can be enabled as described below. The Web pentest mode is being reworked and is currently unavailable.
You need uv, Python 3.13 or newer, and Node.js/npm. For local solving, install and sign in to at least one supported Agent CLI. ./init.sh syncs Python dependencies; the first Web start installs and builds the frontend. The shortest manual path is:
git clone https://github.com/FishCodeTech/muteki.git
cd muteki
./init.sh
./run.sh webOpen http://127.0.0.1:3001. The backend defaults to port 8000; both bind to loopback by default. The first frontend build and service initialization may take a while. Use the addresses and readiness output shown in the terminal. Stop the service with Ctrl+C.
- Click Single-task Workers on the home page, or open Worker settings from the lower-left corner of the single-task page.
- Under Agent credentials, check an existing CLI login or add a token, API key, or custom Base URL. Use the real connection test.
- Under Worker lineup, enable at least one Worker and bind it to a usable account and model. Run the full self-check or check that Worker individually.
- Under Runtime environment, choose Local for host CLIs or Container for Docker Workers. Container mode needs a Worker image and injectable credentials.
- Under Reasoning models, select and test a Planner endpoint and model. Configure the Titler if needed.
- Save the settings and return to Single task.
Local mode can reuse an Agent CLI login already available on the host. A container cannot inherit that login automatically; bind an injectable account in the credentials page. To get started quickly, use one CLI that already works in your terminal and select Local mode.
On Single task, enter the challenge name, original description, target URL, known facts, and Flag format. Add attachments with the picker, paste, or drag and drop. Check single versus multiple Flags, Web tools, runtime, and advanced options, then click Dispatch swarm (⌘↵ on Mac). If you leave the category empty, Muteki infers it from the challenge. The first run also prepares a workspace and Workers, so watch the status rather than clicking dispatch repeatedly. Defaults are generally a good starting point.
Only run challenges and targets you own or are authorized to test. Workers can execute commands and access target services.
The project has been tested on macOS 26 and Ubuntu 24.04. On Windows, we recommend running Muteki inside an Ubuntu 24.04 VMware virtual machine.
| Runtime | Where Muteki and Workers run | Tool setup |
|---|---|---|
| macOS local | Web and Workers on the Mac | Run ./ctf-tools/setup.sh for native CTF tools |
| macOS + Docker | Web on the Mac; Workers in Docker | Pull the full Worker image, which includes the CTF toolchain |
| Ubuntu 24.04 local | Web and Workers on Ubuntu | Use the installer with --with-ctf-tools, or run ./ctf-tools/setup-ubuntu.sh |
| Windows + VMware | Windows hosts the VM and browser; Muteki and Workers run inside Ubuntu 24.04 | Run the Ubuntu installer in the VM |
Install Homebrew, Node.js/npm, and the Agent CLI you intend to use. Sign in using that vendor's CLI, then run from the repository root:
./init.sh
export PATH="$HOME/.local/bin:$PATH" # If newly installed uv is not yet on PATH
./ctf-tools/setup.sh
./run.sh webThe setup script uses ctf-tools/Brewfile, prepares Python/Ruby tool environments, and creates command entry points. It does not sign in to Agent CLIs. Select Local under Single-task settings → Runtime environment. When ctf-tools/.ready exists, run.sh loads its tool paths automatically. See ctf-tools/README.md for optional large dictionaries and knowledge copied from the Worker image. To refresh only existing tool links, run ./ctf-tools/setup.sh --link-only. macOS native tools and those inside the Kali image are not identical.
Install and start Docker Desktop, then run:
./init.sh
docker pull ghcr.io/fishcodetech/muteki-worker:latest
./run.sh webSelect Container under Single-task settings → Runtime environment. Check that Docker and the Worker image are available, then bind injectable credentials to the selected Workers. Host CLI logins are not automatically copied into containers. The full image already contains the CTF toolchain, so ctf-tools/setup.sh is not needed on the Mac for this mode. The Web app stays on the Mac; Docker starts Workers per run.
git clone https://github.com/FishCodeTech/muteki.git
cd muteki
./scripts/install-ubuntu.sh --backend local --preflight
./scripts/install-ubuntu.sh --backend local --with-ctf-tools
./scripts/install-ubuntu.sh --backend local --checkThe installer prepares application dependencies, Node.js, the Web service, and local Worker CLIs; it calls ctf-tools/setup-ubuntu.sh for the available CTF tools. Agent CLIs still need their own login or configured credentials. By default it creates muteki-web.service, stores the Web password in ~/.config/muteki/ubuntu.env, and serves the UI on port 3001. Optional apt packages may be unavailable on some mirrors or architectures; the script reports those individually.
If Muteki is already installed and running, you can install just the local CTF tools with ./ctf-tools/setup-ubuntu.sh after ensuring uv is available. That script has only been tested on Ubuntu 24.04; other Linux distributions may need changes. Once ctf-tools/.ready exists, ./run.sh web loads the tool paths.
Create an Ubuntu 24.04 VM, clone the repository inside it, and run the same installer commands there:
git clone https://github.com/FishCodeTech/muteki.git
cd muteki
./scripts/install-ubuntu.sh --backend local --with-ctf-tools
./scripts/install-ubuntu.sh --backend local --checkMuteki, Agent CLIs, CTF tools, and run data stay inside Ubuntu. Open http://<VM-IP>:3001 from Windows and use the Web password saved in the VM's ~/.config/muteki/ubuntu.env. If VMware NAT prevents access to the VM IP, configure port forwarding or use a suitable virtual network. The backend API still listens only on the VM loopback address by default.
| Location | What it controls | First action |
|---|---|---|
| Single-task settings → Agent credentials | Engine detection, host login, tokens/keys, custom endpoints, connection tests | Confirm at least one usable account; never paste keys into a challenge or commit them |
| Single-task settings → Worker lineup | Worker engines, models, reasoning effort, enablement, and per-seat concurrency | Enable one Worker that passes preflight, then add others |
| Single-task settings → Runtime environment | Local/container mode, container scope, network, image, optional VPN | Usually keep the default |
| Single-task settings → Scheduling and budgets | Automatic versus fixed dispatch, concurrency, time, Worker count, cost cap | Keep defaults initially; set limits when needed |
| Single-task settings → Reasoning models | Planner and Titler endpoints, models, generation parameters | Test a real Planner request |
| Settings → Appearance → Workspace mode | Visibility of conversation and competition workspaces | Enable them to try the additional workspaces |
Workers and the Planner are configured separately. Worker CLIs execute the task. The Planner makes coordination decisions and proposes Steps. Local Workers may reuse host logins, while the Planner needs a working model endpoint. Add a compatible custom endpoint in credentials/model settings, then bind it to the appropriate Worker or Planner. Saving settings affects future runs; it does not rebuild a run already in progress.
The UI can configure Claude Code, Codex, Cursor, Pi, OMP, Kimi, Grok, and OpenCode. Whether each one works depends on installation, vendor login, and the chosen runtime. DeepSeek Harness is registered but cannot currently be selected as a Worker because its CLI does not provide the required structured provider capabilities. Follow each vendor's installation and login instructions.
If the page says Needs self-check or No usable credential, test the account under Agent credentials, then check the Worker's account, model, and runtime binding. Self-check sends a real model request and may incur usage.
- Description: Provide the original prompt, category, URL or host and port, Flag format, and facts you have verified. Do not present guesses as confirmed facts.
- Attachments: Upload challenge files. The run workspace keeps inputs and Worker artifacts for later review.
- One or several Flags: A single-Flag run can stop at a qualified candidate. A multi-Flag run keeps collecting; enter the expected count if known, or watch the timeout, budget, and manual stop controls.
- Web tools: Control the Agent's WebSearch and WebFetch capabilities so unrelated external searches do not pull it away from the target.
- Advanced: Set the Flag format, allow operator-input requests, and bound time, Worker count, and cost for this run. Match the competition's Flag rules.
The run page shows Coordinator messages, Worker state, activity, evidence, and candidate Flags. Target all Workers or one Worker from the selector above the composer. Controls change with run state:
| Action | Effect | When to use it |
|---|---|---|
| Type and press Enter / Send | Sends a hint to the selected target; does not create a Step | Add a clue, correct the challenge text, or share a known result |
| Dispatch | Creates a new Step from the input; a URL in the text may update the target | Ask for a specific new direction or action |
| Ask progress | Summarizes verified findings, active directions, and blockers in the conversation | Check status without changing the task |
| Pause / Resume | Pauses and resumes swarm scheduling | Inspect the challenge or wait for operator input |
| Freeze / Unfreeze | Immediately freezes active Workers and releases them later | Stronger intervention in running work |
| Stop | Stops this run and its Workers while keeping the record | The target changed, the run should end, or a limit was reached |
A hint differs from Dispatch. Typing “The /admin directory is confirmed” and pressing Enter adds information. Clicking Dispatch creates a new Step and may launch a Worker. Avoid repeatedly clicking controls while the run is updating.
If a Worker asks for input, answer in the pending card, provide the needed resource, or reject/correct the request so the run can continue.
- Copy the Flag and verify it on the challenge platform.
- Mark false positive: If the platform rejects a candidate, mark that specific Flag as a false positive (
×on its row does the same). For multi-Flag runs, select the candidate first. This reopens solving. - Continue solving: Restart the swarm on the finished run while preserving its existing evidence.
- Ask / Generate writeup: Ask the solving Worker a follow-up or generate a report. On success, the report appears in the conversation and in
sessions/<run-id>/workspace/writeup.md(or under your customMUTEKI_SESSIONS_ROOT). Generation summarizes the run's full context and can take time.
The home page shows only single-task solving by default. In Settings → Appearance → Workspace mode, enable Show conversation and competition modes to reveal Conversation and Competition on the home page, navigation, and search. These workspaces are still being tested; please report issues.
Open Conversation to work with an external Agent across multiple turns and review tool calls, approvals, artifacts, and history. Configure its Agent runtime, credentials, and permissions under Settings → Agents first. Live guidance depends on the connected transport. Its controls are separate from those for a single CTF run.
Create and test a platform connection, then register the remote competition ID. The workspace can sync challenges, schedule single-task runs, inspect candidates, and track remote submission verdicts. CTFd, rCTF, and GZCTF connection options are available; actual capability depends on the probe result. Test account and submission behavior in a practice competition first.
Open Settings → Extensions (or Extensions on the home page when all workspaces are enabled). Sources may be a local directory, archive, Git, HTTP, or Catalog. An extension root needs an Agent Plugins 1.0.0 plugin.json. Generate an install preview, review its source, digest, and permissions, then install. Installed plugins can be enabled, configured, upgraded, rolled back, or removed. Install only extensions you trust.
./run.sh web # API :8000 and UI :3001
./run.sh web --backend-only # API only
./run.sh web --ui-port 3002 # Different UI port
./run.sh web --rebuild-ui # Rebuild after changing UI sourceCopy .env.example to .env at the repository root for MUTEKI_* settings; exported shell variables take precedence. Keep secrets only in ignored locations such as .env and state/_secrets/. If binding Web to a non-loopback address, set MUTEKI_WEB_PASSWORD first or the backend will refuse to start.
Compose starts the FastAPI and Next.js control plane. The Docker daemon starts Workers on demand. Prepare Docker and the Worker image, then set an absolute host data path and Web password:
docker pull ghcr.io/fishcodetech/muteki-worker:latest
MUTEKI_HOST_DATA_ROOT=/opt/muteki/data \
MUTEKI_WEB_PASSWORD='replace-with-a-strong-password' \
docker compose up --buildThe UI remains at http://localhost:3001 by default. MUTEKI_HOST_DATA_ROOT must be an absolute host path shared by the control plane and Workers. The macOS Docker Desktop path has been exercised; Windows Compose has not been verified end to end on real hardware. See .env.example and SECURITY.md for more settings and runtime boundaries.
./run.sh upgrade --check # Check the latest stable release
./run.sh install # Set up a managed installation
muteki upgrade v0.4.0 # Install this release
muteki rollback # Return to the previous installed release
muteki version # Show version and installation kindSingle-task settings → System update provides the same operations. Preserve .env, sessions/, and state/ during upgrades. Do not commit run records or credentials.
| Symptom | Check first |
|---|---|
| Page does not open | The address and ports printed by ./run.sh web; Node/npm, first UI build, and state/_logs/backend.log / state/_logs/ui.log |
| No usable Worker | CLI installation and login, Agent credential test, enabled Worker lineup, account and model binding |
| Container run fails to start | Docker service, Worker image, container credentials, network/VPN, and Runtime environment settings |
| Planner test fails | Endpoint, model, key/Base URL, and network; a working Worker CLI does not configure the Planner |
| Conversation, Competition, or Extensions missing | Enable their visibility in Settings → Appearance → Workspace mode |
| Platform rejects a Flag | Mark that candidate as a false positive, continue solving, and check the challenge text and Flag format |
| Multi-Flag run keeps going | Enter the expected count; otherwise stop manually or use a time/worker/cost limit |
- Code map:
muteki/holds the backend core;apps/web/is FastAPI;apps/web/ui/is Next.js;docker/contains image configuration. - Issues: Report bugs and documentation problems in GitHub Issues, with reproduction steps, version, and redacted logs. Report security vulnerabilities privately as described in SECURITY.md.
- License: GNU AGPL-3.0. External Agent CLIs and model services have their own licenses, terms, and charges.
- Rework pentest mode
- Add a source-code vulnerability discovery mode
- Expand the integrated Agent chat experience
- Make the product fully plugin-based
Thanks to c3 for the Yunjing range account. I used a lot of grit and made plenty of popcorn along the way.
Thanks to l4n for the inspiration; the reviewer brought a major improvement to solving efficiency.
Thanks to 陈橘墨 for range resources and writeups used in testing and tuning.
Thanks to Sam Altman for not banning my account. It got banned; I will remember his name.
Thanks to Dario Amodei for not banning my account. It got banned too; I will remember his name.
This project's design and evaluation drew on the following work:
- NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security — Minghao Shao, Sofija Jancheska, Meet Udeshi, Brendan Dolan-Gavitt, et al. NeurIPS 2024 Datasets & Benchmarks Track. arXiv:2406.05590
- Teams of LLM Agents can Exploit Zero-Day Vulnerabilities — Richard Fang, Rohan Bindu, Akul Gupta, Daniel Kang. EACL 2026. Paper
- D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security — Chenhui Zhang, et al. 2025. arXiv:2502.10931
- HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing — Lajos Muzsai, David Imolai, András Lukács. 2024. arXiv:2412.01778
- CTFAgent: An LLM-powered Agent for CTF Challenge Solving — Jiaze Sun, et al. Computers & Security, 2025. ScienceDirect
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents — Jiahao Zhu, et al. 2025. arXiv:2602.02164
- Related project article












