Capability-based compiler/runner for reproducible agent scenarios
-
Updated
May 22, 2026 - Rust
Capability-based compiler/runner for reproducible agent scenarios
A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.
Business Process AI Worker · τ²-Bench #1 globally (3/3, 100%) · CRMArenaPro Run 8 · Reflexive Agent Architecture
Code translator green agent (Judge) of an agent-as-a-judge programming languages translation.
Fault-injecting OpenEnv training environment for vibe-coded SaaS incidents. 30 scenarios grounded in 2025-26 production failures. Drop-in OpenClaw-RL pool server. Claude Code skill included.
Deterministic offline ComtradeBench judge for evaluating agent robustness under pagination, retries, duplicates, page drift, and totals traps.
Baseline purple agent for the ComtradeBench benchmark: UN Comtrade tool-use under adversarial API conditions.
Winner of the tau2-bench competition on AgentBeats (Berkeley AgentX) at the competition close: a customer-service purple agent reaching 82.5% in the telecom domain. A plain A2A service with no agent framework; the edge is distilled per-domain policy playbooks injected ahead of the task policy.
A2A green-agent orchestrator for evaluating agents on the AppWorld benchmark, built on the AgentBeats SDK
AI-powered clinical triage simulation using Manchester Triage System (MTS). OpenEnv Challenge 2026 entry with A2A protocol support.
Leaderboard infrastructure for the ComtradeBench / AgentBeats agent-evaluation benchmark: task definitions, submission flow, and scoring.
To associate your repository with the agentbeats topic, visit your repo's landing page and select "manage topics."