Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
-
Updated
Sep 11, 2026 - TypeScript
Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
Meta-harness optimization loop wired onto Islo sandboxes. POC: 0/5→5/5 in four proposer steps. Built on islo.dev.
Observe an agent run and GroundEval drafts the policy and diagram for you, no hand-written policy required, then scores what it checked, what it skipped, and what it wasn't allowed to touch.
Scenario Testing for AI Agents
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
AI-operated company. Building agent-friend: universal tool adapter for AI agents. @tool → OpenAI, Claude, Gemini, MCP. Live 24/7 on Twitch.
Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.
Agent 测试转型手册:资深自动化测试工程师的 AI QE 学习路径 + 练手代码
Project page for Meta-harness on Islo (POC). https://zozo123.github.io/meta-harness-on-islo-page/
A reasoning benchmark runner for comparing LLMs as OpenClaw agents use them. 52 prompts, 3 eval sets, 11 traps, LLM-as-judge, tier-based leaderboard.
Transcript-first evaluation tool for comparing coding-agent sessions across Codex, Claude Code, and Pi.
Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, trigger quality, functional quality, regression protection, baseline value, model variance, rollout safety. Never gradients.
Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilot CLI, Codex CLI).
PandaProbe harness turns agent failures into fixes
A curated list of benchmarks, harnesses, leaderboards, and tools for evaluating AI coding agents.
检测 AI Agent 代理指标与真实业务结果背离的审计工具,面向增长与运营团队
A durable, long-running agent that improves and evaluates other agents.
开源通用 AI Agent 真实任务评测 · 同 Prompt、客观开奖、评分细则全公开 | Open-source evaluation of general-purpose AI Agents on real-world tasks with verifiable outcomes — by PingWest / 硅星人
Documented, reproducible finding: VulcanBench declarative grader mis-scored all functional tasks as 0.0 due to a repo-root pytest-cov addopts leak. Filed upstream issue #79.
Turn agent telemetry into eval jobs — outcome, quality, and spend — correlated with business outcomes per agent run.
To associate your repository with the agent-eval topic, visit your repo's landing page and select "manage topics."