Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
-
Updated
Jul 16, 2025 - Python
Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
Evidence-grounded AI agent for Java code repair with LLM-guided patches and human-in-the-loop review.
Exploring and improving the quality of ChatGPT-generated code for LeetCode programming tasks.
基于 AI Agent 服务自动化修复系统:Agent 自动读取错误日志,定位 Bug,生成补丁,运行测试,提交 PR,并通知开发者 Review。AI-powered auto-fix agent for web services: analyzes logs, patches code, runs tests, and creates pull requests automatically.
Repository-level automated code repair agent using SWE-Bench dataset
Trusted autonomy T&E runtime that links mission needs, hazards, scenarios, telemetry, evidence, verification reports, and hash-chained ledgers so AI/autonomous decisions can be reviewed instead of merely trusted.
Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
A reliability layer for AI-built systems: detect failures (tests or runtime drift), reproduce, repair one ticket at a time behind an approval gate, and prove the fix. The safety boundaries most AI agents skip.
Gymnasium RL environment for training LLM agents to autonomously debug and fix Python code with secure sandboxing and test-driven feedback.
A minimal lab for improving programs, agents and model parameters. Real demos, frozen evaluations, inspectable evidence. Zero runtime dependencies.
PyPatch— OpenEnv RL environment where AI agents debug & fix buggy Python code across 3 difficulty levels.
🦑 CT 11 — Secure Self-Learning Repair Agent. Droste Fusion. 4 Engines. Reflection Engine. 9/10 Benchmark.
Run broken Python code → it fixes itself. Local-first Python runtime repair.
When Free Executors Cost More: The Free-Executor Paradox in Iterative LLM Code-Repair Loops (paper + reproducibility kit)
AI proposes. Humans decide. Source-available AI assurance and live authority control plane for governed AI-produced code change: MCP/API pre-tool enforcement, authenticated agents, scoped authorization, evidence-conditioned authority, human-gated execution, revision-bound verification, chained receipts, and auditable replay.
5-agent LangGraph system that autonomously resolves GitHub issues, benchmarked across single-agent baseline, paid API, and open-weight HPC configurations on SWE-bench Lite
Self-healing code reasoning engine. Detective → QA → Patcher closed loop on SWE-bench, with an RL layer that turns every reasoning trace into DPO training data.
Execution-grounded reinforcement learning for software repair on SWE-bench
OpenEnv-based reinforcement learning environment for automated Python code debugging and repair.
To associate your repository with the code-repair topic, visit your repo's landing page and select "manage topics."