Self-Evolving Code Repair & Evaluation Agent based on LLMs
Published:
Overview
Developed an advanced, self-evolving code repair and automated evaluation agentic system. The system automates the full lifecycle of detecting, localizing, patching, and testing bugs in complex software repositories.
Key Technical Contributions
- Agentic Workflow & Multi-Stage Closed Loop: Built a core code repair agent on Google Gemini with a multi-stage closed-loop: Intent Analysis → Dependency Localization → Patch Planning → Sandboxed Execution → Error Reflection & Retry. Achieved a 100% tool invocation success rate and an 11.40% end-to-end task success rate in complex long-context environments.
- Benchmark Evaluation & SOTA Performance: Conducted extensive ablation and benchmark comparisons against mainstream baselines (e.g., Agentless, OpenHands); built an automated evaluation framework achieving execution verification metrics of EV-Micro 8.50% and EV-Macro 23.03%, significantly outperforming open-source SOTA baselines.
- High-Availability Sandboxed System Engineering: Built an asynchronous orchestration hub with Django REST Framework, PostgreSQL, Celery, and Redis; containerized multi-agent execution environments with Docker and automated Pytest feedback loops, averaging only ~0.35M tokens per task to optimize evaluation precision and compute cost.
Tech Stack & Tools
- Agent Frameworks: ReAct, Agentic Workflow, Tool Use / Function Calling, Self-Reflection
- Infrastructure & Sandboxing: Docker Containerization, Celery, Redis, PostgreSQL, Django REST Framework
- Testing & Verification: Pytest automated test feedback loops, SWE-bench style evaluation
