Self-Evolving Code Repair & Evaluation Agent based on LLMs

Published:

Overview

Developed an advanced, self-evolving code repair and automated evaluation agentic system. The system automates the full lifecycle of detecting, localizing, patching, and testing bugs in complex software repositories.

Key Technical Contributions

  • Agentic Workflow & Multi-Stage Closed Loop: Built a core code repair agent on Google Gemini with a multi-stage closed-loop: Intent Analysis → Dependency Localization → Patch Planning → Sandboxed Execution → Error Reflection & Retry. Achieved a 100% tool invocation success rate and an 11.40% end-to-end task success rate in complex long-context environments.
  • Benchmark Evaluation & SOTA Performance: Conducted extensive ablation and benchmark comparisons against mainstream baselines (e.g., Agentless, OpenHands); built an automated evaluation framework achieving execution verification metrics of EV-Micro 8.50% and EV-Macro 23.03%, significantly outperforming open-source SOTA baselines.
  • High-Availability Sandboxed System Engineering: Built an asynchronous orchestration hub with Django REST Framework, PostgreSQL, Celery, and Redis; containerized multi-agent execution environments with Docker and automated Pytest feedback loops, averaging only ~0.35M tokens per task to optimize evaluation precision and compute cost.

Tech Stack & Tools

  • Agent Frameworks: ReAct, Agentic Workflow, Tool Use / Function Calling, Self-Reflection
  • Infrastructure & Sandboxing: Docker Containerization, Celery, Redis, PostgreSQL, Django REST Framework
  • Testing & Verification: Pytest automated test feedback loops, SWE-bench style evaluation