1. Anatomy of Repository Coding Agents
4 lessonsUnderstand how coding agents interact with repositories, tools, and development environments to execute multi-step tasks.
2. Building Controlled Evaluation Environments
4 lessonsSet up isolated, reproducible repository states where agents can execute tasks without production risk.
3. Multi-Signal Grading Systems
5 lessonsDesign graders that combine functional tests, static analysis, and custom checks to evaluate agent output quality.
4. Task Stratification and Risk Frameworks
4 lessonsCreate policies that assign different trust levels and review requirements based on change type and repository area.
5. Eval Suite Construction from Real Work
4 lessonsBuild evaluation datasets from historical issues, pull requests, and production incidents to test agent capabilities.
6. Integration with Development Pipelines
4 lessonsEmbed agent evaluation into CI/CD workflows, pull request automation, and continuous delivery systems.
7. Governance, Auditability, and Regression Tracking
5 lessonsEstablish systems for logging agent decisions, tracking performance over time, and maintaining compliance requirements.
Want the full course when it launches? Join the waitlist and we will notify you.