Project
Coach-Crew Service: Multi-Agent Self-Evaluation Framework Using AI Coaches
Objective
1. To build an AI-driven evaluation system where AI agents assess and improve the performance of other AI agents. 2. To implement a multi-layer coaching pipeline (Primary Generator → Lightweight Coach → Heavy Coach) for scalable, reliable, self-evaluating agent workflows. 3. To enable confidence-based bypass, fallback evaluation paths, and revision loops for improved output quality. 4. To provide a FastAPI-based interface that exposes endpoints such as /ask, /evaluate, and /batch-evaluate for developers and researchers. 5. To integrate Redis-based memory management and fallback mechanisms for tracking agent history and intermediate reasoning.
Outcome
The project is expected to deliver a functional AI coaching framework capable of generating refined answers along with confidence scores and intermediate reasoning steps. Users interacting with the service will receive evaluated, validated, and improved responses rather than raw model outputs. The system will demonstrate how hierarchical coaching layers enhance quality and correctness through automated review loops. Upon completion, it will serve as a foundational model for future AI governance and agent-based evaluation research.
Apply By Date 01 Oct 2025
Students 0 / 1
Duration 3 months
Mentor Aakriti Aggarwal
Tools-Technologies
Data Science Experience, Watson APIs, WatsonX.ai, WatsonX.governance
Platform
1 ) WatsonX
College
1. IIT Bombay



Aakriti Aggarwal' Comment

This project is well-aligned with current advancements in AI governance and agent-based systems, offering students an opportunity to explore how artificial intelligence can evaluate itself through structured coaching mechanisms. The multi-layer review pipeline provides valuable insights into improving generative quality, reducing errors, and strengthening reliability across different tasks. As part of the work, the team will also compare and benchmark the performance of our agent against other existing AI evaluation frameworks to understand strengths, weaknesses, and possible optimizations. Through this initiative, students will gain practical exposure to designing intelligent review loops and confidence-based workflows. Overall, the project reflects a forward-looking vision of how AI evaluation systems will evolve in the near future.