What's Happening?
The Center for Employment Opportunities (CEO), the largest reentry employment organization in the United States, has significantly improved its AI-powered mock interview coach, ACE. This AI coach assists formerly incarcerated individuals in preparing
for job interviews across more than 30 U.S. cities. Initially, while participants liked ACE, there was no systematic way to evaluate its effectiveness. Researchers from The Bike Shop, an AI research lab, partnered with CEO to develop an evaluation system using a second AI, an 'LLM-as-judge,' to score coaching calls against a rubric of good coaching. This system identified key weaknesses in the original ACE, such as a lack of dialogue, low participant practice rates, and over-praising weak answers. Following these findings, a new version of the AI coach was designed and tested. An A/B test revealed substantial improvements: 96% of participants now attempt an improved answer after feedback, compared to 46% previously; the AI coach successfully limited feedback load 98% of the time, up from 44%; and it helped participants build answers from their own experience 68% of the time, a significant increase from 10%. The new coach also fostered dialogue 60% of the time, a feature previously absent.
Why It's Important?
This development is crucial for U.S. social impact initiatives, particularly in criminal justice reform and workforce development. The improved AI coach addresses a critical need for accessible, effective job interview preparation for formerly incarcerated individuals, a population that faces significant barriers to employment. By augmenting human coaches, the AI tool can provide consistent, scalable practice opportunities that human coaches cannot offer around the clock. This can lead to better interview performance, increased employment rates, and ultimately, reduced recidivism. The systematic evaluation approach, using an 'LLM-as-judge,' sets a precedent for ensuring AI tools in sensitive social sectors are not only well-received but genuinely effective and safe. It highlights the importance of rigorous testing and continuous improvement for AI applications, especially when serving vulnerable populations where errors could have severe consequences, such as inadvertently encouraging the sharing of sensitive personal information during interviews.
What's Next?
For CEO, the evaluation suite is not the final step but rather infrastructure for continuous improvement of the AI coach. The organization plans to use this system to further refine the coach's capabilities and, more importantly, to answer whether better AI coaching translates into participants being more prepared, confident, and effective in telling their stories during actual job interviews. To maintain the system's effectiveness, coaches will need to regularly label a sample of production calls. If the AI coach is applied to different populations, transcripts will need to be re-labeled to ensure the judges remain aligned with human evaluations in new contexts. The process of adding new quality dimensions, such as an 'annoyingness' judge, will also require the same rigorous definition and iteration process. This ongoing work aims to ensure the AI tool remains a reliable and impactful resource, protecting users and maximizing its utility in the reentry process.
Beyond the Headlines
The project highlights deeper implications regarding the ethical deployment and continuous oversight of AI in social services. The initial challenge of human coaches disagreeing on evaluation criteria underscores the complexity of defining 'good' coaching and the need for precise, measurable rubrics when developing AI. The iterative process of refining these definitions and training AI judges reveals the significant human effort required to make AI truly effective and safe, particularly in contexts involving vulnerable populations. The emphasis on protecting users from potential harm, such as oversharing sensitive information, points to the critical ethical responsibility in AI development for social impact. This initiative also demonstrates a shift from simply deploying AI to building robust evaluation frameworks that ensure AI tools augment human capabilities rather than replace them, fostering a more nuanced understanding of human-AI collaboration in addressing complex societal challenges.











