Webinar Recap: Scaling Quality with AI Evaluation Models

Design Executive Council × Dscout | Webinar Recording

Webinar Recap: Scaling Quality with AI Evaluation Models

Webinar Recap: Scaling Quality with AI Evaluation Models

Design Executive Council × Dscout | Webinar Recording

August 5, 2026
Webinar Recap: Scaling Quality with AI Evaluation Models
Summary

When AI can generate more than teams can review, how do you define what "good enough to ship" actually means?

AI is forcing organizations to answer a question many have quietly deferred: what does quality actually look like, and who is accountable when the work moves through a chain a researcher only touched lightly? In this session, design and research leaders from Dscout, The Cigna Group, Datadog, and TD Bank shared how they are encoding standards into AI systems, drawing the line on human involvement, and thinking about what agentic research unlocks beyond faster output.

Hosted in partnership with Dscout, the conversation featured Gordon Ching (Design Executive Council), Michael Winnick (Dscout), Christina Vallery (The Cigna Group), Catherine Walker (The Cigna Group), Anshuman Kumar (Datadog), and Christian Rohrer (TD Bank Group).

This session is Track 4 of Modernizing Research with AI: Customer Centricity at Scale, a four-part series with Dscout and the Design Executive Council. It closes the arc by turning to the layer that holds the rest together: how teams define quality in non-deterministic systems, where humans stay accountable, and what evaluation and governance make possible in an agentic future.

What we covered in the webinar

  1. Aligning and encoding quality in AI systems: How teams are arriving at shared definitions of quality, translating those definitions into machine-readable standards, and using AI as a forcing function to close the gap between tribal knowledge and institutional insight. 
  2. Review frameworks and accountability at scale: How leaders are deciding which decisions deserve human judgment, why "human in the loop" is being reframed toward stronger ownership language, and where accountability sits when AI does more of the work. 
  3. Evaluation and governance in the agentic era: What agentic research unlocks beyond speed, including the potential to move past persona-level compression toward understanding human variability, and why governance is the condition that makes that shift responsible. 

Key takeaways

Scaling quality with AI is less about new tooling and more about codifying judgment, repositioning humans as owners rather than reviewers, and using the evaluation mindset to elevate research's role in the business.

  • From tribal knowledge to machine-readable standards: AI is forcing teams to make "what good looks like" explicit and encodable. Panelists described moving voice, tone, design system, content, and accessibility standards into rules, rubrics, and prompt libraries that agents can reference, turning previously insular expertise into shared infrastructure the whole organization can build on.  
  • Human at the helm, not in the loop: Leaders pushed back on framings that cast humans as bystanders. Emerging frameworks separate human-defined governance, human judgment, and human authority, clarifying which decisions AI can operate within, which require a person to make the call, and which stay entirely with humans who answer for the outcome. 
  • Evals as a UX opportunity: Rubric design, LLM-as-judge evaluations, and iterative prompt refinement map closely to existing UX skills. Teams that lean into evaluation are earning executive interest and a stronger seat at the table, because leadership wants to understand how these systems are performing for real users, not just whether they ship. 
  • Toward post-persona research: Several panelists pointed to a horizon where AI helps research see human variability at a scale traditional methods compress, moving past segments and personas toward responding to individuals in context. That shift depends on governance being built alongside capability, not bolted on afterward. 

Path forward

The panelists kept returning to a shift in posture. Quality is no longer something inspected at the end of a process; it is something encoded at the start. That reframe puts research and design teams in a different position, closer to the standards, rubrics, and guardrails that decide what an AI system is allowed to produce. The work of making judgment explicit, articulating it precisely enough that a machine can act on it, is where much of the value now sits.

The longer horizon is more ambitious. If evaluation and governance hold, AI-assisted research can begin to surface the depth of human variability that traditional methods have had to compress. That is a meaningful unlock for accessibility, ethics, and personalization, but only for teams that treat governance as the enabling condition rather than a downstream check. The teams making progress are the ones building both capabilities in parallel.

Continue the series

This is Track 4 of a 4-part series. Each track examines a different dimension of the shift: 

Explore the premier membership network for design executives

DXC provides design executives a trusted network to create impact alongside their peers, access agenda-shaping research, and amplify the strategic role and value of design leadership at the C-Suite and Board-level of the world’s largest companies.

Express interest

Download the report

Thank you! Download the report here:
Download PDF
Oops! Something went wrong while submitting the form.