When AI can generate more than teams can review, how do you define what "good enough to ship" actually means?
AI is forcing organizations to answer a question many have quietly deferred: what does quality actually look like, and who is accountable when the work moves through a chain a researcher only touched lightly? In this session, design and research leaders from Dscout, The Cigna Group, Datadog, and TD Bank shared how they are encoding standards into AI systems, drawing the line on human involvement, and thinking about what agentic research unlocks beyond faster output.
Hosted in partnership with Dscout, the conversation featured Gordon Ching (Design Executive Council), Michael Winnick (Dscout), Christina Vallery (The Cigna Group), Catherine Walker (The Cigna Group), Anshuman Kumar (Datadog), and Christian Rohrer (TD Bank Group).
This session is Track 4 of Modernizing Research with AI: Customer Centricity at Scale, a four-part series with Dscout and the Design Executive Council. It closes the arc by turning to the layer that holds the rest together: how teams define quality in non-deterministic systems, where humans stay accountable, and what evaluation and governance make possible in an agentic future.
What we covered in the webinar
- Aligning and encoding quality in AI systems: How teams are arriving at shared definitions of quality, translating those definitions into machine-readable standards, and using AI as a forcing function to close the gap between tribal knowledge and institutional insight.
- Review frameworks and accountability at scale: How leaders are deciding which decisions deserve human judgment, why "human in the loop" is being reframed toward stronger ownership language, and where accountability sits when AI does more of the work.
- Evaluation and governance in the agentic era: What agentic research unlocks beyond speed, including the potential to move past persona-level compression toward understanding human variability, and why governance is the condition that makes that shift responsible.
Key takeaways
Scaling quality with AI is less about new tooling and more about codifying judgment, repositioning humans as owners rather than reviewers, and using the evaluation mindset to elevate research's role in the business.
- From tribal knowledge to machine-readable standards: AI is forcing teams to make "what good looks like" explicit and encodable. Panelists described moving voice, tone, design system, content, and accessibility standards into rules, rubrics, and prompt libraries that agents can reference, turning previously insular expertise into shared infrastructure the whole organization can build on.
- Human at the helm, not in the loop: Leaders pushed back on framings that cast humans as bystanders. Emerging frameworks separate human-defined governance, human judgment, and human authority, clarifying which decisions AI can operate within, which require a person to make the call, and which stay entirely with humans who answer for the outcome.
- Evals as a UX opportunity: Rubric design, LLM-as-judge evaluations, and iterative prompt refinement map closely to existing UX skills. Teams that lean into evaluation are earning executive interest and a stronger seat at the table, because leadership wants to understand how these systems are performing for real users, not just whether they ship.
- Toward post-persona research: Several panelists pointed to a horizon where AI helps research see human variability at a scale traditional methods compress, moving past segments and personas toward responding to individuals in context. That shift depends on governance being built alongside capability, not bolted on afterward.
Path forward
The panelists kept returning to a shift in posture. Quality is no longer something inspected at the end of a process; it is something encoded at the start. That reframe puts research and design teams in a different position, closer to the standards, rubrics, and guardrails that decide what an AI system is allowed to produce. The work of making judgment explicit, articulating it precisely enough that a machine can act on it, is where much of the value now sits.
The longer horizon is more ambitious. If evaluation and governance hold, AI-assisted research can begin to surface the depth of human variability that traditional methods have had to compress. That is a meaningful unlock for accessibility, ethics, and personalization, but only for teams that treat governance as the enabling condition rather than a downstream check. The teams making progress are the ones building both capabilities in parallel.
Continue the series
This is Track 4 of a 4-part series. Each track examines a different dimension of the shift:
- Track 1: Building the research team of tomorrow
How leaders are rebuilding teams, redefining performance, and steering change
- Track 2: Rethinking capacity with AI in the workflow
How leaders are reshaping where teams spend their time
- Track 3: Retaining customer-centricity when everyone builds
Protecting user truth when research is democratized








