Evals as the New PRD: Expedia’s AI Chief on Guardrails, Agents, and the Governance Wave

AI governance is shifting from chasing perfect models to building robust evaluation frameworks. At the VB Transform 2026 event, Expedia Group’s first chief AI and data officer reframed the traditional product requirement document as a living spine of evaluation—evals that encode what a product should do, including red-teaming and security checks, long before a single line of code is written. The idea is simple but ambitious: the thinking behind the product lives in the evals, and AI-assisted code will push that approach further, turning the entire design process into an eval-driven cycle.

Expedia’s approach unfolds across three guardrails: principles, processes, and automation. Principles set the high-level decision framework, processes formalize how decisions are checked across teams, and automation implements them. The toll gates align evaluation rounds, red-teaming, and security reviews to each agent’s risk level, with checks becoming mandatory as risk rises. This is not about locking down innovation but about shaping a feedback loop that remains flexible as the system scales and threats evolve.

Architecturally, Expedia favors a composition of specialized agents over a single monolithic model. Tools combine into skills, skills assemble into sub‑agents, and sub‑agents orchestrate the whole system. A core design principle is to fix the user’s agency: the agent can recommend and discuss, but the user must make the final click to act. Latency-driven decisions mix retrieval-augmented generation with direct API tool calls—real-time answers for pricing or availability, balanced by context from Expedia’s own review corpus rather than surface-level supplier claims.

In parallel, teams outside travel are wrestling with ROI and governance in real time. Atlassian researchers argue that ROI comes from teams, not individuals, and that context graphs, workflow redesign, and explicit AI working agreements drive performance. The broader picture across the industry includes OpenAI urging enterprises to use its scorecard to measure AI worth, while China’s Qwen and other low-cost models broaden supplier choices for enterprises. The conversation touches policy and public trust too—covering AI datacentres, elections, and the societal impact of automation.

Beyond business value, the AI era is exposing political, labor, and infrastructure questions. Reports and commentary point to AI’s ripple effects—from AI-dense media layoffs in Australia to concerns about datacentre cooling and water use in the UK, from election guidance risks to the reliability of AI-assisted civic processes. As organizations ship more agents and deploy more coordinated ecosystems, guardrails will need to move left in the design process and stay lean as feedback loops accelerate. In short, governance calibrated to risk, continuous monitoring, and human agency remain the keystones of responsible AI at scale.

Sources

  1. Evals are the new PRD, Expedia’s AI chief tells VB Transform 2026
  2. OpenAI Urges Enterprises to Use Its Scorecard to Measure Worth of AI
  3. Alibaba Qwen 3.8 Max Shows China Closing in on U.S. Models
  4. Headaches for Silicon Valley as China chips away at the US’s lead in the AI race
  5. Gen Z is living in an intimacy economy, where connection is commodified
  6. Tell us your experiences of living near AI datacentres in the UK
  7. Atlassian: Research shows organizations should approach AI at the team level, not the individual level, to achieve true ROI
  8. Man of his word: Pope Leo speeches declared human-authored by Australian AI detection tool
  9. Nine to axe 30 jobs at the Age and SMH due to ‘extreme’ AI disruption
  10. Election voting advice from AI chatbots ‘inaccurate and unreliable’
  11. Not enough water for UK’s datacentre plans, trade body says
You may also like

Related posts

Write a comment
Your email address will not be published. Required fields are marked *

Scroll
wpChatIcon
wpChatIcon