Revolutionizing Product Development: The Role of AI Evals at Expedia
Expedia's AI chief, Xavi Amatriain, shares insights on how AI evaluations are changing the landscape of product development, replacing traditional PRDs with dynamic evals that prioritize security and user experience.

In the rapidly evolving landscape of technology, product development processes are undergoing significant transformations. One of the most groundbreaking changes comes from the travel industry giant, Expedia, where AI evaluations are being hailed as the new Product Requirements Document (PRD). This shift, articulated by Xavi Amatriain, Expedia's Chief AI and Data Officer, emphasizes the importance of integrating AI evaluations into the core of product design and development.
During his recent address at the VB Transform 2026 conference, Amatriain outlined how evals—comprehensive assessments of AI systems—are superseding traditional PRDs. By embedding security requirements and operational principles into these evaluations from the outset, companies can streamline their development processes and enhance product reliability. This article explores this paradigm shift, its implications for product development in the tech industry, and the broader context of AI governance.
Understanding the Shift from PRDs to Evals
Traditionally, Product Requirements Documents (PRDs) have served as a blueprint for product development, detailing the necessary features, functionality, and user interactions. However, as AI technologies become increasingly complex and integral to product offerings, the limitations of PRDs have become apparent.
Amatriain suggests that evals encapsulate not just what a product should do but also how it should function concerning security and user feedback. This dynamic approach allows for a more flexible and adaptive product development cycle, where the requirements evolve based on real-time performance data and user interactions.
The Role of Red Teaming in Evals
One of the critical components of AI evals is red teaming, a process where independent groups assess the security and functionality of AI systems by actively attempting to compromise or misuse them. This proactive approach ensures that potential vulnerabilities are identified and addressed before products are deployed in a live environment.
Amatriain asserts that integrating red teaming into the eval process helps organizations uncover security gaps and biases that could affect user experience and trust. By prioritizing these assessments, companies can build more robust systems that stand up to real-world challenges.

The Necessity of Governance in AI Systems
As organizations increasingly rely on AI for decision-making, the need for effective governance structures becomes paramount. Amatriain argues that governance should correlate with the level of risk associated with specific AI applications. For low-risk applications, minimal governance may suffice, while high-risk scenarios necessitate stricter oversight.
Expedia's governance model operates on three levels:
- Principles: High-level guidelines that articulate decision-making processes.
- Processes and Tools: Specific methodologies and tools designed to enforce these principles.
- Automation: Automated systems that facilitate compliance with governance frameworks.
This layered approach allows Expedia to adapt its governance strategies as the risk landscape evolves. Amatriain emphasizes the importance of embedding these principles into the company culture, ensuring that all employees understand and adhere to them.

Specialized Agents vs. Monolithic Intelligence
Amatriain's vision for AI architecture diverges from the traditional monolithic intelligence models. He advocates for the use of specialized agents—discrete AI systems optimized for specific tasks. This compositional approach allows for greater flexibility and security, as each agent can be individually evaluated and secured before integration into a larger system.
For instance, if a travel booking system consists of multiple specialized agents—such as one for hotel bookings, another for flight comparisons, and yet another for customer feedback—each can be developed, tested, and secured independently. This modular design not only enhances the overall security of the system but also streamlines updates and improvements.
Design Considerations for AI Agents
When developing these specialized agents, Amatriain stresses the importance of thoughtful design. Each agent should have clear guidelines regarding its tone, user interaction, and context management. This systemic design approach ensures that the agents work harmoniously and effectively within the broader system.

Empowering Users in Decision-Making
In the travel industry, where information changes rapidly, it is critical to empower users with agency in their decision-making processes. Amatriain highlights Expedia's approach of providing recommendations without executing transactions on behalf of the user. By maintaining this boundary, users are encouraged to engage actively with the system, fostering a sense of control and responsibility.
For example, when a user queries about hotel prices, the AI agent can provide immediate, relevant information, cross-referencing user-generated reviews and real-time supplier data. However, the final decision to book remains with the user, ensuring they are fully aware of their choices and the implications of those choices.
Addressing Security in AI Development
As AI systems become more prevalent, the threats they face are also evolving. Amatriain warns that security considerations must be integrated into the design process rather than treated as an afterthought. He argues that organizations must anticipate and prepare for potential threats, including attacks from other AI systems.
Amatriain's insights are underscored by recent research showing that a significant percentage of enterprises have experienced security incidents involving AI agents. In fact, over half of surveyed companies reported at least one incident, highlighting the urgent need for robust security measures in AI development.
The key to effective security lies in establishing a feedback loop that allows organizations to monitor AI performance continuously and respond to incidents swiftly. This proactive approach ensures that vulnerabilities are addressed before they can be exploited.

Key Takeaways
- AI evaluations (evals) are replacing traditional PRDs as the foundation for product development, emphasizing security and adaptability.
- Governance structures should align with risk levels, allowing for flexible oversight based on the potential impact of AI applications.
- Specialized agents offer a more secure and efficient alternative to monolithic AI systems, enabling focused development and testing.
- Empowering users in decision-making processes enhances engagement and accountability in AI-driven environments.
- Proactive security measures and continuous monitoring are essential to safeguard AI systems against evolving threats.
Frequently Asked Questions
What are AI evaluations and how do they differ from traditional PRDs?
AI evaluations (evals) are comprehensive assessments that focus on both functionality and security of AI systems, while traditional Product Requirements Documents (PRDs) primarily outline feature sets and user interactions. Evals are designed to adapt based on real-time performance and user feedback, making them more dynamic than static PRDs.
How does Expedia manage AI governance?
Expedia employs a three-tiered governance model that includes high-level principles, specific processes and tools, and automated systems. This approach allows the company to align governance with the risk levels of different AI applications, ensuring that oversight is appropriate to the potential impact of the technology.
Why are specialized agents preferred over monolithic AI systems?
Specialized agents are preferred because they allow for focused development, testing, and security assessments. By creating discrete AI systems optimized for specific tasks, organizations can improve overall system security and adaptability, reducing the risks associated with larger, more complex monolithic models.
How can organizations enhance the security of their AI systems?
Organizations can enhance AI security by integrating security considerations into the design process, establishing continuous feedback loops for performance monitoring, and preparing for potential threats from both human actors and other AI systems. Proactive measures are essential to safeguard against vulnerabilities and ensure the integrity of AI-driven applications.
Comments
Google's Gemini 3.6 Flash: Revolutionizing AI Efficiency and Cost-Effectiveness
Google's latest Gemini models, including the 3.6 Flash, offer significant reductions in AI token costs, enhancing efficiency for long-horizon engineering tasks. With pricing strategies designed to empower enterprises, these new models mark a pivotal shift in the AI landscape.

Related articles
Popular in AI Tools
- SpaceX's Grok 4.5: Disruption in AI Coding at Unmatched Prices
- Gaming Data: The Future of Training AI for General Intelligence
- OpenAI's GPT-5.6: A New Era for Microsoft Copilot and Beyond
- The AI Deployment Dilemma: Balancing Autonomy and Governance
- Kimi 3: A New Frontier in Open Source AI and Its Global Implications