STAI CDT PhD student Israel Mason-Williams was selected to participate in one of the world’s first academic programmes dedicated to AI evaluation. Below, he reflects on the experience and discusses why robust AI evaluation is increasingly important in today’s rapidly evolving landscape.
Why evaluating AI matters more than ever
Over the last few years, AI systems have been increasingly deployed across society. But as these technologies become increasingly capable, an important question arises: how do we know when AI is performing well, and when it isn’t?
Understanding these systems’ capabilities and limitations through rigorous evaluation is crucial to ensure that AI is adopted in the correct context and that when it is adopted, there is a strong understanding of the tasks it can and cannot perform.
As AI is being adopted so quickly, researchers face the challenge of developing consistent ways to assess these systems. This means bringing together evaluation methods from different scientific fields and agreeing on shared standards, so that AI systems can be tested accurately and their performance can be understood and compared reliably.
Joining a global community of AI Evaluators
Earlier this year, I was selected, as 1 of 40 fellows, from more than 600 applicants, to join the world’s first academic programme dedicated to AI Evaluation. The programme was sponsored by the Valencian Graduate School and Research Network of Artificial Intelligence (ValgrAI) and fully funded by Coefficient Giving.
The programme was created to respond to an existing shortage of experts in AI Evaluation in AI Safety Institutes and research labs, and to position AI evaluation as an academic discipline in its own right.
Over five months, the fellowship brought together a global cohort of researchers and practitioners interested in understanding how we can better test, monitor and guide the development of AI systems.
The programme included sessions from experts across industry, academia and government, including the UK AI Safety Institute (AISI), EU AI Office, FAR AI, RAND, Epoch AI, Apollo Research, Redwood Research, Microsoft Research, and Google DeepMind. Their insights highlighted the importance of rigorous evaluation in ensuring AI systems are both useful and trustworthy.
The value of an interdisciplinary approach to AI evaluation
One unique aspect of the programme is that it brings together insights from disciplines that are often studied separately, including psychometrics, statistics, artificial intelligence and human-computer interaction, creating a more holistic approach to AI evaluation.
Among the highlights for me were the sessions exploring how we can better understand, test and evaluate AI systems. These covered topics ranging from measuring performance and reliability to assessing safety, alignment and societal impact. Combined with reflective assignments, the programme offered a holistic framework for evaluating AI systems and understanding where they perform well, where they struggle and what risks may emerge as they become more capable.
Capstone week in Valencia
The programme concluded with a capstone week in Valencia, where fellows presented their own research on AI evaluation.
I presented research from the AgentGuard project, a collaboration with Microsoft Research. The project investigated how differences in the behavioural tendencies of large AI models can lead to varying levels of performance across reasoning tasks. In other words,
The research demonstrated that evaluating advanced AI systems requires more than measuring outcomes alone. Understanding how a model behaves and responds to different prompts can provide important insights into its strengths, limitations and suitability for particular tasks. As AI systems become more sophisticated, evaluation approaches must move beyond one-size-fits-all benchmarks and take these behavioural differences into account.
The research contributes to the wider AgentGuard initiative, which aims to develop tools that help organisations deploy AI agents more safely and reliably in complex real-world settings.
Alongside presenting our research, we took part in workshops and discussions focused on AI’s wider impact on society, how emerging governance frameworks are taking shape and how advanced systems can be stress-tested through red-teaming exercises.
Being part of a global cohort was one of the most rewarding aspects of the programme, exposing me to a wide range of perspectives on AI evaluation and highlighting that different risks matter to different stakeholders.

Key learnings from the programme
As a PhD student researching the fundamental properties of AI systems, the programme provided valuable new perspectives on my work and introduced different ways of thinking about how AI capabilities and behaviours can be assessed. The skills and knowledge gained have strengthened my ability to investigate key research questions and critically evaluate the theories used to explain how AI systems work.
The programme also encouraged me to think more deeply about AI evaluation itself. It provided a new lens through which to assess research findings, helping me better understand whether evaluation methods are truly measuring what they intend to measure and where existing approaches could be improved. This perspective is particularly valuable when reviewing and assessing research, where rigorous evaluation plays a crucial role in ensuring reliable and meaningful conclusions.
A key learning for me is that AI evaluation is truly relative to the questions being asked, and while frameworks exist to answer specific questions, each evaluation requires task specific context to ensure that evaluations are accurate and insights are appropriately communicated. Overall, to ensure appropriate AI evaluation, we must move away from a one-size-fits-all evaluation – this is a change that should be embraced by all scientists so that we can adapt to the new risk landscape posed by advanced AI.
