Thought Leadership Aug 12, 2026

What is the relationship between AI safety and AI evaluations?

A look at why both are important to fund

Mala Kumar

With contributions from Theo Skeadas, Kristi Arbogast, and Julie Hollek

If we posit that AI will play increasingly important roles in everything we do, then it stands to reason that AI should not make the world worse and at least some use of AI should serve a greater social good. One of the most important principles when using technology to serve a greater social good is first to “Do No Harm.” Even with the best intentions to improve the world, there are very real and very complicated power dynamics, cultural nuances, and learning curves that can make the presence of a product, program, project, or initiative a net negative for a given context or population. 

Among tech for social good and responsible tech practitioners, part of a job is to question what could go wrong if a new digital product is introduced in a given context, sometimes before there is a popular name for the phenomenon. What if a social media group intended to help connect and empower teenagers in a village is infiltrated by an imposter to trick a vulnerable person into submission? Or what if an SMS campaign for maternal health check-ins is hijacked to con an expectant mother out of money? While we now know these as examples of catfishing and phishing, the possible negative repercussions of these “for good” initiatives were much less known 15 or 20 years ago. 

We find ourselves in a similar position now with AI, where the consequences of its use are often unknown. As we discuss in our LinkedIn Learning course, generative AI is general purpose, robust, and pervasive, which means it can negatively affect anyone if used irresponsibly. It’s along this realm of “Do No Harm” that the idea of AI evaluations and AI safety have emerged, though not just for vulnerable populations, but for all individuals and populations. In this blog post, we’ll discuss some key concepts and how both are positioned to ensure that AI does no harm and can instead serve a greater social good.

Who does what?

An original definition of AI safety is a catch-all term for all disciplines that ensure AI does not cause harm or present hazards to individuals or societies. By this definition, AI evaluations are a part of AI safety. While the boundaries around what constitutes AI safety have changed several times in the past few years, a subset of AI safety initiatives have been dubbed “frontier AI safety research,” which refers to research and mitigation efforts that are directly incorporated into a generative AI model. AI evaluations, on the other hand, refer to investigation and mitigation efforts of consumer-facing generative AI models, with or without other AI system building blocks. Therefore, frontier AI safety research focuses on the design of the generative AI model, whereas AI evaluations focus on the use of the generative AI model or AI system. 

Frontier AI safety research is usually conducted by frontier model companies or by well-resourced AI safety research labs due to the significant compute cost and research capacity required. Open weights of frontier models alone can be several terabytes in size and require one or more GPUs to access and modify; the frontier model training data is orders of magnitude larger. Meaningful analysis of such quantities of data requires specialized data scientists and AI / ML engineers, in addition to the compute and hardware requirements. 

The technical requirements, a long lead time to see results, and the high cost of frontier AI safety research present significant barriers to entry for most organizations. Therefore, AI evaluations have gained popularity in international civil society, governments, and among small and medium businesses that are trying to investigate how an AI model or system behaves or performs on a case-by-case basis. 

For example, a frontier model company may decide that a critical piece of AI safety research is how to prevent all of its models from enabling financial fraud in major American and European banks, as that is seen as a universal and constant need. Therefore, they may invest money into research to identify known vulnerabilities and implement guardrails directly in the generative AI model to prevent this type of financial fraud before releasing the model. However, a less common demand may be preventing financial fraud in sub-Saharan African informal markets. An INGO working in Kenya may choose to conduct AI evaluations to ensure its AI system that uses the same frontier generative AI model is not vulnerable to informal African market fraud risk. Both the AI safety work of the frontier model company and the AI evaluation work of the INGO are consistent to the principle of “Do No Harm,” albeit at very different stages of the AI software development lifecycle. Both are important to responsibly use the generative AI model in financial AI for social good solutions.

Risk – Time, Severity and Likelihood

In actuarial sciences used in insurance industries, risk is often thought of in terms of time and how likely or how often an event of a certain severity may happen. Weather insurance, for example, refers to natural disasters that are projected to happen 1 in 5, 1 in 10, or 1 in 100 years. Catastrophic natural disasters may be projected to happen 1 in 500 years. Car insurance companies look at individual factors such as age, gender and driving history to determine how likely someone may be to get into an accident before calculating a driver’s monthly premium.

Likewise, AI safety and AI evaluations can be thought of along three axes: time, event severity, and likelihood. AI evaluations, including the evaluations we do at Humane Intelligence, mostly focus on current and near future risks that are highly likely to happen – things like (algorithmic) bias, discrimination, factual inaccuracies, or hallucinations in one-off hiring, healthcare, financial, or economic decisions that are placed on or informed by AI systems. Increasingly, AI evaluations are also looking at decisions or tools as they evolve in the longer term, such as an AI model or system disproportionately affecting a population’s economic livelihood, or agentic AI systems replacing entire government service delivery programs.

AI Evaluations and AI Safety – Event Severity, Time, and Likelihood*

*Note that this is for demonstrative purposes and not based on empirical research. The axes assume a consistent context or operating environment.

Frontier model research and other types of AI safety initiatives have fractured and reunited over the years along medium, high and catastrophic risk. A subset of AI safety commonly dubbed “X-Risk” has focused entirely on catastrophic risks, such as preventing frontier AI models and systems from enabling mass biological or nuclear weapons development and deploying all the weapons at the same time. As shown above, these “catastrophic risk” scenarios are further in the future and are relatively unlikely to happen, but if they did happen, could result in the destruction of the entire human race. Again, both AI evaluations and AI safety follow the principle of “Do No Harm” in terms of time, event severity, and likelihood of occurring, though often focused on different timelines and tranches of risk.

Which one is more important?

Both AI evaluations and other AI safety initiatives are important to enable a future in which AI makes human lives better. As the capabilities of AI models and systems increase, so too do the vectors of harm and risk, and it’s important not to summarily dismiss those possibilities. Ignoring what is relatively likely to happen in favor of only catastrophic risk scenarios, however, is equally problematic. If global income inequality continues to rise, and large portions of a population are economically disenfranchised, plunged into poverty, or stripped of basic health and human services due to current or near term risks posed by AI, the world may be greatly destabilized far before any “catastrophic risk” materializes. At Humane Intelligence, we therefore advocate funding both AI evaluations and other AI safety initiatives, albeit in ways and with distributions that actually enable a greater social good. To do this, we must “Do No Harm” by getting ahead of a wide spectrum of possible negative consequences, even if we don’t yet have names for everything that is happening.

Interested in discussing these topics further? Want to hire us? Please reach out to info@humane-intelligence.org if you are interested in an OpEd, podcast or news segment on your publication. Read more about our programs and services here.

Sign up for our newsletter
Sign up for our newsletter