Our Programs and Services

Our Evaluation Overview

Of four established types of AI evaluations, Humane Intelligence provides three as programs and services: AI red teaming, bias bounties, and contextual (bespoke) evaluations. AI red teaming is a great first point of entry to AI evaluations for people and organizations with subject matter expertise in a given domain or subject matter area. Data outputs from red teaming events can then be used to create bias bounty challenges. AI contextual evals are the best fit for organizations wishing to do a comprehensive analysis of its AI model or system, and can include our new knowledge graph / ontology based methodology. Note that AI evaluations are often iterative and non-linear. The chart to the right / below is intentionally reductive to demonstrate relationships and refers to how Humane Intelligence approaches these evaluation types. Other organizations may define these differently.

Our Types of Evaluations

AI Red Teaming

AI red teaming is a semi-structured testing approach to assess and improve the safety and effectiveness of AI models and systems by identifying vulnerabilities, limitations, and potential areas for improvement. Humane Intelligence offers red teaming events as a paid service using our own software, which will are releasing under an open source software license in 2026.

Participants seated and listening to Rumman present at the IMDA Singapore event

AI Contextual Evaluations

AI Contextual Evaluations are rigorous, mixed-method, bespoke evaluations designed to give a comprehensive analysis of an AI model or system’s performance for a specified problem space. In 2026, we are rapidly developing our knowledge graph / ontological AI problem space mapping and contextual AI evaluations. Humane Intelligence designs and runs contextual evals as a paid service.

Bias Bounty

Bias bounties are collaboratively designed sets of challenges that bring together researchers, impacted communities, and domain experts to rigorously examine and improve AI / ML systems, models, and datasets. Our final self-hosted bias bounty closed in November 2025. In 2026, we are working to move our bias bounty program onto Zindi, a global data science challenge platform.

A simplified view

New to AI Evaluations?

If you recently received philanthropic funding or are otherwise tasked with implementing an AI model or system for the first time, read our AI Evaluation Overview for a simplified explanation of what we provide. You can download this as a PDF or send this link to others.

In Terms of Cookies…

Bias Bounties vs Red Teaming

We are often asked about examples of red teaming and bias bounties. Imagine if we did a red teaming exercise for cookies – yes, cookies. A red teaming cookie event could involve participants of any background or skill level identifying the basics of the situation: does the cookie taste good or bad? Is the cookie too hard? Was it baked correctly or does it fall apart when it’s picked up? A cookie bias bounty goes deeper and would involve people with more skill or knowledge to identify what went wrong. In this case, someone who works in the bakery might realize that salt was used instead of sugar because of a mislabelled sugar container. Likewise, red teaming is the best starting point to identify potential issues with AI / ML models or systems. Bias bounties then go further to explain the problem and can result in mitigations or solutions.
What red teaming or a bias bounty could tell us about our cookies
Let’s work together

Want to hire us?

We have worked with education companies, international civil society, industry, and governments to design and run red teaming events, bias bounties, and bespoke contextual evaluations.

Sign up for our newsletter