Strategy Mar 19, 2026

Learn about our new knowledge graph-based methodology for AI contextual evaluations

We’re developing a new methodology built on statistical rigor and domain expertise

Mala Kumar

Co-authors: Mala Kumar, Annie Brown, Julie Hollek, Theo Skeadas

Updated: June 1, 2026


One of the most common questions our clients ask is, “How do I know if an AI model or system performs well in a particular use case?” When we unpack this statement, we immediately observe several complexities. For comparison, testing if an AI model or system performs well in a given use case often requires testing multiple AI models or systems. Asking if an AI model or system “performs well” can refer to a range of harms, including bias, factuality, or misdirection. One of the most complicated parts of the question is “particular use case,” as this implies knowing what exactly is in or out of scope of an evaluation.  

"/

In fact, claims about an AI model or system evaluation is often a major point of contention. A claim may be that a certain LLM or GenAI chatbot performs the best in terms of misdirection in education. In reality, the evaluation testing coverage may have exclusively focused on higher education in the United States, and not education more broadly. Since drilling down into an AI evaluation’s coverage is often difficult, it can be hard to verify or refute the claim.

Our new knowledge graph / ontology approach

To address these complexities, Humane Intelligence has recently launched a new contextual AI evaluation methodology using knowledge graphs and ontologies. As we detailed in our post about the Sustainable Development Goals (SDGs), ontologies can offer much richer information than taxonomies using the same data. Ontologies can account for the strength, proximity, and clustering of relationships. Likewise, our shift into ontological-based AI evaluations gives us richer information.

"/

Our new methodology is enabling us to capture and embrace the complexities that arise in any given topic using statistically rigorous, scientific methods. Key to our work is that elusive problem mentioned above – problem space coverage. With knowledge graphs and ontologies, we can better understand and document what is in and out of scope in an evaluation, identify evaluation coverage gaps, and make more accurate claims. An ontological approach also enables traceability in evaluation prompt creation, as we can find each metric in terms of its elements, giving us much more specific insight into what data actually shifts model behavior. This helps optimize compute, cost and resources, and helps create reproducible evaluations, in terms of testing coverage. Ultimately, our new methodology will help our clients make better product go/no-go decisions for complex situations and will allow them to offer more transparency to their users, clients, and stakeholders.

How does this work?

At the core of Humane Intelligence’s work is human subject matter expertise and lived experience. That hasn’t changed with our new methodology. 

We start by co-designing and building a formal “map” of the main problem space – an ontology – with our clients and their stakeholders. The ontology is the structure and governing rules of the problem space. We then populate the ontology with details specific to the use case, which produces a knowledge graph that captures key actors, events, geographies, and other relevant attributes. Depending on the problem space, we can add additional layers (or facets) to our ontology and knowledge graph. 

From this knowledge graph, we create a collection of realistic scenarios that represent the problem space. Depending on the breadth and depth of coverage, we then use a mix of human experts and LLMs to construct prompts from those scenarios that can be used for AI red teaming workshops or other AI evaluation types. Each scenario and prompt is tagged with metadata that orients where it sits in the knowledge graph and encodes the key relationships. 

At this stage, we do quality assurance to discard nonsensical or impossible scenarios and prompts. We can stratify our final set of prompts based on the compute and cost constraints of the client. For example, if a client has enough budget to test 5,000 prompts (or let’s say 70k tokens), we could stratify the prompts so that they cover and represent key clusters, known failure points, or un(der)explored relationships. 

Finally, we use the prompts in our AI red teaming app, annotate the responses, and create the final analysis and recommendations. Because prompts and scenarios are traceable in our knowledge graph, we can quickly identify possible reasons why a vulnerability or harm might have surfaced. For example, our stratification might show that all prompts about education in a particular country, regardless of phrasing, tone, or exploit technique, returned unacceptable responses. This makes it much faster and easier than with a traditional taxonomical approach to identify what needs to be strengthened or mitigated. Taken together, this rigorous ontological approach and mathematical stratification mean we can make more confident claims about the evaluation problem space coverage — something traditional evaluation approaches have struggled to achieve, especially in complex, high-stakes domains. 

Click here to read a more detailed example

What’s next?

We are currently developing our knowledge graph / ontology methodology with several clients. Over the coming year, we will be exploring new statistical and mathematical models, improving our ontology coverage type (topologies of ontologies), experimenting with system prompts, and standardizing our software build so that clients can easily access their evaluations and evaluation inputs.

If you are interested in working with us on a knowledge graph AI evaluation, please contact us. We’re excited to collaborate with you!

Taxonomies vs. Ontologies: Choosing The Right Approach

Which approach is better – taxonomies or ontologies? Both! Like most complicated things, there is no single answer. By no means are we abandoning our taxonomical-based AI evaluations, as there are several reasons why they may be a better fit for a client. Here’s our simple guide:

Taxonomies are better for:

  • Smaller AI evaluation budgets
  • AI red teaming workshops that are primarily for educational purposes
  • Rapid, one-off evaluations
  • Well-defined and well documented problem spaces

Knowledge graphs / ontologies are better for:

  • Evaluations that build or expand over time
  • Reproducibility and comparability 
  • Teams with a high degree of subject matter expertise
  • Complex or high-stakes problem spaces that are constantly evolving
Sign up for our newsletter
Sign up for our newsletter