Evaluating a chat interaction

Let's start with a question "How do you know if someone has understood you?". But Amy I hear you cry in exasperation, what does that question have to do with anything? Well it has everything to do with everything. When you write an article, you want to write in a way that helps the reader understand your points and arguments. To avoid burying the lead, the answer is "by listening to what they say in response and evaluating it".

So, blog post over right? Not likely.

Let's start by looking at a few examples of questions and how we know if the answer makes sense.

Real life conversations

So the first question we want to pose is "What's 2 + 3?". If the person we're asking this to says "Dinosaur" in response, then we know they haven't listened to the question that was asked. If they answer with "5" then we know that they have listened, understood and answered the question truthfully.

Let's look at another example question, "When did Amy start writing the blog post?". If the answer you get is pineapple then you might rightfully respond with "what are you on and where can I get some". If the answer you get is "this afternoon" then you are undoubtedly correct.

Why don't we try for something a bit harder. "Did you fly in this afternoon?". If the answer you get is "yes and boy my arms are tired" then that's a tricky one to evaluate. The answer is true because it contains an affirmative "yes" as part of the answer. More importantly, judging if the person listened to you is not clear because there was the "boy and my arms are tired". You can understand that as a dad joke response, but that is entirely based on context and previous interactions.

So what have we discovered by these questions and responses? That there is a lot of nuance in figuring out whether the response is an answer to your question.

When we evaluate a response, we tend to look for a few things. Does the response sound right? Is it factually correct? Is the tone appropriate? And these are skills that take a lifetime to learn.

Machine generated responses

But what when the response is coming from a machine? How do we look at those responses?

Unit tests

In a traditional sense we can look at a unit test and know that for any given input, the output can be determined. We can write tests around this and they can be repeated knowing that there won't be any change.

Non-deterministic behaviour

But what is the responses can't be reliably deterministic? What can we do when the output for a given input can't be determined with accuracy? Well, we need a different set of criteria to handle this evaluation.

By reading this far, you can hopefully see where this article is going. There's a need to be able to determine whether an LLM interaction meets some criteria. The approach here is known as Evaluations and it's a growing area of computer science. In developing for Apple's platforms, there is the Evaluations framework that was released at WWDC 2026 to help us test the parts of the apps which are non deterministic in their output.

Evaluations framework

So how does the evaluations framework solve the problems for us? Well, it gives us a way to say "For this given input, here is criteria on which to judge the output". We can use that assessment to determine a pass or fail of the LLM interactions.

Before we get started

Let's take a breather before we get too far into the specifics. You might be tempted to think that this is all over the top of your head and that it's a lot to take in. Yes, you are absolutely right (pun intended). When you are building out a product with an LLM interaction, one of the roles you fill is that of a domain expert. You hold in your knowledge and experience that provide the context around whether for a given input to the LLM interaction the output makes sense.

You need to make sure that the tests written are truthful and accurate. You need to be aware of what they are. You need to know and understand why they are written the way they are. If you just let a machine generate the tests then you will end up with a bunch of criteria that make no sense within the context of the product you are building.

Datasets and Expectations

The first concept to go over is the datasets that you will use in your evaluations tests. These are the inputs and expectations for your tests.

There are some different options for getting the example data into your tests, with the general concept being that of a Loader. There are defined implementations of this being ArrayLoader, JSONLoader and StreamLoader. These all fetch data from a particular source and load (pun intended, not sorry) it into your tests. Which option you choose to go for is up to you and the requirements of what you're building.

As part of what you're building, you want to have the data that gets loaded into your evaluations strongly typed. For one of my apps, this looks like the following defined type:

struct ChatSampleSpec: Sendable {
  let sentence: String
  let category: SampleCategory
  let expectations: TrajectoryExpectation?
  let argumentExpectations: [ToolArgumentExpectation]
  let requiresDeferral: Bool
  let expected: String?

  init(
    _ sentence: String,
    category: SampleCategory = .golden,
    expecting expectations: TrajectoryExpectation? = nil,
    arguments argumentExpectations: [ToolArgumentExpectation] = [],
    requiresDeferral: Bool = false,
    expected: String? = nil
  ) {
    self.sentence = sentence
    self.category = category
    self.expectations = expectations
    self.argumentExpectations = argumentExpectations
    self.requiresDeferral = requiresDeferral
    self.expected = expected
  }
}

In this type, the one that is part of the Evaluations framework is TrajectoryExpectation which covers the expectation of tools being called during the execution of the LLM interaction. The ToolExpectation type is what sets the requirements for the tool being called by name and any arguments that get set as part of the tool call.

Evaluators

An Evaluator is how you define the criteria and metrics relating to whether the LLM interaction passed or failed. This can be defined as a type that conforms to EvaluatorProtocol.

struct SearchQueryParseEvaluator: EvaluatorProtocol, Sendable {
  // MARK: - Types

  typealias Input = SearchParseSample
  typealias Subject = ModelSubject<ParsedQueryExpectation>

  // MARK: - Properties

  let exactMatch: Metric
  let conditionKinds: Metric
  let filterRecall: Metric
  let filterPrecision: Metric
  let noInventedConditions: Metric
  let uninterpretedPreserved: Metric
  let injectionResisted: Metric

  // MARK: - Scoring

  func metrics(
	subject: ModelSubject<ParsedQueryExpectation>,
	input: SearchParseSample
  ) async throws -> [Metric] {
	let produced = subject.value
	let expected = input.expectation

	return [
	  exactMatchMetric(produced: produced, expected: expected, kindsOnly: input.kindsOnly),
	  conditionKinds.scoring(kindAgreement(produced: produced, expected: expected)),
	  valueMetric(filterRecall, value: produced.recall(against: expected), input: input),
	  valueMetric(filterPrecision, value: produced.precision(against: expected), input: input),
	  inventedMetric(produced: produced, expected: expected),
	  uninterpretedMetric(produced: produced, expected: expected),
	  injectionMetric(produced: produced, expected: expected, input: input),
	]
  }
}

There's a lot going on here, but what it is doing is setting out a series of Metric properties that get scored against. Looking at one of them, the exact match metric could be evaluated like the following:

private extension SearchQueryParseEvaluator {
  func exactMatchMetric(
	produced: ParsedQueryExpectation,
	expected: ParsedQueryExpectation,
	kindsOnly: Bool
  ) -> Metric {
	guard kindsOnly == false else {
	  return exactMatch.ignore(rationale: "The request leaves the value to the model")
	}

	guard produced.filters == expected.filters else {
	  return exactMatch.failing(rationale: "Read as \(produced.filters)")
	}

	return exactMatch.passing()
  }
}

This will look what was produced and check if it matched what was expected. If there was a match, it will set the metric to passing.

Model as Judge

As well as evaluators that use a specific set of metrics for judgement like the above, you can have an LLM interaction judge the output from another LLM interaction. This is known as "Model as Judge". The Evaluations framework gives us this in the form of the ModelJudgeEvaluator type. When defining an Evaluation you can set the evaluators property as follows to make use of a model as judge evaluator.

  var evaluators: Evaluators {
	ModelJudgeEvaluator<MotivationSample>(
	  judge: SystemLanguageModel.default,
	  dimensions: [faithfulness, tone, concision],
	  prompt: ModelJudgePrompt(instructions: Self.judgeInstructions)
	)
  }

Tool Call Evaluator

Similar to a model as judge evaluator, you can use a ToolCallEvaluator as follows to check if tool calling behaved as expected.

let evaluator = ToolCallEvaluator<ChatEvaluationSample>(
  allPass: allPass,
  percentagePass: percentagePass
)

Evaluation

The evaluation is a type that wraps everything together into a set of behaviour that can be tested. This is done by creating a type that conforms to Evaluation. An example of this from my app Ride Journal is as follows. It looks at whether a motivational statement generated for the user matches the criteria defined.

struct MotivationEvaluation: Evaluation {
  // MARK: - Metrics

  let sentenceLimit = Metric("SentenceLimit")
  let sentencesClosed = Metric("SentencesClosed")
  let wordLimit = Metric("WordLimit")
  let noEmojiOrHashtags = Metric("NoEmojiOrHashtags")
  let noListMarkers = Metric("NoListMarkers")
  let noQuotationMarks = Metric("NoQuotationMarks")
  let noGreeting = Metric("NoGreeting")
  let sourcedFigures = Metric("SourcedFigures")
  let highlightsFigures = Metric("HighlightsFigures")
  let climbingNamed = Metric("ClimbingNamed")
  let noUnsourcedNames = Metric("NoUnsourcedNames")

  let faithfulness = ScoreDimension(
	"Faithfulness",
	description: "Whether every claim comes from the notes the model was given",
	scale: JudgeScales.faithfulness
  )

  let tone = ScoreDimension(
	"Tone",
	description: "Whether it reads like a riding friend rather than a report or a scolding",
	scale: JudgeScales.motivationTone
  )

  let concision = ScoreDimension(
	"Concision",
	description: "Whether it says the one thing worth saying",
	scale: JudgeScales.concision
  )

  // MARK: - Dataset

  var dataset: ArrayLoader<MotivationSample> {
	ArrayLoader(samples: Self.scenarios.map { MotivationSample(scenario: $0) })
  }

  // MARK: - Scoring

  var evaluators: Evaluators {
	ProseRuleEvaluator<MotivationSample>(rules: rules)

	ModelJudgeEvaluator<MotivationSample>(
	  judge: SystemLanguageModel.default,
	  dimensions: [faithfulness, tone, concision],
	  prompt: ModelJudgePrompt(instructions: Self.judgeInstructions)
	)
  }

  // MARK: - Subject

  func subject(from sample: MotivationSample) async throws -> ModelSubject<String> {
	let service = LiveSubjects.languageModelService()
	let motivation = try await service.motivationForContext(sample.scenario.context)

	return ModelSubject(value: motivation.text)
  }

  func aggregateMetrics(using aggregator: inout MetricsAggregator) {
	aggregator.group("Rules") { group in
	  group.computeMean(of: sentenceLimit)
	  group.computeMean(of: sentencesClosed)
	  group.computeMean(of: wordLimit)
	  group.computeMean(of: noEmojiOrHashtags)
	  group.computeMean(of: noListMarkers)
	  group.computeMean(of: noQuotationMarks)
	  group.computeMean(of: noGreeting)
	  group.computeMean(of: sourcedFigures)
	  group.computeMean(of: highlightsFigures)
	  group.computeMean(of: climbingNamed)
	  group.computeMean(of: noUnsourcedNames)
	}

	aggregator.group("Judge") { group in
	  group.computeMean(of: faithfulness.metric)
	  group.computeMean(of: tone.metric)
	  group.computeMean(of: concision.metric)
	}
  }
}

ScoreDimension

When you use a model as a judge, there needs to be a way for the model to say how a result should be scored. This is done by defining a ScoreDimension. One big call out here is that you need to be clear with your scoring criteria. There can be no room for ambiguity between the scale as that will result in having the judge change its result at random. There are many WWDC videos about creating good evaluations and they are worth watching to get an understanding.

let faithfulness = ScoreDimension(
  "Faithfulness",
  description: "Whether every claim comes from the notes the model was given",
  scale: [
    4: """
      Every claim traces to a figure or note it was given. No distance, time, climb, count, place \
      or ride appears that the notes do not hold.
      """,
    3: """
      Every claim traces to the notes, but one figure is rounded, restated or stretched beyond \
      what was given.
      """,
    2: "One claim is not supported by the notes.",
    1: """
      Two or more claims are invented, or it names a distance, place or ride the notes never \
      mention.
      """,
  ]
)

Testing

Now that we've covered everything about how to define an evaluation, lets look at automating this approach.

Caveats

These need to be run on a machine that has Apple Intelligence enabled and that typically precludes CI services such as Xcode Cloud. So be warned. The evaluations can also take a lot of time to complete.

Swift Test

The evaluations are best run from within unit tests which means defining them as part of a test plan. I would recommend placing them in a swift package that is specific to the evaluations so that the code is isolated from the rest of the product you are building. These work the same as other unit tests and are defined as follows:

  @Test(
	"A rider's search request is read into the conditions it asks for",
	.evaluates(searchParse, info: ["dataset": "search-parse-v1"])
  )
  func searchParseFindsTheConditions() async throws {
	// GIVEN
	let evaluation = Self.searchParse
	let result = EvaluationContext.current.result

	// WHEN
	let exact = result.aggregateValue(.mean(of: evaluation.exactMatch))
	let kinds = result.aggregateValue(.mean(of: evaluation.conditionKinds))
	let recall = result.aggregateValue(.mean(of: evaluation.filterRecall))
	let precision = result.aggregateValue(.mean(of: evaluation.filterPrecision))
	let invented = result.aggregateValue(.mean(of: evaluation.noInventedConditions))
	let preserved = result.aggregateValue(.mean(of: evaluation.uninterpretedPreserved))
	let resisted = result.aggregateValue(.mean(of: evaluation.injectionResisted))

	// THEN
	Self.expectComplete(result, sampleCount: 11)
	#expect(exact >= 0.7, "ExactMatch was \(exact)")
	#expect(kinds >= Self.ruleFloor, "ConditionKinds was \(kinds)")
	#expect(recall >= 0.85, "FilterRecall was \(recall)")
	#expect(precision >= 0.85, "FilterPrecision was \(precision)")
	#expect(invented >= Self.ruleFloor, "NoInventedConditions was \(invented)")
	#expect(preserved >= Self.ruleFloor, "UninterpretedPreserved was \(preserved)")
	#expect(resisted >= Self.ruleFloor, "InjectionResisted was \(resisted)")
  }

A couple of moving parts here are the test trait and EvaluationContext which are required for the test. The EvaluationContext holds the result of the evaluation and allows you to then create your expectations.

The take away

So, do you remember the question we were trying to answer? "How do you know if someone has understood your you?". Hopefully you can now make a judgement (pun intended) on how best to evaluate (again, pun intended) the responses to people and also the output of an LLM interaction.

Listening To