Using TypeSafe’s Jev for evals in Datadog Agent Observability
Fouad Wahabi
Software Engineering Lead
Alex Barksdale
Senior Software Engineer
Miguel Tulla Lizardi
Software Engineer
TypeSafe AI released Jev in September 2026 to do one thing: make decisions. Give it a state (a string or a JSON object) plus a set of typed questions, and it returns typed answers with probabilities. It never explains itself, and that constraint is the whole idea.
Evaluation pipelines have spent the last two years asking text generators for yes/no verdicts, wrapping the reply in a JSON schema, and paying generation prices for what amounts to a single bit. Jev is built for that bit, which makes it a good fit for the two evaluation surfaces in Datadog Agent Observability: online evals, which score production spans as they arrive, and experiments, which score a dataset offline. The same rubric can drive both.
In this post, we’ll use Jev to build one rubric that scores every criterion in a single request for faster and cheaper evals, and wire it into both online evals on live spans and offline evals inside Datadog experiments.
What Jev returns
Each question for Jev is identified by a key, and its answer is returned under that key. Jev supports three question types, and each one returns a different kind of answer:
| Question Type | Returned Signal | Example evaluation |
|---|---|---|
| Noul | Probability that a yes/no proposition is true | Are the reply’s policy claims supported by the retrieved excerpts? |
| Choice | Selected category, probabilities for all categories, and confidence | What is the reply’s main failure mode? |
| Score | A probability-weighted average of rubric levels, plus the level distribution and confidence | How severe is the potential customer impact? |
A Noul probability expresses uncertainty about a proposition; it does not measure how much of the reply is correct. Choice probabilities and Score confidence summarize the spread of their probabilities, though keep in mind that a confident answer can still be wrong.
Questions in the same request are evaluated independently against a shared state. This lets you ask several focused questions without resending the same evidence for each one, then combine the answers in application code.
Putting Jev to the test
We’ll use a support agent for a fictional airline, Vega Air, as an example for using Jev. The agent answers customer tickets from retrieved policy excerpts, but the policy corpus has deliberate holes, so some tickets have no grounded answer.
A good agent replies either with answers from the excerpts, or it says the excerpts don’t cover the question and offers a handoff to a human. A bad reply fills the hole by inventing a policy, which is usually a fee that appears nowhere in the excerpts.
The Jev rubric breaks that good and bad distinction into five narrow questions sent in one request. Two of them are below and are a Noul and a Choice question type. Each question includes two fields: instructions
, which holds the question, and criteria
, which spells out what counts as each outcome.
instructions
can be a plain string or a JSON object. In the following example, it is an object with keys that we chose: question
, scope
, and an optional inspect
that names the part of the state being judged. failure_mode
skips inspect
because its question already names reply
. You can find the other three questions, and a walkthrough of each field, in our Jev rubric notebook.
The unclear
option in failure_mode
is deliberate. A Choice question always returns the option with the highest probability, so Jev never abstains. If an evaluation needs a way to say “cannot judge this one,” that outcome has to exist in the criteria. TypeSafe’s self-consistency cookbook uses the same pattern for moderation.
Here’s a real response. The ticket asked about cancellation compensation when the retrieved policy only covers delays, so the agent declined to answer and offered a handoff:
The interesting part sits inside failure_mode
. Jev picked none
, but none
at 0.46 and partial_answer
at 0.42 are nearly tied, and confidence came back at 0.34. Flattening that to the string none
throws the interesting part away. A near-tie between two categories is a signal in its own right, and a natural trigger for routing the trace to a human reviewer.
The composite verdict stays in application code:
That composite verdict could have been a sixth Jev question. Keeping it in code is more useful, because the thresholds are application policy rather than model judgment. Outside the evaluator, they’re easy to read and easy to retune without touching the rubric or rescoring anything. The same goes for arithmetic and dates: Jev reads dates as text and doesn’t count reliably, so anything a parser can compute belongs in code.
The call itself is one request against a pinned model:
TypeSafeClient
reads TYPESAFE_API_KEY
from the environment, so nothing else needs configuring. The state is three named fields rather than the whole trace. Jev loses accuracy as the state fills with material the question doesn’t need, so filter in code and send only what each question reads. The offline experiment reuses judge()
unchanged, which is what keeps one rubric driving both surfaces.
Scoring live spans with online evals
For online evals, the appeal of Jev is that you pay for the verdict and nothing else. Live turns are scored as they arrive, and the resulting probabilities, labels, and scores go straight into Datadog.
The traced application never imports Jev. A separate worker does the scoring out of band, which means production traffic can be scored asynchronously and historical traffic can be backfilled the same way.
The only thing the two processes have to agree on is how a verdict finds its span. Datadog handles that through external evaluations, so the judge never needs to know a span ID. The application tags each span with a domain key (e.g., turn_id
) and the scorer joins on that tag.
In this example, the scorer reads a JSON Lines file, standing in for whatever queue, table, or store already holds the turns you want judged:
Passing the turn’s own timestamp_ms
rather than the current time keeps reruns idempotent. Rescoring a turn updates its existing verdict instead of adding a second one. Both halves end-to-end can be found in the online evals notebook in our GitHub repo.
Mapping Jev answers to Datadog metrics
LLMObs.submit_evaluation
accepts four metric types: score
, categorical
, boolean
, and json
. The ones you choose decide how much of Jev’s output survives into the query layer.
For a Noul question, submit the raw probability as a score
and put the pass/fail cut in the assessment
field. Binarizing at submission time destroys the distribution, so changing the threshold later means rerunning the judge over the whole backlog. Keeping the probability turns a threshold change into a query change. Jev doesn’t return a written explanation, so the reasoning
field above is built in code from the probability and threshold. Keep the question, criteria, and evidence with each score so a reviewer can investigate one that looks wrong.
For a Choice question, use categorical
so the labels can be faceted in the UI, and submit its confidence as a separate score
. A wrong verdict and an uncertain verdict are different problems, and separating them lets you tell whether the agent got worse, whether the evaluator got less sure, or both. Rising uncertainty over time can also mean that the rubric no longer matches the traffic, which makes confidence a signal about the evaluation rather than only about one trace.
Tag every metric with judge_model
. Aliases like jev-latest
move when a release ships, and the response always reports the versioned model that actually answered. Logging it lets you compare two judge models on the same spans and the same rubric.
Using Jev inside a Datadog experiment for offline evals
A Datadog Agent Observability experiment takes a dataset, a task, and a list of evaluators. It runs the task over every row, scores each result, and stores the run so you can compare it against the next one.
The interface expects one evaluator object per metric. Ported naively, that’s one Jev request per evaluator per row. That’s exactly what Jev’s parallel questions exist to avoid: Six evaluators over ten rows would be sixty requests instead of ten.
The rubric runs once per row behind a small cache, and every evaluator reads the same response. One detail matters in that cache: guard the dictionary, not the request. experiment.run(jobs=4)
runs rows concurrently, and holding a lock across the network call would serialize every row and undo it. Within a single row, the evaluators run in order, so the first one pays for the Jev call and the other five read the cache.
The evaluators themselves are thin:
Wiring it up looks like any other experiment:
The offline surface can also ask something that the online one cannot. Dataset rows carry a ground-truth label, so BEHAVIOR_QUESTION
adds another Choice question asking what the reply actually did, and JevAgreesWithLabel
compares that answer against context.expected_output
. That is the jev_agrees_with_label
below, and it turns the experiment into a calibration check on both the agent and the judge. Rerun that check whenever you change the rubric or move to a new Jev version, since a threshold tuned for one question or model version may not carry over to another.
The cache class, all six evaluators, and the dataset setup are in our third Jev rubric for experiments.
Get started with Jev for your Datadog evals
You can access Jev directly or through AI gateways like OpenRouter or Vercel AI Gateway. TypeSafe provides Python and JavaScript SDKs, as well as a TypeSafe agent skill that gives coding agents the full API context. Like any judge, Jev has known limitations, so measure its agreement with human reviewers and its repeatability on your own traffic before you rely on it.
On the Datadog side, everything runs on the released ddtrace>=v4.5.0
package. LLMObs.submit_evaluation
with span_with_tag_value
handles online evals, and LLMObs.experiment
with BaseEvaluator
subclasses handles offline evals. There’s no preview builds and no private endpoints. Clone the repo, add your keys, and run them in order:
Our GitHub repo contains three notebooks that are a complete, runnable version of everything in this post:
1-jev-rubric.ipynb builds the rubric one question at a time and reads the answers. This notebook needs a TypeSafe key to get started.
2-online-evals.ipynb traces the agent, then scores the spans out of band.
3-experiments.ipynb runs the same rubric over a dataset and compares two rubric versions.
Check out our Agent Observability documentation to learn more about monitoring and evaluating your LLM applications. And read our blog on using evaluation frameworks with Agent Observability.
If you’re new to Datadog, get started with a free 14-day trial.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.