Skip to main content
Blog

Trust used to be implied

Evals aren't just how you fix your agent — they're how you convince people to trust it.

Trust used to be implied

A Head of Data told me he'd found a second use for his evals.

He'd spent weeks getting his agent to answer correctly. Correct, correct, correct — the same questions, day after day, until the results were boring. Then, instead of keeping those scores to himself, he broadcast them across the company. Not to prove the agent worked. To make people trust it before they ever typed a question, over and over again.

I'd come into the conversation expecting to talk about how he built his evals — what they looked like, what they caught, how his context improved. That's the job we all know: evals are how you find what's broken and fix it. His second use surprised me, because it's obvious the moment you hear it. Agents are finicky…and the work to make them answer consistently must play a role not just in the testing but also in the adoption strategy.

Evals tell the data team what to fix. But there's a second job, maybe the bigger one: they show the business the agent is worth trusting.

Dashboards gave us trust for free

Before AI, there were dashboards, the source of truth to the business. Dashboards were encoded, rigorously QA’d, and built by humans, so you felt protected from inaccuracies. Teams built these dashboards, and people used them (most of the time), and the numbers were blessed.

AI breaks that contract, and now users ask whatever they'd like. But this doesn’t reduce the burden of truth from the data team. It expands it. Data teams curate context from all across the business, organizing it for agents and documenting things that lived in people’s heads or a random SQL query.

Now you’re on a razor’s edge. One wrong agent answer away from breaking that trust and business users rejecting the thing you’ve worked so hard to build. You can observe these conversations and even view poor agent responses in places like the Context Studio, but that’s after someone has already had the bad experience.

So how do you build trust in this environment? Let’s set up evals! The data team can run these at regular intervals, catching misses in the context or agent behavior and fixing them. Adoption problem solved? Not quite. If teams only review these evals internally, they are missing the opportunity presented by that Head of Data. Internal reviews are great to understand what your team should work on… they don’t serve you at all in driving adoption. So we have to point them outward as well.

Now you have to earn it over and over

You’re playing both an acquisition and retention game. Even, dare I say, a bit of marketing, to drive people to use the things you’ve built. If you’re a consumer brand with several bad reviews on Amazon, it’s going to be very hard for you to get new customers or win back the ones that you’ve disappointed without a heavy hand. You have to show them how you’ve improved your product, demonstrate that others are loving it as well, and break through the noise to get them to open their wallets again.

This is what data teams must do as well. Not just run the evals and “fix the product,” but market that fix and promote the improvements to win back and gain new users across your organization. You need to earn the baseline trust before you broadcast it; a faulty fix is going to hurt you more than no fix at all. When you get there, it’s not simply shipping the eval results in a Slack channel. It requires regaining the belief they can use this as confidently as the dashboards of yore.

Evals on the surface feel complex; you have to simplify the results and tell the story just like you would any other analysis. Consistent high scores on known-known questions may be boring to a data team that’s worked hard to model that data, but to a business user, that’s real signal they can trust the agent. Promoting that as loud and often as possible in various channels to show people that you’ve:

  1. Done the work to organize the data
  2. Built confidence in a repeatable answer to the most common questions

The Head of Data spoke about not just sharing scores, but the domains and sample questions inside the eval, to give users even more confidence that it’s not just right, but right for the things they care about.

The ROI bill is due

For as long as I've been in data, our impact has been measured in dashboard views and clicks. Did someone open the thing? Did they open it twice? That was the pulse — a thin, almost insulting proxy for months of modeling, curation, and judgment that never showed up in the number.

The chart I want you to picture is different. Plot your context improvements over time. Layer your eval scores on top of them, climbing as the context gets richer. Then lay adoption behind both. For the first time, the work and its results sit in one honest picture — not "did someone click," but whether the business is actually using data to make decisions, and whether your work is accelerating that adoption.

And the timing isn't an accident. Every company is under the same spotlight right now — the AI bill is due, the noise is deafening, and the question has curdled from "isn't this exciting" into "now prove it." Teams everywhere are reaching for evidence and coming up with platitudes about non-determinism. Data teams happen to be holding something that actually produces the proof.

But the chart only bends upward if you run both jobs. The inward one is quiet: catch the gaps, fix the context, keep the agent honest. The outward one is loud: take the proof and put it in front of the people who stopped trusting you after one bad answer. Do only the first and you've built a well-tuned agent that’s sparsely used — the exact dashboard problem you started with.

It isn't enough to build something great and set it loose. The proof was never going to market itself.

This is something we think a lot about at Hex, where we're creating a platform that makes it easy to build and share interactive data products which can help teams be more impactful.

If this is is interesting, click below to get started, or to check out opportunities to join our team.