Chime® built a context workflow and eval set that drove 5K additional insights per month via the Hex Agent
Chime built an eval set to maintain trusted context for every new data domain launched, so 500+ users could self-serve business insights in conversational form with Hex's agent.
About Chime
Chime® (NASDAQ: CHYM) is a financial technology company founded on the premise that core banking services should be helpful, easy, and free. Chime offers a broad range of low-cost banking, payments, lending, and investing products that address the most critical financial needs of everyday people, serving 10M+ members.
Outcome
- 5K additional insights per month via the Hex agent
- 96% of Threads conversations run with no user doubt
- 60% of questions come from non-technical users
- 500+ users self-serve data with Hex’s agent
- 50+ data domains
- 60% reduction domain launch time, from two and a half weeks down to five days
Scale achieved
- Unlocking self-serve insights for the entire organization
- Trust and accuracy for data across the org
- A question flagged and context fixed before anyone had to ask again
The challenge: a fintech serving 10M+ Americans that constantly launches new products & features needed reliable insights
Chime® wanted every team member across marketing, product, operations, and finance to self-serve their own data questions and build a more data-driven organization. As a public company with nearly 1,500 employees, 14 years of history and a growing product lineup, maintaining trustworthy context for both new hires and executives was table stakes. Products like checking, savings, a credit builder card, SpotMe®, MyPay®, investing, and more each carried their own metric definitions and edge cases. With so many evolving complex products and new offerings regularly launched, finding out where to begin was daunting.
The pace Chime operates at means that any context built for Hex’s agent starts going stale almost immediately. A definition that's accurate the day it ships can be out of date the next day, so Chime needed to know exactly how much context Hex’s agent needed for each data domain and to prove it gave accurate answers before launching a domain company-wide.
"We have so many people and products that our context is never stagnant. Even when we launch a new data domain for the first time, it's already stale the next day. Hex’s Context Studio and Evals give us the visibility needed to manage our context at scale." — Jordan Farrer, Director of Data Science & Analytics, Chime
The solution: 96% of Threads conversations run with no user doubt, a tribute to Chime's investment in context and regular evals
Chime bought Hex before agentic analytics existed, betting that one collaborative workspace would outlast the patchwork of tools it replaced. When Hex's agent capabilities matured, that early bet paid off, giving teams a central place to go beyond traditional reporting to true business intelligence in conversational form.
Given the breadth of Chime’s business, there needed to be a strategy to build the right context for Hex’s agent and maintain trust. Jordan Farrer, Director of Data Science and Analytics at Chime, developed a data domain strategy for the context layer — segmenting Chime’s data landscape into logical areas and writing the playbook on how to endorse each for Hex’s agent.
The team’s context workflow included domain evaluations, continuous monitoring, and tests to keep pace with a business that never stops changing. Today, the context layer powers everything from business insights and exploratory analysis to building data visualizations via generative apps.
"We frequently run evals on our context library to maintain accuracy and trust in the answers the Hex agent gives. Having everything on Hex's platform provides visibility, governance, and control over how our users consume data across the organization." - Dori Wilson, Senior Data Scientist, Chime
Chime has a data stack that enables 500+ users to self-serve in natural language
Chime also needed a data stack that could keep up with the company's growth, while being able to maintain control over the minute details. Their data is stored in Snowflake, which is their system of record for raw and transformed data, and flows through dbt. From there, dbt tables get column-level detail through YAML files. Chime then organizes that context into domains, built around products like investing or SpotMe, a workflow like customer onboarding, or an experience like a customer support interaction, and keeps each one separate to reduce overlap and conflict.
On top of those tables, Chime built a semantic view and a guide for each domain. Guides that carry the top-line company-level facts live in the team's GitHub repo, and sync into Hex, while more table-level details stay in the semantic layer. The team also endorses its most-trusted projects, so Hex's agent can pull from vetted projects built across the company, not just guides written for this purpose.
This structure lets Hex understand the nuances of the business. Hex’s agent expresses when there are options on how a metric can be defined and shared when the context is not endorsed.
Chime's retention rate, for example, means something different depending on which team is asking. So instead of picking one definition and guessing, Chime's guides teach the agent to ask which definition applies before answering, the same way a data scientist would.
The result is a data stack that lets non-technical and technical users access data themselves. Hex is ergonomic for external agents and meets Chime's users where they're already working. More than 500 people across Chime, including its senior leadership, now ask data questions in plain language via the Hex UI, Slack integration, and AI platforms via MCP.
Hex serves as Chime’s data hub, providing governance over how context is accessed, observability over agent responses, and security through OAuth-based connections. That level of visibility and control gives Chime’s users access to data without compromising trustworthiness.
“There has always been an infinite number of questions people wanted to ask. Hex is what finally lets us answer them and dive deeper than we could before, making our whole organization more data-informed.” — Jordan Farrer, Director of Data Science, Chime.
Hex’s Evals cut domain launch time by 60%, from two and a half weeks to five days
With agentic analytics, Chime needed the agent to understand where the data lived, how the business measured things, and be able to prove that its understanding was accurate before the team was able to trust it broadly. They wanted each domain that they launched to have well-documented data foundations, a semantic guide, and evaluation set.
With this in mind, Chime wanted to evaluate Hex’s agent like evaluating a data science team member. Connor Macdonald, Senior Data Scientist at Chime, built this three-level framework, separating questions by level of complexity, for the evaluation set:

For the investing data domain, one of Chime's newest products, questions ranged from how many members were investing on the Chime platform, to what characteristics made investing users unique? They didn’t let a domain ship until Hex’s agent passed all three levels consistently.
Chime tested five points during their evals, including:
- Is the response accurate compared to the reference answer?
- Are the right tables, guides, and semantic models being used?
- Is the query structured correctly?
- Does the agent ask clarifying questions as expected?
- If context drift conflicts across previously endorsed domains over time.
Dori Wilson, Senior Data Scientist at Chime, built an evaluation platform outside of Hex to run the eval set for each domain on a regular cadence to flag regressions. This tool submitted questions to Hex one-by-one and used an LLM-as-judge for evaluation. But this required manual tuning and maintenance.
So, when Hex launched its evals through the CLI in beta, Dori rebuilt the process around it, pairing it with Claude to evaluate the questions from the launched data domains. As new domains launch, Dori continues to add questions to the evaluation set.
"Building my own eval tooling was tedious, but Hex's CLI is what let me stop maintaining a second system and focus on the actual questions that judge our context layer." — Dori Wilson, Senior Data Scientist, Chime
Ongoing evals maintain 96% agentic answer without user doubt with 60% of questions coming from non-technical users
Passing Hex’s evals is the gate before a domain launches, but Chime also needed to make sure context didn't drift once a domain was live. Context Studio gave the team a running signal on what people were asking and where to build or fix context next, whether it was a column needing a clearer description or a product detail in a guide that needed updating.
Today, Connor manages that workflow, prioritizing the context gaps that Hex’s Suggestions flags. For example, when a senior executive got an uncertain answer to a question about SpotMe, Context Studio appropriately flagged exactly why: an "active SpotMe member" column wasn't documented clearly enough. Connor confirmed the fix with the subject matter expert, then pushed the change directly into the context layer.
That same review happens for every suggestion Context Studio raises. Connor checks each one with the product owner or a subject matter expert, then either accepts the fix in Context Studio or pushes a manual update to the team's GitHub repo. Chime tracks grouped suggestion themes worth investigating, and deliberately chooses which ones to investigate based on how many people flag the question.
That discipline shows up in Chime's numbers. Ongoing evaluation work has helped maintain Chime's instance of Threads answers at 96% without user doubt. At least one quarter of all questions are now coming from non-technical users who used to depend on an analyst for every one.
"Context Studio tells us exactly where to look before a bad answer persists. It's turned context maintenance from guesswork into a systematic process." — Connor Macdonald, Senior Data Scientist, Chime
The impact: an AI agent a regulated fintech is willing to trust
Chime's data team has moved from fielding requests one at a time to building infrastructure that lets the rest of the company find its own answers.
With Hex, Chime has been able to:
- Unlock 5K additional insights per month via the Hex agent using trusted context
- Sustain 96% of Threads conversations with no user doubt, starting with semantic models, guides, and evals built for every domain launch, and proactively monitoring context drift via the Context Studio
- Have 60% of questions now come from non-technical users, up from a data team that used to field every request itself
- Enable 500+ users to self-serve in plain language, across product, ops, and finance teams that used to wait in a queue for every question
- Maintain 50+ data domains, each gated by passing all three eval question levels
- Reduce domain launch time by 60%, from two and a half weeks down to five days
Chime continues to add domains to its roadmap, evaluating Hex's newly launched eval suite to run its full evaluation pipeline, and working toward a future where subject matter experts own their domain’s context.
CHIME is a Registered Trademark of Chime Financial, Inc., used with permission.