Research Blog
2 days ago
7min
Hill climbing to glory: using evals to cut AI errors 7x
How 1,800 evals showed us AI feature quality improvements go beyond prompting


Aug 13, 2026
16min
Introducing DataBench
A frontier benchmark for complex data work and analytical reasoning

Jul 8, 2026
7min
Kimi K2.7 in Hex: near-frontier analytics at a fraction of the cost
Open models have finally caught up to their closed counterparts
Jun 9, 2026
7min
We had to build new evals for Fable
Claude Fable 5 is the first model since Opus 4.5 to meaningfully improve at analytical reasoning
Author/ Izzy MillerTags/ Data, ResearchMay 22, 2026
10min
How we built a lab to evaluate data agents
Inside Hex's eval architecture and the synthetic business it runs on.
Author/ Izzy MillerTags/ Engineering, ResearchMay 8, 2026
8min
How we built (and rebuilt) topic discovery into the Hex Context Agent
What are people actually asking your agent?
Author/ Charlene ChamblissTags/ Engineering, Research
Feed [7]
2 days ago
Hill climbing to glory: using evals to cut AI errors 7x
How 1,800 evals showed us AI feature quality improvements go beyond prompting
Aug 13, 2026
A frontier benchmark for complex data work and analytical reasoning
Jul 8, 2026
Kimi K2.7 in Hex: near-frontier analytics at a fraction of the cost
Open models have finally caught up to their closed counterparts
Jun 9, 2026
We had to build new evals for Fable
Claude Fable 5 is the first model since Opus 4.5 to meaningfully improve at analytical reasoning
May 22, 2026
How we built a lab to evaluate data agents
Inside Hex's eval architecture and the synthetic business it runs on.
May 8, 2026
How we built (and rebuilt) topic discovery into the Hex Context Agent
What are people actually asking your agent?