Skip to main content
Blog

The AI data sprawl

Nobody can trace it. The data team still owns the mess.

The AI data sprawl

The story of the last year in data: AI has gotten really good at writing code, and people are using it for analytics.

“Vibe coded” dashboards, queries, reports, and apps are everywhere. Millions of “non-technical” people are writing billions of lines of code, extremely quickly. At this point, it may actually be the case that more "non-technical" people are writing data code than using legacy BI tools.

And for good reason! The alternative was filing a ticket and waiting two weeks for a chart that answered a slightly different question. Now, you can ask in plain language and get exactly what you want in moments. Nobody is going back.

But this is all very… messy. And the problem isn’t that it’s all code, rather it’s how it’s generated, and where it lives. It’s built on random context and porous permissions, then scattered across one-off chats, HTML files pasted into Slack, and someone’s laptop. It runs once. Nobody reviews it. No one can trace it, reproduce it, or vouch for it.

Are any of these answers actually right?

Let’s be honest: you have no idea, and no way to find out. There’s no record of what context the agent got, which table it grabbed, or which filter it silently dropped before the answer landed in a board deck. The numbers look great – formatted, charted, and confidently narrated. That’s the problem! You can’t evaluate quality you can’t observe, or propagate a metric definition into a thousand one-off chats. The only feedback loop is noticing that people disagree.

Traditional BI tools can’t help – they’re built around proprietary specs and rigid dashboards. Bolting on a chat sidebar just gives people one more thing to vibe-code their way around.

This is AI Data Sprawl: a growing pile of plausible-looking numbers, one-off apps, and millions of lines of ungoverned code. Nobody can trace it. The data team still owns the mess, though!

And so data platform owners face an unhappy choice: let people generate whatever code they want and give up on governance – or lock them into a BI tool they trust precisely because it can't do very much.

This is the big problem everyone is feeling – and no one has really cracked the code.

We're working on this at Hex. Generative Data Apps are a big step – making it easy to use agents with context, controls, and collaboration (see the latest here).

But this story is still incomplete! And filling that out is what we're focused on right now, with the largest and most ambitious project we've ever taken on.

We'll share more soon (and much more at Prompt) – we can't wait to show you.

This is something we think a lot about at Hex, where we're creating a platform that makes it easy to build and share interactive data products which can help teams be more impactful.

If this is is interesting, click below to get started, or to check out opportunities to join our team.