Benchouse blog

Every agent aces the easy questions

Issue 1 of the Benchouse newsletter, sent to subscribers on 29 August 2026 and republished here as sent — the figures below are pulled from the same run artifacts the email used. Email-only plumbing (unsubscribe, forwarding notes) is omitted.

Hey, you who care deeply about analytics agents,

I’m really excited to share the first set of benchmarked agents with you all. What follows is the first published result from the Benchouse leaderboard. I asked every agent the same 300 questions on a fully equivalent semantic layer against the same data.

I originally had hoped and promised to publish fifteen agents by September 1st. Alas, I got to seven. Each integration takes a lot of days: an adapter, context-layer equivalency and a data warehouse. And I preferred to get these seven out rather than postpone.

On September 1st the results will also be published at benchouse.ai/benchmark.

The results

In the first table you can find the top-level results; later on I’ll get into each agent in more detail.

ConfigurationAccuracyCompletenessRestraintMeasured
cortex-agents · claude-opus-4-8 · snowflake semantic view84%100%100%300 of 300
hex · gpt-5.6 (Hex) · hex semantic model (dbt)74%100%100%267 of 300
nao · claude-sonnet-5 · nao table docs70%100%100%293 of 300
supersimple · LLM not disclosed · supersimple semantic model70%98%100%300 of 300
genie · LLM not disclosed · genie space instructions66%68%80%271 of 300
cortex-analyst · claude-sonnet-4-6 · snowflake semantic view55%81%100%300 of 300
lightdash · claude-sonnet-4-6 · lightdash semantic model54%95%100%267 of 300

Question bank seed-v1, judged by claude-opus-4-8-r4, k=1, seed 42. Judge agreement is not yet measured for this judge version. "Measured" counts the questions that reached a metric's denominator — cells lost to adapter or judge errors sit in none of them.

Accuracy: of the questions it did answer, how many were right?

Completeness: of the questions it had the data to answer, how many did it answer?

Restraint: some questions cannot be answered from the data, and the correct response is to say so.

Nobody paid to be in this table, and there is no price for a better place on it. No vendor saw the questions before the run. None saw the results before you did.

I give this away because what I sell is the same measurement pointed at your own agent. That only works if you can trust this table, so keeping it clean is how I stay in business.

What is your agent getting wrong?

As you can see, agents are not perfect (yet). When your team asks your agent a question, how often is it right? And how do you know?

We will tell you. We help you be in control of your agent. Using the same tech that publishes the public leaderboard, we built the perfect control center for your analytics agents.

Take the next slot → — interview + waitlist access only. Onboarding each customer one at a time to guarantee results.

Detailed Results

This is the juicy part. When I first started developing the benchmark I mostly asked questions like “How many orders were there in January?”, and honestly: the benchmark was boring. Most agents answer these questions really well. Where it gets interesting is when you ask “Why did my orders for the fishtank category tank?” (no pun intended), or even better “Which marketing ads should I invest more money in to improve my fishtank sales?”

Configuration · tierAccuracyCompletenessMeasured
cortex-agents · claude-opus-4-8 · snowflake semantic view
T1 · descriptive99%99%100
T2 · diagnostic81%100%100
T3 · prescriptive72%100%100
hex · gpt-5.6 (Hex) · hex semantic model (dbt)
T1 · descriptive98%100%97
T2 · diagnostic69%100%83
T3 · prescriptive55%100%87
nao · claude-sonnet-5 · nao table docs
T1 · descriptive96%100%98
T2 · diagnostic68%100%99
T3 · prescriptive47%100%96
supersimple · LLM not disclosed · supersimple semantic model
T1 · descriptive98%100%100
T2 · diagnostic73%99%100
T3 · prescriptive39%96%100
genie · LLM not disclosed · genie space instructions
T1 · descriptive96%99%96
T2 · diagnostic43%94%81
T3 · prescriptive0%16%94
cortex-analyst · claude-sonnet-4-6 · snowflake semantic view
T1 · descriptive95%99%100
T2 · diagnostic46%92%100
T3 · prescriptive0%53%100
lightdash · claude-sonnet-4-6 · lightdash semantic model
T1 · descriptive97%99%96
T2 · diagnostic39%93%83
T3 · prescriptive20%94%88

How the questions are graded

Descriptive asks what happened. How many orders were there in total? A single number, checkable against the warehouse.

Diagnostic asks what happened and why. It needs the agent to join the right things and reason about the result, not merely retrieve it.

Prescriptive asks what happened, why, and what to do about it. This is the work an analyst is actually paid for, and it is where the table above falls off a cliff.

Finally

I hope you enjoyed this as much as I enjoyed building it. This is not a static benchmark; it will be extended over the coming months with more agents. Let me know which ones you think should definitely be up there.

— André Baaij, founder of Benchouse

André Baaij
André Baaij

Founder of Benchouse, the control center for your analytics agents. Runs the public benchmark on analytics agents. Former Head of Data at Babylist.