Every agent aces the easy questions
29 August 2026
Issue 1 of the Benchouse newsletter, sent to subscribers on 29 August 2026 and republished here as sent — the figures below are pulled from the same run artifacts the email used. Email-only plumbing (unsubscribe, forwarding notes) is omitted.
Hey, you who care deeply about analytics agents,
I’m really excited to share the first set of benchmarked agents with you all. What follows is the first published result from the Benchouse leaderboard. I asked every agent the same 300 questions on a fully equivalent semantic layer against the same data.
I originally had hoped and promised to publish fifteen agents by September 1st. Alas, I got to seven. Each integration takes a lot of days: an adapter, context-layer equivalency and a data warehouse. And I preferred to get these seven out rather than postpone.
On September 1st the results will also be published at benchouse.ai/benchmark.
The results
In the first table you can find the top-level results; later on I’ll get into each agent in more detail.
| Configuration | Accuracy | Completeness | Restraint | Measured |
|---|---|---|---|---|
| cortex-agents · claude-opus-4-8 · snowflake semantic view | 84% | 100% | 100% | 300 of 300 |
| hex · gpt-5.6 (Hex) · hex semantic model (dbt) | 74% | 100% | 100% | 267 of 300 |
| nao · claude-sonnet-5 · nao table docs | 70% | 100% | 100% | 293 of 300 |
| supersimple · LLM not disclosed · supersimple semantic model | 70% | 98% | 100% | 300 of 300 |
| genie · LLM not disclosed · genie space instructions | 66% | 68% | 80% | 271 of 300 |
| cortex-analyst · claude-sonnet-4-6 · snowflake semantic view | 55% | 81% | 100% | 300 of 300 |
| lightdash · claude-sonnet-4-6 · lightdash semantic model | 54% | 95% | 100% | 267 of 300 |
Question bank seed-v1, judged by claude-opus-4-8-r4, k=1, seed 42. Judge agreement is not yet measured for this judge version. "Measured" counts the questions that reached a metric's denominator — cells lost to adapter or judge errors sit in none of them.
Accuracy: of the questions it did answer, how many were right?
Completeness: of the questions it had the data to answer, how many did it answer?
Restraint: some questions cannot be answered from the data, and the correct response is to say so.
Nobody paid to be in this table, and there is no price for a better place on it. No vendor saw the questions before the run. None saw the results before you did.
I give this away because what I sell is the same measurement pointed at your own agent. That only works if you can trust this table, so keeping it clean is how I stay in business.
What is your agent getting wrong?
As you can see, agents are not perfect (yet). When your team asks your agent a question, how often is it right? And how do you know?
We will tell you. We help you be in control of your agent. Using the same tech that publishes the public leaderboard, we built the perfect control center for your analytics agents.
Take the next slot → — interview + waitlist access only. Onboarding each customer one at a time to guarantee results.
Detailed Results
This is the juicy part. When I first started developing the benchmark I mostly asked questions like “How many orders were there in January?”, and honestly: the benchmark was boring. Most agents answer these questions really well. Where it gets interesting is when you ask “Why did my orders for the fishtank category tank?” (no pun intended), or even better “Which marketing ads should I invest more money in to improve my fishtank sales?”
| Configuration · tier | Accuracy | Completeness | Measured |
|---|---|---|---|
| cortex-agents · claude-opus-4-8 · snowflake semantic view | |||
| T1 · descriptive | 99% | 99% | 100 |
| T2 · diagnostic | 81% | 100% | 100 |
| T3 · prescriptive | 72% | 100% | 100 |
| hex · gpt-5.6 (Hex) · hex semantic model (dbt) | |||
| T1 · descriptive | 98% | 100% | 97 |
| T2 · diagnostic | 69% | 100% | 83 |
| T3 · prescriptive | 55% | 100% | 87 |
| nao · claude-sonnet-5 · nao table docs | |||
| T1 · descriptive | 96% | 100% | 98 |
| T2 · diagnostic | 68% | 100% | 99 |
| T3 · prescriptive | 47% | 100% | 96 |
| supersimple · LLM not disclosed · supersimple semantic model | |||
| T1 · descriptive | 98% | 100% | 100 |
| T2 · diagnostic | 73% | 99% | 100 |
| T3 · prescriptive | 39% | 96% | 100 |
| genie · LLM not disclosed · genie space instructions | |||
| T1 · descriptive | 96% | 99% | 96 |
| T2 · diagnostic | 43% | 94% | 81 |
| T3 · prescriptive | 0% | 16% | 94 |
| cortex-analyst · claude-sonnet-4-6 · snowflake semantic view | |||
| T1 · descriptive | 95% | 99% | 100 |
| T2 · diagnostic | 46% | 92% | 100 |
| T3 · prescriptive | 0% | 53% | 100 |
| lightdash · claude-sonnet-4-6 · lightdash semantic model | |||
| T1 · descriptive | 97% | 99% | 96 |
| T2 · diagnostic | 39% | 93% | 83 |
| T3 · prescriptive | 20% | 94% | 88 |
How the questions are graded
Descriptive asks what happened. How many orders were there in total? A single number, checkable against the warehouse.
Diagnostic asks what happened and why. It needs the agent to join the right things and reason about the result, not merely retrieve it.
Prescriptive asks what happened, why, and what to do about it. This is the work an analyst is actually paid for, and it is where the table above falls off a cliff.
Finally
I hope you enjoyed this as much as I enjoyed building it. This is not a static benchmark; it will be extended over the coming months with more agents. Let me know which ones you think should definitely be up there.
— André Baaij, founder of Benchouse
Figures pulled from the eval-harness leaderboard-2026S2 run artifacts via cmd/newsletterfigures on 2026-09-05, question bank seed-v1, judge claude-opus-4-8-r4 — never transcribed. A new set of results goes out every few months: subscribe in the sidebar.