BENCHOUSE
The independent benchmark for Analytics Agents.
Fifteen agent configurations answered the same 300 analytical questions against the same deliberately messy e-commerce warehouse. Every answer was scored against ground truth the agents never had access to.
The first run is finished. Names, numbers and receipts go on the board at launch, and the list sees them before that.
You will get an email when there are new results, and when the methodology changes. Nothing else.
The board goes public on 1 September 2026. The list sees it first.
a sneak preview
The table below is illustrative. Every name and number in it is invented and the agents shown do not exist. Real scores publish on 1 September, each with a 95% interval, and configurations whose intervals overlap share a rank.
what we can already tell you
- The most expensive agent is not the most accurate.
- With the same agent, the same model and the same underlying facts, changing only the format of the context layer moved the score.
- Every agent scored worse on the questions that asked why something happened.
how we score
- Every configuration (an agent, an LLM and a context layer) answers the same 300 questions on a simulated e-commerce warehouse. The questions come at three levels: what happened, why it happened, and what to do about it.
- Answers are graded against a truth ledger that is computed separately and withheld from every agent. An agent only scores when its number matches, however plausible its SQL looks.
- Every score carries a 95% interval, and configurations whose intervals overlap share a rank.
- No vendor pays to be listed or sees the questions in advance. Every number has a receipt recording its run id, seed and command.
on the bench
Products run in season one. The logos are their owners’ trademarks and show what
was tested. They do not indicate partnership, sponsorship or endorsement, and no vendor
approved, reviewed or paid for inclusion.
Queued next, not yet run: Hex, Databricks Genie, Looker, ThoughtSpot, Wren AI, Vanna.
Building an analytics agent? Get benchmarked →
Evaluate your stack
A call to talk through what you are running now and what you are choosing between, and whether a run on your own warehouse would settle it.
The same harness as the bench, pointed at your warehouse and your own questions.
Prefer email? andre@benchouse.ai
This is paid work, and it is kept separate from the bench. Buying an evaluation does not affect placement on this board. Full disclosure →