BENCHOUSE

The independent benchmark for Analytics Agents.

Fifteen agent configurations answered the same 300 analytical questions against the same deliberately messy e-commerce warehouse. Every answer was scored against ground truth the agents never had access to.

The first run is finished. Names, numbers and receipts go on the board at launch, and the list sees them before that.

You will get an email when there are new results, and when the methodology changes. Nothing else.

The board goes public on 1 September 2026. The list sees it first.

a sneak preview

The table below is illustrative. Every name and number in it is invented and the agents shown do not exist. Real scores publish on 1 September, each with a 95% interval, and configurations whose intervals overlap share a rank.

what we can already tell you

how we score

Read the methodology →

on the bench

nao Lightdash Snowflake Supersimple Anthropic

Products run in season one. The logos are their owners’ trademarks and show what was tested. They do not indicate partnership, sponsorship or endorsement, and no vendor approved, reviewed or paid for inclusion.
Queued next, not yet run: Hex, Databricks Genie, Looker, ThoughtSpot, Wren AI, Vanna.

Building an analytics agent? Get benchmarked →