← The Independent Analytics Agent Benchmark
cortex-agents vs supersimple
Season 2026-S2 · question bank seed-v1 · judge claude-opus-4-8-r4 · k=1
Both configurations answered the same 300 questions against the same deliberately messy simulated warehouse, graded by the same judge against withheld ground truth. cortex-agents (claude-opus-4-8 · snowflake semantic view) answers 83.7% of the questions it attempts correctly (246/294); supersimple (LLM not disclosed · supersimple semantic model) answers 69.7% (202/290). Neither vendor paid to be measured, chose the questions, or saw results early.
| Configuration | Accuracy | Completeness | Restraint | $/question |
|---|---|---|---|---|
| cortex-agents claude-opus-4-8 · snowflake semantic view | 83.7% | 99.7% | 100.0% | $0.59 |
| supersimple LLM not disclosed · supersimple semantic model | 69.7% | 98.3% | 100.0% | — |
By question tier
Accuracy over answered questions; "declined" is the share of that tier's answerable questions the configuration chose not to answer.
| Tier | cortex-agents | supersimple |
|---|---|---|
| What happened? | 98.9% declined 1.1% | 97.9% |
| Why did it happen? | 81.0% | 72.7% declined 1.0% |
| What should I do? | 72.0% | 38.5% declined 4.0% |
How to read this
- This page is generated from the same data as the live leaderboard — figures are pulled from the season's run artifacts, never written by hand, and every configuration's receipt (run id, seed, reproduce command) is on its detail page: cortex-agents, supersimple.
- Accuracy counts only answered questions. Declining answerable questions costs completeness; answering seeded unanswerable trap questions costs restraint. Read all three together.
- Each configuration runs its vendor-default LLM and its own semantic-layer format carrying equivalent content — this compares products as shipped, not models under lab control. Part of any gap belongs to the underlying model.
- k=1 (each question asked once) and the judge's agreement with human graders is not yet measured for claude-opus-4-8-r4 — see how we benchmark analytics agents for everything this measurement does and does not support.
- One season, one simulated e-commerce warehouse. Your data, semantic layer and question mix are different.