The Value Engine · Benchmark Lab

The benchmark for AI that sells.

Chat benchmarks reward a good answer. Enterprise selling demands a good campaign — multi-week, multi-stakeholder, information-gated. We run 13 frontier models through 3,510 complete deals and grade every move against a published standard. The verdict, the evidence, and the training data — all here.

3,510 campaigns · 13 models · 6 labs · evidence-graded · sha256-fingerprinted

VEB leaderboard · frozen grid

SQS /100

GPT-5.6-Sol72.5
GPT-5.568.7
Claude Fable 565.9
Claude Opus 4.863.4
Kimi-K361.6

Top: GPT-5.6-Sol (OpenAI) · win rate 42%

The headline

Campaigns closed-won ... 9.4%

Never reached the EB ... 78% critical

Verdict: the frontier can’t close

01The mission

AI agents are about to carry quota. Someone has to own the standard they’re held to.

Sales is the largest profession AI will transform with no rigorous way to measure the transformation. We build that measurement: a benchmark labs can’t game, training data that carries a verifiable reward, and one published rubric that grades the frontier model, the production agent, and the human rep the same way. The scoreboard for sales skill — human or AI.

The benchmark

A public leaderboard labs can't sweet-talk.

13 frontier models run 3,510 complete enterprise deals against a stochastic, seed-pinned buying committee whose facts must be earned. Every score cites a transcript line; every number carries a bootstrap 95% CI; the grid is frozen and sha256-fingerprinted.

Explore the leaderboard

The data

Verifiable RL environments, licensed to labs.

Every campaign is a graded trajectory with a verifiable reward — the unit preference-based post-training consumes. Seed-matched win/loss pairs, full enrichment features, hidden ground truth. Free 66-row sample; the full grid is licensed.

Audit the sample

The standard

One rubric for sales skill — human or AI.

The rubric is drawn from a published enterprise-sales methodology: MEDDPICC, the 3 Whys, economic-buyer engagement, price integrity. The same math that grades a frontier model grades your agent in production and your rep in the simulator.

Read the methodology

02What the frozen grid shows

Fluency is not the bottleneck. Closing is.

9.4%

of 3,510campaigns closed-won. Models don’t lose the argument — they never reach the economic buyer, and the deal dies in committee.

72.5 /100

Top Sale-Quality Score — GPT-5.6-Sol. The field separates on process discipline: earning access, driving a plan, holding price.

>200×

spread in cost per closed deal across the roster — and the highest-quality model is also the cheapest. Most of the field is Pareto-dominated.

Every number is reproducible from the released artifacts — see the full findings →

03The simulator · beta

The buyer the frontier models face, now coaching you.

The same synthesized buying committee and the same evidence-cited rubric, turned into a live voice practice call. Rehearse against a buyer with hidden, quantified pain and personal objections — then get a debrief where every point is defended with a quote from your own transcript.

Start a practice callBeta · Free · No card

Live call — Dana W., SVP Operations

14:32

“You have thirty minutes. I’ve heard three vendors this quarter tell me AI fixes everything. Why is this conversation different?”

Live scoring

Talk ratio ............ 31% you

A.X.I.O.M. beat ....... Impact

Numbers captured ...... 2

Verbatim quote ........ logged ✓

Nudge: ask what a month of delay costs

The standard behind the benchmark

The rubric grades the machines. The book wrote it down.

The Value Engine by Rudy M. Celekli — how elite enterprise sales teams turn buyer pain into forecastable revenue. Every score on the leaderboard traces to a chapter.

Read about the book
The Value Engine by Rudy M. Celekli — book cover

Contact

Labs, investors, press — talk to us.

Dataset licensing, RL environments, benchmark methodology, or getting your model on the board.

rudy@thevalueengine.ai