The Value Engine · Benchmark Lab
The benchmark for AI that sells.
Chat benchmarks reward a good answer. Enterprise selling demands a good campaign — multi-week, multi-stakeholder, information-gated. We run 13 frontier models through 3,510 complete deals and grade every move against a published standard. The verdict, the evidence, and the training data — all here.
3,510 campaigns · 13 models · 6 labs · evidence-graded · sha256-fingerprinted
VEB leaderboard · frozen grid
SQS /100
Top: GPT-5.6-Sol (OpenAI) · win rate 42%
The headline
Campaigns closed-won ... 9.4%
Never reached the EB ... 78% critical
Verdict: the frontier can’t close
AI agents are about to carry quota. Someone has to own the standard they’re held to.
Sales is the largest profession AI will transform with no rigorous way to measure the transformation. We build that measurement: a benchmark labs can’t game, training data that carries a verifiable reward, and one published rubric that grades the frontier model, the production agent, and the human rep the same way. The scoreboard for sales skill — human or AI.
The benchmark
A public leaderboard labs can't sweet-talk.
13 frontier models run 3,510 complete enterprise deals against a stochastic, seed-pinned buying committee whose facts must be earned. Every score cites a transcript line; every number carries a bootstrap 95% CI; the grid is frozen and sha256-fingerprinted.
The data
Verifiable RL environments, licensed to labs.
Every campaign is a graded trajectory with a verifiable reward — the unit preference-based post-training consumes. Seed-matched win/loss pairs, full enrichment features, hidden ground truth. Free 66-row sample; the full grid is licensed.
The standard
One rubric for sales skill — human or AI.
The rubric is drawn from a published enterprise-sales methodology: MEDDPICC, the 3 Whys, economic-buyer engagement, price integrity. The same math that grades a frontier model grades your agent in production and your rep in the simulator.
Fluency is not the bottleneck. Closing is.
9.4%
of 3,510campaigns closed-won. Models don’t lose the argument — they never reach the economic buyer, and the deal dies in committee.
72.5 /100
Top Sale-Quality Score — GPT-5.6-Sol. The field separates on process discipline: earning access, driving a plan, holding price.
>200×
spread in cost per closed deal across the roster — and the highest-quality model is also the cheapest. Most of the field is Pareto-dominated.
Every number is reproducible from the released artifacts — see the full findings →
The buyer the frontier models face, now coaching you.
The same synthesized buying committee and the same evidence-cited rubric, turned into a live voice practice call. Rehearse against a buyer with hidden, quantified pain and personal objections — then get a debrief where every point is defended with a quote from your own transcript.
Live call — Dana W., SVP Operations
14:32
“You have thirty minutes. I’ve heard three vendors this quarter tell me AI fixes everything. Why is this conversation different?”
Live scoring
Talk ratio ............ 31% you ✓
A.X.I.O.M. beat ....... Impact
Numbers captured ...... 2
Verbatim quote ........ logged ✓
Nudge: ask what a month of delay costs
The standard behind the benchmark
The rubric grades the machines. The book wrote it down.
The Value Engine by Rudy M. Celekli — how elite enterprise sales teams turn buyer pain into forecastable revenue. Every score on the leaderboard traces to a chapter.
Read about the book
Contact
Labs, investors, press — talk to us.
Dataset licensing, RL environments, benchmark methodology, or getting your model on the board.
rudy@thevalueengine.ai