What this is actually for
The expensive mistakes are never the ones somebody flagged. They're the ones everyone in the room agreed on.
What we built for ourselves
We put plans, vendor proposals and designs in front of frontier models from different providers, blind to each other, scored one to a thousand.
Reviewers never see each other's work. The standard multi-agent pattern where they do produces convergence, not review.
A local model with our business context sits on the panel alongside them.
We run it on our own work — including having the local model critique the cloud system and sending that critique to the owner unedited.
What you get
- —Independent scored reviews from multiple providers plus a context-aware local model
- —The full spread reported, with outlier reasoning preserved verbatim
- —A written verdict with the dissent left in, not summarised away
Hear how it got built
We recorded the build of each of these, including the parts that didn't work.
- Episode 8 — Score It One To A Thousandrecording
- Episode 10 — Racing Fuel In A Lawn Mowerrecording
Episodes publish as they're cut. Nothing here is a teaser for something that doesn't exist — the systems are already running.
What this doesn't do
It catches consensus errors. It will not make a bad plan good. And the moment anyone starts reporting a tidy average instead of the spread, the product is ruined — the disagreement is the signal.
What it runs on
Built on NVIDIA hardware and software we own and operate, not a reseller relationship.
Where it's going next: Self-serve intake — submit a document, get a scored verdict back.
A member of the NVIDIA Inception program
Fifteen minutes, no deck.
Bring one specific problem in one sentence. We'll tell you whether it's worth building, roughly what it takes, and sometimes that the answer is don't build it.
Start hereTwo steps before we talk — it takes about five minutes.