Most companies are about to make a six-figure AI decision with no way to check the supplier's numbers. We build a benchmark from your own real examples, measure what the supplier's system actually does on it, and give you a report you can put in front of a board.
The fee is fixed and we are paid the same whatever we find.
An AI supplier tells you their system is 94% accurate. That number was produced by them, on data they chose, against a comparison they chose. It may be completely honest and still tell you nothing about what will happen on your documents, your customers and your edge cases.
The question almost nobody asks: compared to what? A system that scores 94% where a far cheaper approach scores 91% is a very different purchase from one where the cheap approach scores 60%. The comparison is where the value is, and it is the part a supplier has the least incentive to run.
We have spent this year publishing what goes wrong with measurement, using our own mistakes as the material. Four of our published records are corrections to our own earlier results, including one where a confident, statistically significant finding turned out to be an instrument pointed slightly off, and one where our own technique produced a real saving and still lost money once we priced the cost of the measurement itself. That is the expertise being sold here.
A sample of your own real examples, with the correct answers established independently, and held out so that nothing in it was used to build or tune the thing being tested. This is the part that makes the result about you rather than about the supplier's demo.
We run the supplier's system, and we run a sensible cheaper alternative next to it, because a number with nothing to compare it against is not evidence. Results come with the uncertainty attached, so you know which differences are real and which are noise.
Plain English first: what it does well, where it fails, what it would cost you when it fails, and whether the difference justifies the price. Then the full method, so your technical people or theirs can check every step.
We build custom AI models. So an evaluation of somebody else's that concludes "buy ours instead" is exactly what you should be worried about.
Three things make that not how this works, and they are contractual, not aspirational:
Fixed prices, agreed in advance. For context, these are engagements meant to protect a decision that is usually many times larger.
$7,500 · fixed
A desk review of what the supplier has already published or told you. No access to their system needed, and they need not be involved or informed.
The right first step when you are early, or when you need something small enough to approve without a procurement cycle.
$28,000 · fixed
The full thing. A benchmark built from your own data, the supplier's system measured on it against a real baseline, and a report written for the people signing the cheque.
You keep the benchmark. It is reusable for the next supplier, the next version, and for holding whoever you choose to their numbers later.
$4,500 · per day
For acquisitions, procurement disputes, and situations where somebody needs a defensible technical opinion on whether a supplier's claims hold up.
Priced by the day because the scope of these is not knowable in advance, and pretending otherwise would mean padding the estimate.
Prices are for a single system on a single use case. Multiple suppliers or multiple use cases are scoped and quoted before you commit. If an evaluation ends up smaller than we scoped, you are billed the smaller amount.
There are 46 research records on this site, and a good share of them are negative results or corrections to our own earlier work. That is unusual to the point of being nearly unique among people selling AI, and it is the whole credential for this service.
Ask any other supplier to show you the experiments that went against them. The answer, and how long it takes them to produce one, is itself the most useful evaluation you will run.
A procurement decision rarely has a quarter to spare, and an evaluation that arrives after the contract is signed is worth nothing. A Claim Check takes about a week and a full evaluation about three.
That is possible because building this kind of measurement is what the team does continuously as published research, so the method and the tooling already exist. We are not inventing an approach on your engagement and billing you for the learning.
For a Claim Check, no, and they need not know. For a full evaluation we need enough access to run their system on your held-out examples, which is usually arranged through you as part of a trial or pilot. A supplier who refuses reasonable access for an independent test has told you something important, and we will put that in the report.
Common, and not a blocker. Part of the engagement is establishing correct answers for a sample, which is work you need doing anyway and which you keep afterwards. If your data is genuinely not ready, the free Data Readiness Scorecard will tell you that in an afternoon, at no cost.
Yes, and it is a common reason to call. The same method works on a system already in production, and the question becomes whether it is earning its cost and where it is quietly failing. That report is usually what a renewal conversation needs.
Then we will say so, and we will also say plainly that we sell that, so you can weigh the recommendation accordingly. You are free to take the benchmark we built and hand it to anybody. That is the point of you owning it.
If you are not ready to commission anything, the free Data Readiness Scorecard runs on your own machine, never sends us your data, and tells you whether you have what an evaluation would need. And the research record is free to read and is the best evidence of how we work.