Hard questions, answered plainly
Including the one about whether this is a certification. It isn't, and we say so in the same words every time.
Those tools are run by the team that built the AI, on data they chose. Useful for the builder, but self-graded. We are independent: a separate team, our own tests, private held-out cases, and the full method shown.
Every score ships with its working: what was counted, out of how many, on which held-out set, and a breakdown of the failures. If we cannot show how a number was made, we do not publish it.
No. We provide an independent evaluation report and scorecard — an evidence-based opinion at a point in time. It is not a certificate, a guarantee, or an accreditation, and we do not present it as one.
Both. For agents we trace each step and tool call, not just the final answer, so the report reflects how the agent actually behaves across a whole task.
A clear description of what the AI is meant to do, access to run it, the claims being made about it, and examples of the real tasks it should handle. [[finalise the exact intake list]]
[[timeframe]], depending on scope.
You get a ranked list of the specific issues and how to reproduce each one, so the builder can fix them. Then we re-evaluate.
[[state your data-handling and confidentiality policy]]