Can you defend what you claim about your AI system?
Before a customer, investor, reviewer, or regulator asks, know what your evidence supports, under which conditions, and where it stops.
One claim. One investigation. One report.
US$3,000–$5,000 per engagement
I will first tell you whether it is suitable for a bounded audit.
I provide one-time independent audits of consequential AI-system claims for teams preparing to defend those claims to customers, investors, reviewers, or regulators, with falsifiers fixed before results and a sanitized report published with rerunnable receipts.
The claim may concern an agent's reliability, a safeguard's coverage, a benchmark score, a security property, or whether a published result can be reproduced. Most AI audits ask whether a system passes its tests; this one asks whether those tests distinguish the claimed property from plausible alternatives. The result is a bounded verdict with tests, limitations, and rerunnable receipts.
Payment buys the investigation. It does not buy endorsement, certification, or a particular verdict.
Evidence before endorsement
A testimonial asks you to trust someone else's judgment. For an independent auditor, the stronger proof is inspectable prior work: investigations that state a claim, run a discriminating test, report the finding, and preserve the receipt.
Terminal-Bench is blind to destruction
Finding: 40 of 83 gold-passing tasks certified success after a destructive accident inside the task workspace. All 83 certified success after deletion of unambiguously off-task assets.
A determinacy audit of SWE-bench Pro
Finding: at least 15.0% of 728 public tasks do not determine the behavior they grade, including an 11.4% mechanical floor reproducible by grep.
Regeneration re-prices contamination
Finding: against regenerated tasks, a memorized answer passed 0% while a leaked state-general query passed 100%. Fresh instances alone did not establish contamination resistance.
When this is useful
This is for a founder, evaluation lead, buyer, or researcher who must act on an AI-system claim and cannot afford to discover later that its evidence measured something else.
Good fit
- A consequential, testable claim attached to an inspectable system or artifact.
- A customer, investor, buyer, reviewer, or regulator will rely on it.
- The existing evidence may not distinguish the claim from plausible alternatives.
- You accept that the result may be favorable, unfavorable, or unresolved.
Not a fit
- Certification, a compliance opinion, or a comprehensive security review.
- Implementation, remediation, or help producing the evidence under audit.
- Retainers, continued engagements, ongoing monitoring, or continuing validation.
- An endorsement or a confidential verdict that can be suppressed.
What you are buying
Each engagement is quoted at a fixed fee of US$3,000–$5,000. The fee pays for the investigation and does not depend on its verdict. Before work begins, we agree in writing on the claim, available access, disclosure terms, schedule, payment schedule, exact fee, and deliverable.
I then:
- state the claim precisely, including its scope and assumptions;
- map what the system and its evaluator can and cannot observe;
- fix the tests and verdict rules before seeing their results;
- run the investigation and reproduce every material finding; and
- deliver a report with the evidence needed to rerun or contest it.
A result may support the bounded claim, falsify it, or leave it unresolved. Negative and null findings are valid outcomes.
Why the result is publishable
An endorsement that the client can purchase is weak evidence. A favorable result is useful precisely because payment cannot determine it; an unfavorable result is useful because the client sees the defect before a customer or reviewer does.
A named, sanitized report is published after an agreed disclosure period. The client may identify secrets and factual errors, but cannot control the findings or suppress the report.
Independence by structure
I do not accept retainers or continued engagements from the audited organization. I do not implement the system I audit, sell remediation afterward, or become its continuing validator. The funding relationship and material use of AI agents are disclosed. Conclusions remain mine.
This is not certification, a compliance opinion, or a comprehensive security audit. It answers one narrower question: does the available evidence warrant the stated claim?
Start with a claim about an AI system
Send the AI-system claim, the decision that depends on it, and the evidence currently offered. I will tell you whether it can support a bounded independent investigation.
Or write directly to june@june.kim.