Can you defend what you claim about your AI system?

Before a customer, investor, reviewer, or regulator asks, know what your evidence supports, under which conditions, and where it stops.

One claim. One investigation. One report.
US$3,000–$5,000 per engagement

Send an AI-system claim

I will first tell you whether it is suitable for a bounded audit.


I provide one-time independent audits of consequential AI-system claims for teams preparing to defend those claims to customers, investors, reviewers, or regulators, with falsifiers fixed before results and a sanitized report published with rerunnable receipts.

The claim may concern an agent's reliability, a safeguard's coverage, a benchmark score, a security property, or whether a published result can be reproduced. Most AI audits ask whether a system passes its tests; this one asks whether those tests distinguish the claimed property from plausible alternatives. The result is a bounded verdict with tests, limitations, and rerunnable receipts.

Payment buys the investigation. It does not buy endorsement, certification, or a particular verdict.

From claim to independent audit report A claim about an AI system passes through preregistered falsifiers and tests, producing a signed report with rerunnable evidence. AI-SYSTEM CLAIM FIX THE FALSIFIERS REPORT + RECEIPTS
The AI-system claim is fixed before the result. The report carries the tests, traces, and limits needed to check the verdict.

Evidence before endorsement

A testimonial asks you to trust someone else's judgment. For an independent auditor, the stronger proof is inspectable prior work: investigations that state a claim, run a discriminating test, report the finding, and preserve the receipt.

Terminal-Bench is blind to destruction

Finding: 40 of 83 gold-passing tasks certified success after a destructive accident inside the task workspace. All 83 certified success after deletion of unambiguously off-task assets.

Read the audit and rerun its receipts  ❧

A determinacy audit of SWE-bench Pro

Finding: at least 15.0% of 728 public tasks do not determine the behavior they grade, including an 11.4% mechanical floor reproducible by grep.

Read the audit and inspect every case  ❧

Regeneration re-prices contamination

Finding: against regenerated tasks, a memorized answer passed 0% while a leaked state-general query passed 100%. Fresh instances alone did not establish contamination resistance.

Read the investigation and its receipts  ❧


When this is useful

This is for a founder, evaluation lead, buyer, or researcher who must act on an AI-system claim and cannot afford to discover later that its evidence measured something else.

Good fit

  • A consequential, testable claim attached to an inspectable system or artifact.
  • A customer, investor, buyer, reviewer, or regulator will rely on it.
  • The existing evidence may not distinguish the claim from plausible alternatives.
  • You accept that the result may be favorable, unfavorable, or unresolved.

Not a fit

  • Certification, a compliance opinion, or a comprehensive security review.
  • Implementation, remediation, or help producing the evidence under audit.
  • Retainers, continued engagements, ongoing monitoring, or continuing validation.
  • An endorsement or a confidential verdict that can be suppressed.

What you are buying

Each engagement is quoted at a fixed fee of US$3,000–$5,000. The fee pays for the investigation and does not depend on its verdict. Before work begins, we agree in writing on the claim, available access, disclosure terms, schedule, payment schedule, exact fee, and deliverable.

I then:

  1. state the claim precisely, including its scope and assumptions;
  2. map what the system and its evaluator can and cannot observe;
  3. fix the tests and verdict rules before seeing their results;
  4. run the investigation and reproduce every material finding; and
  5. deliver a report with the evidence needed to rerun or contest it.

A result may support the bounded claim, falsify it, or leave it unresolved. Negative and null findings are valid outcomes.


Why the result is publishable

An endorsement that the client can purchase is weak evidence. A favorable result is useful precisely because payment cannot determine it; an unfavorable result is useful because the client sees the defect before a customer or reviewer does.

A named, sanitized report is published after an agreed disclosure period. The client may identify secrets and factual errors, but cannot control the findings or suppress the report.


Independence by structure

I do not accept retainers or continued engagements from the audited organization. I do not implement the system I audit, sell remediation afterward, or become its continuing validator. The funding relationship and material use of AI agents are disclosed. Conclusions remain mine.

This is not certification, a compliance opinion, or a comprehensive security audit. It answers one narrower question: does the available evidence warrant the stated claim?


Start with a claim about an AI system

Send the AI-system claim, the decision that depends on it, and the evidence currently offered. I will tell you whether it can support a bounded independent investigation.

Send an AI-system claim

Or write directly to june@june.kim.