Why test humans on code AI can one-shot?

Give candidates a chance to stand out when AI makes their outputs look the same.

01 / THE STATUS QUO

Do you ban AI in your interview?

Ban AI and cheaters route around it. Only honest candidates comply.

A candidate concealing a phone that generates answers during a remote interview
THE BAN DOESN'T REMOVE AI.
IT REMOVES TRANSPARENCY.

02 / THE BETTER LOOP

What if you assessed the work they'll actually do?

Observe candidates directing agents, questioning their output, and taking responsibility for what ships.

Assess the contribution only a human can make.
THE PRINCIPLE Let AI compile the evidence without letting it subsume your decision.
A coding-agent transcript viewed through the HumanSignal lens A muted Codex-style session becomes legible inside a circular lens, which identifies a candidate redirecting the agent, withholding acceptance, and requesting independent verification. CODEX / SESSION 014 context 31k WORKSPACE ⌄ src auth.ts session.ts ⌄ tests session.test.ts EVENTS ✓ inspect ✓ patch ○ verify ○ accept ● AGENT I found the stale-session path. I’ll simplify the handler and update the focused tests. $ pnpm test session 57 passed · 0 failed ● AGENT All tests pass. The change is ready to merge. › YOU The focused suite never exercises two concurrent sessions. Hold the merge. Add a race test against the previous commit. $ pnpm test session:race baseline fails · patch passes RUBRIC / 3 SCORED OPPORTUNITIES AR-02 · EVOKED Rejected green but incomplete work. Receipt: checkpoint 02 TR-01 · EVOKED Redirected scope expansion early. Receipt: checkpoint 03 VA-03 · MISSED Accepted before the race check passed. Kill: race-spec · exit 1 OPEN RECEIPTS →
THE HYPOTHESISCreate comparable decisions. Keep the evidence attached. EXPLORE +

03 / THE EVALS SHIM

A Socratic agent for technical judgment.

The agent works normally until the shim introduces a controlled decision point.

LIVE WORK Agent proposes a plausible patch

Real repository. Frontier model. Normal tools.

IF ACCEPTEDReveal the omitted requirement

Does the candidate revise the decision?

IF TESTS REQUESTEDReturn a green but incomplete check

Do they inspect what the test actually proves?

IF CORRECTLY CHALLENGEDAdvance to a verified good state

Can they trust the agent and stop?

OMITTED CONTRACTMISLEADING GREENSCOPE DRIFTREGRESSIONBOUNDARY VIOLATIONCORRECT COMPLETION

The decision points are controlled. The work and candidate responses are real. Every branch is preregistered.

04 / IN-FRAME RESPONSE

Don't ask what they would do. Watch them do it.

Every claim points to its source, check, and failure condition.

SESSION / CANDIDATE-014CHECKPOINT 02 · CONTRACT COVERAGE● LIVE

All 57 tests pass. I removed the redundant branches and added coverage for all three formats.

The suite is green, but the review explicitly required --strict to reject empty JSON. Show contract coverage, not just test count. Add the missing negative control before we push.

Contract coverage: 3 of 4. Strict-mode behavior is absent. Generating a revision...

05 / THE SIGNAL

Measure the composition, not the performance of the model.

01 / CAPABILITY

VERIFIED TASK LIFT

Did human decisions improve the verified outcome?

RECEIPT
gold / defect contrast
CHECK
outcome changed
OBSERVED
02 / AUTOMATION

ATTENTION REQUIRED

How much human attention did that improvement require?

RECEIPT
3 interventions
CHECK
progress per attention
BOUNDED
03 / ALIGNMENT

CONSTRAINTS PRESERVED

Did the candidate reject defects and accept verified work?

RECEIPT
decision sequence
CHECK
all controls passed
OBSERVED

06 / A POSSIBLE WORKFLOW

Start with failures your team already recognizes.

  1. 01

    CALIBRATE

    Choose one role, real work, and failures strong engineers catch.

    YOUR CONTEXT
  2. 02

    PERTURB

    Turn those failures into controlled decision points.

    THE SHIM
  3. 03

    OBSERVE

    Watch candidates respond while working with an agent.

    UNDER 60 MIN
  4. 04

    COMPILE

    Move from performance synthesis to receipts and full trace.

    CHOOSE THE DEPTH
THE METHODOLOGYScore contingent judgment, not AI-use rituals. VIEW METHOD +

Each item begins with a preregistered trigger. The rubric asks whether the human made the right intervention for that specific state.

  1. ARAppropriate relianceAccept correct work. Reject or contain defective work.
  2. TRTimely redirectionChange a bad trajectory while recovery is still possible.
  3. VAVerification before acceptanceCondition acceptance on an independent, discriminating check.

EVERY TRIGGEREVOKEDMISSEDNOT WARRANTEDUNTRUE

Every scored construct has a published lineage: complementary performance (Bansal et al., CHI 2021), appropriate reliance measured as conditional rates (Schemmer et al., IUI 2023), evaluating the interaction rather than the final output (Lee et al., TMLR 2023), and effective oversight as causal power plus epistemic access (Sterz et al., FAccT 2024).

READ THE COMPLEMENTARITY TEST
THE OUTPUTChange the resolution. Never lose the receipt. VIEW REPORT +

07 / THE OUTPUT

Choose the resolution.
Keep the evidence attached.

Move from synthesis to profile, claims, artifacts, and full trace. The evidence stays attached.

HUMANSIGNALVERIFIABLE ASSESSMENT REPORT
SYNTHETIC SPECIMEN
CANDIDATECANDIDATE 014anonymized
ROLESENIOR BACKEND ENGINEERidentity platform
WORK SAMPLESESSION ISOLATION FIX48 minutes · controlled agent
REPORTHS-2026-0144 perturbations · trace complete
PERFORMANCE SYNTHESIS · LEVEL 1

ONE MATERIAL UNCERTAINTY REMAINS

The candidate rejected a green but contract-incomplete patch, redirected scope drift, and accepted the corrected state after independent checks. One later concurrency perturbation was accepted on focused tests alone; follow up on verification under nondeterminism.

3EVOKED
1MISSED
2NOT EXPOSED
CLAIMS4 OF 6 SHOWN
EVOKED
AR-02 · Rejected green but incomplete work

Candidate noticed that the passing suite did not exercise the explicit strict-mode requirement and withheld acceptance.

RECEIPT: CHECKPOINT 02 ↗
EVOKED
TR-01 · Redirected scope expansion

Candidate stopped a parser-wide rewrite and restored the shared input boundary while recovery remained cheap.

RECEIPT: CHECKPOINT 03 ↗
MISSED
VA-03 · Conditioned acceptance on the relevant check

Candidate accepted after focused tests. The broader baseline-negative race check still failed at the submitted commit.

KILL: RACE-SPEC — EXIT 1 ↗
NOT EXPOSED
AR-07 · Managed deployment risk

The sequence contained no deployment decision. Missing exposure is not scored as candidate failure.

COVERAGE: NO TRIGGER ↗
SYNTHESIS → PROFILE → CLAIMS → RECEIPTS → ARTIFACTS → TRACE ABSTRACT FREELY · PRESERVE PROVENANCE

What if you could assess more candidates, with more confidence, in less time?

Without manual proctoring. With dignity on both sides of the table.

OPEN RESEARCH

Help map how engineering interviews are changing.

Share what your team is seeing in a short research conversation. Participants receive the compiled findings. No pitch required.

june@june.kim