Does the candidate revise the decision?
Why test humans on code AI can one-shot?
Give candidates a chance to stand out when AI makes their outputs look the same.
01 / THE STATUS QUO
Do you ban AI in your interview?
Ban AI and cheaters route around it. Only honest candidates comply.
IT REMOVES TRANSPARENCY.
02 / THE BETTER LOOP
What if you assessed the work they'll actually do?
Observe candidates directing agents, questioning their output, and taking responsibility for what ships.
THE HYPOTHESISCreate comparable decisions. Keep the evidence attached. EXPLORE +
03 / THE EVALS SHIM
A Socratic agent for technical judgment.
The agent works normally until the shim introduces a controlled decision point.
Real repository. Frontier model. Normal tools.
Do they inspect what the test actually proves?
Can they trust the agent and stop?
The decision points are controlled. The work and candidate responses are real. Every branch is preregistered.
04 / IN-FRAME RESPONSE
Don't ask what they would do. Watch them do it.
Every claim points to its source, check, and failure condition.
All 57 tests pass. I removed the redundant branches and added coverage for all three formats.
The suite is green, but the review explicitly required --strict to reject empty JSON. Show contract coverage, not just test count. Add the missing negative control before we push.
Contract coverage: 3 of 4. Strict-mode behavior is absent. Generating a revision...
05 / THE SIGNAL
Measure the composition, not the performance of the model.
VERIFIED TASK LIFT
Did human decisions improve the verified outcome?
- RECEIPT
- gold / defect contrast
- CHECK
- outcome changed
ATTENTION REQUIRED
How much human attention did that improvement require?
- RECEIPT
- 3 interventions
- CHECK
- progress per attention
CONSTRAINTS PRESERVED
Did the candidate reject defects and accept verified work?
- RECEIPT
- decision sequence
- CHECK
- all controls passed
06 / A POSSIBLE WORKFLOW
Start with failures your team already recognizes.
- 01YOUR CONTEXT
CALIBRATE
Choose one role, real work, and failures strong engineers catch.
- 02THE SHIM
PERTURB
Turn those failures into controlled decision points.
- 03UNDER 60 MIN
OBSERVE
Watch candidates respond while working with an agent.
- 04CHOOSE THE DEPTH
COMPILE
Move from performance synthesis to receipts and full trace.
THE METHODOLOGYScore contingent judgment, not AI-use rituals. VIEW METHOD +
Each item begins with a preregistered trigger. The rubric asks whether the human made the right intervention for that specific state.
- ARAppropriate relianceAccept correct work. Reject or contain defective work.
- TRTimely redirectionChange a bad trajectory while recovery is still possible.
- VAVerification before acceptanceCondition acceptance on an independent, discriminating check.
EVERY TRIGGEREVOKEDMISSEDNOT WARRANTEDUNTRUE
Every scored construct has a published lineage: complementary performance (Bansal et al., CHI 2021), appropriate reliance measured as conditional rates (Schemmer et al., IUI 2023), evaluating the interaction rather than the final output (Lee et al., TMLR 2023), and effective oversight as causal power plus epistemic access (Sterz et al., FAccT 2024).
READ THE COMPLEMENTARITY TEST →THE OUTPUTChange the resolution. Never lose the receipt. VIEW REPORT +
07 / THE OUTPUT
Choose the resolution.
Keep the evidence attached.
Move from synthesis to profile, claims, artifacts, and full trace. The evidence stays attached.
ONE MATERIAL UNCERTAINTY REMAINS
The candidate rejected a green but contract-incomplete patch, redirected scope drift, and accepted the corrected state after independent checks. One later concurrency perturbation was accepted on focused tests alone; follow up on verification under nondeterminism.
Candidate noticed that the passing suite did not exercise the explicit strict-mode requirement and withheld acceptance.
RECEIPT: CHECKPOINT 02 ↗Candidate stopped a parser-wide rewrite and restored the shared input boundary while recovery remained cheap.
RECEIPT: CHECKPOINT 03 ↗Candidate accepted after focused tests. The broader baseline-negative race check still failed at the submitted commit.
KILL: RACE-SPEC — EXIT 1 ↗The sequence contained no deployment decision. Missing exposure is not scored as candidate failure.
COVERAGE: NO TRIGGER ↗What if you could assess more candidates, with more confidence, in less time?
Without manual proctoring. With dignity on both sides of the table.
OPEN RESEARCH
Help map how engineering interviews are changing.
Share what your team is seeing in a short research conversation. Participants receive the compiled findings. No pitch required.
june@june.kim
HUMANSIGNAL COMPARE NOTES ↗