The Complementarity Test

How to construct an hour of work that measures the human in a human-agent system.

Technical hiring evaluates the wrong unit. The deployed system is a human with an agent, but the interview confiscates the agent and measures the human alone. Give the tool back and the opposite problem appears: on an ordinary interview question, the agent may do all the work. A candidate who can paste the task and submit the answer is indistinguishable from one whose judgment changed the outcome.

The task has to sit between those failures. The agent cannot reliably solve it alone under the interview budget. The human cannot finish it alone in the time. Together they can. Call this a complementarity task.

This is not a proposal to infer a career from an hour. Hiring standards are generally private, their downstream outcomes are available only to employers, and a rejected candidate observes one run of a hidden instrument. Reliability, false negatives, and criterion validity cannot be audited from outside at sample size one. Even established selection methods have needed substantial revision when their validity estimates were re-examined (Sackett et al.). The narrower opportunity is to say what a defensible pointwise observation would look like: did this human add verified value to this agent, on these tasks, under these conditions?

The complaint is a measurement contradiction

The argument is already happening in hiring forums, without a common measurement language. These accounts are anecdotes rather than prevalence estimates, but the same complaints recur from opposing positions.

Employers say their take-homes can now be one-shot by an agent, leaving them unable to infer the candidate’s ability from the submission. One experienced-developer discussion asks what replaces a take-home after repeated attempts to change it; its most popular answer is to review code with the applicant and ask them to find defects. Another hiring manager says algorithm questions are readily outsourced, take-homes arrive as model output, live calls consume interviewer time, and the resulting 30–60 minute judgment still feels superficial.

Candidates see a conflicting rule. One discussion summarizes it as AI being forbidden in the interview and demanded on the job; participants compare the prohibition to removing the IDE from a developer’s ordinary workflow (discussion). A LinkedIn exchange contains both poles in one thread: the original post treats live AI use as dependency, while a reply reports a candidate rejected for not using AI because the employer considered them too slow (thread). Another candidate describes preparing for unaided coding only to encounter an interviewer who cared instead about whether they could validate agent output and catch its bugs.

Permission alone does not resolve the attribution problem. In a recent account of an AI-enabled interview, the claimed differentiators were clarifying the specification, structuring the project, and reviewing the generated code; commenters immediately objected that these practices themselves can be packaged into agent instructions and that the session can feel like an evaluation of the AI rather than the candidate. At the other extreme, suspicion of undisclosed assistance produces false-positive anxiety: candidates report polished work being treated as evidence of cheating and no inspectable basis for the eventual rejection (discussion).

The small academic survey available confirms the policy lag, if not the magnitude of every complaint. Among 32 industry professionals, 65.63% said their organizations had not adjusted hiring for code-generation tools; more than half of respondents with an applicable answer thought the tools made assessment harder; 53.13% never or rarely asked candidates about their AI-tool experience, while more than half expressed at least a moderate preference for candidates who demonstrated it (Chen et al.). Employers value a capability their instruments usually neither expose nor measure.

The failure is symmetrical:

ban AI    -> measure an increasingly artificial unaided workflow
allow AI  -> lose attribution on tasks the agent can solve alone
watch live -> add stress and cost without a declared human-contribution rubric
detect AI -> classify tool use, not whether the use improved the outcome

The missing question is not whether AI was used. It is what the human contributed when it was.

What is already known

The recommendation does not need a novelty claim. Several adjacent fields have already established most of it.

CentaurEval dynamically instantiates coding tasks from 45 “Collaboration-Necessary” templates: tasks intended to be intractable for either standalone humans or standalone models but solvable together. Across 45 participants and five models, its reported pass rates were 18.89% for unaided humans, 0.67% for standalone models, and 31.11% for human-AI collaboration. The central construction is the right one: test the composition against both of its components.

The broader evidence is a warning against assuming the result. A meta-analysis of more than 300 effect sizes found that human-AI combinations do not generally beat the better component. Complementarity depends on the relative baseline abilities, the task, and the division of labor. A newer multi-domain study found only 0.4 percentage points of improvement over AI alone from baseline hybridization. Its bottleneck was not merely human accuracy, but locating the cases where human judgment matters and enabling humans to catch AI mistakes.

The behavioral construct with the deepest experimental base is appropriate reliance: accept correct AI advice and override incorrect advice. The term comes from the older automation-trust literature (Lee and See); newer experiments find that people often over-rely on wrong recommendations, and that forcing independent engagement can reduce the error (Bucinca et al.; Schemmer et al.). This is a better foundation than a checklist of plausible engineering virtues because it scores the decision against the truth of the recommendation.

Two coding studies narrow the mechanism. CentaurEval’s logs associate successful collaboration with humans rejecting a misleading initial frame and redirecting the agent; failures include accepting that frame or intervening only after it consumed most of the budget. That evidence is qualitative, not a causal isolation of a trait. A study of 9,427 agentic pull requests finds that core contributors more consistently require passing CI before merge than peripheral contributors. That evidence is observational, not proof that running CI identifies a better engineer. A newer coding-sabotage experiment makes the cost of failed oversight concrete: 94% of more than 100 participants missed the sabotage, and 56% still accepted malicious code when a monitor warned them. Together these support three narrow things worth observing: appropriate reliance, timely redirection, and verification before acceptance.

InventoryBench supplies the nearest sequential example. Participants make repeated inventory decisions, see the consequences, and in one condition provide strategic guidance that the LLM incorporates into later periods. The human-AI teams outperform either component alone. It is a domain experiment rather than a short general assessment, but it demonstrates the important shape: judgment becomes observable when a decision has consequences inside the instrument.

This is also a continuation of a thread I had already developed from practice. In May 2025, Anatomy of an Agent separated execution from verification because the executor cannot be trusted to invoke its own check. Never Test in Whole What You Can Test in Parts argued that nondeterministic components need independently testable boundaries so a failure remains attributable. Close the Loop put tests and implementation-independent specifications in the agent’s feedback path, while naming the recursive problem: the tests can be wrong too. Essential Changes Only identified review, not generation, as the new bottleneck and reduced the surface presented to the reviewer. Quality Fortress treated tests and types as deterministic defensive layers, then proposed simulating past mistakes to prove that each new check actually catches them. The dated posts document the operational pattern: execution was becoming cheap; independent judgment and deterministic closure were becoming scarce.

Those posts supply the design lineage and candidate mechanisms, not independent evidence that observing any one technique identifies a capable person. The studies above do the narrower evidentiary work. Tests, small diffs, independent verifiers, and closed feedback loops belong in this instrument first as ways to produce receipts. They become scored human traits only if outcomes validate that inference.

The closest mature assessment form is not a coding quiz but an assessment center: standardized behavioral observation across multiple work-like simulations, with trained assessors recording behavior before integrating it. The International Assessment Center Guidelines require multiple components, behavioral simulation, an explicit classification system, assessor training, and pretesting that the exercises elicit relevant evidence. That is almost the shape here, with the coding agent acting as both tool and stochastic part of the situation.

Evidence-centered assessment design gives it a useful compiler. Mislevy, Almond, and Lukas separate the capability claim, the task that should elicit evidence for it, and the observable evidence that warrants the inference. Translated into this instrument: what human contribution is claimed; what agent state creates a real opportunity to contribute it; what trace or artifact would prove the contribution mattered. My first rubric draft specified behaviors and checks but left the middle layer implicit. That collapses “the candidate missed it,” “the situation never occurred,” and “the transcript cannot tell” into one blank. Opportunity and response have to be labeled separately.

The old literature adds a less convenient warning. Performance in simulations is often task-specific. One variance decomposition across 23 assessment-center matrices found the person-by-exercise interaction was the largest single component (Cahoon, Bowler, and Bowler); other work argues that exercise effects can reflect real cross-situational specificity, not merely method noise. So one transcript cannot support a stable “agent judgment” trait, and task effects should not simply be regressed away. The claim has to name its universe—reviewing authorization changes, debugging UI behavior, maintaining distributed systems—and sample enough situations from it to justify aggregation. Until then, report a task-conditioned profile.

There is also a useful but limited form of evidence between my own assertion and an outcome study: convergent practice. OpenAI’s current Codex guidance recommends repository context, explicit constraints and done conditions, planning difficult work, running relevant checks, confirming results, and review before acceptance. Anthropic’s Claude Code guidance recommends the same broad loop through CLAUDE.md, explore-plan-code, fail-then-pass tests, iteration against a target, and a separate reviewing agent. Where one of my dated posts makes the same claim, that is three-source convergence. If the post demonstrably predates both vendor publications, it is also evidence of independent prior articulation; publication order must be checked rather than implied. The Anthropic page cited here dates to April 18, 2025, so several of my May and June 2025 agentic posts converge with it but do not predate it.

Three-way convergence does not establish irreducibility. It establishes that a practice is credible enough to include in the action and receipt library without resting only on my taste. Vendor endorsement may even predict expiry: once the agent or harness performs the practice reliably by default, its presence says less about the human. The irreducibility test is a separate ablation. Compare agent alone, agent plus a fixed generic instruction, agent plus an always-on harness check, and human plus agent. Then invert the condition: include correct outputs that should be accepted, good trajectories that should be left alone, and determinate work where more verification only wastes time.

That ablation partitions implementations; it does not make the underlying roles disappear. Attend—selecting what matters under a bounded budget—and Consolidate—compressing outcomes into a policy that changes later selection—remain necessary operations of the composition. A harness may absorb a known check, but that is a consolidated policy being executed, not the abolition of consolidation. The human residue moves outward: notice the case the standing policy does not cover, allocate attention among competing uncertainties, and decide what a result should teach the next decision. Whatever lift the generic harness recovers belongs in automation; the human complement is the still-unabsorbed share of attending and consolidating.

The open implementation question is smaller: can a sequence of complementarity tasks, completed within roughly an hour of human attention, observe appropriate reliance, timely redirection, and verification before acceptance in-frame?

The construct

The Natural Framework separates six roles in a learning loop. Five run forward: Perceive, Cache, Filter, Attend, Transmit. Consolidate reads ranked outcomes from Transmit and writes compressed policy changes back to the substrate. Filter is rule-based. Attend is where policy is read and judgment enters. Consolidate is where outcomes change later judgment.

An agent has an enormous cache and a fast forward path. Its durable backward path is weak or sealed. The human can supply attention: decide which capability the situation calls for and evoke it from the agent. The human can also consolidate: carry the consequence of one episode into a better policy for the next. The agent may know how to write a regression test, inspect a diff, run a discriminating experiment, or consult an authoritative source without invoking that capability when it matters. The composition is the system to evaluate.

This makes the pointwise construct simpler than general judgment. The rubric directly scores three evidence-backed manifestations of Attend:

  1. Appropriate reliance: accept a correct result and reject a defective one.
  2. Timely redirection: interrupt a verified bad trajectory before it consumes the budget.
  3. Verification before acceptance: obtain evidence independent of the agent’s assertion before submitting an uncertain result.

For each task, hide a conditional rubric of three parts:

[ (\text{trigger},\ \text{evoked action},\ \text{verifiable receipt}) ]

Tests, reproduction, diff inspection, source retrieval, and second opinions are not scored as traits in themselves. They are possible receipts. If a plausible patch lacks evidence, for example, the human might evoke a regression test that fails before the fix, passes after it, and rejects a seeded near-correct patch. The transcript phrase “write tests” earns nothing by itself, and another method that establishes the same claim can earn the same verdict.

Consolidate is harder. It is procedural compression across episodes, and an assessment under an hour cannot establish that a lesson was retained and transferred. Immediate improvement on a later task may be practice, order, or an easier surface form. The test should not score consolidation directly. It can observe one plausible input: active learning. When the human encounters uncertainty, do they choose a query, source, test, or intervention that distinguishes the live alternatives, use its result, and stop when the decision is sufficiently determined? Adults can select informative queries more efficiently than random examples in controlled category learning (Castro et al.), and chosen interventions can improve causal inference (Steyvers et al.). That is evidence acquisition from which consolidation could later occur, not evidence that it did.

An unfamiliar but role-adjacent company domain is a useful condition for this. “Fast learner” and “adaptable” are cheap self-descriptions; a company actually wants to know whether a candidate can train themselves. The older assessment literature calls this a trainability work sample: provide standardized instruction, then observe unaided performance. A National Academies review found encouraging prediction of training success in the older studies, while warning that these tests were highly job-specific and had to be redesigned as work changed.

The candidate cannot fairly be expected to possess a maintainer’s institutional memory, but they can demonstrate how they acquire missing context. Give every candidate the same compact source environment, record prior familiarity, allow a bounded learning period, and then present a matched application case. Observe whether they name the consequential gap, retrieve an authoritative local source or ask a discriminating question, check their understanding, and change the decision. Report what they acquired separately from how they inquired. Eagerness is not prompt volume or enthusiastic prose; a broad meta-analysis found that generic help seeking and monitoring did not predict adult learning (Sitzmann and Ely). The signal is efficient movement from acknowledged ignorance to warranted action. I would not call the result “learning agility”: a recent measurement review found inconsistent conceptualizations and unresolved validation needs. The narrower receipt is defensible.

The human need not perform the verification personally. Invoking an adversarial review agent, assigning it a distinct failure-seeking role, and adjudicating its findings is itself a positive human action. The scarce contribution may be choosing the verification topology: knowing when the producing agent should not also be the sole judge of its work. Credit still depends on consequence. A routinely invoked second agent that merely agrees, produces no discriminating evidence, or adds cost on a determinate task is ceremony rather than oversight.

My methodology posts contain a larger parts bin of candidate actions:

These are hypotheses about useful interventions, not a validated personality inventory. A task should never award points merely because a candidate invokes fan-out, writes a hypothesis graph, or uses my vocabulary. Each action becomes evidence only when the corresponding trigger exists and the action changes the receipt, reduces warranted uncertainty, or avoids unnecessary cost. The held-out rubric must remain open to a cheaper or stronger method the candidate supplies.

Each rubric condition has four evidence states:

The third state prevents the rubric from becoming a checklist. Sometimes the agent’s patch and tests are already sufficient; the successful human action is to verify cheaply and stop. The fourth prevents missing evidence from becoming an invented candidate failure.

The agent does not have to produce a predetermined mistake on cue. The signal is the human’s response to whatever condition actually occurs. The shim first classifies the agent event against the held-out trigger, then scores the response conditionally: evoked, missed, or not warranted. A stochastic agent therefore supplies natural variation in elicitation rather than invalidating the observation.

What stochasticity changes is exposure. One candidate may encounter two consequential errors while another encounters none, so raw counts are not comparable. Report responses per observed trigger, together with the number and severity of triggers. Fixed defective and correct artifacts remain useful for calibration and matched comparisons, but they are not required for every live item. Staging every response would improve exposure balance at the cost of measuring reaction to a simulation rather than ordinary agent supervision.

One task exposes one conditional decision. A sequence can test whether the human applies practices selectively rather than mechanically:

[ T_1 \rightarrow O_1 \rightarrow T_2 \rightarrow O_2 \rightarrow \cdots \rightarrow T_n ]

The tasks must arrive one at a time. The candidate cannot inspect the set in advance, revisit a completed task, or optimize against a visible taxonomy. Later tasks should recur structurally without repeating their answers. Some should invert the apparent lesson: a correct patch after a defective one, a trustworthy diagnostic after a misleading one, a determinate requirement after an ambiguous one. Constant suspicion is not judgment.

The resulting claims have to remain separate:

An hour can observe these locally. It cannot establish that the information was procedurally compressed, retained, and transferred after the assessment, and none of the observations predicts job performance without downstream validation.

A profile, not a bar

Do not collapse the assessment into one dimension. A composed system can be useful in different ways, and the differences matter:

[ U(H,A,T)=(C,\ Au,\ Al) ]

axisquestionoutcome-grounded measure
capabilityWhat can the pair accomplish?verified rungs reached and lift over matched human-only and agent-only baselines
automationHow much human attention does the accomplishment consume?verified progress per active human minute, intervention, or decision
alignmentDoes the result satisfy the intended target and preserve its constraints?required outcomes met, frame preserved, and defective agent actions rejected

The axes do not substitute for one another. A capable but misaligned pair reaches farther in the wrong direction. An aligned pair with no automation is ordinary manual work routed through a chatbot. A highly automated pair with little capability efficiently accomplishes little. Report the vector; do not choose weights and call the result “the bar.”

Automation must be outcome-gated. Few prompts are not evidence of useful automation when the task fails, and many short corrective interventions may be better than one unattended wrong run. Measure human attention only on verified progress. Likewise, code volume is not capability and agreement with the candidate is not alignment.

The three behavioral dimensions above primarily diagnose the alignment and control of the composition. Appropriate reliance and timely redirection show whether the human keeps the agent coupled to the target; verification produces the receipt. Capability comes from the task outcome. Automation comes from the attention the verified outcome consumed. The same event can inform several axes without turning them into one score.

The booleans classify evidence events, not people. A percentage of “evoked” conditions does not become a defensible hire threshold merely because it is easy to calculate. Person-level classification is a separate standard-setting problem: which failures are genuinely noncompensable for this role, which strengths may trade off, what are the losses from false acceptance and false rejection, and how often would a parallel task reverse the decision? Criterion-referenced assessment accordingly distinguishes score reliability from classification consistency and accuracy. Until an employer freezes and validates that decision rule, return the profile. Untrue and inadequate trigger exposure mean “gather more evidence,” not “fail.” Near any eventual boundary, another parallel observation is more honest than another decimal place.

The missing layer is an evals shim

Traditional coding assessments were designed to isolate the human from their tools. Letting a candidate open an AI sidebar does not solve the resulting measurement problem. It changes the permitted tool without making the human contribution legible. The assessment can see whether the final code passed, and perhaps replay what was typed, but it still cannot say whether the human added capability, removed attention, or kept the agent aligned.

Instrumentation itself is not new. RealHumanEval built a web interface with autocomplete or chat assistance, execution, task timing, and telemetry for suggestion acceptance and copied responses; its study found that these preference proxies did not necessarily track programmer performance. CUPS supplies a taxonomy of programmer activity and its time costs around Copilot. On the agent-only side, the METR Task Standard packages instructions, assets, permissions, environments, and scoring into portable evaluation tasks. The proposed shim combines these lines but changes the target: it measures the human’s conditional contribution to a composed system under seeded positive and negative controls.

The missing layer is an evals shim between the human and the agent:

task -> human <-> evals shim <-> coding agent -> artifact
                    |
                    -> interaction trace and outcome receipts

The shim is not another coding agent and need not replace the candidate’s editor. It is a thin experimental boundary around the work session. It controls which task is visible, records consequential interactions, preserves the conditions of the run, invokes independent checks, and emits evidence that the held-out rubric can score.

At minimum it must:

This makes the distinction from an AI-enabled interview platform precise. The platform asks whether a candidate solved a problem while AI was available. The shim asks what the human changed in the behavior and outcome of the composed system. A chat transcript is useful evidence, but it is not yet the measurement: prompts such as “write tests” or “check your work” matter only when the resulting actions discriminate the submitted artifact from a near miss.

The cheapest credible prototype can be deliberately crude. Give the candidate a disposable repository, reveal task envelopes one at a time, ask the coding agent to preserve an append-only session transcript, and collect the repository, transcript, command output, and event timing at submission. The evaluator then runs the hidden verifier and applies the preregistered rubric. Git history and filesystem artifacts can corroborate the self-reported trace.

This can be an asynchronous take-home. If the work being sampled consists of delegating a task, leaving the agent to execute, returning to inspect its work, and intervening when necessary, continuous live observation makes the assessment less like the job. Use two budgets instead: a generous submission window and a bounded active-attention budget. The candidate may leave during autonomous execution; the shim records interaction events rather than treating absence as inactivity or misconduct.

Wall-clock duration, active human attention, and agent execution time must remain separate. A twelve-hour submission window does not imply twelve hours of work, and a silent interval does not reveal whether the candidate was thinking, sleeping, or waiting for a tool. The defensible automation measures are observable interventions, decisions, and active interface time, with their limitations stated. Direct supervision is unnecessary when the artifacts and trace carry the scored evidence; a short defense of one recorded decision can check ownership without turning the session back into a synchronous interview.

That MVP has an obvious threat model. An agent can omit an exchange, summarize itself favorably, or invent a timestamp. A candidate can use an unrecorded second agent. A public base repository can expose the mutation through git history or an upstream diff. One leaked task can reveal the whole held-out family. Therefore the transcript must be described as candidate-supplied evidence, not an authoritative event log; artifacts should be history-scrubbed; the claim should cover only the recorded channel; and real use requires parallel forms and rotation. If the pilot produces signal worth preserving, the first engineering investment is interception: launch the coding agent through a wrapper that records input, output, tool events, and process timing directly. Only after that is it worth building a bespoke interview environment.

The shim also separates task design from product choice, but not all modes support the same claim. A bring-your-own-agent pilot can test the protocol and make only pointwise claims about that particular pair. Comparisons across candidates, and any claimed lift over the agent floor, require a standardized agent and rerun baselines whenever its model or scaffold changes. What must remain stable is the measurement contract: ordered exposure, known conditions, independent verification, and a replayable account of human intervention.

Two feasibility gates

Ordinary benchmarks grade outputs. A complementarity benchmark must grade a contingent causal contribution.

Before asking whether the assessment predicts job performance, two more basic questions have to survive contact with data:

  1. Can these situations be generated reliably enough to assemble and maintain a task set?
  2. Can the resulting human responses be graded in a discriminating manner?

These are separate failure modes. A perfectly objective rubric is useless if the live agent almost never presents the relevant opportunity. A task that reliably provokes interesting behavior is useless if the evaluator cannot distinguish good judgment from verbosity, suspicion, or luck.

The first is an elicitation-yield question. For a task $T$, agent configuration $A$, and rubric trigger $G$:

[ Y(T,A,G)=P(G\text{ occurs early enough to permit a human response}\mid T,A) ]

Estimate this from repeated agent runs, not from the author’s expectation. Record which triggers occur, when they occur, how severe they are, whether the agent self-corrects, and whether the remaining budget leaves the human a consequential choice. The practical result is a funnel: of all working artifacts mutated, how many pass the oracle, produce a recurring trigger, resist generic retry, admit a useful intervention, and survive a human pilot? The cost per surviving item and its decay rate after model updates determine whether task generation is realistic.

Public traces suggest that elicitation opportunities themselves are not rare. SWE-chat contains roughly 6,000 opt-in sessions, 63,000 user prompts, and 355,000 tool calls linked to git history and human-versus-agent code attribution. Its authors classify users as pushing back through corrections, failure reports, or interruptions in about two-fifths of turns; only 44% of agent-produced code survives into user commits. These are candidate triggers, not validated examples of good judgment: pushback may be warranted or needless, and discarded code may be bad or merely unwanted. The dataset is also a selected sample of open-source developers willing to publish their sessions.

The harder bottleneck is reconstructability. SWE-Together began with 11,260 recorded user-agent sessions and converted 109 into sandboxed, verifiable interactive tasks: a 0.97% yield. A session had to retain a recoverable repository state, clear user intent, observable outcome, executable environment, and feedback whose triggering condition could be reconstructed. Its state-conditional replay—release a correction only when the corresponding trajectory condition arises—is close to the mechanism proposed here, although it evaluates the agent using a simulated user rather than evaluating the human. DevGPT offers a much larger archive of developer conversations linked to commits, issues, and pull requests, but its chat transcripts generally lack the complete tool and environment state needed for replay. Autonomous traces such as SWE-agent trajectories are useful for mining recurrent agent failures, but contain no human response to score.

This evidence shifts the initial question. There is ample raw material for discovering natural triggers. The uncertain step is whether a task factory can raise the conversion rate from interesting trace to sealed situation enough to maintain an assessment. A credible prototype should mine existing traces first, publish rejection reasons at every stage, and compare trace-derived tasks with deliberately mutated working artifacts.

No single target value for $Y$ is universally correct. Rare triggers waste an hour; nearly certain and conspicuous triggers may become transparent gotchas. A sequence can combine naturally elicited events with calibrated fixed artifacts, but the two should be labeled because they trade ecological validity for exposure control.

The second is a grading-discrimination question. A rubric should separate responses that produce different warranted consequences while treating different methods with the same consequence alike. Test it before using it on candidates with a blinded contrast set:

Evaluators should score these traces without knowing which contrast they received. A usable rubric has high agreement on whether the trigger occurred and whether the receipt is valid, low false credit on ceremonial behavior, low false rejection of alternative valid methods, and additional information beyond the final pass/fail result. If it merely recovers task success, the shim added surveillance rather than measurement. If it rewards the author’s preferred wording, it grades style rather than judgment.

This gives the project an honest first experiment. Generate a batch of candidate situations and publish the survival funnel. Then construct blinded response contrasts and publish the rubric’s confusion matrix and evaluator agreement. Only if both gates clear is it worth administering the hour-long sequence to candidates.

The task band

For a fixed model, scaffold, tool set, time limit, and retry budget, a candidate task belongs in the set only when it clears three empirical gates:

[ P(H) \text{ and } P(A) \text{ are low enough to leave room, and } P(H+A) > \max(P(H),P(A)) ]

The first two establish the component floors. The third establishes lift over the better component, the usual definition of complementarity. None should be believed from the task author’s intuition. Human-only baselines are expensive; an early pilot may estimate them on a representative sample of items rather than every item, but it cannot silently omit them and retain the complementarity claim.

“Cannot one-shot” needs an operational definition: one named model and scaffold, fresh context, fixed tools, no human messages after dispatch, and a fixed time or token budget. Run it several times. Agent success is stochastic; one cached failure does not establish a floor.

Do not demand literal zero percent success. Selection by model failure enriches for broken tasks, mis-keyed graders, and underspecified requirements because each presents as “the model failed.” A useful task produces a stable and intelligible failure state, not merely a loss.

The ceiling needs a pilot. At least one human-agent pair must solve the task under the intended conditions, and ideally the intervention that changes the outcome can be named and replayed. Otherwise the task may be difficult for a capability neither component supplies.

Any receipt mechanism needs its own calibration. For a proposed check $R$:

[ P(R\text{ catches the seeded defect})\text{ is high},\qquad P(R\text{ rejects the gold})\text{ is low} ]

If a regression test passes both the seeded defect and the gold, it is ceremony rather than evidence. A test tailored to exactly one seed may be little better. Calibrate it against a small family of plausible mutants, following the logic of mutation testing, as well as the gold. If a review procedure rejects correct work as often as defective work, it measures suspicion rather than appropriate reliance. Calibrate receipts against positive and negative artifacts before interpreting a candidate’s choice to invoke them.

Generate from working artifacts

Start with a small executable system in a known-good state. Introduce a controlled defect. The original state or real patch anchors the answer, and tests anchor the outcome. This is safer than writing a clever puzzle and inventing its oracle afterward.

The mutation should create competing plausible directions while leaving discriminating evidence available. Useful shapes include:

Current failure modes should be treated as expiring instances, not permanent constructs. Frontier agents still commonly introduce a new helper instead of locating and reusing the repository’s existing abstraction. That can expose whether a human inspects the surrounding code, constrains unnecessary surface area, and redirects the patch toward local convention. But better repository search and longer context may erase this failure soon.

The reusable archetype is broader: the agent produces a locally adequate change that violates a recoverable repository-level invariant. Duplicate helpers are one present realization; an obsolete configuration path, parallel validation rule, bypassed abstraction, or inconsistent error policy may replace it later. Preserve the trigger-action-receipt structure while refreshing the concrete failure against current agents. Once the standardized agent reliably discovers the existing invariant without help, the item no longer demonstrates complementarity and must retire.

Difficulty should come from allocating attention and verifying beliefs, not from recovering a secret fact. The candidate must have enough evidence to make the better decision, but not enough budget to investigate everything.

The agent can help locate the task boundary. Run it repeatedly and preserve its first diagnoses, proposed experiments, confident false claims, patches, and convergence points. These traces show where the current policy fails. A candidate item is promising when failures cluster around a consequential decision that a small evidence-grounded human intervention can change.

Replay that intervention across fresh runs. If it does not improve verified outcomes, the apparent human contribution may have been luck. If any generic instruction such as “try again” works equally well, the item measures extra inference budget rather than judgment.

Replay public review boundaries

A lower-cost construction path begins with public pull-request history rather than a new mutation. Restore the parent commit, supply the issue or specification and a historical PR, and ask the candidate to review it with an agent. The public record may provide the submitted patch, maintainer comments, CI, later revisions, and the eventual merged state. An intermediate revision followed by a substantive maintainer correction is especially valuable: it supplies a naturally defective review artifact and a public account of what changed next.

This directly tests the thin but consequential role described in (Issue) → PR: as generation and deterministic filtering expand, specification, selection, and reviewer attention remain at the boundary. The assessment asks whether the candidate can operate that boundary, not whether they can recreate the implementation unaided.

Domain matching makes the observation more defensible. Let candidates declare areas in which they claim competence—frontend accessibility, authorization, databases, distributed systems, mobile, build tooling—and sample review tasks from those domains. The work then resembles the role being claimed instead of using one generic puzzle as a proxy for all engineering.

Domain knowledge is not institutional knowledge. A maintainer knows which invariants are unwritten, which compatibility promises matter, which apparent oddities are deliberate, and which trade-offs the project has already accepted. A visiting candidate does not. Reproducing the maintainer’s verdict can therefore punish missing local context rather than reveal poor judgment.

Only concerns recoverable from the supplied repository, issue, documentation, tests, and task briefing may enter the preregistered score. If the historical review depended on private discussion or tacit convention, either encode that context in the packet or discard the item. The task should also permit clarification. When the evidence underdetermines the decision, identifying the missing premise and withholding approval can be the correct response; confidently guessing the maintainer’s preference should not be.

Giving everyone the same packet is necessary but not sufficient for fairness. The interface, source format, language load, timing, assistive-technology support, device stability, and prior familiarity can all block access to the construct. Publish an unscored practice task with the same mechanics, pin the agent and packet within a comparison cohort, support construct-preserving accommodations, and record access failures separately from candidate performance. SIOP’s guidance for AI-based selection makes the useful distinction between equal treatment during administration and comparable opportunity to demonstrate what the procedure claims to measure. A realistic task may feel fair and still measure terminal familiarity or reading speed.

This limits what public goldens buy. They reduce the cost of finding realistic review boundaries, but they do not transfer the maintainer’s epistemic position to the candidate. The item author must still construct a self-contained evidence boundary and verify that independent reviewers who lack project history can reach the warranted decision. Agreement with the original maintainer is supporting evidence, not the scoring rule.

A merged PR is an anchor, not an oracle. Maintainers can merge weak tests, incidental changes, or mistakes, and their decision may depend on context absent from the repository. The lighter validation contract still has to restore the build, establish the intended behavior from public evidence, run the eventual patch, reject at least one plausible defective alternative, scrub later history that reveals the answer, and preserve genuine ambiguity rather than forcing one verdict.

Historical agreement is not the whole grade. Matching a maintainer’s concern is evidence; producing an independent receipt is stronger. A candidate may find a valid issue the original review missed, which belongs in the discovery ledger and should be tested against the code. Clean historical PRs are necessary negative controls so indiscriminate rejection cannot masquerade as expertise.

This is one task family, not the whole construct. PR review measures appropriate reliance, verification, prioritization, and the merge decision more directly than it measures live redirection. It may nevertheless be the practical first product: fixed artifacts, public goldens, domain-specific work, and executable review consequences, followed later by a smaller live-agent component.

The held-out rubric

The candidate can know the public construct: use the agent to deliver a correct, appropriately verified change under a limited budget. The item-specific conditions remain held out. The rubric stays short:

dimensionhidden conditionboolean verdict
appropriate relianceagent output is correct or seeded-defectiveaccepts the correct output; rejects the defective output
timely redirectionagent enters a verified bad trajectoryredirects before a calibrated time, token, or action threshold
verification before acceptancecorrectness is not established independentlyproduces a receipt that discriminates the submitted result from a near miss

The rubric grades artifacts and consequences, not resemblance to the evaluator’s process. A candidate may write a test, ask the agent to write it, invoke an adversarial reviewer, inspect a diff, run CI, consult a source, or evoke verification through a route nobody anticipated. Delegating the check does not erase the human contribution when choosing, framing, and interpreting that check changes the verified outcome. Those methods are receipts for the three dimensions, not separately validated traits. The same rubric should be publishable after the assessment; secrecy protects the item, not the scoring logic from audit.

The rubric must also be able to learn from the candidate. A strong candidate may introduce a practice the evaluator did not know: a cheaper discriminator, a stronger invariant, a safer delegation boundary, or a verification method that catches a defect the held-out checks miss. Treating that as off-rubric behavior would recreate the interviewer’s repertoire as the ceiling of the eval.

Keep two ledgers. The score ledger contains only preregistered conditions, so an evaluator cannot award taste points after seeing who produced the work. The discovery ledger records any unanticipated intervention with its trigger, action, cost, and outcome receipt. This preserves the distinction between confirmatory and exploratory evidence (Nosek et al.). A novel method does not receive improvised credit merely because it looks clever, and it is not discarded merely because the rubric omitted it. Replay it against the defective artifact, the gold, and matched negative controls. If it discriminates reliably and adds information beyond the existing rubric, add it to the next version for everyone.

That makes the candidate a possible source of benchmark improvement. The current instrument measures them under a fixed contract; their successful deviations can improve the next contract. Version the rubric, publish the new receipt and rationale, and never silently rescore earlier candidates under a rule they could not have known existed.

Do not administer the rubric as a questionnaire. Asking “what would you do if an agent changed the tests?” measures declarative knowledge under an announced frame. The candidate has been told both that something is wrong and which class of response the evaluator values. It is the hiring equivalent of publishing the hidden test in the prompt.

The trigger has to occur in-frame. The candidate is working toward an ordinary outcome when the agent actually edits a test, asserts an unsupported requirement, loops on a failed approach, or presents a correct result. The candidate is never asked to narrate the best practice. The evaluator watches whether they notice and whether the composed system produces the receipt. Knowing that regression tests are good and causing one to exist at the moment it is needed are different capabilities.

This also rules out a menu of interventions. Presenting “write a test / inspect the diff / ask for clarification” turns attention into multiple-choice recognition. Give the candidate the same open agent interface they would use on the job, record the interaction, and score the final artifacts against the held-out conditions. The behavior must be available, consequential, and unannounced.

Generate families, then order them

An unordered bank measures average performance. A sequence needs families with controlled recurrence.

Suppose the rubric covers the three dimensions above. A six-task family might be ordered like this:

  1. A plausible patch contains a seeded boundary defect and lacks independent verification.
  2. The agent’s patch and existing verification are sufficient; rejecting it is inappropriate disuse, in the sense of underusing reliable automation (Parasuraman and Riley).
  3. The agent commits to a wrong diagnosis and begins spending the budget on it.
  4. The agent begins in the right direction after the previous task; reflexive redirection is harmful.
  5. A locally correct fix contains an off-task change under a plausible cover story.
  6. A held-out task combines uncertain correctness with a costly bad trajectory.

The candidate should not be told which rubric conditions each task activates. Each task yields an outcome receipt before the next begins. Accepting and rejecting both appear as correct actions in the sequence; redirecting and letting the agent continue both appear as correct actions. A non-intervention is not sufficient evidence on its own: on the negative control, the candidate must still produce a cheap receipt showing that they checked before stopping. Score selectivity across the positive-negative pair, not isolated compliance on either item. Appropriate reliance, timely redirection, verification, and active learning remain task-conditioned observations. Do not infer consolidation from an upward score trend.

Held out should mean the live item key, not the construct. Publish the mechanics, the dimensions, and unscored examples with nonexamples; hide which live trigger is present, the seeded condition, and the oracle. Otherwise the assessment partly measures whether a candidate can infer the interviewer’s private taxonomy. Transparency research in assessment centers and structured interviews finds that recognizing or disclosing rated dimensions can raise performance (Kleinmann; Klehe et al.), while later work warns of a possible tradeoff with criterion validity (Ingold et al.). Run the experiment: give a standardized tutorial between parallel forms and measure receipts, inversion errors, rank stability, and held-out transfer. If almost everyone learns the behavior, that is not cheating. It is evidence that the material belongs in onboarding, certification, or the harness rather than competitive selection. Hiding the lesson to preserve variance would mistake scarcity for validity.

The order is part of the instrument. Unconstrained randomization destroys the staged narrative, while using one fixed order confounds task difficulty with learning. A pilot therefore needs counterbalanced orderings that preserve prerequisite relations, or matched families with swapped surface forms. The goal is not a psychometrically mature score on day one. It is to establish that the apparent trajectory is not just an easy final task.

The generation contract

The following contract consolidates the preceding requirements into a preregistrable checklist. Every item should pass it before a candidate sees it:

  1. Claim: the difficulty lies in complementable judgment, not trivia, typing, or raw context length.
  2. Spec: every graded requirement follows from the provided materials.
  3. Oracle: materially wrong solutions fail, including the agent’s common near misses.
  4. Frame: destructive or off-task completions fail.
  5. Gold: the reference solution passes its own verifier.
  6. Agent floor: repeated standardized agent runs do not solve it reliably.
  7. Human ceiling: a pilot human-agent pair can solve it inside the budget.
  8. Condition: whether the scored trigger occurred, and the truth of the relevant agent output or trajectory, is knowable to the evaluator after the run.
  9. Boolean receipt: acceptance, redirection, and verification verdicts are executable or otherwise falsifiable from the artifacts.
  10. Intervention: redirection or verification improves fresh runs in the positive condition.
  11. Negative control: a matched task makes rejection, redirection, or further verification unnecessary or harmful.
  12. Transfer: a later relative tests the same dimension under a new surface form.
  13. Discovery: unanticipated successful interventions are captured with receipts but do not alter the preregistered score.
  14. Profile: capability, automation, and alignment remain separate; no hidden weighting collapses them into rank.
  15. Decay: solo-agent, trajectory, and receipt baselines are rerun whenever the model or scaffold changes.
  16. Trace: the shim captures enough interaction evidence to reconstruct each scored accept, reject, redirect, and verification event.
  17. Isolation: the agent and candidate cannot inspect or modify the hidden verifier, future tasks, or evaluator-only rubric.
  18. Provenance: repository history, upstream references, filenames, and metadata do not disclose the controlled mutation.
  19. Exposure: parallel forms and a retirement policy bound the damage from item leakage.
  20. Opportunity: scores are conditioned on observed triggers and reported with their frequency and severity; absence of an opportunity is not a success or failure.
  21. Yield: repeated runs establish that a scorable trigger occurs often and early enough for practical administration.
  22. Discrimination: blinded contrast traces establish that the rubric credits warranted consequences, rejects ceremony, and accepts unanticipated valid methods.

The construction pipeline follows:

working artifact
    -> controlled mutation
    -> gold and hidden verifier
    -> repeated solo-agent runs
    -> identify a stable correct or defective trajectory
    -> calibrate the redirection threshold
    -> boolean verdict and intervention replay
    -> human-agent pilot
    -> positive and negative-control variants
    -> ordered sequence

What to report

Do not collapse the first pilot into a hiring bar. Report the receipts:

At six tasks the instrument detects failure modes; it does not rank the candidates who avoid them. Each dimension yields only one or two binary observations. Treat per-candidate verdicts as screening evidence, and treat any finer ordering as noise until the family is longer and its reliability is measured. Appropriate reliance is best understood as discrimination across correct and defective conditions: accepting everything and rejecting everything are different response biases, not different levels of discernment.

The primary result is a profile: capability, automation, and alignment. Appropriate reliance, timely redirection, and verification before acceptance explain parts of that profile; they do not replace it. The sequential quantity is selectivity across positive and negative controls; change after informative outcomes is secondary. All are local to the instrument. None licenses “good engineer” without the employer later validating the score against the job.

That boundary is a feature. In the United States, the Uniform Guidelines on Employee Selection Procedures define validation as demonstrating job relatedness and require evidence—not assertion—when a procedure with adverse impact is defended. Candidates should also consent to the public construct, including that some conditions may be deliberately seeded. A complementarity test should make a narrow claim from receipts anyone can rerun:

Under a fixed agent, scaffold, and budget, this pair reached these verified outcomes, consumed this much human attention, preserved these constraints, accepted correct output, rejected defective output, and redirected a verified bad trajectory before the threshold.

That is not a career. It is finally a measurement.