Hire June Kim
I build research systems that can be inspected, rerun, and judged by their artifacts: benchmark audits, agent pipelines, evaluation harnesses, review protocols, and tools that turn AI research questions into working software.
I am an AI systems / research engineer. The depth is independent evaluation: I audit frontier coding benchmarks for construct validity and find where the headline metric measures the wrong thing, every verdict a re-runnable receipt. The breadth is shipping: 10+ years across Google, Loom, and startups, and fixes landed in dozens of open-source projects across Rust, Go, C++, and Python, which is what lets me read any benchmark's test suite in any stack.
The common thread is checkable work. I'd rather hand you something you can verify than something you have to believe.
The best fit is a team that needs someone between research and product: build the system, run it against real workflows, measure what breaks, and turn the result into a better product or protocol.
Forward me for
- Research engineering for AI systems, agents, coding tools, or evaluations.
- One-time independent construct-validity audits of consequential AI-system claims.
- Applied AI work where ambiguous research ideas need to become shipped systems.
- Developer tools, AI infrastructure, or agentic workflow products.
- Founding or early engineering roles around research-adjacent products.
What other people staked on the work
Judge these first. Every item is someone else's action on a public timestamp, not my description of my own work.
- 92 pull requests merged across 76 external organizations in 2026, including hyper, TiDB, Servo, Enzyme, wild, flux, burn, luminal, ag2, and pertpy. A maintainer pressing merge is a cost paid by someone with no reason to flatter me. Full list.
- Archived and regradeable, with DOIs. Every audit ships its receipts to Zenodo, versioned, with a regrade script: Terminal-Bench audit, SWE-bench Pro run, Assurance at the Boundary. Anyone can re-run the grading and disagree with me on the record.
- Fixes filed upstream, not just findings published. harbor#2266 implements the Terminal-Bench frame gate as an opt-in observational check, CI green, currently open. inspect_ai#4462 does the same for Inspect. When I say a benchmark is broken, I send the patch.
Provenance
Twice I have audited a coding benchmark, received no reply, and watched something happen anyway: DeepSWE shipped a re-graded v1.1 eighteen days later in which exactly the four tasks I flagged climbed while the pooled rate held flat, and OpenAI retracted its SWE-bench Pro recommendation twenty-nine days after my determinacy audit went up on Scale's tracker. I claim no causation, only the dates, which sit on other people's servers. What the comparison is actually about.
Proof
- Benchmark and eval auditing: public construct-validity audits of frontier coding benchmarks (ProgramBench, SWE-bench Pro, DeepSWE, Terminal-Bench, SWE-rebench, SWE-bench Verified), plus determinacy, a reusable auditing tool. Each verdict is a grep-checkable receipt anyone can re-run.
- The Hypothesis Graph: a harness-layer data structure for coding agents — testable-claim nodes, refutation-condition edges — that makes an agent's reasoning inspectable and reusable. arXiv-shape preprint with reproducible, DOI-archived artifacts.
- Verifiable Knowledge: a protocol where each agent presents each claim with a falsifiable condition another agent can re-run instead of attesting its own work. Accountable failure outranks unaccountable assertion.
- Slop Slope: tested whether adversarial review loops turn test-passing LLM code from coin-flip drafts into merge-ready artifacts.
- Agentic open-source contribution: pipeline for finding issues, generating patches, submitting PRs, and tracking maintainer outcomes.
- Cognition and epistemology research: abduction, memory, provenance, review loops, and agentic systems as a theory layer for visible tools.
- Also: Sweep & Triage, (PR) → merged, PageLeft, and Runnable Textbooks.
Resume & Contact
Resume page · PDF · Markdown