Hire June Kim

I build research systems that can be inspected, rerun, and judged by their artifacts: benchmark audits, agent pipelines, evaluation harnesses, review protocols, and tools that turn AI research questions into working software.

I am an AI systems / research engineer. The depth is independent evaluation: I audit frontier coding benchmarks for construct validity and find where the headline metric measures the wrong thing, every verdict a re-runnable receipt. The breadth is shipping: 10+ years across Google, Loom, and startups, and fixes landed in dozens of open-source projects across Rust, Go, C++, and Python, which is what lets me read any benchmark's test suite in any stack.

The common thread is checkable work. I'd rather hand you something you can verify than something you have to believe.

The best fit is a team that needs someone between research and product: build the system, run it against real workflows, measure what breaks, and turn the result into a better product or protocol.

Forward me for

What other people staked on the work

Judge these first. Every item is someone else's action on a public timestamp, not my description of my own work.

Provenance

Twice I have audited a coding benchmark, received no reply, and watched something happen anyway: DeepSWE shipped a re-graded v1.1 eighteen days later in which exactly the four tasks I flagged climbed while the pooled rate held flat, and OpenAI retracted its SWE-bench Pro recommendation twenty-nine days after my determinacy audit went up on Scale's tracker. I claim no causation, only the dates, which sit on other people's servers. What the comparison is actually about.

Proof

Resume & Contact

Resume page · PDF · Markdown

june@june.kim · LinkedIn · GitHub