Hire June Kim

I build research systems that can be inspected, rerun, and judged by their artifacts: benchmark audits, agent pipelines, evaluation harnesses, review protocols, and tools that turn AI research questions into working software.

I am an AI systems / research engineer. The depth is independent evaluation: I audit frontier coding benchmarks for construct validity and find where the headline metric measures the wrong thing, every verdict a re-runnable receipt. The breadth is shipping: 10+ years across Google, Loom, and startups, and fixes landed in dozens of open-source projects across Rust, Go, C++, and Python, which is what lets me read any benchmark's test suite in any stack.

The common thread is anti-credence, pro-merit: replace trust in credentials, affiliation, and polished claims with artifacts that earn trust through evidence, provenance, review, execution, and use.

The best fit is a team that needs someone between research and product: build the system, run it against real workflows, measure what breaks, and turn the result into a better product or protocol.

Forward me for

Proof

Resume & Contact

Resume page · PDF · Markdown

june@june.kim · LinkedIn · GitHub