SEC-bench Pro, our benchmark for long-horizon software security tasks, is now available on arXiv.
Hwiwon
Lee
I build and evaluate AI agents for cybersecurity.
Computer Science PhD student at the University of Illinois Urbana-Champaign, advised by Lingming Zhang. My work sits between software security, systems, and AI.
Dispatches
Field log / 2026Agentic Vulnerability Reasoning on COTS Binaries is now available on arXiv.
Recognized as a Microsoft Most Valuable Security Researcher for 2026 Q1.
Selected papers
Reading index / 04SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
A benchmark of 344 validated vulnerabilities across browser engines and the Linux kernel, designed to measure realistic agent bug hunting.
PDF: SEC-bench ProAgentic Vulnerability Reasoning on COTS Binaries
An end-to-end study of autonomous vulnerability discovery and debugger-verified validation on commercial Windows binaries.
PDF: Agentic Vulnerability Reasoning on COTS BinariesSEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
A fully automated framework for constructing and evaluating authentic proof-of-concept generation and vulnerability patching tasks.
PDF: SEC-benchBENZENE: A Practical Root Cause Analysis System with an Under-Constrained State Mutation
A practical, automated crash-diagnosis system that uses under-constrained state mutation to rank root causes efficiently.
PDF: BENZENE