publications

Research on AI agents, software security, and dependable systems.

2026

  1. SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
    Hwiwon Lee, Jiawei Liu, Dongjun Kim, and 4 more authors
    arXiv preprint arXiv:2605.26548, 2026
    Adopted by OpenAI for evaluating latest models + Google VRP $20,000 bounty
  2. Agentic Vulnerability Reasoning on COTS Binaries
    Hwiwon Lee, Jongseong Kim, and Lingming Zhang
    arXiv preprint arXiv:2605.05000, 2026
    39 zero-days (23 assigned CVEs) and $200,000+ in Microsoft bounty awards

2025

  1. SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
    Hwiwon Lee, Ziqi Zhang, Hanxiao Lu, and 1 more author
    In Proceedings of the Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS’25), 2025

2024

  1. BENZENE: A Practical Root Cause Analysis System with an Under-Constrained State Mutation
    Younggi Park, Hwiwon Lee, Jinho Jung, and 2 more authors
    In 45th IEEE Symposium on Security and Privacy (Oakland’24), 2024
    IEEE Symposium on Security and Privacy 2024 Distinguished Paper Award

2022

  1. SoK: Demystifying Cyber Resilience Quantification in Cyber-Physical Systems
    Hwiwon Lee, Sosun Kim, and Huy Kang Kim
    In 2022 IEEE International Conference on Cyber Security and Resilience (CSR), 2022

2020

  1. PhantomFS-v2: Dare You to Avoid This Trap
    Jione Choi, Hwiwon Lee, Younggi Park, and 6 more authors
    IEEE Access, 2020