publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

2026

  1. NeurIPS LCFM
    img_sketchssm.png
    SketchSSM: Write to the Full State, Read from a Compact Sketch
    Omin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer, Sehoon Kim*, and Jae W. Lee*
    NeurIPS 2026 LCFM Workshop; extended version under review
  2. Under Review
    img_herald.png
    HERALD: High Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval
    Omin Kwon, Doyeon Kim, Jongseok Park, Seung Yul Lee, Ion Stoica, and Jae W. Lee
    2026
  3. NeurIPS
    img_mage.jpg
    MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM
    Omin Kwon, Yeonjae Kim, Doyeon Kim, Minseo Kim, Yeonhong Park, and Jae W. Lee
    NeurIPS 2026& ICML 2026 AdaptFM Workshop (Best Paper Award)

2025

  1. ICCAD
    img_star.jpg
    STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware Randomizer
    Omin Kwon*, Kyungjun Oh*, Jaeyong Lee, Myungsuk Kim, and Jihong Kim
    ICCAD 2025
    U.S. Patent Application with SK hynix (No. 19/569899; filed March 17, 2026)
  2. NeurIPS
    img_nestedfp.png
    NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
    Haeun Lee, Omin Kwon, Yeonhong Park, and Jae W. Lee
    NeurIPS 2025
  3. CAL
    img_aide.png
    AiDE: Attention-FFN Disaggregated Execution for Cost-Effective LLM Decoding on CXL-PNM
    KyungSoo Kim, Omin Kwon, Yeonhong Park, and Jae W. Lee
    IEEE CAL, 2025