FuriosaAI presented four accepted papers at ICML 2026 in Seoul this summer. Featured in recent IT press coverage, this research highlights Furiosa’s software stack innovation, demonstrating that high-performance, energy-efficient AI acceleration requires deep co-optimization across model architectures, serving pipelines, and hardware execution. All four papers were written by the company’s AI Research Group, which is part of the Algorithm Team at Furiosa.
ReJump: Tree-jump representation for LLM reasoning
Overview: Models LLM reasoning processes as a tree structure with non-adjacent transitions ("jumps") to capture complex execution behaviors such as backtracking, verification, and calculation. By systematically analyzing these execution paths, ReJump improves Best-of-N selection strategies by up to 9.1% on reasoning tasks.
Why It Matters: Moves LLM evaluation beyond simple final-accuracy metrics, providing a structured framework to diagnose how reasoning models explore, overthink, or verify intermediate steps
Furiosa Co-Authors: Wonjun Kang (AI Research Engineer) and Heeju Kim (Software Engineer)
Collaborating Organizations: UW-Madison, Microsoft Research, SNU, and KRAFTON
LoSA: Locality-aware sparse attention for block-wise diffusion language models
Overview Solves the key "KV inflation" memory bottleneck in diffusion language models (DLMs) by identifying active versus stable tokens across denoising steps. It reuses cached prefix attention for stable tokens and applies sparse attention exclusively to active tokens.
Why It Matters: Enables non-autoregressive, block-wise diffusion models to scale to long-context generation without memory-bound speed degradation, unlocking practical non-autoregressive serving alternatives
Furiosa Co-Authors: Minjae Lee (AI Research Engineer) and Wonjun Kang (AI Research Engineer)
Collaborating Organizations: UC Berkeley and UT Austin
AsyncOPD: How stale can on-policy distillation be?
Overview: Establishes an asynchronous training pipeline for on-policy distillation (OPD) by decoupling rollout generation from learner updates. By analyzing staleness dynamics and optimizing reverse-KL estimators under finite teacher-score caches, it achieves 1.6×–3.8× higher training throughput over strict synchronous baselines.
Why It Matters: Eliminates synchronous idle bottlenecks during teacher-student distillation, dramatically speeding up post-training pipelines for reasoning models
Furiosa Co-Authors: Wonjun Kang (AI Research Engineer), Kevin Galim (AI Research Engineer), Seunghyuk Oh (AI Research Engineer), Donghoon Kim (AI Research Engineer), Minjae Lee (AI Research Engineer), Minseo Kim (AI Research Engineer), and Hyung Il Koo (Chief Algorithm Officer)
Collaborating Organizations: Ajou University, UC Berkeley, Microsoft Research, KRAFTON, and Ludo Robotics
EfficientRollout: System-aware self-speculative decoding for RL rollouts
Overview:Accelerates reinforcement learning (RL) post-training through system-aware speculative decoding. It combines a 4-bit quantized self-drafting model, a system-aware speculative decoding toggle, and dynamic adaptation of draft lengths to evolving target policies, thereby speeding up rollouts by up to 19.6% and end-to-end training by 12.7% losslessly.
Why It Matters: Overcomes dynamic policy drift and batch-shrinking bottlenecks during RL, making large-scale post-training significantly faster and more compute-efficient
Furiosa Co-Authors: Minseo Kim (AI Research Engineer), Minjae Lee (AI Research Engineer), Seunghyuk Oh (AI Research Engineer), Kevin Galim (AI Research Engineer), Donghoon Kim (AI Research Engineer), Hyung Il Koo (Chief Algorithm Officer), and Wonjun Kang (AI Research Engineer)
Collaborating Organizations: UC Berkeley
Happy hour with ICML attendees
During ICML week, Furiosa welcomed the community to our Gangnam HQ for a casual happy hour with visiting researchers, system architects, engineers, and collaborators. It was exciting to open our doors to more than 100 attendees and create a space for meaningful conversations.

As we continue building our full-stack ecosystem, we are committed to bridging foundational AI research with production-grade execution. Events like these are an important part of fostering collaboration across academic and industry, and we’re grateful to everyone who joined us to exchange ideas and build new connections.
We look forward to future research conferences and community gatherings as we continue working toward our mission of making AI truly sustainable.




Written by
The Furiosa Team




