FuriosaAI at ICML 2026: Advancing full-stack software efficiency

News
August 6, 2026

Summary

Written by

The Furiosa Team

Share this article

No items found.

FuriosaAI presented four accepted papers at ICML 2026 in Seoul this summer. Featured in recent IT press coverage, this research highlights Furiosa’s software stack innovation, demonstrating that high-performance, energy-efficient AI acceleration requires deep co-optimization across model architectures, serving pipelines, and hardware execution. All four papers were written by the company’s AI Research Group, which is part of the Algorithm Team at Furiosa. 

ReJump: Tree-jump representation for LLM reasoning

Overview: Models LLM reasoning processes as a tree structure with non-adjacent transitions ("jumps") to capture complex execution behaviors such as backtracking, verification, and calculation. By systematically analyzing these execution paths, ReJump improves Best-of-N selection strategies by up to 9.1% on reasoning tasks.

Why It Matters: Moves LLM evaluation beyond simple final-accuracy metrics, providing a structured framework to diagnose how reasoning models explore, overthink, or verify intermediate steps

Furiosa Co-Authors: Wonjun Kang (AI Research Engineer) and Heeju Kim (Software Engineer)

Collaborating Organizations: UW-Madison, Microsoft Research, SNU, and KRAFTON

Read the paper

LoSA: Locality-aware sparse attention for block-wise diffusion language models

Overview Solves the key "KV inflation" memory bottleneck in diffusion language models (DLMs) by identifying active versus stable tokens across denoising steps. It reuses cached prefix attention for stable tokens and applies sparse attention exclusively to active tokens.

Why It Matters: Enables non-autoregressive, block-wise diffusion models to scale to long-context generation without memory-bound speed degradation, unlocking practical non-autoregressive serving alternatives

Furiosa Co-Authors: Minjae Lee (AI Research Engineer) and Wonjun Kang (AI Research Engineer)

Collaborating Organizations: UC Berkeley and UT Austin

Read the paper

AsyncOPD: How stale can on-policy distillation be?

Overview: Establishes an asynchronous training pipeline for on-policy distillation (OPD) by decoupling rollout generation from learner updates. By analyzing staleness dynamics and optimizing reverse-KL estimators under finite teacher-score caches, it achieves 1.6×–3.8× higher training throughput over strict synchronous baselines.

Why It Matters: Eliminates synchronous idle bottlenecks during teacher-student distillation, dramatically speeding up post-training pipelines for reasoning models

Furiosa Co-Authors: Wonjun Kang (AI Research Engineer), Kevin Galim (AI Research Engineer), Seunghyuk Oh (AI Research Engineer), Donghoon Kim (AI Research Engineer), Minjae Lee (AI Research Engineer), Minseo Kim (AI Research Engineer), and Hyung Il Koo (Chief Algorithm Officer)

Collaborating Organizations: Ajou University, UC Berkeley, Microsoft Research, KRAFTON, and Ludo Robotics

Read the paper

EfficientRollout: System-aware self-speculative decoding for RL rollouts

Overview:Accelerates reinforcement learning (RL) post-training through system-aware speculative decoding. It combines a 4-bit quantized self-drafting model, a system-aware speculative decoding toggle, and dynamic adaptation of draft lengths to evolving target policies, thereby speeding up rollouts by up to 19.6% and end-to-end training by 12.7% losslessly.

Why It Matters: Overcomes dynamic policy drift and batch-shrinking bottlenecks during RL, making large-scale post-training significantly faster and more compute-efficient

Furiosa Co-Authors: Minseo Kim (AI Research Engineer), Minjae Lee (AI Research Engineer), Seunghyuk Oh (AI Research Engineer), Kevin Galim (AI Research Engineer), Donghoon Kim (AI Research Engineer), Hyung Il Koo (Chief Algorithm Officer), and Wonjun Kang (AI Research Engineer)

Collaborating Organizations: UC Berkeley

Read the paper

Happy hour with ICML attendees

During ICML week, Furiosa welcomed the community to our Gangnam HQ for  a casual happy hour with visiting researchers, system architects, engineers, and collaborators. It was exciting to open our doors to more than 100 attendees and create a space for meaningful conversations. 

As we continue building our full-stack ecosystem, we are committed to bridging foundational AI research with production-grade execution. Events like these are an important part of fostering collaboration across academic and industry, and we’re grateful to everyone who joined us to exchange ideas and build new connections. 

We look forward to future research conferences and community gatherings as we continue working toward our mission of making AI truly sustainable. 

Written by

The Furiosa Team

Share this article

white dot background graphic

Get the latest updates on FuriosaAI

Thanks for submitting the form.