COLM
2026

Scaling Test-Time Compute for DLLMs via Parallel Search

Saarth Gaonkar, Sebastian Zhao, Minjae Lee, Coleman Hooper, Wonjun Kang, Michael W. Mahoney, Kurt Keutzer, Chenfeng Xu, Amir Gholami
Workshop
Accepted
diffusion-llm
search
parallel-decoding

Abstract

Diffusion Large Language Models (DLLMs) have emerged as a promising paradigm for parallel and any-order text generation. Unlike autoregressive (AR) LLMs, which decode one token at a time from left to right, DLLMs can reveal multiple token positions in arbitrary order, opening a broad design space for training-free test-time scaling. In this work, we study how to scale parallel search for DLLMs under practical accuracy–throughput constraints by jointly considering algorithmic accuracy and system efficiency. We first formalize design decisions for DLLM parallel search, including search granularity, scoring, compute allocation, unmasking policy, block size, and parallel denoising, and evaluate existing parallel search methods applied to DLLMs. Our results show that standard confidence-based unmasking can reduce trajectory diversity, while diversity-aware unmasking preserves trajectory diversity and can improve accuracy in model-dependent settings. Finally, we empirically demonstrate how test-time compute scaling behaves across the dimensions under joint algorithm–system constraints, pushing the DLLM test-time scaling frontier under realistic serving conditions.

‍

Related Publications

Scaling Test-Time Compute for DLLMs via Parallel Search

COLM
2026
diffusion-llm
search
parallel-decoding
​
View Job

Characterizing High Bandwidth Flash for LLM Serving

arXiv
2026
kv-cache
long-context
​
View Job

AsyncOPD: How Stale Can On-Policy Distillation Be?

NeurIPS
2026
reasoning
distillation
asynchronous-execution
​
View Job