Scaling Test-Time Compute for DLLMs via Parallel Search
Abstract
Diffusion Large Language Models (DLLMs) have emerged as a promising paradigm for parallel and any-order text generation. Unlike autoregressive (AR) LLMs, which decode one token at a time from left to right, DLLMs can reveal multiple token positions in arbitrary order, opening a broad design space for training-free test-time scaling. In this work, we study how to scale parallel search for DLLMs under practical accuracy–throughput constraints by jointly considering algorithmic accuracy and system efficiency. We first formalize design decisions for DLLM parallel search, including search granularity, scoring, compute allocation, unmasking policy, block size, and parallel denoising, and evaluate existing parallel search methods applied to DLLMs. Our results show that standard confidence-based unmasking can reduce trajectory diversity, while diversity-aware unmasking preserves trajectory diversity and can improve accuracy in model-dependent settings. Finally, we empirically demonstrate how test-time compute scaling behaves across the dimensions under joint algorithm–system constraints, pushing the DLLM test-time scaling frontier under realistic serving conditions.