Furiosa Algorithms Research

We work at the intersection of AI and hardware to make AI computing sustainable

Grid pattern with red, yellow, white dots and squares surrounding a large black irregular shape center-right.
RL Post-Training
Deploy the most capable models with strong latency and throughput.
Abstract digital art of colored dots and squares in red, blue, white, and black forming curved shapes.
Non-autoregressive Generation
Lower total cost of ownership with less energy, fewer racks, and air-cooled data centers of today.
Pixelated close-up of a red, white, black, and green Nike swoosh logo pattern.
Inference-time Algorithms
Stay future-proof for tomorrow’s models and transition with ease.

Publications

Scaling Test-Time Compute for DLLMs via Parallel Search

COLM
2026
Accepted
Workshop
diffusion-llm
search
parallel-decoding
​
View Job

Characterizing High Bandwidth Flash for LLM Serving

arXiv
2026
Published
kv-cache
long-context
​
View Job

AsyncOPD: How Stale Can On-Policy Distillation Be?

NeurIPS
2026
Accepted
Poster
reasoning
distillation
asynchronous-execution
​
View Job

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

NeurIPS
2026
Accepted
Poster
speculative-decoding
reinforcement-learning
rollout-generation
​
View Job

Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback

TMLR
2026
search
transformer
​
View Job

TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs

EACL
2026
speculative-decoding
vision-lanuage
​
View Job

LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models

ICML
2026
diffusion-llm
long-context
sparse-attention
​
View Job

Draft-based Approximate Inference for LLMs

ICLR
2026
speculative-decoding
kv-cache
​
View Job

ParallelBench: Understanding the Tradeoffs of Parallel Decoding in Diffusion LLMs

ICLR
2026
parallel-decoding
benchmark
​
View Job

Inference-Aligned SFT for Diffusion LLMs via Group-based Trajectory Sampling

ICLR
2026
Workshop
diffusion-llm
discrete-diffusion
sft
​
View Job

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization

2026
Preprint
kv-cache
quantization
​
View Job

Counting Guidance for High Fidelity Text-to-Image Synthesis

WACV
2025
text-to-image
diffusion
​
View Job

ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning

ICML
2025
reasoning
search
interpretability
​
View Job

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

ICML
2025
Oral
reward-model
reasoning
​
View Job

Parameter-Efficient Fine-Tuning of State Space Models

ICML
2025
ssm
peft
​
View Job

State-offset Tuning: State-based Parameter-Efficient Fine-Tuning for State Space Models

ACL
2025
ssm
peft
​
View Job

Eta Inversion: Designing an Optimal Eta Function for Diffusion-based Real Image Editing

ECCV
2024
diffusion
image-editing
​
View Job

Can MLLMs Perform Multimodal In-Context Learning for Text-to-Image Generation?

COLM
2024
text-to-image
in-context-learning
​
View Job
Long-term collaboration partners on LLM efficiency, quantization, parallel decoding, and other advanced research areas for efficient inference.
Logo of Bair Artificial Intelligence Research
Logo of University of Wisconsin Madison

Join our team

​
See Open Roles