White Paper: Benchmarking RNGD on Backend.AI for 1.3–1.5x greater efficiency in enterprise AI inference

Technical Updates
August 6, 2026

Summary

Written by

Share this article

No items found.

FuriosaAI has partnered with AI infrastructure platform provider Lablup to publish a new white paper evaluating an enterprise AI inference stack that combines Furiosa's RNGD inference accelerator with Lablup's Backend.AI orchestration platform.

The white paper, titled "RNGD meets Backend.AI," provides detailed benchmark data comparing four RNGD accelerators on Backend.AI against a control setup of four NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs running on bare metal.

This chart compares throughput per Watt when running Qwen3-32B.

The evaluation focuses on critical real-world serving metrics, including throughput, latency, power efficiency, and scalability. Testing with the Qwen3-32B (FP8) model demonstrated that RNGD delivers competitive throughput—reaching 95% of the comparison setup at peak concurrency (256 requests)—while drawing 30% to 44% less power. 

The Backend.AI portal.

Across all measured concurrency levels, RNGD achieved 1.3x to 1.5x higher throughput per watt. RNGD also excelled in user responsiveness, keeping Time to First Token (TTFT) under 1 second through concurrency 32 (compared to over 2.9 seconds for the control group).

Optimizing AI infrastructure through hardware and software integration

Beyond benchmark performance, the white paper explores how RNGD and Backend.AI work together to simplify the deployment and management of production-scale AI inference workloads. 

Backend.AI is designed to manage heterogeneous accelerated computing environments, including GPUs, NPUs, and other AI accelerators through a unified platform. The white paper addresses pressing operational challenges, including air-cooling constraints, rising datacenter power density, and Sovereign AI compliance. It outlines practical deployment strategies for managing mixed LLM workloads (such as prefill and decode phases) using Backend.AI’s session execution model and Sokovan orchestrator.

Access the white paper

Read the full findings and deployment guide by downloading the white paper on the Backend.AI site here. A Korean version of the report is also available here. We would like to thank the Backend.AI team for their collaboration and expertise in preparing the report.

About Backend.AI

Backend.AI is a vendor-agnostic accelerated workload hosting platform based on our own home-grown orchestration and job scheduler, running on top of either cloud or on-premises (air-gapped) clusters.

Written by

Share this article

white dot background graphic

Get the latest updates on FuriosaAI

Thanks for submitting the form.