Build at any layer, 
from model to silicon

​
Go to Furiosa Docs

Serving

FURIOSA-LLM

furiosa-llm, drop-in vLLM replacement with peak performance

Point the client you already have at a new endpoint and serve any of pre-compiled models — no rewrite, no proprietary API to learn.
​
Go to Furiosa-LLM Docs

shell

pip install furiosa-llm
furiosa-llm serve [model]
# Serving on :8000 (OpenAI-compatible)
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1"
)
client.chat.completions.create(model="[model]", ...)
from furiosa_llm import LLM

llm = LLM(
    model="[model]",
    dtype="[precision]",
)
out = llm.generate(prompts)
​
Go to Furiosa-LLM Docs
Model Zoo

Start from a model that already runs

OPENAI
Model zoo icon

gpt-oss

4B · 8B · 32B

FP8

Text generation

QWEN

Qwen3

4B · 8B · 32B

FP8

Text generation

UPSTAGE
star icon

Solar Open

4B · 8B · 32B

FP8

Text generation

Qwen3

4B · 8B · 32B

FP8

TEXT GENERATION

Qwen3-Coder

30B-A3B

FP8

CODE GENERATION

Qwen3-Embedding

0.6· 4B· 8B

BF16

TEXT GENERATION

gpt-oss

20B· 120B

MXFP4

TEXT GENERATION

Solar Open

100B

NVFPA16

TEXT GENERATION

EXAONE 4.5

33B

FP8

Vision-Language

EXAONE 3.5

7.8B· 32B

BF16

TEXT GENERATION

Llama 3.3

70B

FP8· INT8· FP16

Text Generation

DeepSeek-R1-Distill

7B· 8B· 14B· 32B· 70B

BF16

Reasoning

Qwen2.5

0.5B· 7B· 14B· 32B

BF16

TEXT GENERATION

BGE-M3

0.6B

BF16

TEXT GENERATION

E5-Mistral

7B

FP16

EMBEDDING

Qwen3 MoE

30B -A3B

FP8

TEXT GENERATION

Qwen3-VL

2B · 4B · 32B

BF16

Vision-Language

Qwen3-Reranker

0.6B· 4B· 8B

BF16

Reranking

Gemma 4

31B

FP8

TEXT GENERATION

K-EXAONE

236B-A23B

NVFPA16

TEXT GENERATION

EXAONE 4.0

32B

FP8

Text Generation

Mistral NeMo

12B

FP8

TEXT GENERATION

Llama 3.1

8B

FP8· BF16

TEXT GENERATION

QwQ

32B

BF16

Reasoning

Qwen2.5-Coder

&B· 14B· 32B

BF16

CODE GENERATION

BGE-Reranker-v2-M3

0.6B

BF16

RERANKING

Harrier-OSS-v1

0.6B

BF16

RERANKING

DEPLOYING

EASY DEPLOYMENT

Pick any deployment path that fits your needs

Gray square icon with three horizontal server-like bars and the text 'Bare Metal' underneath on a black background with a subtle grid pattern.
BARE METAL

A dedicated NXT RNGD Server, hosted at a data center near you.

Gray square block with a cloud icon embossed above the text 'API Endpoint' on a black grid background.
API ENDPOINT

Call models served on RNGD through our cloud partners. Nothing to install.

PROGRAMMING

FURIOSA-TORCH

Move your model on RNGD through furiosa-torch

Bring a model that is not in the zoo. Furiosa-Torch is a native PyTorch backend: move the model to the ‘rngd’ device, or compile it with torch.compile, and the Furiosa compiler takes it from the graph down.
​
Learn more
① Move to the device

python

import torch, furiosa.torch

device = torch.device("furiosa:0")
model = MyModel().to(device)
② Run it your way — three paths, one backend

Eager

torch.compile

Ahead-of-time

out = model(x.to(device))                  # eager: ops run as dispatched

compiled = torch.compile(model, backend=furiosa.torch.backend)
out = compiled(x.to(device))               # whole-graph compile

runnable = CompileModule.from_module(model, inputs)  # AOT → Runnable
​
Learn more
FURIOSA-TCL  /  FURIOSA-VISA

Apply kernel programming at two altitudes

COMING Q4 2026

TCL

Tensor Contraction Language

Declare what to compute over named symbolic axes. Kernels are generic over shape; the compiler owns tiling, layout and scheduling.

python

import furiosa.tcl as tcl

@tcl.kernel
def linear(x: tcl.bf16[M, K], w: tcl.bf16[N, K]) -> tcl.f32[M, N]:
    # axes missing from the output are contracted away
    return tcl.operation.contract(x, w, out=[M, N])
COMING Q4 2026

vISA

TCP Virtual ISA

Hand-optimize the hot path. Reason in tensors while directly managing memory placement and Tensor Unit scheduling.

rust

#[device(chip=1)]
fn gemm(a: DmTensor<bf16, m![M, K]>,
        b: DmTensor<bf16, m![K, N]>) -> DmTensor<f32, m![M, N]> {
    // you choose the memory tier and the schedule
    a.fetch().contract(b).commit()
}

Dive Deeper

​
TCP Paper, presented at ISCA 2024
Download the peer-reviewed architecture paper.
Read article
​
Keynote video by CTO and CRO
Watch a video where our engineering leads go in-depth of our HW and SW design philosophy.
Read article
​
Hot Chips Presentation by CEO June Paik
Watch the full presentation by our CEO at Hot Chips 2025.
Read article