
A Computer architecture built around how AI models compute
Why the abstraction layer matters
Inference efficiency is a hardware-software problem

When hardware and model don’t match, data spends more time moving than being computed on, and capex buys idle silicon. The fix starts at the abstraction layer between software and hardware.
The abstraction layer defines the trade-offs

The abstraction is a contract: it decides what the silicon optimizes on its own and what the compiler has to arrange across the whole graph. For modern models, that contract has to handle multi-dimensional data layouts.
Balancing this contract has forced a zero-sum compromise across three vertices:
- Generality: supports any compute pattern, at the cost of control logic and data movement.
- Performance: Maximizing throughput for a static point in time, but inducing architectural obsolescence as models evolve.
- Efficiency: optimizes individual hardware blocks, but pushes tiling and scheduling complexity onto the compiler.
Forcing multi-dimensional tensors into fixed 2D grids costs data movement at every layer. Reaching the balanced point at the center of the diagram takes a primitive designed for the shape of AI compute.
Tensor contraction, the right abstraction for AI compute

Matrix-multiply accelerators map these tensors onto fixed 2D grids, so the compiler has to flatten, slice and permute them in memory. That adds layout work and extra data movement at every step.
The Tensor Contraction Processor (TCP) resolves this compromise by raising the hardware execution primitive. By accelerating the full tensor contractions as a single, unified hardware operation, the TCP matches the mathematical expression of the workload directly in silicon—treating the tensor operation as a first-class citizen.

THE GEOMETRY OF AI COMPUTE
Less data movement, higher utilization

The TCP runs multi-dimensional operations directly on-chip, keeps data in local SRAM, and minimizes movement to and from memory.
Driving global compiler optimization

On the TCP, execution is structured and predictable, so parallelism and data routing can be planned at compile time. The Furiosa Compiler treats the whole network graph as one optimization problem and minimizes data movement end to end.
Automatic optimization across shapes, fusions and chip generations


The Furiosa Compiler searches layout, tiling and fusion automatically, so most models run well without hand-written kernels and carry over to new chip generations. When you want finer control, TCL and vISA give you kernel-level access.
Flexible dataflow adaptation for diverse tensor shapes
The TCP redefines generality through micro-architectural plasticity, deploying flexible execution units anchored by fluid dot-product primitives. By integrating these primitives directly inside the compute blocks, the hardware exposes multiple physical pathways for data flow. The compiler can dynamically configure these lanes to modify data-reuse vectors in real time, maintaining peak hardware utilization across diverse tensor layouts.

The sweet spot for inference workloads

The Tensor Contraction Processor sits between the two: programmable enough for changing models and agentic workloads, and efficient because it runs tensor contraction natively. By mapping the higher-dimensional geometry of the tensor directly onto the silicon, the TCP unifies software-programmable dataflow generality with the deterministic power efficiency of dedicated primitives.

MEET RNGD, BUILT ON TCP.
Blog
.avif)
FuriosaAI and Samsung SDS Launch Korea’s First Domestic NPUaaS to Expand Enterprise AI Access

Experience RENEGADE Summit 2026
