Furiosa RNGD AI accelerator card in red with specifications highlighting TFLOPS, HBM3 memory, SRAM, and memory bandwidth.
Tensor Contraction Processor (TCP)

A Computer architecture built around how AI models compute

​
Download TCP Paper

Why the abstraction layer matters

Inference efficiency is a hardware-software problem

Diagram highlighting the intersection of AI algorithms, software, and hardware, demonstrating FuriosaAI’s co-design approach to maximizing inference performance, efficiency, and scalability.
Inference-time scaling means more tokens per request and more varied workloads. At that scale, throughput and cost depend less on raw compute and more on how well the model, the compiler and the hardware work together.

When hardware and model don’t match, data spends more time moving than being computed on, and capex buys idle silicon. The fix starts at the abstraction layer between software and hardware.

The abstraction layer defines the trade-offs

AI architecture diagram illustrating FuriosaAI’s approach to balancing performance, efficiency, and programmability through optimal abstraction and hardware-software-algorithm co-design.

The abstraction is a contract: it decides what the silicon optimizes on its own and what the compiler has to arrange across the whole graph. For modern models, that contract has to handle multi-dimensional data layouts.

‍
Balancing this contract has forced a zero-sum compromise across three vertices:

  • Generality: supports any compute pattern, at the cost of control logic and data movement.
  • Performance: Maximizing throughput for a static point in time, but inducing architectural obsolescence as models evolve.
  • Efficiency: optimizes individual hardware blocks, but pushes tiling and scheduling complexity onto the compiler.

Forcing multi-dimensional tensors into fixed 2D grids costs data movement at every layer. Reaching the balanced point at the center of the diagram takes a primitive designed for the shape of AI compute.

Tensor contraction, the right abstraction for AI compute

AI architecture diagram showing tensor contraction as the primary computational workload in transformer models, linking tensor data structures, neural network layers, and weighted activation functions used in modern AI inference.
AI architecture diagram showing tensor contraction as the primary computational workload in transformer models, linking tensor data structures, neural network layers, and weighted activation functions used in modern AI inference.
Deep learning operations do not occur in flat, two-dimensional structures. In transformers, weights and activations are multi-dimensional tensors, and the dominant operation is tensor contraction: matrix multiplication generalized to higher dimensions.

Matrix-multiply accelerators map these tensors onto fixed 2D grids, so the compiler has to flatten, slice and permute them in memory. That adds layout work and extra data movement at every step.

The Tensor Contraction Processor (TCP) resolves this compromise by raising the hardware execution primitive. By accelerating the full tensor contractions as a single, unified hardware operation, the TCP matches the mathematical expression of the workload directly in silicon—treating the tensor operation as a first-class citizen.
Close-up of a computer chip with blue, black, and pink circuit details visible.

THE GEOMETRY OF AI COMPUTE

Less data movement, higher utilization

Architectural comparison between Furiosa TCP and traditional GPU designs, highlighting how reduced data movement, shared SRAM access, and coordinated execution improve AI inference efficiency and performance.
Architectural comparison between Furiosa TCP and traditional GPU designs, highlighting how reduced data movement, shared SRAM access, and coordinated execution improve AI inference efficiency and performance.
In production inference, data movement is the biggest lever on performance per watt. On GPUs, fixed 2D execution grids mean the software stack keeps reshaping and shuffling data across caches and register files, which adds latency and power.

The TCP runs multi-dimensional operations directly on-chip, keeps data in local SRAM, and minimizes movement to and from memory.

Driving global compiler optimization

Visualization of Furiosa’s compiler technology that maps tensors across distributed SRAM and compute units, optimizing data layouts and execution paths to improve AI inference performance, utilization, and power efficiency.
Eliminating runtime data orchestration overhead requires an architectural contract that substitutes dynamic hardware scheduling with static, compile-time layout optimization. Legacy architectures introduce hardware-controlled caching and dynamic thread scheduling, creating non-deterministic execution behaviors that force the compiler into reactive mitigation.

On the TCP, execution is structured and predictable, so parallelism and data routing can be planned at compile time. The Furiosa Compiler treats the whole network graph as one optimization problem and minimizes data movement end to end.

Automatic optimization across shapes, fusions and chip generations

Operators
Diagram illustrating the large number of AI operators and optimization categories that must be managed across modern machine learning frameworks.
​
Tensor shapes
Abstract visualization of operator fusion, combining multiple AI operations into optimized execution blocks.
​
Fusion
Technical diagram showing examples of operator fusion patterns used to optimize AI model execution.
​
Generations
Diagram showing four generations of AI accelerator hardware, highlighting ongoing architectural evolution and optimization requirements.
The legacy inference software ecosystem is bottlenecked by fragile, hand-crafted operator libraries that rely on manual kernel tuning for isolated tensor geometries and model variations. This heuristic approach fails at scale due to the multivariate combinatorial expansion of dynamic tensor shapes, operator fusions, and advancing micro-architectural revisions. Manual optimization cannot sustain hardware utilization across this multi-dimensional variant space.

The Furiosa Compiler searches layout, tiling and fusion automatically, so most models run well without hand-written kernels and carry over to new chip generations. When you want finer control, TCL and vISA give you kernel-level access.

Flexible dataflow adaptation for diverse tensor shapes

Diagram showing TCP architecture with distributed memory and parallel compute units processing data simultaneously.
Diagram showing TCP reconfiguring compute resources to efficiently process larger tensor operations through parallel dot products.
Diagram showing a TPU’s fixed 128×128 systolic array optimized for structured matrix multiplication workloads.
Fixed-size systolic arrays such as Tensor Processing Units (TPUs) deliver high density but are bound to rigid, fixed-dimension physical systolic arrays built for uniform matrix multiplication. When subjected to the highly variable, asymmetric tensor shapes characteristic of production inference, these rigid grids suffer from massive structural under-utilization and bubble injection.

The TCP redefines generality through micro-architectural plasticity, deploying flexible execution units anchored by fluid dot-product primitives. By integrating these primitives directly inside the compute blocks, the hardware exposes multiple physical pathways for data flow. The compiler can dynamically configure these lanes to modify data-reuse vectors in real time, maintaining peak hardware utilization across diverse tensor layouts.
Two 3D cubes, one light and one dark, separated by a multiplication sign on a black background.

The sweet spot for inference workloads

Diagram comparing GPU, TPU, and Furiosa TCP architectures, highlighting how Furiosa’s inference-optimized Tensor Contraction Processor delivers both high power efficiency and flexible dataflow for AI inference workloads.
Chip design has long traded one thing for another: GPUs are general but spend energy on control and data movement; fixed-function arrays are efficient but rigid.

The Tensor Contraction Processor sits between the two: programmable enough for changing models and agentic workloads, and efficient because it runs tensor contraction natively. By mapping the higher-dimensional geometry of the tensor directly onto the silicon, the TCP unifies software-programmable dataflow generality with the deterministic power efficiency of dedicated primitives.

MEET RNGD, BUILT ON TCP.

​
Explore RNGD

Blog

​
See all posts

FuriosaAI and Samsung SDS Launch Korea’s First Domestic NPUaaS to Expand Enterprise AI Access

News
​
FuriosaAI and Samsung SDS Launch Korea’s First Domestic NPUaaS to Expand Enterprise AI Access

Experience RENEGADE Summit 2026

News
​
Experience RENEGADE Summit 2026

RNGD outperforms RTX Pro 6000 with the latest SDK

Technical Updates
​
RNGD outperforms RTX Pro 6000 with the latest SDK