Abstract grayscale digital art depicting flowing, curved lines and scattered small squares creating wave-like patterns on a black background.

Five EXAONE Generations, One Validated Path: at ~2.3× the Power Efficiency of A100

September 2026
SHARE THIS ARTICLE

Summary

Partner / Customer

LG AI Research

Industry

AI Research Institute

Workload / model

EXAONE 3.0→4.5 & K-EXAONE 236B-A23B (MoE)

Deployment

2× NXT RNGD Server (16 RNGD cards)

Running a domestic foundation model entirely on foreign GPUs leaves the supply chain only half domestic — and for an always-on service, that dependency compounds in power and floor cost every day it stays in production.

power efficiency vs. A100

2.3×

EXAONE generations validated

5

issue-to-fix response cycle

<24h

Power Efficiency

2.3x

A100 server

2.3×

Performance per watt (TPS/W) versus A100, 32B-class FP8 serving

LG AI Research builds the EXAONE family of large language models and operates ChatEXAONE, a group-wide generative AI assistant used across LG's business units. As each EXAONE generation shipped — 3.0 through 4.5, and the 236B-parameter K-EXAONE mixture-of-experts model — every version needed to serve on RNGD without accuracy loss, at a measurable efficiency advantage over its A100 baseline.FuriosaAI validated the full generational lineage on 2× NXT RNGD Server, then carried the same validation assets — NVFP4 artifacts, parser configuration, monitoring templates — into LG U+, LG CNS, and three external consortium engagements, turning one validation into a reference architecture for the whole EXAONE family.

Challenge

A Domestic Model, Still Running on Foreign Silicon

ChatEXAONE is an always-on internal service, so A100-based inference accumulates power and floor cost linearly — the longer the assistant stays live, the more expensive the gap between the model's origin and its infrastructure becomes.LG AI Research needed every EXAONE generation served without accuracy loss, needed the 236B-parameter K-EXAONE MoE model covered alongside the smaller dense models, and needed a quantified efficiency delta against the A100 baseline before committing production traffic to RNGD.

PRECISION
FP8 (32B-class); NVFP4A16 for K-EXAONE 236B-A23B
PARALLELISM
K-EXAONE: DP=1 / TP=32 / PP=1 across 4 RNGD cards
LIVE SERVING
~40.2 tok/s single-request generation, K-EXAONE 236B (low-load condition)
GROUP EXPANSION
LG U+ national AI program; LG CNS OpenStack + Kubernetes PoC

‍

Solution

A Validated, Reproduction-Grade Launch Configuration

FuriosaAI built a serving configuration precise enough that any of the five EXAONE generations — dense or MoE — could be relaunched from the same recipe, with quantization, parallelism, and parser configuration locked down in advance.

Quantization and Parallelism

K-EXAONE 236B-A23B serves in NVFP4A16 (4-bit weights, 16-bit activations), a pre-built .fxb artifact that fits the 236B MoE model inside the 4-card, 192GB HBM3 budget of one RNGD server. Parallelism is DP=1 / TP=32 / PP=1 — all 32 PEs across 4 cards bound as a single tensor-parallel group, with a second server running an independent identical instance for redundancy. The 32B-class generations (3.0 through 4.5) run in FP8 with TP=4 within a single server.

Parsers, Observability, and Generation Transitions

An exaone4 reasoning parser and a hermes tool-call parser, plus auto tool choice, handle reasoning-token separation and function-call schema parsing server-side — no client changes needed for agent integrations.

Prometheus and Grafana track KV-cache usage, prefix-cache hit rate, and wire-pipeline hit rate as the operating standard. Every generation transition — most recently EXAONE 4.5-33B entering SDK 2026.4 — triggers the same sequence: re-run the PTQ pipeline, run a quality-regression check, and re-release the artifact.

Target Workload

Internal Assistant to Agentic MoE

The 32B-class generations (3.0 through 4.5) serve FP8 internal-assistant traffic — summarization, QA, coding, retrieval — at TP=4 within a single server.

K-EXAONE 236B-A23B is a different workload class: an MoE model running reasoning and tool calling simultaneously, sized for agent workloads rather than single-turn chat.

Economics

One Validation, Reused Across Five Engagements

The FP8 serving path for the 32B-class EXAONE generations measures at roughly 2.3× throughput per watt versus the A100 baseline — the efficiency delta LG AI Research needed before committing ChatEXAONE traffic to RNGD.

Because the validation assets were built once and reused as-is, the same engineering investment carried directly into the Aramco consortium, ESTsecurity, and Alphacode proofs of concept — compounding the return on a single validation effort.

Development

From ChatEXAONE to the LG Group

LG U+'s national "AI for Everyone" program is sizing RNGD card counts now, against concrete targets: GPT-OSS-120B and EXAONE 4.5-33B in FP8, 30–35 TPS/user interactivity, 256 concurrency, and a 4K aggregate output-TPS target.

LG CNS has completed its first PoC — RNGD integrated into OpenStack (Kolla) and Kubernetes (Kubespray) at 4 cards per node, with VM PCIe passthrough, NUMA, and hugepages configured. A third-party validator confirmed NUMA-aligned VM performance matches native Kubernetes Pod performance.

VISION

A Reference Architecture for the EXAONE Family

What began as validating one model family on domestic silicon has become the standard onboarding path for every new EXAONE generation — and the template other consortium partners now build on.
Each new generation, from 3.0 through the 236B-parameter K-EXAONE MoE, extends the same reference architecture rather than starting from zero.

SHARE THIS ARTICLE
​
Download story PDF