Abstract grayscale digital art depicting flowing, curved lines and scattered small squares creating wave-like patterns on a black background.

Serving a 100B-Class Model at 128K Context: Production Traffic on One of Korea's Largest Portals

September 2026
SHARE THIS ARTICLE

Summary

Partner / Customer

Upstage

Industry

AI — Foundation Models

Workload / model

Solar Pro 3 (100B-class, 128K context, NVFP4)

Deployment

3× NXT RNGD Server (Daum production)

For a foundation-model company, inference cost is COGS. Expanding Solar's API and service footprint meant finding an alternative to GPU serving economics — and proving a 100B-class model could hold production quality at 128K context on domestic silicon.

power efficiency vs. H200

1.3×

scale-up requested post-launch

10×

context length in production

128K

Power Efficiency

1.3×

H200 server

1.3x

Performance per watt (TPS/W) versus H200, Solar Pro 3 production serving

Upstage builds the Solar family of foundation models and serves them through its API and through partner products like Daum AI Summary, a search-result summarization feature on one of Korea's largest consumer portals. Expanding that footprint meant proving a 100B-class model — Solar Pro 3, at 128K context — could hold production quality and cost discipline on RNGD.FuriosaAI validated the serving configuration on furiosa-llm, shared an offline throughput profile before go-live so expected performance was agreed in advance, and exposed an OpenAI-compatible endpoint so Upstage's application integrated by changing only the endpoint URL.

Model / artifact
solar-enkoja-100b-128k-3.1.0-dpo.5-nvfp4
Serving config
TP=4 / DP=2 across 8 RNGD cards
API compatibility
OpenAI-compatible /v1/chat/completions, endpoint-only integration
Family expansion
Solar Embedding supported; Solar Mini 4 targeted late August

Challenge

Can a 100B-Class Model Hold Production Quality at 128K?

Inference cost is COGS for a foundation-model company. Expanding Solar's API and service footprint required an alternative to H200 serving economics, and the "Korean full stack" position — domestic model plus domestic silicon — is a real differentiator in public-sector and enterprise deals.

The technical gate was concrete: can a 100B-class model at 128K context be served at production quality, on infrastructure Upstage could stand behind for a consumer-scale portal?

Solution

128K Context, Weights Compressed to Make Room for KV Cache

Solar Pro 3's 100B-class weights compress to NVFP4 to fit TP=4 within the 48GB-HBM3-per-card budget — freeing headroom that gets allocated to the real bottleneck at 128K context: KV cache.

Quantization, Memory, and Serving Topology

100B-class weights compress to NVFP4 to fit TP=4 (4 cards) within the 48GB-HBM3-per-card budget. One 8-card server runs as TP=4 × DP=2 — two model replicas for throughput scaling plus per-instance fault isolation. Validation was shared as an offline-throughput profile (CSV) on furiosa-llm 2026.3.0rc3 before go-live, so expected performance was agreed in advance.

API Compatibility and SDK Co-Optimization

An OpenAI-compatible /v1/chat/completions endpoint is exposed as-is — Upstage's application integrates by changing only the endpoint URL, with reasoning-content format and usage fields (including cached_tokens) verified compatible.
Solar Pro 3 entered SDK 2026.4's performance-optimization targets, so each SDK release flows directly into service performance; Solar Embedding is newly supported, and a Solar Mini 4 lightweight variant is targeted for late-August integration.

Target Workload

API Serving and Portal-Scale Summarization

Case 1
Solar Pro 3, 100B-class, 128K context, responses including reasoning tokens — API product requirements.
Case 2
Daum AI Summary analyzes web documents to generate search-result summaries. Portal search means high concurrency, short TTFT, and steep hour-by-hour traffic swings.

Economics

Production Traffic, Then a 10× Scale-Up Request

Commercial traffic from one of Korea's largest consumer portals now runs on a domestic model plus domestic NPU — validation well beyond a proof of concept, measured at roughly 1.3× power efficiency versus H200.

After operational stabilization, Upstage requested a 10× scale-up — entering capacity planning is itself the quantitative trust signal. Additional server allocation (3–4 units) is being evaluated against traffic growth.

Development

Built for the 24-Hour Power Curve

Portal traffic spends long periods at low load, and RNGD's low-load power behavior — low-100s of watts per card near idle — contributes materially to cumulative TCO. This workload is sold on the 24-hour power curve, not peak performance.

With 128K long-context requests mixed in, KV-pool management is the discipline: prefix-cache hit rate is tracked as an operations metric.

VISION

A Three-Party Model for the Korean AI Stack

Infrastructure, model, and application — presented together in a public joint webinar — is the story Upstage and FuriosaAI are building toward: a domestic full stack that doesn't ask customers to trade capability for sovereignty.
Day-N RNGD support for new Solar releases is now standing process, with parallel discussions underway on scaled serving through Samsung SDS's NPUaaS.

SHARE THIS ARTICLE
​
Download story PDF