arrow_backBack to research feed
efficiencyPublished: July 17, 2026

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

By Yuchen Yang, Yifan Zhao, Anisha Dasgupta, Sasa Misailovic

Research TL;DR

"Dynamically quantizes MoE expert weights at runtime, balancing precision and KV cache memory to improve serving throughput by up to 1.94x while maintaining FP16 accuracy."

Abstract

Mixture-of-Experts (MoE) is a popular class of large language models (LLMs), offering high efficiency and accuracy. However, in KV-cache-intensive serving scenarios, MoEs often exhibit a tension between the GPU memory requirements of the model weights and the growing KV cache. We propose PagedWeight, a novel management method for MoE LLM serving that dynamically quantizes MoE model's weights at runtime and balances expert-weight precision with the KV cache sizes. PagedWeight exposes and effectively navigates the complex tradeoff between the model's task accuracy, memory consumption, and throughput/latency. Across several memory-sensitive MoE serving scenarios, PagedWeight improves the quality-memory tradeoff over several existing quantization baselines. PagedWeight achieves FP16-equivalent accuracy with up to 72.0% GPU memory savings and 1.94$\times$ throughput improvement, and improves quality over quantization methods by up to 39.3% at a similar memory budget with at most 4.1% throughput loss.

Technical Analysis & Implementation

Overview§

PagedWeight is a runtime weight quantization method for Mixture-of-Experts (MoE) LLMs that dynamically adjusts precision per expert based on its importance during serving. It aims to reduce GPU memory pressure from model weights to accommodate the growing KV cache, thereby improving throughput without sacrificing quality.

Core Methodology§

PagedWeight operates at the granularity of individual experts in an MoE layer. During inference, it selects a target total memory budget $B$ for weights. The key idea is to assign higher precision to more important experts and lower precision to less important ones, all while keeping the total weight memory under $B$.

Quality-Aware Quantization§

The importance of an expert $e$ is measured by its contribution to the model's output quality. PagedWeight uses a lightweight quality proxy: the average routing probability over a calibration dataset. Let $p_e$ be the average routing score for expert $e$. Then the allocated precision $b_e$ (bits) is determined by solving an optimization problem:

$$ \min_{b_e} \sum_e p_e \cdot L(b_e) \quad \text{s.t.} \quad \sum_e b_e \cdot W_e \leq B, $$

where $L(b_e)$ is the expected quantization loss (e.g., mean squared error) for precision $b_e$, and $W_e$ is the number of weight elements in expert $e$. The solution leads to higher $b_e$ for experts with large $p_e$ or high sensitivity.

Dynamic Paging§

To enable runtime adjustments, PagedWeight introduces a paging mechanism similar to virtual memory. The GPU memory is divided into pages, and expert weights are quantized and stored in pages of varying bit-widths (e.g., 8-bit, 4-bit, 3-bit). During a forward pass, the router determines which experts are needed. PagedWeight fetches the corresponding pages, dequantizes them on-the-fly, and keeps only the active experts in full precision temporarily. Inactive expert pages can be swapped out to CPU memory or compressed further.

Implementation Details§

  • Quantization Scheme: Supports uniform affine quantization with per-expert scaling factors. For a given bit-width $b$, the quantization operation is:

$$Q(w) = \text{round}\left(\frac{w}{s}\right) \cdot s, \quad s = \frac{\max(|w|)}{2^{b-1} - 1}.$$

  • Calibration: Importance scores $p_e$ are computed once from a small calibration set. A calibration set can be reused across serving runs.
  • Memory Management: Uses a page table mapping expert IDs to physical page addresses. The scheduler dynamically adjusts page sizes based on current KV cache occupancy.

Code Illustration§

import torch

class PagedWeightExpert(torch.nn.Module):
    def __init__(self, expert_weight, importance_score):
        super().__init__()
        self.importance = importance_score
        # Store original weight for potential fine-tuning
        self.register_buffer('weight_fp16', expert_weight)
        self.page_size = 0  # in bytes, determined by allocator

def compute_budgeted_precision(importances, weights_size, budget_bytes):
    # Simplified: allocate bits proportional to importance
    total_importance = importances.sum()
    allocation = budget_bytes * importances / total_importance
    bits = allocation / weights_size
    # Clamp to available bit-widths
    bit_options = torch.tensor([3, 4, 8])
    idx = torch.bucketize(bits, bit_options)
    return bit_options[idx.clamp(max=2)]

# Usage:
# budgets = compute_budgeted_precision(importances, expert_params, target_memory)
# Then quantize each expert accordingly.

Experimental Results§

  • Memory Savings: Up to 72% GPU memory reduction vs. FP16 with equivalent accuracy on WikiText-2 perplexity.
  • Throughput: Up to 1.94x improvement over FP16 serving.
  • Quality vs. Memory: Outperforms uniform quantization baselines (e.g., GPTQ, SmoothQuant) by up to 39.3% in perplexity at similar memory budgets.

Summary§

PagedWeight dynamically trades off expert precision for KV cache space, enabling larger effective batch sizes and higher throughput without retraining. It is orthogonal to other compression techniques and can be combined with them.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window1,050,000 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.003720
Output Cost (Generated):$0.022320
Total Est. Cost:$0.026040
Context Window Capacity0.0118%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
GPT-5.5 Pro$30.00$180.00
o3 Mini$1.10$4.40
MiniMax M1$0.55$2.20
GPT-4o-mini$0.15$0.60
GLM 4.7 Flash$0.06$0.40
Gemini 2.5 Flash$0.30$2.50
GPT-5.2-Codex$1.75$14.00
o3 Pro$20.00$80.00
DeepSeek V3.1$0.25$0.95
GPT-4$30.00$60.00
Mistral Medium 3.1$0.40$2.00
Qwen Plus 0728 (thinking)$0.26$0.78
Claude Opus 4$15.00$75.00
o4 Mini$1.10$4.40
GPT-4.1 Mini$0.40$1.60
Claude Sonnet 5$2.00$10.00
Claude Sonnet 4.5$3.00$15.00
Claude Opus 4.7 (Fast)$30.00$150.00
Claude Opus 4.5$5.00$25.00
o1$15.00$60.00
GLM 4.5V$0.60$1.80
GPT-4o (2024-11-20)$2.50$10.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT-5 Chat$1.25$10.00
GPT-5 Nano$0.05$0.40
Mistral Large 2407$2.00$6.00
GPT Chat Latest$5.00$30.00
Claude Sonnet 4.6$3.00$15.00
gpt-oss-120b$0.04$0.17
GPT-5.3-Codex$1.75$14.00
MoonshotAI Kimi Latest$3.00$15.00
Qwen2.5 7B Instruct$0.04$0.10
Llama 3.2 3B Instruct$0.05$0.34
Gemini 3.1 Pro Preview$2.00$12.00
Qwen3.5 Plus 2026-02-15$0.26$1.56
Google Gemini Flash Latest$1.50$9.00
Claude Haiku 4.5$1.00$5.00
GPT-5 Mini$0.25$2.00
Mistral Nemo$0.02$0.03
Gemini 3.1 Flash$0.25$1.50
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-5.6 Luna Pro$1.00$6.00
GPT-5.6 Luna$1.00$6.00
Qwen2.5 72B Instruct$0.36$0.40
Command R (08-2024)$0.15$0.60
KAT-Coder-Air V2.5$0.15$0.60
GPT-4o (2024-05-13)$5.00$15.00
Llama 3.2 11B Vision$0.34$0.34
Llama 4 Maverick$0.20$0.80
KAT-Coder-Pro V2.5$0.74$2.96
Mixtral 8x22B Instruct$2.00$6.00
Mistral Large$2.00$6.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Llama 3 8B Instruct$0.14$0.14
Kimi K2.6$0.68$3.42
Llama 4 Scout$0.10$0.30
GPT-4 Turbo Preview$10.00$30.00
Claude 3 Haiku$0.25$1.25
GLM 5$0.95$2.55
Qwen3 30B A3B Instruct 2507$0.10$0.30
Gemini 2.0 Flash$0.10$0.40
MiniMax M2.7$0.25$1.00
GPT-5.4 Nano$0.20$1.25
GLM 4.5 Air$0.13$0.85
Qwen3 Coder 480B A35B$0.30$1.00
UI-TARS 7B$0.10$0.20
GPT-5.5$5.00$30.00
Mistral Small 4$0.15$0.60
GLM 5 Turbo$1.20$4.00
Qwen3 Coder Next$0.11$0.80
GPT-4o$2.50$10.00
Mistral Small 3$0.10$0.30
Command R+$2.50$10.00
Qwen3 Max Thinking$0.78$3.90
MiniMax M2-her$0.30$1.20
Gemini 2.5 Pro Preview 06-05$1.25$10.00
Claude Fable 5$10.00$50.00
Qwen3 30B A3B Thinking 2507$0.13$1.56
Qwen3.7 Plus$0.32$1.28
Mistral Small 3.2 24B$0.10$0.30
Claude Opus 4.8$5.00$25.00
Claude Opus 4.7$5.00$25.00
Gemma 3n 4B$0.06$0.12
DeepSeek V3.1 Terminus$0.27$1.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
gpt-oss-20b$0.03$0.13
Llama 3.1 405B$0.80$0.80
Claude Opus 4.1$15.00$75.00
MiniMax M3$0.30$1.20
Step 3.7 Flash$0.20$1.15
o3$2.00$8.00
Grok 4.20$1.25$2.50
Qwen3.7 Max$1.48$4.42
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
Llama 3.1 8B$0.04$0.04
DeepSeek V3.2$0.27$0.40
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
GPT-5.1$1.25$10.00
Gemini 3.5 Flash$1.50$9.00
GLM 5V Turbo$1.20$4.00
GPT-5 Image Mini$2.50$2.00
Grok 4.20 Multi-Agent$1.25$2.50
Qwen3 8B$0.12$0.46
Mistral Large 3$0.50$1.50
DeepSeek V4 Pro$0.43$0.87
Llama 3.3 70B Instruct$0.13$0.40
Grok 4.20$1.25$2.50
ERNIE 4.0$1.20$2.40
Mistral Large 2$0.60$1.80
Qwen3.6 Flash$0.19$1.13
DeepSeek V3 0324$0.27$1.12
o1-pro$150.00$600.00
Qwen Plus 0728$0.26$0.78
Yi-Lightning$0.15$0.30
GPT-5.4 Mini$0.75$4.50
Seed-2.0-Mini$0.10$0.40
Qwen3 235B A22B Thinking 2507$0.30$3.00
Qwen3.5-122B-A10B$0.26$2.08
Qwen3.5-Flash$0.07$0.26
GPT Audio$2.50$10.00
GPT Audio Mini$0.60$2.40
Qwen3.5 397B A17B$0.39$2.34
Command R$0.15$0.60
MiniMax M2.5$0.15$0.90
GPT-5.1 Chat$1.25$10.00
Grok 4.3$1.25$2.50
Seed-2.0-Lite$0.25$2.00
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Mistral Small 3.1 24B$0.35$0.56
GPT-5.6 Sol Pro$5.00$30.00
GPT-5$1.25$10.00
Mistral Medium 3$0.40$2.00
GPT-5.6 Sol$5.00$30.00
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Claude Opus 4.6$5.00$25.00
GPT-5.1-Codex-Max$1.25$10.00
Ministral 3 14B 2512$0.20$0.20
Grok 4.5$2.00$6.00
Seed 1.6 Flash$0.07$0.30
Qwen3 VL 8B Instruct$0.12$0.46
Gemini 3.1 Pro$2.00$12.00
Google Gemini Pro Latest$2.00$12.00
Gemini 2.5 Pro$1.25$10.00
MiniMax M2$0.30$1.20
GPT-4.1 Nano$0.10$0.40
Qwen3 VL 32B Instruct$0.10$0.42
Llama 4 Maverick$0.20$0.80
Qwen3.6 35B A3B$0.14$1.00
Qwen3.6 Max Preview$1.04$6.24
GPT-5 Image$10.00$10.00
Hy3 preview$0.06$0.21
GPT-5.4 Image 2$8.00$15.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Claude Opus Latest$5.00$25.00
Qwen3.5-35B-A3B$0.14$1.00
Ministral 3 8B 2512$0.15$0.15
o3 Deep Research$10.00$40.00
o4 Mini Deep Research$2.00$8.00
GLM 5.1$0.97$3.04
Gemma 4 26B A4B$0.07$0.34
DeepSeek V4 Flash$0.10$0.20
Ministral 3 3B 2512$0.10$0.10
Gemma 3 4B$0.05$0.10
Gemma 4 31B$0.22$0.55
R1 0528$0.50$2.15
Qwen3.5-9B$0.10$0.15
Qwen 2.5-Coder 32B$0.35$0.70
Qwen3.6 Plus$0.33$1.95
Llama Guard 4 12B$0.18$0.18
GLM 4.7$0.40$1.75
Gemini 3 Flash Preview$0.50$3.00
Qwen3 30B A3B$0.13$0.52
GPT-5.4 Pro$30.00$180.00
GPT-5.4$2.50$15.00
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.13$1.56
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Doubao Pro$0.80$1.60
Qwen3.5-27B$0.26$2.60
GPT-5.6 Terra Pro$2.50$15.00
GPT-5.6 Terra$2.50$15.00
GLM 5.2$0.98$3.07
Claude Opus 4.8 (Fast)$10.00$50.00
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
Hunyuan A13B Instruct$0.14$0.57
Mixtral 8x22B$0.50$1.00
GPT-5.3 Chat$1.75$14.00
Gemini 3.1 Flash Lite Preview$0.25$1.50
Hy3$0.20$0.80
Codestral 2508$0.30$0.90
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
Claude Fable Latest$10.00$50.00
Seed 1.6$0.25$2.00
GPT-5.2 Pro$21.00$168.00
GLM 4.6V$0.30$0.90
Qwen3 VL 30B A3B Instruct$0.13$0.52
Qwen3 Coder 30B A3B Instruct$0.07$0.27
KAT-Coder-Pro V2$0.30$1.20
GPT-4.1$2.00$8.00
Kimi K2 0905$0.60$2.50
Anthropic Claude Sonnet Latest$2.00$10.00
Hunyuan Pro$0.60$1.20
Qwen3.5 Plus 2026-04-20$0.30$1.80
Qwen3.6 27B$0.45$2.70
GPT-5 Pro$15.00$120.00
Grok Build 0.1$1.00$2.00
Mistral Medium 3.5$1.50$7.50
Anthropic Claude Haiku Latest$1.00$5.00
DeepSeek V3.2 Exp$0.27$0.41
Kimi K2.7 Code$0.85$3.80
Lyria 3 Pro Preview$0.00$0.00
GLM 4.5$0.60$2.20
Gemini 2.5 Flash Lite$0.10$0.40
Gemma 3 12B$0.05$0.15
Command A$2.50$10.00
DeepSeek R1$0.70$2.50
Qwen 2.5 72B$0.40$0.80
GPT-4o-mini Search Preview$0.15$0.60
GPT-5.2$1.75$14.00
GPT-4o Search Preview$2.50$10.00
Devstral 2 2512$0.40$2.00
Qwen3 235B A22B Instruct 2507$0.09$0.55
Qwen3 32B$0.08$0.28
DeepSeek V3$0.20$0.80
Claude 3.5 Sonnet v2$3.00$15.00
MiniMax M2.1$0.30$1.20
Command R7B (12-2024)$0.04$0.15
Llama 3.3 70B Instruct$0.13$0.40
GPT-5.2 Chat$1.75$14.00
GPT-5.1-Codex-Mini$0.25$2.00
Qwen-Plus$0.26$0.78
Qwen2.5 Coder 32B Instruct$0.66$1.00
GPT-4o (2024-08-06)$2.50$10.00
Llama 3.1 8B Instruct$0.05$0.08
Llama 3.1 70B Instruct$0.40$0.40
Kimi K3$3.00$15.00
Muse Spark 1.1$1.25$4.25
Kimi K2 Thinking$0.60$2.50
Voxtral Small 24B 2507$0.10$0.30
gpt-oss-safeguard-20b$0.07$0.30
Qwen3 VL 8B Thinking$0.12$1.36
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
Qwen3 Next 80B A3B Thinking$0.10$0.78
Qwen2.5 VL 72B Instruct$0.80$1.00
R1 Distill Llama 70B$0.80$0.80
R1$0.70$2.50
MiniMax-01$0.20$1.10
Qwen3 Next 80B A3B Instruct$0.10$1.10
Lyria 3 Clip Preview$0.00$0.00
Qwen3 VL 235B A22B Thinking$0.26$2.60
Qwen3 VL 235B A22B Instruct$0.21$1.90
ERNIE 4.5 VL 424B A47B$0.42$1.25
Claude Sonnet 4$3.00$15.00
Qwen3 14B$0.12$0.24
Qwen3 235B A22B$0.46$1.82
o4 Mini High$1.10$4.40
Gemma 2 27B$0.65$0.65
Mistral Large 3 2512$0.50$1.50
GPT-5 Codex$1.25$10.00
Gemma 3 27B$0.10$0.30
Saba$0.20$0.60
Llama 3.2 11B Vision Instruct$0.34$0.34
o3 Mini High$1.10$4.40
Llama 3.2 1B Instruct$0.03$0.20
GPT-4 Turbo$10.00$30.00
GPT-3.5 Turbo Instruct$1.50$2.00
GPT-3.5 Turbo 16k$3.00$4.00
GPT-3.5 Turbo$0.50$1.50
SHARE RESEARCH: