vllm.models.glm5next.common.model
¶
_LOGIT_SCALE = 1.0
module-attribute
¶
Output logit scale. A GLM-5.3-Flash trained value that neither the
checkpoint nor Glm5NextTextConfig carries.
_MHC_POST_MULT_VALUE = 2.0
module-attribute
¶
mHC post-multiplier. A GLM-5.3-Flash trained value that neither the
checkpoint nor Glm5NextTextConfig carries.
_MHC_TAU = 0.05
module-attribute
¶
mHC routing temperature. A GLM-5.3-Flash trained value that neither the
checkpoint nor Glm5NextTextConfig carries.
_VISION_RMS_NORM_EPS = 1e-06
module-attribute
¶
Vision tower RMSNorm epsilon.
GLM-5.3-Flash checkpoints ship vision_config.rms_norm_eps = 1e-5, but the
vision tower was trained with 1e-6. Serving with 1e-5 drifts the RMSNorm and
produces repetitive/degraded image descriptions, so force the trained value
regardless of the checkpoint field.
_dequant_fp8_block(weight_fp8, scale_inv, block_size=128)
¶
Dequantize a block-FP8 (e4m3) weight with per-block scale to BF16.
Unlike scaled_dequantize this tolerates a non-divisible (partial last
block) shape by zero-padding to a multiple of block_size before the
scale broadcast and trimming back afterwards (e.g. kv_a_proj_with_mqa is
576 rows = 4*128 + 64).
Source code in vllm/models/glm5next/common/model.py
_fused_shared_expert_name(name, n_routed_experts)
¶
Point a checkpoint mlp.shared_experts.* tensor at the fused MoE's
shared-expert slot, which follows the routed experts; other names are
returned unchanged.
Source code in vllm/models/glm5next/common/model.py
_fused_shared_experts_tuned(parallel_config)
¶
AITER has fused-MoE configs tuned for the fused shared-expert shape (one more expert and one more top-k slot than the routed MoE) only on gfx950, with every expert on each rank and its weights split by TP4 or TP8. Data, prefill context and expert parallelism change that split, so any other GPU or parallel layout would run untuned fallback kernels.
Source code in vllm/models/glm5next/common/model.py
_num_fused_shared_experts(n_shared_experts, enabled)
¶
Expert slots the fused MoE appends for the shared expert; must match the
num_fused_shared_experts that FusedMoE allocates.
Source code in vllm/models/glm5next/common/model.py
_try_load_fp8_attn_proj(name, tensor, buf, params_dict, loaded_params, kv_a_pad_size)
¶
Dequantize FP8 q_a_proj / kv_a_proj_with_mqa / o_proj to BF16 on load.
The FP8 checkpoint stores these as block-FP8 (weight + weight_scale_inv),
but the model holds them in BF16 (fused_qkv_a_proj is always BF16 via
DeepSeekV2FusedQkvAProjLinear; o_proj is excluded by
modules_to_not_convert). When the model target is BF16 (no
weight_scale_inv param) we dequantize; otherwise we return False so the
normal stacked/direct path loads the FP8 tensor as-is.
Source code in vllm/models/glm5next/common/model.py
1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 | |
_validate_supported_config(config)
¶
Reject checkpoints using config options this implementation lacks.
The kpool indexer kernels always keep the incomplete trailing pool, so a checkpoint asking otherwise would be served silently wrong.