nano-vllm/nanovllm/models/__init__.py at 29e102720b3bbcb53f789805bd8949a005f9ab61 - nano-vllm - Gitea: Git with a cup of tea

zijie-tian/nano-vllm

Files

Zijie Tian 726e4b58cf ✨ feat: add GLM-4-9B-Chat-1M model support

Add support for GLM-4 model architecture with the following changes:

- Add glm4.py with ChatGLMForCausalLM, GLM4Model, GLM4Attention, GLM4MLP
- Add GLM4RotaryEmbedding with interleaved partial rotation (rotary_dim = head_dim // 2)
- Add apply_rotary_emb_interleaved function for GLM-4 style RoPE
- Add GLM-4 weight name conversion and loading in loader.py
- Add GLM-4 chat template conversion in test_ruler.py
- Add trust_remote_code=True for GLM-4 config loading

Key GLM-4 specific adaptations:
- QKV bias enabled (add_qkv_bias: true)
- RoPE with rope_ratio scaling (base = 10000 * rope_ratio)
- Interleaved RoPE (pairs adjacent elements, not first/second half)
- Partial rotation (only half of head_dim is rotated)
- Uses multi_query_group_num instead of num_key_value_heads
- Uses kv_channels instead of head_dim
- Uses ffn_hidden_size instead of intermediate_size

Tested with RULER niah_single_1 (5 samples): 100% accuracy
Both GPU-only and CPU offload modes verified

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

2026-01-28 13:15:57 +08:00

11 lines

343 B

Python

Raw Blame History

 """Model registry and model implementations."""
 from nanovllm.models.registry import register_model, get_model_class, MODEL_REGISTRY
 # Import models to trigger registration
 from nanovllm.models import qwen3
 from nanovllm.models import llama
 from nanovllm.models import glm4
 __all__ = ["register_model", "get_model_class", "MODEL_REGISTRY"]