nano-vllm/nanovllm/config.py at 4484a1482cbb66d26123077debabf0f5c0419919

Files

Zijie Tian 29e102720b 🐛 fix: support multiple EOS tokens for GLM-4

GLM-4 uses multiple EOS tokens [151329, 151336, 151338] where 151336
(<|user|>) should also stop generation. Previously only the first EOS
from tokenizer was used, causing generation to always hit max_tokens.

Changes:
- config.py: Change eos type to int | list[int]
- llm_engine.py: Read eos_token_id from hf_config (contains full list)
- scheduler.py: Use set for efficient multi-EOS lookup

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

2026-01-28 13:23:53 +08:00

3.5 KiB

Raw Blame History

View Raw

3.5 KiB Raw Blame History

3.5 KiB

Raw Blame History