reasoning: swap to dense Qwen3-32B-bnb-4bit for reliable tool calling
DeepSeek-R1-distill's tool_choice=auto narration bug needed a real fix, not a workaround -- Qwen3's native tool-call format (hermes-compatible chat template) solves it at the source instead of parsing around it. Dense Qwen3-32B avoids the MoE arch/quantization pitfalls hit by the two prior swap attempts (Kimi-distilled Qwen3.6 MoE, AWQ Qwen3-30B-A3B) -- same bnb-4bit path already proven working on this sm70 (V100) node.
This commit is contained in:
@@ -12,7 +12,18 @@ spec:
|
|||||||
predictor:
|
predictor:
|
||||||
containers:
|
containers:
|
||||||
- args:
|
- args:
|
||||||
- --model=unsloth/DeepSeek-R1-Distill-Qwen-32B-bnb-4bit
|
# DeepSeek-R1-Distill-32B retired: tool_choice="auto" (what pi sends)
|
||||||
|
# hit a documented vLLM/R1-family conflict -- the model narrated fake
|
||||||
|
# tool_calls in its <think> block instead of emitting real ones,
|
||||||
|
# regardless of parser. Tried swapping to a Kimi-distilled Qwen3.6
|
||||||
|
# MoE checkpoint and an AWQ-quantized Qwen3-30B-A3B first -- both
|
||||||
|
# failed on real, separate blockers (unrecognized model_type; then
|
||||||
|
# marlin INT4 kernels needing compute capability 80+, but worker-1's
|
||||||
|
# GPU is sm70/V100). Landed on dense Qwen3-32B instead: native Qwen3
|
||||||
|
# tool-call format (no narration bug), bnb-4bit works fine on sm70
|
||||||
|
# (proven by the old DeepSeek config already), and no MoE
|
||||||
|
# arch/quantization risk this time.
|
||||||
|
- --model=unsloth/Qwen3-32B-bnb-4bit
|
||||||
- --served-model-name=reasoning
|
- --served-model-name=reasoning
|
||||||
- --quantization=bitsandbytes
|
- --quantization=bitsandbytes
|
||||||
- --dtype=float16
|
- --dtype=float16
|
||||||
@@ -23,19 +34,11 @@ spec:
|
|||||||
- --max-num-seqs=4
|
- --max-num-seqs=4
|
||||||
- --enable-chunked-prefill
|
- --enable-chunked-prefill
|
||||||
- --enable-prefix-caching
|
- --enable-prefix-caching
|
||||||
- --reasoning-parser=deepseek_r1
|
# qwen3 is vLLM's dedicated reasoning parser for this family's <think>
|
||||||
# Without these, any client sending tool_choice="auto" (pi does, for
|
# blocks.
|
||||||
# Read/Bash/etc.) gets a 400: "auto" tool choice requires
|
- --reasoning-parser=qwen3
|
||||||
# --enable-auto-tool-choice and --tool-call-parser to be set.
|
# hermes is the documented tool-call parser for general (non-Coder)
|
||||||
# deepseek_v3 (matching --reasoning-parser above) 400s here --
|
# Qwen3 models -- native chat template support, not narrated text.
|
||||||
# "DeepSeek-V3 Tool parser could not locate tool call start/end
|
|
||||||
# tokens in the tokenizer" -- this checkpoint is a Qwen2.5-32B base
|
|
||||||
# distilled on R1 traces, so its tokenizer never got DeepSeek-V3's
|
|
||||||
# own special tool-call tokens registered even though it kept R1's
|
|
||||||
# <think> reasoning format. hermes parses tool calls from plain text
|
|
||||||
# patterns instead of special tokens, so it works against the
|
|
||||||
# underlying Qwen tokenizer regardless. Verified live: deepseek_v3
|
|
||||||
# 400s, hermes returns a real tool_calls response.
|
|
||||||
- --enable-auto-tool-choice
|
- --enable-auto-tool-choice
|
||||||
- --tool-call-parser=hermes
|
- --tool-call-parser=hermes
|
||||||
# vLLM 0.11.0's native OffloadingConnector -- spills KV cache blocks
|
# vLLM 0.11.0's native OffloadingConnector -- spills KV cache blocks
|
||||||
|
|||||||
Reference in New Issue
Block a user