32 blocks was a ~1GB safety-valve leftover from the num_cpu_blocks=2000 hang incident, not meaningful offload capacity. This model's KV cache is ~32MB/128-token block (64 layers, 8 KV heads x 128 head_dim, fp16) -- 256 blocks gives ~8GB of real DRAM offload (32,768 tokens), comfortably under the pod's 36Gi limit alongside the ~20GB bnb-4bit weights.