vLLM prefix caching for phonellm (nemotron_h): bug report + first fix upstream

#1
by kushalpatil - opened

Heads-up for the phonellm team: we benchmarked phonellm-alpha-1 for a multi-turn voice workload and found vLLM's prefix caching is a no-op for nemotron_h hybrids — a byte-identical resend re-prefills the whole transcript every turn (v0.28.0 and nightly, all cache modes/KV dtypes), which capped concurrency ~5x below an equivalent-activation attention model on one L40S. We root-caused the CPU-backend half and sent a fix upstream, with the details and a no-GPU repro here:

The GPU-path no-op on the v0.28 images is still open in the issue. Since turn-2+ TTFT for phone calls rides on prefix reuse, you probably want this fixed more than anyone — a nudge/review from your side on the issue would help. Happy to share the benchmark harness details.

Sign up or log in to comment