tue 03:14 · berlin
The work happens once.
$ claude
> long-ctx vllm OOM after fp16 prefix cache;
> raising max_model_len blew the kv budget
[debug session · 47 min]
> root cause: prefix cache page size too small
> fix: --kv-cache-dtype=fp8 + --block-size=32
[folklore · auto-saved trace]
signed · @you · ed25519 3a4b…