A single recorded GPU sequence (a CUDA Graph) cut a 7B model's draft latency from 112ms to 25ms — type0 | type0