Record successful 2048-token CUDA Graph retest and preserve failure samples
This commit is contained in:
@@ -62,3 +62,8 @@
|
||||
|
||||
修正依据:
|
||||
https://github.com/blazux/qwen3.8-Flash-DGX/blob/main/src/patch_mamba_block_size.py
|
||||
|
||||
## 后续 2048 预算复测
|
||||
|
||||
上述 CUDA Graph 未完成验收是首次实验的状态。后续 graph+前缀缓存组合已通过三轮复测,
|
||||
短回答中位速度提升约 4.4%,默认仍保留 eager。详见 [复测报告](cuda-graph-retest.md)。
|
||||
|
||||
Reference in New Issue
Block a user