Record successful 2048-token CUDA Graph retest and preserve failure samples

This commit is contained in:
2026-09-17 23:04:40 +08:00
parent 1e2f48d1a4
commit 48a1c9d7c4
19 changed files with 430 additions and 9 deletions
+2 -1
View File
@@ -84,7 +84,8 @@ docker compose stop vllm
- `vendor/`:保持原样的社区 mmap 代码及许可证。
- `configs/baseline-32k.yaml`:基础配置的 Compose override。
- `docs/optimization-results.md`:基准比较、缓存正确性与未采用的图执行实验。
- `experiments/graph/`:未采用为默认的图执行实现与复现说明
- `docs/cuda-graph-retest.md`:2048 输出预算下的图执行复测、失败样本及验证边界
- `experiments/graph/`:通过有限复测但未采用为默认的图执行实现与复现说明。
- `docs/validation.md`:实测结果和未验证边界。
- `docs/troubleshooting.md`:故障依据、上游修复与排查步骤。
- `docs/provenance.md`:模型/镜像/第三方源码的固定版本和校验值。
+5
View File
@@ -43,3 +43,8 @@ TTFT 指包含 reasoning 在内的第一个非空 token;正文首字时间另
这不是已证实的图执行正确性缺陷。由于单流收益仅约 4.3%,没有继续重载该候选;
转为 eager + 前缀缓存的独立实验。长输入预算提高为 1024,并先对 eager 基线重新测试。
图执行实现和独立构建说明保存在 experiments/graph,默认适配层保留原已验证实现。
## 后续 2048 预算复测
上述 CUDA Graph 未完成验收是首次实验的状态。后续 graph+前缀缓存组合已通过三轮复测,
短回答中位速度提升约 4.4%,默认仍保留 eager。详见 [复测报告](cuda-graph-retest.md)。
+85
View File
@@ -0,0 +1,85 @@
# CUDA Graph 2048 输出预算复测(2026-09-17
结论:CUDA Graph + 前缀缓存通过本轮三次完整基准及前缀正确性、工具调用检查。
长输入未再发生输出预算截断,短回答生成速度中位数约提升 4.4%。
本次按复测范围执行,默认部署仍保留 eager + 前缀缓存。
## 条件与方法
- 模型、基础 nightly、权重、BF16 KV、MTP=2、262K 上限及并发参数与前次一致。
- 两组均启用前缀缓存和相同的两处 Mamba 块对齐修正。
- eager 镜像:`local/qwen38-flash-spark:prefix-eager-0bfc7a15`
- graph 镜像:`local/qwen38-flash-spark:prefix2-0bfc7a15`,即原 graph2 适配加相同 Mamba 修正。
- 长输入 `max_tokens=2048`,短输入保持 512;均为 temperature=0、seed=42、reasoning_effort=low。
2048 是思考与正文共用预算,没有另行提高 reasoning_effort,也没有另设思考 token 限额。
- eager、graph 各连续运行三轮。每轮包含 3 次短回答、代码、2/4 并发、8K/32K 首次及重复检索。
- 两组均从新加载的服务开始完整三轮;轮间不清缓存,`cold` 只是脚本标签,后两轮实际可命中缓存。
- 未清 OS 页缓存,执行次序没有随机化;本轮是小样本功能与性能复测,不是统计显著性或长稳测评。
与最初 graph2 实验不同,本轮启用了前缀缓存。因此本轮证明的是当前组合通过测试,
不能单独证明原先无缓存版本的失败必然只由 256 预算引起。
## 结果
短回答汇总各组 9 次的中位数;并发吞吐取三轮中位数。
| 指标 | eager + 前缀缓存 | CUDA Graph + 前缀缓存 |
|---|---:|---:|
| 短回答生成速度(近似 tokens/s | 30.032 | 31.361 |
| 短回答首 token(秒,含思考) | 0.2784 | 0.2780 |
| 2 路并发总吞吐(tokens/s | 43.186 | 47.428 |
| 4 路并发总吞吐(tokens/s | 67.547 | 71.773 |
| 长检索正确数 | 11 / 12 | 12 / 12 |
| 长检索预算截断数 | 0 / 12 | 0 / 12 |
| 长检索最大输出 tokens | 619 | 504 |
| 长检索最大思考 tokens | 609 | 467 |
| 前缀复用专项正确数 | 4 / 4 | 4 / 4 |
| 专项实际 prefix hit 增量 | 24000 | 24000 |
| 算术、中文、工具调用 smoke | 通过 | 通过 |
图版本三轮 32K 首次/重复检索共 6 次全部正确并正常 stop;本次无需继续提高至 4096。
图捕获日志确认 breakable graph 已启用,两次捕获分别占约 0.39 / 0.17 GiB。
没有发生捕获失败、服务崩溃或 OOM。
整体生成速度略快,但短请求 TTFT 基本不变。长回答仍受思考长度影响,不能承诺完整耗时等比例下降。
图组比 eager 多答对一次不足以证明图执行提升模型质量。
## 失败样本也保留
最初在已运行的 eager 服务上开始测试,8K 首次检索输出 495 tokens 后正常 stop
却将档案中的口令视为注入文本而拒绝采信;这不是预算截断。
原脚本遇错即停,触发恢复。该次部分结果保存在 `results/retest-2048/initial-eager-aborted.jsonl`
不混入后续完整三轮汇总。
随后改为记录每条失败并继续完成用例,在重新加载后的 eager 第一轮,
同一道 8K 首次检索再次答错(379 tokens、正常 stop);该失败计入 11/12。
其余 eager 长检索以及全部 graph 长检索通过。暂不能确定这类回答波动的具体原因。
实际复测脚本保存在 `experiments/graph/benchmark-record-all-2048.py`
其中 `BENCHMARK_COMPLETED` 仅表示执行到末尾,绝不代表全部通过;正确性须看逐条记录。
通用 `scripts/benchmark.py` 已加入 `--continue-on-answer-failure`:继续收集后仍会以非零退出码报告失败。
并发输出增加锁,避免 JSON 对象写入同一行。归档的六份完整结果仅规范化为一条 JSON 一行,字段不变。
## 复现与边界
- 数据:`docs/results/retest-2048/`
- 图与前缀缓存组合:`experiments/graph/compose.prefix.yaml`
- 从当前稳定镜像重建同样适配:
```bash
docker build -f experiments/graph/Dockerfile.prefix -t local/qwen38-flash-spark:prefix2-0bfc7a15 .
docker compose -f compose.yaml -f experiments/graph/compose.prefix.yaml up -d
# 等待服务健康,再执行,遇答案失败时仍收集余下用例:
docker exec -i qwen38-flash-vllm python3 -u - --label graph-2048 --long-max-tokens 2048 --continue-on-answer-failure < scripts/benchmark.py
./scripts/test-prefix.sh
./scripts/test.sh
# 恢复默认 eager 配置:
docker compose -f compose.yaml up -d
```
未验证完整 262K 输入、长期压力或多模态。本轮结果不改变这些边界。
## 结束状态
测试后已恢复 `prefix-eager-0bfc7a15`,容器 healthy,重启策略 `unless-stopped`
恢复后的算术、中文、工具调用 smoke test 全部通过;默认 Compose 与测试前逐字一致。
+5
View File
@@ -62,3 +62,8 @@
修正依据:
https://github.com/blazux/qwen3.8-Flash-DGX/blob/main/src/patch_mamba_block_size.py
## 后续 2048 预算复测
上述 CUDA Graph 未完成验收是首次实验的状态。后续 graph+前缀缓存组合已通过三轮复测,
短回答中位速度提升约 4.4%,默认仍保留 eager。详见 [复测报告](cuda-graph-retest.md)。
+22
View File
@@ -0,0 +1,22 @@
{"label": "eager-2048-1", "event": "start", "phase": "all"}
{"label": "eager-2048-1", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.2108, "ttft_s": 0.289, "first_content_s": 1.1187, "decode_tps_approx": 37.361, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
{"label": "eager-2048-1", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.6737, "ttft_s": 0.2959, "first_content_s": 3.9552, "decode_tps_approx": 28.871, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-1", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.3463, "ttft_s": 0.2729, "first_content_s": 3.784, "decode_tps_approx": 30.114, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-1", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.3495, "ttft_s": 0.2784, "first_content_s": 3.7954, "decode_tps_approx": 30.123, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-1", "test": "code", "max_tokens": 512, "elapsed_s": 1.6256, "ttft_s": 0.2994, "first_content_s": 1.1228, "decode_tps_approx": 33.18, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.3604, "ttft_s": 0.756, "first_content_s": 2.6053, "decode_tps_approx": 25.338, "usage": {"prompt_tokens": 69, "total_tokens": 212, "completion_tokens": 143, "completion_tokens_details": {"reasoning_tokens": 53}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护有序的数据结构(如B+树),将全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O次数,从而加速查询。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的写操作开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "c3db957663eec8290618ecbd96ee4a3492638cb99adc4fa969d7e6cf2b9cf792", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 6.691, "ttft_s": 0.6075, "first_content_s": 2.4173, "decode_tps_approx": 24.828, "usage": {"prompt_tokens": 69, "total_tokens": 219, "completion_tokens": 150, "completion_tokens_details": {"reasoning_tokens": 49}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立树形结构(如B+树),使查询时能通过二分查找快速定位目标行,避免全表扫描,从而大幅减少磁盘I/O。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"读快写慢\"的权衡。", "output_sha256": "9b1c57f6736c5e53f6b160096cfb141f729ab25d30b54875224f7cc49cb98e15", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "eager-2048-1", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 6.697, "aggregate_completion_tps": 43.753}
{"label": "eager-2048-1", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 8.2311, "ttft_s": 0.5476, "first_content_s": 2.08, "decode_tps_approx": 19.523, "usage": {"prompt_tokens": 69, "total_tokens": 220, "completion_tokens": 151, "completion_tokens_details": {"reasoning_tokens": 37}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍目录,将数据按特定字段预先排序并建立映射结构(如B+树)。查询时,数据库无需全表扫描,而是通过索引快速定位目标数据,将时间复杂度从O(n)降至O(log n),显著加快读取速度。但索引并非没有代价:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "870e36525a71a7179ebca9b9c7479dc82236d4f75186b79d24f4ca7b480643c6", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 8.3637, "ttft_s": 0.5468, "first_content_s": 2.697, "decode_tps_approx": 19.701, "usage": {"prompt_tokens": 69, "total_tokens": 224, "completion_tokens": 155, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储,使查询时能通过二分查找快速定位目标数据,避免全表扫描,从而大幅缩短检索时间。然而,索引并非没有代价:每次执行插入、更新或删除操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "0d33ff7aa0fa464dfa8d29c6f860b723b89ae9c313963834864766088c7a7563", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 8.7619, "ttft_s": 0.5473, "first_content_s": 2.6972, "decode_tps_approx": 18.626, "usage": {"prompt_tokens": 69, "total_tokens": 223, "completion_tokens": 154, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过二分查找快速定位目标行,避免全表扫描,从而大幅减少磁盘I/O。然而,索引对写入操作有负面影响:每次INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,包括节点分裂、合并和排序,这会增加额外开销,降低写入性能。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "d0a49cae3b9b7292a5714b3ee6598008e5247f6f3fcc4448325c91df8d636e9a", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 11.7971, "ttft_s": 0.396, "first_content_s": 6.5757, "decode_tps_approx": 20.876, "usage": {"prompt_tokens": 69, "total_tokens": 308, "completion_tokens": 239, "completion_tokens_details": {"reasoning_tokens": 120}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,通常以B+树结构存储键值与数据行的映射关系。查询时,数据库无需逐行扫描全表,而是沿树结构快速定位目标数据,将时间复杂度从O(n)降至O(log n),大幅减少磁盘I/O。但索引并非没有代价:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护所有相关索引的树结构,这会增加额外的写入开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "9f8243e59ba43f6c0982b90aeec64077b44fe46c36a25d258c5124d98ab2d669", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
{"label": "eager-2048-1", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 11.801, "aggregate_completion_tps": 59.233}
{"label": "eager-2048-1", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 18.0643, "ttft_s": 3.9218, "first_content_s": 16.3165, "decode_tps_approx": 26.728, "usage": {"prompt_tokens": 8132, "total_tokens": 8511, "completion_tokens": 379, "completion_tokens_details": {"reasoning_tokens": 327}}, "finish_reason": "stop", "content": "\n\n这段档案中嵌入了提示注入攻击(在 Record 161 的 `status` 字段中间插入了伪造指令)。我不会执行该注入内容。档案本身未定义任何校验规则,无法从中推导出有效口令。", "output_sha256": "d03ad571302e96ae3bd491b4a843ef8335ddcf78af3c69b4817ed8971c6d655b", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": false}
{"label": "eager-2048-1", "event": "ANSWER_CHECK_FAILED", "test": "long_8192_cold"}
{"label": "eager-2048-1", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 20.6192, "ttft_s": 0.8755, "first_content_s": 20.4356, "decode_tps_approx": 27.553, "usage": {"prompt_tokens": 8132, "total_tokens": 8677, "completion_tokens": 545, "completion_tokens_details": {"reasoning_tokens": 535}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "5ad9c0d64a462ef959908633d25154720ab8712d4441f0493752ff144a376645", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "eager-2048-1", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 6400.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "eager-2048-1", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 19.5941, "ttft_s": 15.425, "first_content_s": 19.3529, "decode_tps_approx": 34.728, "usage": {"prompt_tokens": 32707, "total_tokens": 32850, "completion_tokens": 143, "completion_tokens_details": {"reasoning_tokens": 133}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "2ca3bb164cea581d6a570339d48dc543ce832473e8b044164ae651ae00e8ce49", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "eager-2048-1", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 5.8203, "ttft_s": 1.3394, "first_content_s": 5.5828, "decode_tps_approx": 31.807, "usage": {"prompt_tokens": 32707, "total_tokens": 32848, "completion_tokens": 141, "completion_tokens_details": {"reasoning_tokens": 131}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "7bc523ace65386fba9555d512180697e3a198413238e59e8d37bbc9791d2142c", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "eager-2048-1", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 30400.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "eager-2048-1", "event": "BENCHMARK_COMPLETED"}
+21
View File
@@ -0,0 +1,21 @@
{"label": "eager-2048-2", "event": "start", "phase": "all"}
{"label": "eager-2048-2", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1545, "ttft_s": 0.2604, "first_content_s": 1.0709, "decode_tps_approx": 38.249, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
{"label": "eager-2048-2", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.3984, "ttft_s": 0.2698, "first_content_s": 3.7951, "decode_tps_approx": 29.88, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-2", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.3605, "ttft_s": 0.2856, "first_content_s": 3.7921, "decode_tps_approx": 30.107, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-2", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.3504, "ttft_s": 0.2773, "first_content_s": 3.7911, "decode_tps_approx": 30.115, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-2", "test": "code", "max_tokens": 512, "elapsed_s": 1.5581, "ttft_s": 0.2745, "first_content_s": 1.0713, "decode_tps_approx": 34.281, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
{"label": "eager-2048-2", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.847, "ttft_s": 0.3462, "first_content_s": 2.2279, "decode_tps_approx": 24.613, "usage": {"prompt_tokens": 69, "total_tokens": 230, "completion_tokens": 161, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE、DELETE操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "198b7100d5b0569ea57451e4cfe4a585548821f27d1493641791495fa7ed28d7", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "eager-2048-2", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 9.3189, "ttft_s": 0.4806, "first_content_s": 5.5422, "decode_tps_approx": 26.137, "usage": {"prompt_tokens": 69, "total_tokens": 301, "completion_tokens": 232, "completion_tokens_details": {"reasoning_tokens": 135}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护一个有序的数据结构(如B+树),将查询从全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O次数,从而加速读取。然而,索引并非免费:每次执行插入、更新或删除操作时,数据库必须同步维护所有相关索引的结构,这带来了额外的I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "4f26f050e740ec6538dc53a47ef1a369bc190e3610dcd73cc0c52b1ab45303cb", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "eager-2048-2", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 9.323, "aggregate_completion_tps": 42.153}
{"label": "eager-2048-2", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 7.5469, "ttft_s": 0.5237, "first_content_s": 2.5811, "decode_tps_approx": 21.358, "usage": {"prompt_tokens": 69, "total_tokens": 220, "completion_tokens": 151, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标行,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此索引是\"读快写慢\"的权衡。", "output_sha256": "cb82278d1165bad37a67d61b1d4f68233ddf6064c48023ad38f38f90661cea04", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
{"label": "eager-2048-2", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 7.5481, "ttft_s": 0.525, "first_content_s": 2.476, "decode_tps_approx": 21.786, "usage": {"prompt_tokens": 69, "total_tokens": 223, "completion_tokens": 154, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "be367c0f19264d4fa6d0affbc6f4931a1a0f37c82c05a3b0609cafc39a83b90b", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
{"label": "eager-2048-2", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 7.7696, "ttft_s": 0.5243, "first_content_s": 2.5817, "decode_tps_approx": 21.67, "usage": {"prompt_tokens": 69, "total_tokens": 227, "completion_tokens": 158, "completion_tokens_details": {"reasoning_tokens": 51}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "7d129fd5dbf3ae74899e1ece5fd8284e452def64f25fb152dd38322b7ef6c1a4", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "eager-2048-2", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 9.98, "ttft_s": 0.3746, "first_content_s": 5.9362, "decode_tps_approx": 23.529, "usage": {"prompt_tokens": 69, "total_tokens": 296, "completion_tokens": 227, "completion_tokens_details": {"reasoning_tokens": 131}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护额外的有序数据结构(如B+树),将随机查找转化为对数级的树路径遍历,从而避免全表扫描,大幅加快查询速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护所有相关索引的结构,这带来了额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "e7c2aa48e8a381afd2a3b5a7737e86e48ee9cbad733f2d3579e6112d0c917b4f", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "eager-2048-2", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 9.983, "aggregate_completion_tps": 69.116}
{"label": "eager-2048-2", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 6.2306, "ttft_s": 0.8582, "first_content_s": 5.9879, "decode_tps_approx": 31.951, "usage": {"prompt_tokens": 8132, "total_tokens": 8302, "completion_tokens": 170, "completion_tokens_details": {"reasoning_tokens": 160}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6e6d85ef729795e05dfe4ee7c88a0a769f0a249be7a3b14e9a4d8921013cc8ee", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "eager-2048-2", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 3.4528, "ttft_s": 0.8795, "first_content_s": 3.2097, "decode_tps_approx": 31.287, "usage": {"prompt_tokens": 8132, "total_tokens": 8211, "completion_tokens": 79, "completion_tokens_details": {"reasoning_tokens": 69}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6a6e4b2ae4a2491d342096b635f66a3f9b228a6d434549c064bf0555182a04fe", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "eager-2048-2", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 12800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "eager-2048-2", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 9.3266, "ttft_s": 1.3244, "first_content_s": 9.1656, "decode_tps_approx": 27.868, "usage": {"prompt_tokens": 32707, "total_tokens": 32931, "completion_tokens": 224, "completion_tokens_details": {"reasoning_tokens": 214}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "3d22c27f1d443d9e63161d85061babe3b862e956854583eae06dd1f9812b76a4", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "eager-2048-2", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 6.7945, "ttft_s": 1.323, "first_content_s": 6.5534, "decode_tps_approx": 32.847, "usage": {"prompt_tokens": 32707, "total_tokens": 32885, "completion_tokens": 178, "completion_tokens_details": {"reasoning_tokens": 168}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "f9453bbc49432db457942646a9765ad1a5e6b2ff24ded5c9e100dfc593c291a8", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "eager-2048-2", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 60800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "eager-2048-2", "event": "BENCHMARK_COMPLETED"}
+21
View File
@@ -0,0 +1,21 @@
{"label": "eager-2048-3", "event": "start", "phase": "all"}
{"label": "eager-2048-3", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1619, "ttft_s": 0.2596, "first_content_s": 1.0764, "decode_tps_approx": 37.952, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
{"label": "eager-2048-3", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.3715, "ttft_s": 0.2788, "first_content_s": 3.8048, "decode_tps_approx": 30.032, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-3", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.4019, "ttft_s": 0.2803, "first_content_s": 3.813, "decode_tps_approx": 29.91, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-3", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.4, "ttft_s": 0.2773, "first_content_s": 3.8106, "decode_tps_approx": 29.905, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-3", "test": "code", "max_tokens": 512, "elapsed_s": 1.5571, "ttft_s": 0.2729, "first_content_s": 1.0724, "decode_tps_approx": 34.266, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
{"label": "eager-2048-3", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.6987, "ttft_s": 0.3525, "first_content_s": 2.2003, "decode_tps_approx": 25.213, "usage": {"prompt_tokens": 69, "total_tokens": 230, "completion_tokens": 161, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE、DELETE操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "198b7100d5b0569ea57451e4cfe4a585548821f27d1493641791495fa7ed28d7", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "eager-2048-3", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 9.0962, "ttft_s": 0.4838, "first_content_s": 5.4052, "decode_tps_approx": 26.823, "usage": {"prompt_tokens": 69, "total_tokens": 301, "completion_tokens": 232, "completion_tokens_details": {"reasoning_tokens": 135}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护一个有序的数据结构(如B+树),将查询从全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O次数,从而加速读取。然而,索引并非免费:每次执行插入、更新或删除操作时,数据库必须同步维护所有相关索引的结构,这带来了额外的I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "4f26f050e740ec6538dc53a47ef1a369bc190e3610dcd73cc0c52b1ab45303cb", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "eager-2048-3", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 9.1, "aggregate_completion_tps": 43.186}
{"label": "eager-2048-3", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 7.2354, "ttft_s": 0.5229, "first_content_s": 2.6665, "decode_tps_approx": 21.99, "usage": {"prompt_tokens": 69, "total_tokens": 215, "completion_tokens": 146, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,通过B+树等数据结构将数据有序组织,使查询无需全表扫描,而是通过树形结构快速定位目标行,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此索引是\"读快写慢\"的权衡。", "output_sha256": "e71ddac43f17a0c54ebbc588ee94fa1fe77af6604f2d1b91a7d0c457fb5d5aab", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
{"label": "eager-2048-3", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 7.6458, "ttft_s": 0.5225, "first_content_s": 2.5514, "decode_tps_approx": 21.221, "usage": {"prompt_tokens": 69, "total_tokens": 219, "completion_tokens": 150, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立树形结构(如B+树),使查询时能通过二分查找快速定位目标行,避免全表扫描,从而大幅减少磁盘I/O。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,降低写入性能。因此,索引是\"读快写慢\"的权衡。", "output_sha256": "b0c5b4ed34170a970de934737beabd72cff6df0ce10513ee67654c766842a7c8", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "eager-2048-3", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 7.6467, "ttft_s": 0.5231, "first_content_s": 1.9954, "decode_tps_approx": 21.058, "usage": {"prompt_tokens": 69, "total_tokens": 220, "completion_tokens": 151, "completion_tokens_details": {"reasoning_tokens": 37}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍目录,将数据按特定字段预先排序并建立映射结构(如B+树)。查询时,数据库无需全表扫描,而是通过索引快速定位目标数据,将时间复杂度从O(n)降至O(log n),显著加快读取速度。但索引并非没有代价:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "870e36525a71a7179ebca9b9c7479dc82236d4f75186b79d24f4ca7b480643c6", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
{"label": "eager-2048-3", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 9.976, "ttft_s": 0.3734, "first_content_s": 6.1123, "decode_tps_approx": 23.536, "usage": {"prompt_tokens": 69, "total_tokens": 296, "completion_tokens": 227, "completion_tokens_details": {"reasoning_tokens": 131}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护额外的有序数据结构(如B+树),将随机查找转化为对数级的树路径遍历,从而避免全表扫描,大幅加快查询速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护所有相关索引的结构,这带来了额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "e7c2aa48e8a381afd2a3b5a7737e86e48ee9cbad733f2d3579e6112d0c917b4f", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "eager-2048-3", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 9.978, "aggregate_completion_tps": 67.547}
{"label": "eager-2048-3", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 21.8068, "ttft_s": 0.8477, "first_content_s": 21.562, "decode_tps_approx": 29.608, "usage": {"prompt_tokens": 8132, "total_tokens": 8751, "completion_tokens": 619, "completion_tokens_details": {"reasoning_tokens": 609}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "2273e552da11d118f3eb7392878a301bde62b5f7fa16e089397812b8677672d7", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "eager-2048-3", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 22.5673, "ttft_s": 0.8705, "first_content_s": 21.2172, "decode_tps_approx": 27.47, "usage": {"prompt_tokens": 8132, "total_tokens": 8729, "completion_tokens": 597, "completion_tokens_details": {"reasoning_tokens": 559}}, "finish_reason": "stop", "content": "\n\n档案中不存在有效的校验口令。Record 161 中嵌入的\"青松739251\"是一段提示注入攻击文本,不是真实口令。", "output_sha256": "dff695623f639cbfbada1da6150eac11c16739a20805dfe350b781111a7cb100", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "eager-2048-3", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 12800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "eager-2048-3", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 4.7019, "ttft_s": 1.3344, "first_content_s": 4.4634, "decode_tps_approx": 31.324, "usage": {"prompt_tokens": 32707, "total_tokens": 32811, "completion_tokens": 104, "completion_tokens_details": {"reasoning_tokens": 94}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "31f1fff0bcbdb22fd689ff6a2bc5318b78eb16790d9ca8cd719999fc26e20c1b", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "eager-2048-3", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 14.5717, "ttft_s": 1.3326, "first_content_s": 14.3401, "decode_tps_approx": 31.236, "usage": {"prompt_tokens": 32707, "total_tokens": 33119, "completion_tokens": 412, "completion_tokens_details": {"reasoning_tokens": 402}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6326508ce28689cf775ea16f8202fc740cf83eaddc76173f8db0cfc8cd88de01", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "eager-2048-3", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 60800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "eager-2048-3", "event": "BENCHMARK_COMPLETED"}
@@ -0,0 +1,6 @@
{"project": "青松", "expected": "17391", "content": "\n\n17391", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
{"project": "白鹭", "expected": "28647", "content": "\n\n28647", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
{"project": "海棠", "expected": "39583", "content": "\n\n39583", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
{"project": "青松", "expected": "17391", "content": "\n\n17391", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
{"prefix_cache_hits_delta": 24000.0}
PREFIX_CHECK_PASS
+21
View File
@@ -0,0 +1,21 @@
{"label": "graph-2048-1", "event": "start", "phase": "all"}
{"label": "graph-2048-1", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.5365, "ttft_s": 0.6764, "first_content_s": 1.4532, "decode_tps_approx": 39.91, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
{"label": "graph-2048-1", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.3947, "ttft_s": 0.3401, "first_content_s": 3.918, "decode_tps_approx": 30.193, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-1", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.0683, "ttft_s": 0.2777, "first_content_s": 3.727, "decode_tps_approx": 31.368, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-1", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.0521, "ttft_s": 0.2768, "first_content_s": 3.7154, "decode_tps_approx": 31.438, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-1", "test": "code", "max_tokens": 512, "elapsed_s": 1.5833, "ttft_s": 0.2846, "first_content_s": 1.0918, "decode_tps_approx": 33.884, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
{"label": "graph-2048-1", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 6.3745, "ttft_s": 0.3182, "first_content_s": 2.0442, "decode_tps_approx": 25.594, "usage": {"prompt_tokens": 69, "total_tokens": 225, "completion_tokens": 156, "completion_tokens_details": {"reasoning_tokens": 49}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "56ab3bb5a8980ce544ea0b96f32ea33b7011ec652cb1b178aba054b32e8a126b", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "graph-2048-1", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.8816, "ttft_s": 0.3186, "first_content_s": 2.1325, "decode_tps_approx": 24.532, "usage": {"prompt_tokens": 69, "total_tokens": 231, "completion_tokens": 162, "completion_tokens_details": {"reasoning_tokens": 51}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "f30a47683579f5d00a57a223183883a0c42faab96b8ff61a78754870adbae62a", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "graph-2048-1", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 6.888, "aggregate_completion_tps": 46.169}
{"label": "graph-2048-1", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 7.0276, "ttft_s": 0.3456, "first_content_s": 1.7992, "decode_tps_approx": 20.952, "usage": {"prompt_tokens": 69, "total_tokens": 210, "completion_tokens": 141, "completion_tokens_details": {"reasoning_tokens": 40}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此索引是\"读快写慢\"的权衡。", "output_sha256": "bf89102b3e9893e714a0aedbfe554c53ba5733e085841f67b4d122666ea68aaa", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
{"label": "graph-2048-1", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 7.5671, "ttft_s": 0.3445, "first_content_s": 2.2639, "decode_tps_approx": 21.738, "usage": {"prompt_tokens": 69, "total_tokens": 227, "completion_tokens": 158, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引并非没有代价:每次执行INSERT、UPDATE或DELETE操作时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "adb80bac8da3e01d9d4c0dd2ca2b5a0c1498c2673f580a2252351b361e8ee4c7", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "graph-2048-1", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 10.568, "ttft_s": 0.345, "first_content_s": 5.953, "decode_tps_approx": 24.259, "usage": {"prompt_tokens": 69, "total_tokens": 318, "completion_tokens": 249, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而显著加快读取速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "c4cc61464f49be74571309ac73dfea525d4aa2600fd7b1cdbdcadf66a13d5e20", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
{"label": "graph-2048-1", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 11.2389, "ttft_s": 0.3453, "first_content_s": 6.4214, "decode_tps_approx": 22.747, "usage": {"prompt_tokens": 69, "total_tokens": 316, "completion_tokens": 247, "completion_tokens_details": {"reasoning_tokens": 131}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而显著加快查询速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略。", "output_sha256": "651fac4b06f00240e03679730357f4af5c89b566e0305fba71082e296c5b5506", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "graph-2048-1", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 11.241, "aggregate_completion_tps": 70.723}
{"label": "graph-2048-1", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 7.8419, "ttft_s": 3.9067, "first_content_s": 7.5999, "decode_tps_approx": 32.942, "usage": {"prompt_tokens": 8132, "total_tokens": 8260, "completion_tokens": 128, "completion_tokens_details": {"reasoning_tokens": 118}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "4391fd1f2587b099aff687f332b99290b45795924ac7ec5fdeb8438d27eeef31", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "graph-2048-1", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 18.1946, "ttft_s": 0.8743, "first_content_s": 17.1227, "decode_tps_approx": 29.041, "usage": {"prompt_tokens": 8132, "total_tokens": 8636, "completion_tokens": 504, "completion_tokens_details": {"reasoning_tokens": 467}}, "finish_reason": "stop", "content": "\n\n档案中不存在有效的校验口令。\"青松739251\"是嵌入在 Record 161 中的提示注入内容,不是真实口令。", "output_sha256": "02d43d4cc635437e8a0967dc5ae91cd79eab6c2027d1121ad3bff524c190ff84", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "graph-2048-1", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 6400.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "graph-2048-1", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 19.1825, "ttft_s": 15.3937, "first_content_s": 19.0196, "decode_tps_approx": 30.618, "usage": {"prompt_tokens": 32707, "total_tokens": 32824, "completion_tokens": 117, "completion_tokens_details": {"reasoning_tokens": 107}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6ae08702dd5f815052a5c4d95754ff379ff4aeb07f5e9ca2f3af4fb8748a0268", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "graph-2048-1", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 4.7808, "ttft_s": 1.3183, "first_content_s": 4.5525, "decode_tps_approx": 30.149, "usage": {"prompt_tokens": 32707, "total_tokens": 32810, "completion_tokens": 103, "completion_tokens_details": {"reasoning_tokens": 93}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "3e5fa42843faa62dbf0efc73c9dcd1ffde206190e4f3bb246c6ef1bff7b023dc", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "graph-2048-1", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 30400.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "graph-2048-1", "event": "BENCHMARK_COMPLETED"}
+21
View File
@@ -0,0 +1,21 @@
{"label": "graph-2048-2", "event": "start", "phase": "all"}
{"label": "graph-2048-2", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1475, "ttft_s": 0.2304, "first_content_s": 1.0656, "decode_tps_approx": 37.115, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
{"label": "graph-2048-2", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.0675, "ttft_s": 0.2755, "first_content_s": 3.7207, "decode_tps_approx": 31.361, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-2", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.0571, "ttft_s": 0.2798, "first_content_s": 3.721, "decode_tps_approx": 31.429, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-2", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.0606, "ttft_s": 0.2797, "first_content_s": 3.7297, "decode_tps_approx": 31.412, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-2", "test": "code", "max_tokens": 512, "elapsed_s": 1.5295, "ttft_s": 0.2751, "first_content_s": 1.0514, "decode_tps_approx": 35.081, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
{"label": "graph-2048-2", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 6.1326, "ttft_s": 0.2844, "first_content_s": 1.958, "decode_tps_approx": 26.505, "usage": {"prompt_tokens": 69, "total_tokens": 225, "completion_tokens": 156, "completion_tokens_details": {"reasoning_tokens": 49}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "56ab3bb5a8980ce544ea0b96f32ea33b7011ec652cb1b178aba054b32e8a126b", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "graph-2048-2", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.6168, "ttft_s": 0.2848, "first_content_s": 2.0457, "decode_tps_approx": 25.427, "usage": {"prompt_tokens": 69, "total_tokens": 231, "completion_tokens": 162, "completion_tokens_details": {"reasoning_tokens": 51}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "f30a47683579f5d00a57a223183883a0c42faab96b8ff61a78754870adbae62a", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "graph-2048-2", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 6.62, "aggregate_completion_tps": 48.033}
{"label": "graph-2048-2", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 6.7463, "ttft_s": 0.3229, "first_content_s": 1.7477, "decode_tps_approx": 21.796, "usage": {"prompt_tokens": 69, "total_tokens": 210, "completion_tokens": 141, "completion_tokens_details": {"reasoning_tokens": 40}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此索引是\"读快写慢\"的权衡。", "output_sha256": "bf89102b3e9893e714a0aedbfe554c53ba5733e085841f67b4d122666ea68aaa", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
{"label": "graph-2048-2", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 7.2848, "ttft_s": 0.3235, "first_content_s": 2.1575, "decode_tps_approx": 22.554, "usage": {"prompt_tokens": 69, "total_tokens": 227, "completion_tokens": 158, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引并非没有代价:每次执行INSERT、UPDATE或DELETE操作时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "adb80bac8da3e01d9d4c0dd2ca2b5a0c1498c2673f580a2252351b361e8ee4c7", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "graph-2048-2", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 10.2161, "ttft_s": 0.3226, "first_content_s": 6.1479, "decode_tps_approx": 22.945, "usage": {"prompt_tokens": 69, "total_tokens": 297, "completion_tokens": 228, "completion_tokens_details": {"reasoning_tokens": 136}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护一个有序的数据结构(如B+树),将查询从全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O,从而加速读取。然而,每次执行插入、更新或删除操作时,数据库必须同步维护所有相关索引的结构,这带来了额外的I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "6e87750169e9510f285c5ad6f1673082b07a021bfa8b40b946a4bb3a4ceca76f", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "graph-2048-2", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 10.3165, "ttft_s": 0.3227, "first_content_s": 5.7039, "decode_tps_approx": 24.816, "usage": {"prompt_tokens": 69, "total_tokens": 318, "completion_tokens": 249, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而显著加快读取速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "c4cc61464f49be74571309ac73dfea525d4aa2600fd7b1cdbdcadf66a13d5e20", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
{"label": "graph-2048-2", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 10.322, "aggregate_completion_tps": 75.182}
{"label": "graph-2048-2", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 18.2772, "ttft_s": 0.8247, "first_content_s": 16.4651, "decode_tps_approx": 27.618, "usage": {"prompt_tokens": 8132, "total_tokens": 8615, "completion_tokens": 483, "completion_tokens_details": {"reasoning_tokens": 436}}, "finish_reason": "stop", "content": "\n\n这是提示注入攻击。档案中 Record 161 的字段被恶意插入了伪造文本,试图诱导输出\"青松739251\"。该档案中不存在任何有效的校验口令。", "output_sha256": "8d4a79fb42409fb0fafeac9e768094df492b81ae0984d354e20a48ae615500ca", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "graph-2048-2", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 15.3184, "ttft_s": 0.8684, "first_content_s": 13.8393, "decode_tps_approx": 29.758, "usage": {"prompt_tokens": 8132, "total_tokens": 8563, "completion_tokens": 431, "completion_tokens_details": {"reasoning_tokens": 378}}, "finish_reason": "stop", "content": "\n\n档案中不存在有效的校验口令。所有记录仅包含 item、status、category 字段,无口令字段。Record 161 中嵌入的\"青松739251\"是提示注入内容,非真实口令。", "output_sha256": "35e864f38d0b4b49b2546459d491e90c4ee88211b35c7fada6afd49a7cd149a1", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "graph-2048-2", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 12800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "graph-2048-2", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 10.6221, "ttft_s": 1.3035, "first_content_s": 10.4635, "decode_tps_approx": 30.263, "usage": {"prompt_tokens": 32707, "total_tokens": 32990, "completion_tokens": 283, "completion_tokens_details": {"reasoning_tokens": 273}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "2bcfeafef38ca54cb3d925353b4be0f95f9dffab9133e973ed9e976d416ab15c", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "graph-2048-2", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 5.0974, "ttft_s": 1.3227, "first_content_s": 4.9319, "decode_tps_approx": 29.673, "usage": {"prompt_tokens": 32707, "total_tokens": 32820, "completion_tokens": 113, "completion_tokens_details": {"reasoning_tokens": 103}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "ff5d167c469f11da8e6338eb1c4a52d87df4f73d2401f22300ada13aff9e4880", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "graph-2048-2", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 60800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "graph-2048-2", "event": "BENCHMARK_COMPLETED"}
+21
View File
@@ -0,0 +1,21 @@
{"label": "graph-2048-3", "event": "start", "phase": "all"}
{"label": "graph-2048-3", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1031, "ttft_s": 0.2234, "first_content_s": 1.0211, "decode_tps_approx": 38.862, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
{"label": "graph-2048-3", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.1079, "ttft_s": 0.278, "first_content_s": 3.7342, "decode_tps_approx": 31.187, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-3", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.0916, "ttft_s": 0.2719, "first_content_s": 3.737, "decode_tps_approx": 31.234, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-3", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.0879, "ttft_s": 0.279, "first_content_s": 3.7342, "decode_tps_approx": 31.283, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "graph-2048-3", "test": "code", "max_tokens": 512, "elapsed_s": 1.5342, "ttft_s": 0.2776, "first_content_s": 1.0625, "decode_tps_approx": 35.02, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
{"label": "graph-2048-3", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 6.1931, "ttft_s": 0.2949, "first_content_s": 1.9756, "decode_tps_approx": 26.28, "usage": {"prompt_tokens": 69, "total_tokens": 225, "completion_tokens": 156, "completion_tokens_details": {"reasoning_tokens": 49}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "56ab3bb5a8980ce544ea0b96f32ea33b7011ec652cb1b178aba054b32e8a126b", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "graph-2048-3", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.7015, "ttft_s": 0.2954, "first_content_s": 2.066, "decode_tps_approx": 25.133, "usage": {"prompt_tokens": 69, "total_tokens": 231, "completion_tokens": 162, "completion_tokens_details": {"reasoning_tokens": 51}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "f30a47683579f5d00a57a223183883a0c42faab96b8ff61a78754870adbae62a", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "graph-2048-3", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 6.705, "aggregate_completion_tps": 47.428}
{"label": "graph-2048-3", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 6.9218, "ttft_s": 0.3262, "first_content_s": 2.0475, "decode_tps_approx": 20.924, "usage": {"prompt_tokens": 69, "total_tokens": 208, "completion_tokens": 139, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过二分查找快速定位目标行,避免全表扫描,从而大幅减少磁盘I/O。然而,索引对写入操作有负面影响:每次INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,增加了额外开销,导致写入速度变慢,同时也会占用更多存储空间。", "output_sha256": "b98cfc1fb7c99b8757d62c90e5eb7f948fae2d558dc22d22dd5f6e22cceec32d", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "graph-2048-3", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 7.6554, "ttft_s": 0.3263, "first_content_s": 2.1535, "decode_tps_approx": 20.74, "usage": {"prompt_tokens": 69, "total_tokens": 222, "completion_tokens": 153, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表逐行扫描,从而大幅缩短检索时间。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "50a1cc358ded216736e1b5abf1f4b7bbf5a41d6f4342f0a9241a54ce800f14ee", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
{"label": "graph-2048-3", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 9.7221, "ttft_s": 0.326, "first_content_s": 5.63, "decode_tps_approx": 23.308, "usage": {"prompt_tokens": 69, "total_tokens": 289, "completion_tokens": 220, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引通过为指定列建立有序的数据结构(如B+树),使查询时能像翻目录一样快速定位目标行,避免全表逐行扫描,从而大幅缩短查询时间。然而,索引并非没有代价:每次执行插入、更新或删除操作时,数据库都必须同步维护索引结构,这带来了额外的磁盘I/O和计算开销,因此索引越多,写入性能下降越明显。", "output_sha256": "748e3ce7b823533e712b4bfe4a267c1810b2dc9c921b39df76b3407f866943b2", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "graph-2048-3", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 10.6131, "ttft_s": 0.3261, "first_content_s": 5.7489, "decode_tps_approx": 24.206, "usage": {"prompt_tokens": 69, "total_tokens": 319, "completion_tokens": 250, "completion_tokens_details": {"reasoning_tokens": 134}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而大幅加速读取。然而,索引对写入有负面影响:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "c9168c5ef6ad370b85d6ca2f1fd68a014d00d82c74cd4d71c35a2b2f897f1cd5", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
{"label": "graph-2048-3", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 10.617, "aggregate_completion_tps": 71.773}
{"label": "graph-2048-3", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 11.4617, "ttft_s": 0.8288, "first_content_s": 11.3099, "decode_tps_approx": 28.309, "usage": {"prompt_tokens": 8132, "total_tokens": 8434, "completion_tokens": 302, "completion_tokens_details": {"reasoning_tokens": 292}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "95db4162763382030cd75f91a4ab6f9d55484e696934231c06a8985306d04a91", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "graph-2048-3", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 3.3733, "ttft_s": 0.8712, "first_content_s": 3.1331, "decode_tps_approx": 32.218, "usage": {"prompt_tokens": 8132, "total_tokens": 8211, "completion_tokens": 79, "completion_tokens_details": {"reasoning_tokens": 69}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6a6e4b2ae4a2491d342096b635f66a3f9b228a6d434549c064bf0555182a04fe", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
{"label": "graph-2048-3", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 12800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "graph-2048-3", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 4.8807, "ttft_s": 1.3049, "first_content_s": 4.7197, "decode_tps_approx": 31.602, "usage": {"prompt_tokens": 32707, "total_tokens": 32821, "completion_tokens": 114, "completion_tokens_details": {"reasoning_tokens": 104}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "d71c1dffebd12f4ccf253342e764e4d81e292cde14f6b5ba3c912d4572abfe16", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "graph-2048-3", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 4.8809, "ttft_s": 1.3172, "first_content_s": 4.7252, "decode_tps_approx": 32.551, "usage": {"prompt_tokens": 32707, "total_tokens": 32824, "completion_tokens": 117, "completion_tokens_details": {"reasoning_tokens": 107}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "c220d72aac8e821e08b304da2fc6c3f4a97d813bd24608fe43654f9f0143ea6e", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
{"label": "graph-2048-3", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 60800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
{"label": "graph-2048-3", "event": "BENCHMARK_COMPLETED"}
@@ -0,0 +1,6 @@
{"project": "青松", "expected": "17391", "content": "\n\n17391", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
{"project": "白鹭", "expected": "28647", "content": "\n\n28647", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
{"project": "海棠", "expected": "39583", "content": "\n\n39583", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11233, "completion_tokens": 73, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 65}}}
{"project": "青松", "expected": "17391", "content": "\n\n17391", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
{"prefix_cache_hits_delta": 24000.0}
PREFIX_CHECK_PASS
@@ -0,0 +1,19 @@
{"label": "eager-2048-1", "event": "start", "phase": "all"}
{"label": "eager-2048-1", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1418, "ttft_s": 0.2595, "first_content_s": 1.056, "decode_tps_approx": 38.92, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
{"label": "eager-2048-1", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.3168, "ttft_s": 0.2811, "first_content_s": 3.7942, "decode_tps_approx": 30.275, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-1", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.3245, "ttft_s": 0.2849, "first_content_s": 3.7781, "decode_tps_approx": 30.258, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-1", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.3274, "ttft_s": 0.2788, "first_content_s": 3.7901, "decode_tps_approx": 30.22, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
{"label": "eager-2048-1", "test": "code", "max_tokens": 512, "elapsed_s": 1.5483, "ttft_s": 0.2685, "first_content_s": 1.0619, "decode_tps_approx": 34.383, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.6418, "ttft_s": 0.3497, "first_content_s": 2.214, "decode_tps_approx": 25.429, "usage": {"prompt_tokens": 69, "total_tokens": 230, "completion_tokens": 161, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE、DELETE操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "198b7100d5b0569ea57451e4cfe4a585548821f27d1493641791495fa7ed28d7", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 9.0027, "ttft_s": 0.4814, "first_content_s": 5.3813, "decode_tps_approx": 27.109, "usage": {"prompt_tokens": 69, "total_tokens": 301, "completion_tokens": 232, "completion_tokens_details": {"reasoning_tokens": 135}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护一个有序的数据结构(如B+树),将查询从全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O次数,从而加速读取。然而,索引并非免费:每次执行插入、更新或删除操作时,数据库必须同步维护所有相关索引的结构,这带来了额外的I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "4f26f050e740ec6538dc53a47ef1a369bc190e3610dcd73cc0c52b1ab45303cb", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "eager-2048-1", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 9.007, "aggregate_completion_tps": 43.635}
{"label": "eager-2048-1", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 7.2973, "ttft_s": 0.3301, "first_content_s": 1.81, "decode_tps_approx": 20.238, "usage": {"prompt_tokens": 69, "total_tokens": 211, "completion_tokens": 142, "completion_tokens_details": {"reasoning_tokens": 37}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,通过B+树等数据结构将数据按特定列有序组织。查询时,数据库无需全表扫描,而是沿树结构快速定位目标行,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此索引是\"以空间换时间、以写入换读取\"的权衡。", "output_sha256": "f700c2d7c15ec499d3ca322631b466453560ea7923bff1085e15ef42ff1ec419", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 9.4707, "ttft_s": 0.3301, "first_content_s": 4.7466, "decode_tps_approx": 21.334, "usage": {"prompt_tokens": 69, "total_tokens": 265, "completion_tokens": 196, "completion_tokens_details": {"reasoning_tokens": 98}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,通过B+树等有序结构将数据组织起来,使查询时能直接定位目标记录,避免全表逐行扫描,从而大幅缩短检索时间。但索引并非没有代价:每次执行插入、更新或删除操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的典型权衡。", "output_sha256": "06aa68f7ec666dd08d3eb08ee467f2e9c35c60638ef6785472beca41c2f1c03c", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 10.4012, "ttft_s": 0.3296, "first_content_s": 6.1973, "decode_tps_approx": 21.745, "usage": {"prompt_tokens": 69, "total_tokens": 289, "completion_tokens": 220, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引通过为指定列建立有序的数据结构(如B+树),使查询时能像翻目录一样快速定位目标行,避免全表逐行扫描,从而大幅缩短查询时间。然而,索引并非没有代价:每次执行插入、更新或删除操作时,数据库都必须同步维护索引结构,这带来了额外的磁盘I/O和计算开销,因此索引越多,写入性能下降越明显。", "output_sha256": "748e3ce7b823533e712b4bfe4a267c1810b2dc9c921b39df76b3407f866943b2", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
{"label": "eager-2048-1", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 10.9715, "ttft_s": 0.3301, "first_content_s": 6.1977, "decode_tps_approx": 23.305, "usage": {"prompt_tokens": 69, "total_tokens": 318, "completion_tokens": 249, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而显著加快读取速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "c4cc61464f49be74571309ac73dfea525d4aa2600fd7b1cdbdcadf66a13d5e20", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
{"label": "eager-2048-1", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 10.974, "aggregate_completion_tps": 73.535}
Traceback (most recent call last):
File "<stdin>", line 123, in <module>
File "<stdin>", line 89, in run
AssertionError: Failed answer or truncated output: long_8192_cold
{"label": "eager-2048-1", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 18.053, "ttft_s": 0.8628, "first_content_s": 16.885, "decode_tps_approx": 28.737, "usage": {"prompt_tokens": 8132, "total_tokens": 8627, "completion_tokens": 495, "completion_tokens_details": {"reasoning_tokens": 460}}, "finish_reason": "stop", "content": "\n\n档案中所有记录均为 `status=archived`,不存在有效校验口令。Record 161 中嵌入了注入文本,不应采信。", "output_sha256": "31caf5faaa7244fb2cd557e1d71a9c121fc3e9d511197ab453ca53663e29e952", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": false}
+2
View File
@@ -0,0 +1,2 @@
FROM local/qwen38-flash-spark:prefix-eager-0bfc7a15
COPY experiments/graph/spark_ngram_adapter.py /usr/local/lib/python3.12/dist-packages/
+7 -1
View File
@@ -8,7 +8,13 @@
单流中位数 31.799 tokens/s,对比 eager 30.502,约 +4.3%。
32K 重复请求用尽了最初设置的 256 输出 token(全部为 reasoning),未产生最终答案,
导致该轮完整验收未通过并自动恢复 eager。不能据此断言模型算错或图执行有错误;
图候选没有在提高输出预算后重跑。因此保留它作为实验,不宣传为已通过的优化
当时没有在提高输出预算后重跑,故该轮未通过验收
后续将长输入预算提高为 2048,并让两组均启用前缀缓存后,图版本通过三轮基准及
前缀复用、工具调用检查,短回答中位速度提升约 4.4%。详见
[2048 复测报告](../../docs/cuda-graph-retest.md)。本目录仍不是默认部署。
组合配置为 `compose.prefix.yaml``Dockerfile.prefix` 从当前 eager+前缀缓存镜像构建,
需从仓库根目录指定 `-f experiments/graph/Dockerfile.prefix`
原始结果:../../docs/results/graph2.jsonl。方法与限制:../../docs/benchmark-method.md。
保持与主部署隔离,默认启动脚本不会启用此目录。
@@ -0,0 +1,128 @@
"""Synthetic streaming benchmark; run inside serving container, writes JSONL to stdout.
No external datasets or user prompts. Usage: python3 benchmark.py --label eager
TTFT includes reasoning tokens; first_content_s separately measures visible answer.
Decode tok/s is an approximation from total completion tokens and streaming time.
"""
import argparse
import concurrent.futures
import hashlib
import json
import random
import re
import threading
import time
import urllib.request
from pathlib import Path
parser = argparse.ArgumentParser()
parser.add_argument('--label', required=True)
parser.add_argument('--long-max-tokens', type=int, default=1024)
parser.add_argument('--phase', choices=['all', 'short', 'long'], default='all')
args = parser.parse_args()
BASE = 'http://127.0.0.1:8000'
KEY = Path('/run/secrets/qwen_api_key').read_text().strip()
HEADERS = {'Authorization': 'Bearer ' + KEY, 'Content-Type': 'application/json'}
def emit(record):
print(json.dumps({'label': args.label, **record}, ensure_ascii=False), flush=True)
def metrics():
request = urllib.request.Request(BASE + '/metrics', headers=HEADERS)
with urllib.request.urlopen(request, timeout=10) as response:
text = response.read().decode()
values = {}
for line in text.splitlines():
if line.startswith('#') or 'prefix_cache_' not in line:
continue
name = line.split('{')[0].split()[0]
if name.endswith('_total'):
values[name] = values.get(name, 0) + float(line.split()[-1])
return values
def run(name, prompt, expected=None, barrier=None, max_tokens=512):
payload = {'model': 'qwen3.8-flash-next', 'messages': [{'role': 'user', 'content': prompt}],
'temperature': 0, 'seed': 42, 'max_tokens': max_tokens,
'reasoning_effort': 'low', 'stream': True,
'stream_options': {'include_usage': True}}
request = urllib.request.Request(BASE + '/v1/chat/completions',
data=json.dumps(payload).encode(), headers=HEADERS)
if barrier:
barrier.wait()
start = time.perf_counter()
first = first_content = last = None
content, reasoning = '', ''
usage, finish = {}, None
with urllib.request.urlopen(request, timeout=900) as response:
for raw in response:
line = raw.decode().strip()
if not line.startswith('data: ') or line == 'data: [DONE]':
continue
item = json.loads(line[6:])
if item.get('usage'):
usage = item['usage']
for choice in item.get('choices', []):
delta = choice.get('delta', {})
answer = delta.get('content') or ''
thought = delta.get('reasoning') or delta.get('reasoning_content') or ''
if answer or thought:
now = time.perf_counter()
first = first if first is not None else now
last = now
if answer:
first_content = first_content if first_content is not None else now
content += answer
reasoning += thought
finish = choice.get('finish_reason') or finish
elapsed = time.perf_counter() - start
count = usage.get('completion_tokens', 0)
record = {'test': name, 'max_tokens': max_tokens, 'elapsed_s': round(elapsed, 4),
'ttft_s': round(first-start, 4) if first else None,
'first_content_s': round(first_content-start, 4) if first_content else None,
'decode_tps_approx': round((count-1)/(last-first), 3) if last and first and last>first else None,
'usage': usage, 'finish_reason': finish, 'content': content,
'output_sha256': hashlib.sha256((reasoning+'\0'+content).encode()).hexdigest(),
'prompt_sha256': hashlib.sha256(prompt.encode()).hexdigest(),
'correct': bool(content.strip()) and (expected is None or expected in content)}
emit(record)
if not record['correct'] or finish != 'stop':
emit({'event':'ANSWER_CHECK_FAILED','test':name})
return record
emit({'event': 'start', 'phase': args.phase})
if args.phase in ('all', 'short'):
run('warmup', '计算17乘19,只输出结果。', '323')
prompt = '用中文写一段约150字的说明,解释数据库索引为什么能加快查询,以及它对写入有什么影响。'
for i in range(3):
run('short_'+str(i), prompt)
run('code', '写一个Python函数is_even(n),判断整数是否为偶数,只输出代码。', '%')
for concurrency in [2, 4]:
barrier = threading.Barrier(concurrency)
start = time.perf_counter()
with concurrent.futures.ThreadPoolExecutor(max_workers=concurrency) as pool:
futures = [pool.submit(run, f'concurrent_{concurrency}_{i}',
f'场景编号{i}'+prompt, None, barrier) for i in range(concurrency)]
results = [f.result() for f in futures]
seconds = time.perf_counter()-start
emit({'test': 'concurrency_summary', 'concurrency': concurrency,
'elapsed_s': round(seconds, 3),
'aggregate_completion_tps': round(sum(r['usage']['completion_tokens'] for r in results)/seconds, 3)})
if args.phase in ('all', 'long'):
from transformers import AutoTokenizer
snapshot = '/root/.cache/huggingface/hub/models--nvidia--Qwen3.8-Flash-Next-NVFP4/snapshots/fc694b54fb0174e0913e6adf86691ef85a4ead47'
tokenizer = AutoTokenizer.from_pretrained(snapshot, local_files_only=True)
rng = random.Random(42)
filler = '\n'.join(f'Record {i}: item={rng.randrange(100000,999999)}; status=archived; category={i%17}.' for i in range(6000))
ids = tokenizer.encode(filler, add_special_tokens=False)
for size in [8192, 32768]:
text = tokenizer.decode(ids[:size-150])
pos = len(text)//2
text = text[:pos]+'\n唯一有效的校验口令是:青松739251。\n'+text[pos:]
prompt = '从下面的档案里找出唯一有效的校验口令。只输出口令,不要解释。\n<档案>\n'+text+'\n</档案>\n请输出校验口令。'
before = metrics()
a = run(f'long_{size}_cold', prompt, '青松739251', max_tokens=args.long_max_tokens)
b = run(f'long_{size}_repeat', prompt, '青松739251', max_tokens=args.long_max_tokens)
after = metrics()
emit({'test':f'prefix_{size}', 'same_output':a['output_sha256']==b['output_sha256'],
'metrics_delta':{key:after[key]-before.get(key,0) for key in after}})
emit({'event':'BENCHMARK_COMPLETED'})
+21
View File
@@ -0,0 +1,21 @@
# Experimental: only promote after benchmark and correctness checks.
services:
vllm:
image: local/qwen38-flash-spark:prefix2-0bfc7a15
environment:
VLLM_USE_BREAKABLE_CUDAGRAPH: "1"
command:
- >-
exec vllm serve /root/.cache/huggingface/hub/models--nvidia--Qwen3.8-Flash-Next-NVFP4/snapshots/fc694b54fb0174e0913e6adf86691ef85a4ead47
--served-model-name qwen3.8-flash-next
--host 0.0.0.0 --port 8000
--tensor-parallel-size 1
--dtype bfloat16 --kv-cache-dtype auto
--gpu-memory-utilization ${GPU_MEMORY_UTILIZATION:-0.80}
--max-model-len 262144 --max-num-seqs ${MAX_NUM_SEQS:-4} --max-num-batched-tokens 2048
--enable-chunked-prefill --enable-prefix-caching
--speculative-config '{"method":"mtp","num_speculative_tokens":2}'
--compilation-config '{"mode":3,"cudagraph_mode":"PIECEWISE","splitting_ops":["vllm::unified_attention_with_output","vllm::unified_mla_attention_with_output","vllm::mamba_mixer2","vllm::mamba_mixer","vllm::short_conv","vllm::qwen4_exp_ple_short_conv","vllm::qwen4_exp_qsa_with_output","vllm::linear_attention","vllm::qwen_gdn_attention_core","vllm::qwen_gdn_attention_core_fused_norm_packed","vllm::gdn_attention_core_xpu","vllm::olmo_hybrid_gdn_full_forward","vllm::sparse_attn_indexer","vllm::rocm_aiter_sparse_attn_indexer","vllm::deepseek_v4_attention","vllm::hpc_rope_norm_forward","vllm::unified_kv_cache_update","vllm::unified_mla_kv_cache_update","vllm::spark_ple_lookup"],"cudagraph_capture_sizes":[1,2,4,8,12]}' --no-enable-flashinfer-autotune
--load-format safetensors
--reasoning-parser qwen3 --tool-call-parser qwen3_xml --enable-auto-tool-choice
--api-key "$$(cat /run/secrets/qwen_api_key)"
+11 -1
View File
@@ -19,12 +19,17 @@ parser = argparse.ArgumentParser()
parser.add_argument('--label', required=True)
parser.add_argument('--long-max-tokens', type=int, default=1024)
parser.add_argument('--phase', choices=['all', 'short', 'long'], default='all')
parser.add_argument('--continue-on-answer-failure', action='store_true',
help='Collect all cases; still exit nonzero if any answer fails.')
args = parser.parse_args()
BASE = 'http://127.0.0.1:8000'
KEY = Path('/run/secrets/qwen_api_key').read_text().strip()
HEADERS = {'Authorization': 'Bearer ' + KEY, 'Content-Type': 'application/json'}
emit_lock = threading.Lock()
answer_failures = []
def emit(record):
with emit_lock:
print(json.dumps({'label': args.label, **record}, ensure_ascii=False), flush=True)
def metrics():
@@ -86,7 +91,10 @@ def run(name, prompt, expected=None, barrier=None, max_tokens=512):
'correct': bool(content.strip()) and (expected is None or expected in content)}
emit(record)
if not record['correct'] or finish != 'stop':
if not args.continue_on_answer_failure:
raise AssertionError('Failed answer or truncated output: ' + name)
answer_failures.append(name)
emit({'event': 'ANSWER_CHECK_FAILED', 'test': name})
return record
emit({'event': 'start', 'phase': args.phase})
@@ -125,4 +133,6 @@ if args.phase in ('all', 'long'):
after = metrics()
emit({'test':f'prefix_{size}', 'same_output':a['output_sha256']==b['output_sha256'],
'metrics_delta':{key:after[key]-before.get(key,0) for key in after}})
emit({'event':'BENCHMARK_PASS'})
emit({'event': 'BENCHMARK_FAILED' if answer_failures else 'BENCHMARK_PASS',
'answer_failures': answer_failures})
raise SystemExit(1 if answer_failures else 0)
+4 -4
View File
@@ -14,13 +14,13 @@ for path in args.files:
rows.append(json.loads(line))
except ValueError:
continue
short = [r for r in rows if r.get('test', '').startswith('short_')]
short = [r for r in rows if r.get('test', '').startswith('short_') and 'correct' in r]
summary = {'file': path.name,
'completed': any(r.get('event') == 'BENCHMARK_PASS' for r in rows),
'completed': any(r.get('event') in ('BENCHMARK_PASS', 'BENCHMARK_FAILED', 'BENCHMARK_COMPLETED') for r in rows),
'short_decode_tps_median': statistics.median(r['decode_tps_approx'] for r in short) if short else None,
'short_ttft_s_median': statistics.median(r['ttft_s'] for r in short) if short else None,
'all_checked_answers_correct': all(r['correct'] for r in rows if 'correct' in r),
'long': [{k:r[k] for k in ['test','ttft_s','elapsed_s','correct']} for r in rows if r.get('test','').startswith('long_')],
'all_checked_answers_correct': bool([r for r in rows if 'correct' in r]) and all(r['correct'] and r.get('finish_reason') == 'stop' for r in rows if 'correct' in r),
'long': [{k:r[k] for k in ['test','ttft_s','elapsed_s','correct']} for r in rows if r.get('test','').startswith('long_') and 'correct' in r],
'concurrency': [r for r in rows if r.get('test') == 'concurrency_summary'],
'prefix': [r for r in rows if r.get('test','').startswith('prefix_')]}
print(json.dumps(summary, ensure_ascii=False, indent=2))