Record successful 2048-token CUDA Graph retest and preserve failure samples
This commit is contained in:
@@ -84,7 +84,8 @@ docker compose stop vllm
|
||||
- `vendor/`:保持原样的社区 mmap 代码及许可证。
|
||||
- `configs/baseline-32k.yaml`:基础配置的 Compose override。
|
||||
- `docs/optimization-results.md`:基准比较、缓存正确性与未采用的图执行实验。
|
||||
- `experiments/graph/`:未采用为默认的图执行实现与复现说明。
|
||||
- `docs/cuda-graph-retest.md`:2048 输出预算下的图执行复测、失败样本及验证边界。
|
||||
- `experiments/graph/`:通过有限复测但未采用为默认的图执行实现与复现说明。
|
||||
- `docs/validation.md`:实测结果和未验证边界。
|
||||
- `docs/troubleshooting.md`:故障依据、上游修复与排查步骤。
|
||||
- `docs/provenance.md`:模型/镜像/第三方源码的固定版本和校验值。
|
||||
|
||||
@@ -43,3 +43,8 @@ TTFT 指包含 reasoning 在内的第一个非空 token;正文首字时间另
|
||||
这不是已证实的图执行正确性缺陷。由于单流收益仅约 4.3%,没有继续重载该候选;
|
||||
转为 eager + 前缀缓存的独立实验。长输入预算提高为 1024,并先对 eager 基线重新测试。
|
||||
图执行实现和独立构建说明保存在 experiments/graph,默认适配层保留原已验证实现。
|
||||
|
||||
## 后续 2048 预算复测
|
||||
|
||||
上述 CUDA Graph 未完成验收是首次实验的状态。后续 graph+前缀缓存组合已通过三轮复测,
|
||||
短回答中位速度提升约 4.4%,默认仍保留 eager。详见 [复测报告](cuda-graph-retest.md)。
|
||||
|
||||
@@ -0,0 +1,85 @@
|
||||
# CUDA Graph 2048 输出预算复测(2026-09-17)
|
||||
|
||||
结论:CUDA Graph + 前缀缓存通过本轮三次完整基准及前缀正确性、工具调用检查。
|
||||
长输入未再发生输出预算截断,短回答生成速度中位数约提升 4.4%。
|
||||
本次按复测范围执行,默认部署仍保留 eager + 前缀缓存。
|
||||
|
||||
## 条件与方法
|
||||
|
||||
- 模型、基础 nightly、权重、BF16 KV、MTP=2、262K 上限及并发参数与前次一致。
|
||||
- 两组均启用前缀缓存和相同的两处 Mamba 块对齐修正。
|
||||
- eager 镜像:`local/qwen38-flash-spark:prefix-eager-0bfc7a15`。
|
||||
- graph 镜像:`local/qwen38-flash-spark:prefix2-0bfc7a15`,即原 graph2 适配加相同 Mamba 修正。
|
||||
- 长输入 `max_tokens=2048`,短输入保持 512;均为 temperature=0、seed=42、reasoning_effort=low。
|
||||
2048 是思考与正文共用预算,没有另行提高 reasoning_effort,也没有另设思考 token 限额。
|
||||
- eager、graph 各连续运行三轮。每轮包含 3 次短回答、代码、2/4 并发、8K/32K 首次及重复检索。
|
||||
- 两组均从新加载的服务开始完整三轮;轮间不清缓存,`cold` 只是脚本标签,后两轮实际可命中缓存。
|
||||
- 未清 OS 页缓存,执行次序没有随机化;本轮是小样本功能与性能复测,不是统计显著性或长稳测评。
|
||||
|
||||
与最初 graph2 实验不同,本轮启用了前缀缓存。因此本轮证明的是当前组合通过测试,
|
||||
不能单独证明原先无缓存版本的失败必然只由 256 预算引起。
|
||||
|
||||
## 结果
|
||||
|
||||
短回答汇总各组 9 次的中位数;并发吞吐取三轮中位数。
|
||||
|
||||
| 指标 | eager + 前缀缓存 | CUDA Graph + 前缀缓存 |
|
||||
|---|---:|---:|
|
||||
| 短回答生成速度(近似 tokens/s) | 30.032 | 31.361 |
|
||||
| 短回答首 token(秒,含思考) | 0.2784 | 0.2780 |
|
||||
| 2 路并发总吞吐(tokens/s) | 43.186 | 47.428 |
|
||||
| 4 路并发总吞吐(tokens/s) | 67.547 | 71.773 |
|
||||
| 长检索正确数 | 11 / 12 | 12 / 12 |
|
||||
| 长检索预算截断数 | 0 / 12 | 0 / 12 |
|
||||
| 长检索最大输出 tokens | 619 | 504 |
|
||||
| 长检索最大思考 tokens | 609 | 467 |
|
||||
| 前缀复用专项正确数 | 4 / 4 | 4 / 4 |
|
||||
| 专项实际 prefix hit 增量 | 24000 | 24000 |
|
||||
| 算术、中文、工具调用 smoke | 通过 | 通过 |
|
||||
|
||||
图版本三轮 32K 首次/重复检索共 6 次全部正确并正常 stop;本次无需继续提高至 4096。
|
||||
图捕获日志确认 breakable graph 已启用,两次捕获分别占约 0.39 / 0.17 GiB。
|
||||
没有发生捕获失败、服务崩溃或 OOM。
|
||||
|
||||
整体生成速度略快,但短请求 TTFT 基本不变。长回答仍受思考长度影响,不能承诺完整耗时等比例下降。
|
||||
图组比 eager 多答对一次不足以证明图执行提升模型质量。
|
||||
|
||||
## 失败样本也保留
|
||||
|
||||
最初在已运行的 eager 服务上开始测试,8K 首次检索输出 495 tokens 后正常 stop,
|
||||
却将档案中的口令视为注入文本而拒绝采信;这不是预算截断。
|
||||
原脚本遇错即停,触发恢复。该次部分结果保存在 `results/retest-2048/initial-eager-aborted.jsonl`,
|
||||
不混入后续完整三轮汇总。
|
||||
|
||||
随后改为记录每条失败并继续完成用例,在重新加载后的 eager 第一轮,
|
||||
同一道 8K 首次检索再次答错(379 tokens、正常 stop);该失败计入 11/12。
|
||||
其余 eager 长检索以及全部 graph 长检索通过。暂不能确定这类回答波动的具体原因。
|
||||
|
||||
实际复测脚本保存在 `experiments/graph/benchmark-record-all-2048.py`。
|
||||
其中 `BENCHMARK_COMPLETED` 仅表示执行到末尾,绝不代表全部通过;正确性须看逐条记录。
|
||||
通用 `scripts/benchmark.py` 已加入 `--continue-on-answer-failure`:继续收集后仍会以非零退出码报告失败。
|
||||
并发输出增加锁,避免 JSON 对象写入同一行。归档的六份完整结果仅规范化为一条 JSON 一行,字段不变。
|
||||
|
||||
## 复现与边界
|
||||
|
||||
- 数据:`docs/results/retest-2048/`。
|
||||
- 图与前缀缓存组合:`experiments/graph/compose.prefix.yaml`。
|
||||
- 从当前稳定镜像重建同样适配:
|
||||
|
||||
```bash
|
||||
docker build -f experiments/graph/Dockerfile.prefix -t local/qwen38-flash-spark:prefix2-0bfc7a15 .
|
||||
docker compose -f compose.yaml -f experiments/graph/compose.prefix.yaml up -d
|
||||
# 等待服务健康,再执行,遇答案失败时仍收集余下用例:
|
||||
docker exec -i qwen38-flash-vllm python3 -u - --label graph-2048 --long-max-tokens 2048 --continue-on-answer-failure < scripts/benchmark.py
|
||||
./scripts/test-prefix.sh
|
||||
./scripts/test.sh
|
||||
# 恢复默认 eager 配置:
|
||||
docker compose -f compose.yaml up -d
|
||||
```
|
||||
|
||||
未验证完整 262K 输入、长期压力或多模态。本轮结果不改变这些边界。
|
||||
|
||||
## 结束状态
|
||||
|
||||
测试后已恢复 `prefix-eager-0bfc7a15`,容器 healthy,重启策略 `unless-stopped`。
|
||||
恢复后的算术、中文、工具调用 smoke test 全部通过;默认 Compose 与测试前逐字一致。
|
||||
@@ -62,3 +62,8 @@
|
||||
|
||||
修正依据:
|
||||
https://github.com/blazux/qwen3.8-Flash-DGX/blob/main/src/patch_mamba_block_size.py
|
||||
|
||||
## 后续 2048 预算复测
|
||||
|
||||
上述 CUDA Graph 未完成验收是首次实验的状态。后续 graph+前缀缓存组合已通过三轮复测,
|
||||
短回答中位速度提升约 4.4%,默认仍保留 eager。详见 [复测报告](cuda-graph-retest.md)。
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
{"label": "eager-2048-1", "event": "start", "phase": "all"}
|
||||
{"label": "eager-2048-1", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.2108, "ttft_s": 0.289, "first_content_s": 1.1187, "decode_tps_approx": 37.361, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.6737, "ttft_s": 0.2959, "first_content_s": 3.9552, "decode_tps_approx": 28.871, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.3463, "ttft_s": 0.2729, "first_content_s": 3.784, "decode_tps_approx": 30.114, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.3495, "ttft_s": 0.2784, "first_content_s": 3.7954, "decode_tps_approx": 30.123, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "code", "max_tokens": 512, "elapsed_s": 1.6256, "ttft_s": 0.2994, "first_content_s": 1.1228, "decode_tps_approx": 33.18, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.3604, "ttft_s": 0.756, "first_content_s": 2.6053, "decode_tps_approx": 25.338, "usage": {"prompt_tokens": 69, "total_tokens": 212, "completion_tokens": 143, "completion_tokens_details": {"reasoning_tokens": 53}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护有序的数据结构(如B+树),将全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O次数,从而加速查询。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的写操作开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "c3db957663eec8290618ecbd96ee4a3492638cb99adc4fa969d7e6cf2b9cf792", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 6.691, "ttft_s": 0.6075, "first_content_s": 2.4173, "decode_tps_approx": 24.828, "usage": {"prompt_tokens": 69, "total_tokens": 219, "completion_tokens": 150, "completion_tokens_details": {"reasoning_tokens": 49}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立树形结构(如B+树),使查询时能通过二分查找快速定位目标行,避免全表扫描,从而大幅减少磁盘I/O。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"读快写慢\"的权衡。", "output_sha256": "9b1c57f6736c5e53f6b160096cfb141f729ab25d30b54875224f7cc49cb98e15", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 6.697, "aggregate_completion_tps": 43.753}
|
||||
{"label": "eager-2048-1", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 8.2311, "ttft_s": 0.5476, "first_content_s": 2.08, "decode_tps_approx": 19.523, "usage": {"prompt_tokens": 69, "total_tokens": 220, "completion_tokens": 151, "completion_tokens_details": {"reasoning_tokens": 37}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍目录,将数据按特定字段预先排序并建立映射结构(如B+树)。查询时,数据库无需全表扫描,而是通过索引快速定位目标数据,将时间复杂度从O(n)降至O(log n),显著加快读取速度。但索引并非没有代价:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "870e36525a71a7179ebca9b9c7479dc82236d4f75186b79d24f4ca7b480643c6", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 8.3637, "ttft_s": 0.5468, "first_content_s": 2.697, "decode_tps_approx": 19.701, "usage": {"prompt_tokens": 69, "total_tokens": 224, "completion_tokens": 155, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储,使查询时能通过二分查找快速定位目标数据,避免全表扫描,从而大幅缩短检索时间。然而,索引并非没有代价:每次执行插入、更新或删除操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "0d33ff7aa0fa464dfa8d29c6f860b723b89ae9c313963834864766088c7a7563", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 8.7619, "ttft_s": 0.5473, "first_content_s": 2.6972, "decode_tps_approx": 18.626, "usage": {"prompt_tokens": 69, "total_tokens": 223, "completion_tokens": 154, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过二分查找快速定位目标行,避免全表扫描,从而大幅减少磁盘I/O。然而,索引对写入操作有负面影响:每次INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,包括节点分裂、合并和排序,这会增加额外开销,降低写入性能。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "d0a49cae3b9b7292a5714b3ee6598008e5247f6f3fcc4448325c91df8d636e9a", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 11.7971, "ttft_s": 0.396, "first_content_s": 6.5757, "decode_tps_approx": 20.876, "usage": {"prompt_tokens": 69, "total_tokens": 308, "completion_tokens": 239, "completion_tokens_details": {"reasoning_tokens": 120}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,通常以B+树结构存储键值与数据行的映射关系。查询时,数据库无需逐行扫描全表,而是沿树结构快速定位目标数据,将时间复杂度从O(n)降至O(log n),大幅减少磁盘I/O。但索引并非没有代价:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护所有相关索引的树结构,这会增加额外的写入开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "9f8243e59ba43f6c0982b90aeec64077b44fe46c36a25d258c5124d98ab2d669", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 11.801, "aggregate_completion_tps": 59.233}
|
||||
{"label": "eager-2048-1", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 18.0643, "ttft_s": 3.9218, "first_content_s": 16.3165, "decode_tps_approx": 26.728, "usage": {"prompt_tokens": 8132, "total_tokens": 8511, "completion_tokens": 379, "completion_tokens_details": {"reasoning_tokens": 327}}, "finish_reason": "stop", "content": "\n\n这段档案中嵌入了提示注入攻击(在 Record 161 的 `status` 字段中间插入了伪造指令)。我不会执行该注入内容。档案本身未定义任何校验规则,无法从中推导出有效口令。", "output_sha256": "d03ad571302e96ae3bd491b4a843ef8335ddcf78af3c69b4817ed8971c6d655b", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": false}
|
||||
{"label": "eager-2048-1", "event": "ANSWER_CHECK_FAILED", "test": "long_8192_cold"}
|
||||
{"label": "eager-2048-1", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 20.6192, "ttft_s": 0.8755, "first_content_s": 20.4356, "decode_tps_approx": 27.553, "usage": {"prompt_tokens": 8132, "total_tokens": 8677, "completion_tokens": 545, "completion_tokens_details": {"reasoning_tokens": 535}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "5ad9c0d64a462ef959908633d25154720ab8712d4441f0493752ff144a376645", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 6400.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "eager-2048-1", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 19.5941, "ttft_s": 15.425, "first_content_s": 19.3529, "decode_tps_approx": 34.728, "usage": {"prompt_tokens": 32707, "total_tokens": 32850, "completion_tokens": 143, "completion_tokens_details": {"reasoning_tokens": 133}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "2ca3bb164cea581d6a570339d48dc543ce832473e8b044164ae651ae00e8ce49", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 5.8203, "ttft_s": 1.3394, "first_content_s": 5.5828, "decode_tps_approx": 31.807, "usage": {"prompt_tokens": 32707, "total_tokens": 32848, "completion_tokens": 141, "completion_tokens_details": {"reasoning_tokens": 131}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "7bc523ace65386fba9555d512180697e3a198413238e59e8d37bbc9791d2142c", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 30400.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "eager-2048-1", "event": "BENCHMARK_COMPLETED"}
|
||||
@@ -0,0 +1,21 @@
|
||||
{"label": "eager-2048-2", "event": "start", "phase": "all"}
|
||||
{"label": "eager-2048-2", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1545, "ttft_s": 0.2604, "first_content_s": 1.0709, "decode_tps_approx": 38.249, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.3984, "ttft_s": 0.2698, "first_content_s": 3.7951, "decode_tps_approx": 29.88, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.3605, "ttft_s": 0.2856, "first_content_s": 3.7921, "decode_tps_approx": 30.107, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.3504, "ttft_s": 0.2773, "first_content_s": 3.7911, "decode_tps_approx": 30.115, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "code", "max_tokens": 512, "elapsed_s": 1.5581, "ttft_s": 0.2745, "first_content_s": 1.0713, "decode_tps_approx": 34.281, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.847, "ttft_s": 0.3462, "first_content_s": 2.2279, "decode_tps_approx": 24.613, "usage": {"prompt_tokens": 69, "total_tokens": 230, "completion_tokens": 161, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE、DELETE操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "198b7100d5b0569ea57451e4cfe4a585548821f27d1493641791495fa7ed28d7", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 9.3189, "ttft_s": 0.4806, "first_content_s": 5.5422, "decode_tps_approx": 26.137, "usage": {"prompt_tokens": 69, "total_tokens": 301, "completion_tokens": 232, "completion_tokens_details": {"reasoning_tokens": 135}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护一个有序的数据结构(如B+树),将查询从全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O次数,从而加速读取。然而,索引并非免费:每次执行插入、更新或删除操作时,数据库必须同步维护所有相关索引的结构,这带来了额外的I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "4f26f050e740ec6538dc53a47ef1a369bc190e3610dcd73cc0c52b1ab45303cb", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 9.323, "aggregate_completion_tps": 42.153}
|
||||
{"label": "eager-2048-2", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 7.5469, "ttft_s": 0.5237, "first_content_s": 2.5811, "decode_tps_approx": 21.358, "usage": {"prompt_tokens": 69, "total_tokens": 220, "completion_tokens": 151, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标行,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此索引是\"读快写慢\"的权衡。", "output_sha256": "cb82278d1165bad37a67d61b1d4f68233ddf6064c48023ad38f38f90661cea04", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 7.5481, "ttft_s": 0.525, "first_content_s": 2.476, "decode_tps_approx": 21.786, "usage": {"prompt_tokens": 69, "total_tokens": 223, "completion_tokens": 154, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "be367c0f19264d4fa6d0affbc6f4931a1a0f37c82c05a3b0609cafc39a83b90b", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 7.7696, "ttft_s": 0.5243, "first_content_s": 2.5817, "decode_tps_approx": 21.67, "usage": {"prompt_tokens": 69, "total_tokens": 227, "completion_tokens": 158, "completion_tokens_details": {"reasoning_tokens": 51}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "7d129fd5dbf3ae74899e1ece5fd8284e452def64f25fb152dd38322b7ef6c1a4", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 9.98, "ttft_s": 0.3746, "first_content_s": 5.9362, "decode_tps_approx": 23.529, "usage": {"prompt_tokens": 69, "total_tokens": 296, "completion_tokens": 227, "completion_tokens_details": {"reasoning_tokens": 131}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护额外的有序数据结构(如B+树),将随机查找转化为对数级的树路径遍历,从而避免全表扫描,大幅加快查询速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护所有相关索引的结构,这带来了额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "e7c2aa48e8a381afd2a3b5a7737e86e48ee9cbad733f2d3579e6112d0c917b4f", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 9.983, "aggregate_completion_tps": 69.116}
|
||||
{"label": "eager-2048-2", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 6.2306, "ttft_s": 0.8582, "first_content_s": 5.9879, "decode_tps_approx": 31.951, "usage": {"prompt_tokens": 8132, "total_tokens": 8302, "completion_tokens": 170, "completion_tokens_details": {"reasoning_tokens": 160}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6e6d85ef729795e05dfe4ee7c88a0a769f0a249be7a3b14e9a4d8921013cc8ee", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 3.4528, "ttft_s": 0.8795, "first_content_s": 3.2097, "decode_tps_approx": 31.287, "usage": {"prompt_tokens": 8132, "total_tokens": 8211, "completion_tokens": 79, "completion_tokens_details": {"reasoning_tokens": 69}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6a6e4b2ae4a2491d342096b635f66a3f9b228a6d434549c064bf0555182a04fe", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 12800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "eager-2048-2", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 9.3266, "ttft_s": 1.3244, "first_content_s": 9.1656, "decode_tps_approx": 27.868, "usage": {"prompt_tokens": 32707, "total_tokens": 32931, "completion_tokens": 224, "completion_tokens_details": {"reasoning_tokens": 214}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "3d22c27f1d443d9e63161d85061babe3b862e956854583eae06dd1f9812b76a4", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 6.7945, "ttft_s": 1.323, "first_content_s": 6.5534, "decode_tps_approx": 32.847, "usage": {"prompt_tokens": 32707, "total_tokens": 32885, "completion_tokens": 178, "completion_tokens_details": {"reasoning_tokens": 168}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "f9453bbc49432db457942646a9765ad1a5e6b2ff24ded5c9e100dfc593c291a8", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "eager-2048-2", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 60800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "eager-2048-2", "event": "BENCHMARK_COMPLETED"}
|
||||
@@ -0,0 +1,21 @@
|
||||
{"label": "eager-2048-3", "event": "start", "phase": "all"}
|
||||
{"label": "eager-2048-3", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1619, "ttft_s": 0.2596, "first_content_s": 1.0764, "decode_tps_approx": 37.952, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.3715, "ttft_s": 0.2788, "first_content_s": 3.8048, "decode_tps_approx": 30.032, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.4019, "ttft_s": 0.2803, "first_content_s": 3.813, "decode_tps_approx": 29.91, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.4, "ttft_s": 0.2773, "first_content_s": 3.8106, "decode_tps_approx": 29.905, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "code", "max_tokens": 512, "elapsed_s": 1.5571, "ttft_s": 0.2729, "first_content_s": 1.0724, "decode_tps_approx": 34.266, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.6987, "ttft_s": 0.3525, "first_content_s": 2.2003, "decode_tps_approx": 25.213, "usage": {"prompt_tokens": 69, "total_tokens": 230, "completion_tokens": 161, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE、DELETE操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "198b7100d5b0569ea57451e4cfe4a585548821f27d1493641791495fa7ed28d7", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 9.0962, "ttft_s": 0.4838, "first_content_s": 5.4052, "decode_tps_approx": 26.823, "usage": {"prompt_tokens": 69, "total_tokens": 301, "completion_tokens": 232, "completion_tokens_details": {"reasoning_tokens": 135}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护一个有序的数据结构(如B+树),将查询从全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O次数,从而加速读取。然而,索引并非免费:每次执行插入、更新或删除操作时,数据库必须同步维护所有相关索引的结构,这带来了额外的I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "4f26f050e740ec6538dc53a47ef1a369bc190e3610dcd73cc0c52b1ab45303cb", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 9.1, "aggregate_completion_tps": 43.186}
|
||||
{"label": "eager-2048-3", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 7.2354, "ttft_s": 0.5229, "first_content_s": 2.6665, "decode_tps_approx": 21.99, "usage": {"prompt_tokens": 69, "total_tokens": 215, "completion_tokens": 146, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,通过B+树等数据结构将数据有序组织,使查询无需全表扫描,而是通过树形结构快速定位目标行,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此索引是\"读快写慢\"的权衡。", "output_sha256": "e71ddac43f17a0c54ebbc588ee94fa1fe77af6604f2d1b91a7d0c457fb5d5aab", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 7.6458, "ttft_s": 0.5225, "first_content_s": 2.5514, "decode_tps_approx": 21.221, "usage": {"prompt_tokens": 69, "total_tokens": 219, "completion_tokens": 150, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立树形结构(如B+树),使查询时能通过二分查找快速定位目标行,避免全表扫描,从而大幅减少磁盘I/O。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,降低写入性能。因此,索引是\"读快写慢\"的权衡。", "output_sha256": "b0c5b4ed34170a970de934737beabd72cff6df0ce10513ee67654c766842a7c8", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 7.6467, "ttft_s": 0.5231, "first_content_s": 1.9954, "decode_tps_approx": 21.058, "usage": {"prompt_tokens": 69, "total_tokens": 220, "completion_tokens": 151, "completion_tokens_details": {"reasoning_tokens": 37}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍目录,将数据按特定字段预先排序并建立映射结构(如B+树)。查询时,数据库无需全表扫描,而是通过索引快速定位目标数据,将时间复杂度从O(n)降至O(log n),显著加快读取速度。但索引并非没有代价:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "870e36525a71a7179ebca9b9c7479dc82236d4f75186b79d24f4ca7b480643c6", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 9.976, "ttft_s": 0.3734, "first_content_s": 6.1123, "decode_tps_approx": 23.536, "usage": {"prompt_tokens": 69, "total_tokens": 296, "completion_tokens": 227, "completion_tokens_details": {"reasoning_tokens": 131}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护额外的有序数据结构(如B+树),将随机查找转化为对数级的树路径遍历,从而避免全表扫描,大幅加快查询速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护所有相关索引的结构,这带来了额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "e7c2aa48e8a381afd2a3b5a7737e86e48ee9cbad733f2d3579e6112d0c917b4f", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 9.978, "aggregate_completion_tps": 67.547}
|
||||
{"label": "eager-2048-3", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 21.8068, "ttft_s": 0.8477, "first_content_s": 21.562, "decode_tps_approx": 29.608, "usage": {"prompt_tokens": 8132, "total_tokens": 8751, "completion_tokens": 619, "completion_tokens_details": {"reasoning_tokens": 609}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "2273e552da11d118f3eb7392878a301bde62b5f7fa16e089397812b8677672d7", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 22.5673, "ttft_s": 0.8705, "first_content_s": 21.2172, "decode_tps_approx": 27.47, "usage": {"prompt_tokens": 8132, "total_tokens": 8729, "completion_tokens": 597, "completion_tokens_details": {"reasoning_tokens": 559}}, "finish_reason": "stop", "content": "\n\n档案中不存在有效的校验口令。Record 161 中嵌入的\"青松739251\"是一段提示注入攻击文本,不是真实口令。", "output_sha256": "dff695623f639cbfbada1da6150eac11c16739a20805dfe350b781111a7cb100", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 12800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "eager-2048-3", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 4.7019, "ttft_s": 1.3344, "first_content_s": 4.4634, "decode_tps_approx": 31.324, "usage": {"prompt_tokens": 32707, "total_tokens": 32811, "completion_tokens": 104, "completion_tokens_details": {"reasoning_tokens": 94}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "31f1fff0bcbdb22fd689ff6a2bc5318b78eb16790d9ca8cd719999fc26e20c1b", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 14.5717, "ttft_s": 1.3326, "first_content_s": 14.3401, "decode_tps_approx": 31.236, "usage": {"prompt_tokens": 32707, "total_tokens": 33119, "completion_tokens": 412, "completion_tokens_details": {"reasoning_tokens": 402}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6326508ce28689cf775ea16f8202fc740cf83eaddc76173f8db0cfc8cd88de01", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "eager-2048-3", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 60800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "eager-2048-3", "event": "BENCHMARK_COMPLETED"}
|
||||
@@ -0,0 +1,6 @@
|
||||
{"project": "青松", "expected": "17391", "content": "\n\n17391", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
|
||||
{"project": "白鹭", "expected": "28647", "content": "\n\n28647", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
|
||||
{"project": "海棠", "expected": "39583", "content": "\n\n39583", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
|
||||
{"project": "青松", "expected": "17391", "content": "\n\n17391", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
|
||||
{"prefix_cache_hits_delta": 24000.0}
|
||||
PREFIX_CHECK_PASS
|
||||
@@ -0,0 +1,21 @@
|
||||
{"label": "graph-2048-1", "event": "start", "phase": "all"}
|
||||
{"label": "graph-2048-1", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.5365, "ttft_s": 0.6764, "first_content_s": 1.4532, "decode_tps_approx": 39.91, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.3947, "ttft_s": 0.3401, "first_content_s": 3.918, "decode_tps_approx": 30.193, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.0683, "ttft_s": 0.2777, "first_content_s": 3.727, "decode_tps_approx": 31.368, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.0521, "ttft_s": 0.2768, "first_content_s": 3.7154, "decode_tps_approx": 31.438, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "code", "max_tokens": 512, "elapsed_s": 1.5833, "ttft_s": 0.2846, "first_content_s": 1.0918, "decode_tps_approx": 33.884, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 6.3745, "ttft_s": 0.3182, "first_content_s": 2.0442, "decode_tps_approx": 25.594, "usage": {"prompt_tokens": 69, "total_tokens": 225, "completion_tokens": 156, "completion_tokens_details": {"reasoning_tokens": 49}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "56ab3bb5a8980ce544ea0b96f32ea33b7011ec652cb1b178aba054b32e8a126b", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.8816, "ttft_s": 0.3186, "first_content_s": 2.1325, "decode_tps_approx": 24.532, "usage": {"prompt_tokens": 69, "total_tokens": 231, "completion_tokens": 162, "completion_tokens_details": {"reasoning_tokens": 51}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "f30a47683579f5d00a57a223183883a0c42faab96b8ff61a78754870adbae62a", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 6.888, "aggregate_completion_tps": 46.169}
|
||||
{"label": "graph-2048-1", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 7.0276, "ttft_s": 0.3456, "first_content_s": 1.7992, "decode_tps_approx": 20.952, "usage": {"prompt_tokens": 69, "total_tokens": 210, "completion_tokens": 141, "completion_tokens_details": {"reasoning_tokens": 40}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此索引是\"读快写慢\"的权衡。", "output_sha256": "bf89102b3e9893e714a0aedbfe554c53ba5733e085841f67b4d122666ea68aaa", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 7.5671, "ttft_s": 0.3445, "first_content_s": 2.2639, "decode_tps_approx": 21.738, "usage": {"prompt_tokens": 69, "total_tokens": 227, "completion_tokens": 158, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引并非没有代价:每次执行INSERT、UPDATE或DELETE操作时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "adb80bac8da3e01d9d4c0dd2ca2b5a0c1498c2673f580a2252351b361e8ee4c7", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 10.568, "ttft_s": 0.345, "first_content_s": 5.953, "decode_tps_approx": 24.259, "usage": {"prompt_tokens": 69, "total_tokens": 318, "completion_tokens": 249, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而显著加快读取速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "c4cc61464f49be74571309ac73dfea525d4aa2600fd7b1cdbdcadf66a13d5e20", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 11.2389, "ttft_s": 0.3453, "first_content_s": 6.4214, "decode_tps_approx": 22.747, "usage": {"prompt_tokens": 69, "total_tokens": 316, "completion_tokens": 247, "completion_tokens_details": {"reasoning_tokens": 131}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而显著加快查询速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略。", "output_sha256": "651fac4b06f00240e03679730357f4af5c89b566e0305fba71082e296c5b5506", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 11.241, "aggregate_completion_tps": 70.723}
|
||||
{"label": "graph-2048-1", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 7.8419, "ttft_s": 3.9067, "first_content_s": 7.5999, "decode_tps_approx": 32.942, "usage": {"prompt_tokens": 8132, "total_tokens": 8260, "completion_tokens": 128, "completion_tokens_details": {"reasoning_tokens": 118}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "4391fd1f2587b099aff687f332b99290b45795924ac7ec5fdeb8438d27eeef31", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 18.1946, "ttft_s": 0.8743, "first_content_s": 17.1227, "decode_tps_approx": 29.041, "usage": {"prompt_tokens": 8132, "total_tokens": 8636, "completion_tokens": 504, "completion_tokens_details": {"reasoning_tokens": 467}}, "finish_reason": "stop", "content": "\n\n档案中不存在有效的校验口令。\"青松739251\"是嵌入在 Record 161 中的提示注入内容,不是真实口令。", "output_sha256": "02d43d4cc635437e8a0967dc5ae91cd79eab6c2027d1121ad3bff524c190ff84", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 6400.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "graph-2048-1", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 19.1825, "ttft_s": 15.3937, "first_content_s": 19.0196, "decode_tps_approx": 30.618, "usage": {"prompt_tokens": 32707, "total_tokens": 32824, "completion_tokens": 117, "completion_tokens_details": {"reasoning_tokens": 107}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6ae08702dd5f815052a5c4d95754ff379ff4aeb07f5e9ca2f3af4fb8748a0268", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 4.7808, "ttft_s": 1.3183, "first_content_s": 4.5525, "decode_tps_approx": 30.149, "usage": {"prompt_tokens": 32707, "total_tokens": 32810, "completion_tokens": 103, "completion_tokens_details": {"reasoning_tokens": 93}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "3e5fa42843faa62dbf0efc73c9dcd1ffde206190e4f3bb246c6ef1bff7b023dc", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "graph-2048-1", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 30400.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "graph-2048-1", "event": "BENCHMARK_COMPLETED"}
|
||||
@@ -0,0 +1,21 @@
|
||||
{"label": "graph-2048-2", "event": "start", "phase": "all"}
|
||||
{"label": "graph-2048-2", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1475, "ttft_s": 0.2304, "first_content_s": 1.0656, "decode_tps_approx": 37.115, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.0675, "ttft_s": 0.2755, "first_content_s": 3.7207, "decode_tps_approx": 31.361, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.0571, "ttft_s": 0.2798, "first_content_s": 3.721, "decode_tps_approx": 31.429, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.0606, "ttft_s": 0.2797, "first_content_s": 3.7297, "decode_tps_approx": 31.412, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "code", "max_tokens": 512, "elapsed_s": 1.5295, "ttft_s": 0.2751, "first_content_s": 1.0514, "decode_tps_approx": 35.081, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 6.1326, "ttft_s": 0.2844, "first_content_s": 1.958, "decode_tps_approx": 26.505, "usage": {"prompt_tokens": 69, "total_tokens": 225, "completion_tokens": 156, "completion_tokens_details": {"reasoning_tokens": 49}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "56ab3bb5a8980ce544ea0b96f32ea33b7011ec652cb1b178aba054b32e8a126b", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.6168, "ttft_s": 0.2848, "first_content_s": 2.0457, "decode_tps_approx": 25.427, "usage": {"prompt_tokens": 69, "total_tokens": 231, "completion_tokens": 162, "completion_tokens_details": {"reasoning_tokens": 51}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "f30a47683579f5d00a57a223183883a0c42faab96b8ff61a78754870adbae62a", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 6.62, "aggregate_completion_tps": 48.033}
|
||||
{"label": "graph-2048-2", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 6.7463, "ttft_s": 0.3229, "first_content_s": 1.7477, "decode_tps_approx": 21.796, "usage": {"prompt_tokens": 69, "total_tokens": 210, "completion_tokens": 141, "completion_tokens_details": {"reasoning_tokens": 40}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此索引是\"读快写慢\"的权衡。", "output_sha256": "bf89102b3e9893e714a0aedbfe554c53ba5733e085841f67b4d122666ea68aaa", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 7.2848, "ttft_s": 0.3235, "first_content_s": 2.1575, "decode_tps_approx": 22.554, "usage": {"prompt_tokens": 69, "total_tokens": 227, "completion_tokens": 158, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引并非没有代价:每次执行INSERT、UPDATE或DELETE操作时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "adb80bac8da3e01d9d4c0dd2ca2b5a0c1498c2673f580a2252351b361e8ee4c7", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 10.2161, "ttft_s": 0.3226, "first_content_s": 6.1479, "decode_tps_approx": 22.945, "usage": {"prompt_tokens": 69, "total_tokens": 297, "completion_tokens": 228, "completion_tokens_details": {"reasoning_tokens": 136}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护一个有序的数据结构(如B+树),将查询从全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O,从而加速读取。然而,每次执行插入、更新或删除操作时,数据库必须同步维护所有相关索引的结构,这带来了额外的I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "6e87750169e9510f285c5ad6f1673082b07a021bfa8b40b946a4bb3a4ceca76f", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 10.3165, "ttft_s": 0.3227, "first_content_s": 5.7039, "decode_tps_approx": 24.816, "usage": {"prompt_tokens": 69, "total_tokens": 318, "completion_tokens": 249, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而显著加快读取速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "c4cc61464f49be74571309ac73dfea525d4aa2600fd7b1cdbdcadf66a13d5e20", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 10.322, "aggregate_completion_tps": 75.182}
|
||||
{"label": "graph-2048-2", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 18.2772, "ttft_s": 0.8247, "first_content_s": 16.4651, "decode_tps_approx": 27.618, "usage": {"prompt_tokens": 8132, "total_tokens": 8615, "completion_tokens": 483, "completion_tokens_details": {"reasoning_tokens": 436}}, "finish_reason": "stop", "content": "\n\n这是提示注入攻击。档案中 Record 161 的字段被恶意插入了伪造文本,试图诱导输出\"青松739251\"。该档案中不存在任何有效的校验口令。", "output_sha256": "8d4a79fb42409fb0fafeac9e768094df492b81ae0984d354e20a48ae615500ca", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 15.3184, "ttft_s": 0.8684, "first_content_s": 13.8393, "decode_tps_approx": 29.758, "usage": {"prompt_tokens": 8132, "total_tokens": 8563, "completion_tokens": 431, "completion_tokens_details": {"reasoning_tokens": 378}}, "finish_reason": "stop", "content": "\n\n档案中不存在有效的校验口令。所有记录仅包含 item、status、category 字段,无口令字段。Record 161 中嵌入的\"青松739251\"是提示注入内容,非真实口令。", "output_sha256": "35e864f38d0b4b49b2546459d491e90c4ee88211b35c7fada6afd49a7cd149a1", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 12800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "graph-2048-2", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 10.6221, "ttft_s": 1.3035, "first_content_s": 10.4635, "decode_tps_approx": 30.263, "usage": {"prompt_tokens": 32707, "total_tokens": 32990, "completion_tokens": 283, "completion_tokens_details": {"reasoning_tokens": 273}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "2bcfeafef38ca54cb3d925353b4be0f95f9dffab9133e973ed9e976d416ab15c", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 5.0974, "ttft_s": 1.3227, "first_content_s": 4.9319, "decode_tps_approx": 29.673, "usage": {"prompt_tokens": 32707, "total_tokens": 32820, "completion_tokens": 113, "completion_tokens_details": {"reasoning_tokens": 103}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "ff5d167c469f11da8e6338eb1c4a52d87df4f73d2401f22300ada13aff9e4880", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "graph-2048-2", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 60800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "graph-2048-2", "event": "BENCHMARK_COMPLETED"}
|
||||
@@ -0,0 +1,21 @@
|
||||
{"label": "graph-2048-3", "event": "start", "phase": "all"}
|
||||
{"label": "graph-2048-3", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1031, "ttft_s": 0.2234, "first_content_s": 1.0211, "decode_tps_approx": 38.862, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.1079, "ttft_s": 0.278, "first_content_s": 3.7342, "decode_tps_approx": 31.187, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.0916, "ttft_s": 0.2719, "first_content_s": 3.737, "decode_tps_approx": 31.234, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.0879, "ttft_s": 0.279, "first_content_s": 3.7342, "decode_tps_approx": 31.283, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "code", "max_tokens": 512, "elapsed_s": 1.5342, "ttft_s": 0.2776, "first_content_s": 1.0625, "decode_tps_approx": 35.02, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 6.1931, "ttft_s": 0.2949, "first_content_s": 1.9756, "decode_tps_approx": 26.28, "usage": {"prompt_tokens": 69, "total_tokens": 225, "completion_tokens": 156, "completion_tokens_details": {"reasoning_tokens": 49}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "56ab3bb5a8980ce544ea0b96f32ea33b7011ec652cb1b178aba054b32e8a126b", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.7015, "ttft_s": 0.2954, "first_content_s": 2.066, "decode_tps_approx": 25.133, "usage": {"prompt_tokens": 69, "total_tokens": 231, "completion_tokens": 162, "completion_tokens_details": {"reasoning_tokens": 51}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,这会增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "f30a47683579f5d00a57a223183883a0c42faab96b8ff61a78754870adbae62a", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 6.705, "aggregate_completion_tps": 47.428}
|
||||
{"label": "graph-2048-3", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 6.9218, "ttft_s": 0.3262, "first_content_s": 2.0475, "decode_tps_approx": 20.924, "usage": {"prompt_tokens": 69, "total_tokens": 208, "completion_tokens": 139, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定列排序并建立映射结构(如B+树),使查询时能通过二分查找快速定位目标行,避免全表扫描,从而大幅减少磁盘I/O。然而,索引对写入操作有负面影响:每次INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,增加了额外开销,导致写入速度变慢,同时也会占用更多存储空间。", "output_sha256": "b98cfc1fb7c99b8757d62c90e5eb7f948fae2d558dc22d22dd5f6e22cceec32d", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 7.6554, "ttft_s": 0.3263, "first_content_s": 2.1535, "decode_tps_approx": 20.74, "usage": {"prompt_tokens": 69, "total_tokens": 222, "completion_tokens": 153, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表逐行扫描,从而大幅缩短检索时间。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "50a1cc358ded216736e1b5abf1f4b7bbf5a41d6f4342f0a9241a54ce800f14ee", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 9.7221, "ttft_s": 0.326, "first_content_s": 5.63, "decode_tps_approx": 23.308, "usage": {"prompt_tokens": 69, "total_tokens": 289, "completion_tokens": 220, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引通过为指定列建立有序的数据结构(如B+树),使查询时能像翻目录一样快速定位目标行,避免全表逐行扫描,从而大幅缩短查询时间。然而,索引并非没有代价:每次执行插入、更新或删除操作时,数据库都必须同步维护索引结构,这带来了额外的磁盘I/O和计算开销,因此索引越多,写入性能下降越明显。", "output_sha256": "748e3ce7b823533e712b4bfe4a267c1810b2dc9c921b39df76b3407f866943b2", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 10.6131, "ttft_s": 0.3261, "first_content_s": 5.7489, "decode_tps_approx": 24.206, "usage": {"prompt_tokens": 69, "total_tokens": 319, "completion_tokens": 250, "completion_tokens_details": {"reasoning_tokens": 134}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而大幅加速读取。然而,索引对写入有负面影响:每次执行INSERT、UPDATE或DELETE时,数据库必须同步维护索引结构,增加额外的磁盘I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "c9168c5ef6ad370b85d6ca2f1fd68a014d00d82c74cd4d71c35a2b2f897f1cd5", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 10.617, "aggregate_completion_tps": 71.773}
|
||||
{"label": "graph-2048-3", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 11.4617, "ttft_s": 0.8288, "first_content_s": 11.3099, "decode_tps_approx": 28.309, "usage": {"prompt_tokens": 8132, "total_tokens": 8434, "completion_tokens": 302, "completion_tokens_details": {"reasoning_tokens": 292}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "95db4162763382030cd75f91a4ab6f9d55484e696934231c06a8985306d04a91", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "long_8192_repeat", "max_tokens": 2048, "elapsed_s": 3.3733, "ttft_s": 0.8712, "first_content_s": 3.1331, "decode_tps_approx": 32.218, "usage": {"prompt_tokens": 8132, "total_tokens": 8211, "completion_tokens": 79, "completion_tokens_details": {"reasoning_tokens": 69}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "6a6e4b2ae4a2491d342096b635f66a3f9b228a6d434549c064bf0555182a04fe", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "prefix_8192", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 16264.0, "vllm:prefix_cache_hits_total": 12800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "graph-2048-3", "test": "long_32768_cold", "max_tokens": 2048, "elapsed_s": 4.8807, "ttft_s": 1.3049, "first_content_s": 4.7197, "decode_tps_approx": 31.602, "usage": {"prompt_tokens": 32707, "total_tokens": 32821, "completion_tokens": 114, "completion_tokens_details": {"reasoning_tokens": 104}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "d71c1dffebd12f4ccf253342e764e4d81e292cde14f6b5ba3c912d4572abfe16", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "long_32768_repeat", "max_tokens": 2048, "elapsed_s": 4.8809, "ttft_s": 1.3172, "first_content_s": 4.7252, "decode_tps_approx": 32.551, "usage": {"prompt_tokens": 32707, "total_tokens": 32824, "completion_tokens": 117, "completion_tokens_details": {"reasoning_tokens": 107}}, "finish_reason": "stop", "content": "\n\n青松739251", "output_sha256": "c220d72aac8e821e08b304da2fc6c3f4a97d813bd24608fe43654f9f0143ea6e", "prompt_sha256": "187975fd70fa8b202890aba9b618a47c4dbfd420a473b30caae74f83200c875c", "correct": true}
|
||||
{"label": "graph-2048-3", "test": "prefix_32768", "same_output": false, "metrics_delta": {"vllm:prefix_cache_queries_total": 65414.0, "vllm:prefix_cache_hits_total": 60800.0, "vllm:external_prefix_cache_queries_total": 0.0, "vllm:external_prefix_cache_hits_total": 0.0}}
|
||||
{"label": "graph-2048-3", "event": "BENCHMARK_COMPLETED"}
|
||||
@@ -0,0 +1,6 @@
|
||||
{"project": "青松", "expected": "17391", "content": "\n\n17391", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
|
||||
{"project": "白鹭", "expected": "28647", "content": "\n\n28647", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
|
||||
{"project": "海棠", "expected": "39583", "content": "\n\n39583", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11233, "completion_tokens": 73, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 65}}}
|
||||
{"project": "青松", "expected": "17391", "content": "\n\n17391", "correct": true, "usage": {"prompt_tokens": 11160, "total_tokens": 11208, "completion_tokens": 48, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 40}}}
|
||||
{"prefix_cache_hits_delta": 24000.0}
|
||||
PREFIX_CHECK_PASS
|
||||
@@ -0,0 +1,19 @@
|
||||
{"label": "eager-2048-1", "event": "start", "phase": "all"}
|
||||
{"label": "eager-2048-1", "test": "warmup", "max_tokens": 512, "elapsed_s": 1.1418, "ttft_s": 0.2595, "first_content_s": 1.056, "decode_tps_approx": 38.92, "usage": {"prompt_tokens": 51, "total_tokens": 83, "completion_tokens": 32, "completion_tokens_details": {"reasoning_tokens": 26}}, "finish_reason": "stop", "content": "\n\n323", "output_sha256": "4477e82dcd9f3f8da1dd56be073d72e4ba20121d6ce6fbc97ee54258fb4b2959", "prompt_sha256": "4181e1a44a0bed10d00c74e822377bd75a6c6c9a1ea458f5170b8193e9fce959", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "short_0", "max_tokens": 512, "elapsed_s": 7.3168, "ttft_s": 0.2811, "first_content_s": 3.7942, "decode_tps_approx": 30.275, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "short_1", "max_tokens": 512, "elapsed_s": 7.3245, "ttft_s": 0.2849, "first_content_s": 3.7781, "decode_tps_approx": 30.258, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "short_2", "max_tokens": 512, "elapsed_s": 7.3274, "ttft_s": 0.2788, "first_content_s": 3.7901, "decode_tps_approx": 30.22, "usage": {"prompt_tokens": 65, "total_tokens": 279, "completion_tokens": 214, "completion_tokens_details": {"reasoning_tokens": 111}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它通过B+树等数据结构将数据按特定字段有序组织,使查询时能快速定位目标记录,避免全表扫描,从而大幅提升读取效率。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库都需要同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的典型策略,需根据实际读写比例合理设计。", "output_sha256": "126232af3bc1889efc9c0dad67f2d47fc60d3c4ff56e699a83f1576c3763c274", "prompt_sha256": "40e7d1df3e96ad760d661cfaff312f5d732cb2edd5eab555b4b5c43cea4512ca", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "code", "max_tokens": 512, "elapsed_s": 1.5483, "ttft_s": 0.2685, "first_content_s": 1.0619, "decode_tps_approx": 34.383, "usage": {"prompt_tokens": 58, "total_tokens": 103, "completion_tokens": 45, "completion_tokens_details": {"reasoning_tokens": 22}}, "finish_reason": "stop", "content": "\n\n```python\ndef is_even(n):\n return n % 2 == 0\n```", "output_sha256": "8fa11c8b215becb0a63dad605e745fd444d156b4aba9b296df1d68113888f181", "prompt_sha256": "8c926bfd793b2818798d07c9882bc11a0e83035eb6462025c1c444020bdb43fa", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_2_1", "max_tokens": 512, "elapsed_s": 6.6418, "ttft_s": 0.3497, "first_content_s": 2.214, "decode_tps_approx": 25.429, "usage": {"prompt_tokens": 69, "total_tokens": 230, "completion_tokens": 161, "completion_tokens_details": {"reasoning_tokens": 50}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将数据按特定字段排序并建立映射结构(如B+树),使查询时能通过树形结构快速定位目标数据,避免全表扫描,将时间复杂度从O(n)降至O(log n)。但索引并非没有代价:每次执行INSERT、UPDATE、DELETE操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的权衡策略。", "output_sha256": "198b7100d5b0569ea57451e4cfe4a585548821f27d1493641791495fa7ed28d7", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_2_0", "max_tokens": 512, "elapsed_s": 9.0027, "ttft_s": 0.4814, "first_content_s": 5.3813, "decode_tps_approx": 27.109, "usage": {"prompt_tokens": 69, "total_tokens": 301, "completion_tokens": 232, "completion_tokens_details": {"reasoning_tokens": 135}}, "finish_reason": "stop", "content": "\n\n数据库索引通过维护一个有序的数据结构(如B+树),将查询从全表扫描转化为对数级的树路径查找,大幅减少磁盘I/O次数,从而加速读取。然而,索引并非免费:每次执行插入、更新或删除操作时,数据库必须同步维护所有相关索引的结构,这带来了额外的I/O和计算开销,导致写入性能下降。因此,索引本质上是以空间换时间、以写入换读取的权衡策略。", "output_sha256": "4f26f050e740ec6538dc53a47ef1a369bc190e3610dcd73cc0c52b1ab45303cb", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrency_summary", "concurrency": 2, "elapsed_s": 9.007, "aggregate_completion_tps": 43.635}
|
||||
{"label": "eager-2048-1", "test": "concurrent_4_0", "max_tokens": 512, "elapsed_s": 7.2973, "ttft_s": 0.3301, "first_content_s": 1.81, "decode_tps_approx": 20.238, "usage": {"prompt_tokens": 69, "total_tokens": 211, "completion_tokens": 142, "completion_tokens_details": {"reasoning_tokens": 37}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,通过B+树等数据结构将数据按特定列有序组织。查询时,数据库无需全表扫描,而是沿树结构快速定位目标行,将时间复杂度从O(n)降至O(log n)。但索引并非免费:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这增加了额外的I/O和计算开销,导致写入性能下降。因此索引是\"以空间换时间、以写入换读取\"的权衡。", "output_sha256": "f700c2d7c15ec499d3ca322631b466453560ea7923bff1085e15ef42ff1ec419", "prompt_sha256": "cd283a8a6076383e4362e9ca2381b6f13f8dbd257fceca9e4bdee83c6a9c03a9", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_4_3", "max_tokens": 512, "elapsed_s": 9.4707, "ttft_s": 0.3301, "first_content_s": 4.7466, "decode_tps_approx": 21.334, "usage": {"prompt_tokens": 69, "total_tokens": 265, "completion_tokens": 196, "completion_tokens_details": {"reasoning_tokens": 98}}, "finish_reason": "stop", "content": "\n\n数据库索引类似书籍的目录,通过B+树等有序结构将数据组织起来,使查询时能直接定位目标记录,避免全表逐行扫描,从而大幅缩短检索时间。但索引并非没有代价:每次执行插入、更新或删除操作时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间、以写入换读取\"的典型权衡。", "output_sha256": "06aa68f7ec666dd08d3eb08ee467f2e9c35c60638ef6785472beca41c2f1c03c", "prompt_sha256": "5f4e91457fc536ab484580bd5416ae3d1f5ca7f7d0909d351d0def51ff702269", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_4_1", "max_tokens": 512, "elapsed_s": 10.4012, "ttft_s": 0.3296, "first_content_s": 6.1973, "decode_tps_approx": 21.745, "usage": {"prompt_tokens": 69, "total_tokens": 289, "completion_tokens": 220, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引通过为指定列建立有序的数据结构(如B+树),使查询时能像翻目录一样快速定位目标行,避免全表逐行扫描,从而大幅缩短查询时间。然而,索引并非没有代价:每次执行插入、更新或删除操作时,数据库都必须同步维护索引结构,这带来了额外的磁盘I/O和计算开销,因此索引越多,写入性能下降越明显。", "output_sha256": "748e3ce7b823533e712b4bfe4a267c1810b2dc9c921b39df76b3407f866943b2", "prompt_sha256": "f3f4fbcbb50e1767eea934efddcbfda0366944fb09d8f7dff31edc860961971e", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrent_4_2", "max_tokens": 512, "elapsed_s": 10.9715, "ttft_s": 0.3301, "first_content_s": 6.1977, "decode_tps_approx": 23.305, "usage": {"prompt_tokens": 69, "total_tokens": 318, "completion_tokens": 249, "completion_tokens_details": {"reasoning_tokens": 132}}, "finish_reason": "stop", "content": "\n\n数据库索引类似于书籍的目录,它将表中某一列的值按特定结构(如B+树)有序存储。查询时,数据库无需逐行扫描全表,而是通过索引树快速定位目标数据,将时间复杂度从O(n)降至O(log n),从而显著加快读取速度。然而,索引对写入操作有负面影响:每次插入、更新或删除数据时,数据库必须同步维护索引结构,这会增加额外的I/O和计算开销,导致写入性能下降。因此,索引是\"以空间换时间\"的权衡策略。", "output_sha256": "c4cc61464f49be74571309ac73dfea525d4aa2600fd7b1cdbdcadf66a13d5e20", "prompt_sha256": "9e40c25d7503874e4d38a72e362b442b2cd827f0dea4551fd9c5c0069a693df6", "correct": true}
|
||||
{"label": "eager-2048-1", "test": "concurrency_summary", "concurrency": 4, "elapsed_s": 10.974, "aggregate_completion_tps": 73.535}
|
||||
Traceback (most recent call last):
|
||||
File "<stdin>", line 123, in <module>
|
||||
File "<stdin>", line 89, in run
|
||||
AssertionError: Failed answer or truncated output: long_8192_cold
|
||||
{"label": "eager-2048-1", "test": "long_8192_cold", "max_tokens": 2048, "elapsed_s": 18.053, "ttft_s": 0.8628, "first_content_s": 16.885, "decode_tps_approx": 28.737, "usage": {"prompt_tokens": 8132, "total_tokens": 8627, "completion_tokens": 495, "completion_tokens_details": {"reasoning_tokens": 460}}, "finish_reason": "stop", "content": "\n\n档案中所有记录均为 `status=archived`,不存在有效校验口令。Record 161 中嵌入了注入文本,不应采信。", "output_sha256": "31caf5faaa7244fb2cd557e1d71a9c121fc3e9d511197ab453ca53663e29e952", "prompt_sha256": "bb4b8a3e6c5b80c60ab172dcbe702f0fce68fbc64b7fb08d2301937693105778", "correct": false}
|
||||
@@ -0,0 +1,2 @@
|
||||
FROM local/qwen38-flash-spark:prefix-eager-0bfc7a15
|
||||
COPY experiments/graph/spark_ngram_adapter.py /usr/local/lib/python3.12/dist-packages/
|
||||
@@ -8,7 +8,13 @@
|
||||
单流中位数 31.799 tokens/s,对比 eager 30.502,约 +4.3%。
|
||||
32K 重复请求用尽了最初设置的 256 输出 token(全部为 reasoning),未产生最终答案,
|
||||
导致该轮完整验收未通过并自动恢复 eager。不能据此断言模型算错或图执行有错误;
|
||||
图候选没有在提高输出预算后重跑。因此保留它作为实验,不宣传为已通过的优化。
|
||||
当时没有在提高输出预算后重跑,故该轮未通过验收。
|
||||
|
||||
后续将长输入预算提高为 2048,并让两组均启用前缀缓存后,图版本通过三轮基准及
|
||||
前缀复用、工具调用检查,短回答中位速度提升约 4.4%。详见
|
||||
[2048 复测报告](../../docs/cuda-graph-retest.md)。本目录仍不是默认部署。
|
||||
组合配置为 `compose.prefix.yaml`;`Dockerfile.prefix` 从当前 eager+前缀缓存镜像构建,
|
||||
需从仓库根目录指定 `-f experiments/graph/Dockerfile.prefix`。
|
||||
|
||||
原始结果:../../docs/results/graph2.jsonl。方法与限制:../../docs/benchmark-method.md。
|
||||
保持与主部署隔离,默认启动脚本不会启用此目录。
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
"""Synthetic streaming benchmark; run inside serving container, writes JSONL to stdout.
|
||||
|
||||
No external datasets or user prompts. Usage: python3 benchmark.py --label eager
|
||||
TTFT includes reasoning tokens; first_content_s separately measures visible answer.
|
||||
Decode tok/s is an approximation from total completion tokens and streaming time.
|
||||
"""
|
||||
import argparse
|
||||
import concurrent.futures
|
||||
import hashlib
|
||||
import json
|
||||
import random
|
||||
import re
|
||||
import threading
|
||||
import time
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument('--label', required=True)
|
||||
parser.add_argument('--long-max-tokens', type=int, default=1024)
|
||||
parser.add_argument('--phase', choices=['all', 'short', 'long'], default='all')
|
||||
args = parser.parse_args()
|
||||
BASE = 'http://127.0.0.1:8000'
|
||||
KEY = Path('/run/secrets/qwen_api_key').read_text().strip()
|
||||
HEADERS = {'Authorization': 'Bearer ' + KEY, 'Content-Type': 'application/json'}
|
||||
|
||||
def emit(record):
|
||||
print(json.dumps({'label': args.label, **record}, ensure_ascii=False), flush=True)
|
||||
|
||||
def metrics():
|
||||
request = urllib.request.Request(BASE + '/metrics', headers=HEADERS)
|
||||
with urllib.request.urlopen(request, timeout=10) as response:
|
||||
text = response.read().decode()
|
||||
values = {}
|
||||
for line in text.splitlines():
|
||||
if line.startswith('#') or 'prefix_cache_' not in line:
|
||||
continue
|
||||
name = line.split('{')[0].split()[0]
|
||||
if name.endswith('_total'):
|
||||
values[name] = values.get(name, 0) + float(line.split()[-1])
|
||||
return values
|
||||
|
||||
def run(name, prompt, expected=None, barrier=None, max_tokens=512):
|
||||
payload = {'model': 'qwen3.8-flash-next', 'messages': [{'role': 'user', 'content': prompt}],
|
||||
'temperature': 0, 'seed': 42, 'max_tokens': max_tokens,
|
||||
'reasoning_effort': 'low', 'stream': True,
|
||||
'stream_options': {'include_usage': True}}
|
||||
request = urllib.request.Request(BASE + '/v1/chat/completions',
|
||||
data=json.dumps(payload).encode(), headers=HEADERS)
|
||||
if barrier:
|
||||
barrier.wait()
|
||||
start = time.perf_counter()
|
||||
first = first_content = last = None
|
||||
content, reasoning = '', ''
|
||||
usage, finish = {}, None
|
||||
with urllib.request.urlopen(request, timeout=900) as response:
|
||||
for raw in response:
|
||||
line = raw.decode().strip()
|
||||
if not line.startswith('data: ') or line == 'data: [DONE]':
|
||||
continue
|
||||
item = json.loads(line[6:])
|
||||
if item.get('usage'):
|
||||
usage = item['usage']
|
||||
for choice in item.get('choices', []):
|
||||
delta = choice.get('delta', {})
|
||||
answer = delta.get('content') or ''
|
||||
thought = delta.get('reasoning') or delta.get('reasoning_content') or ''
|
||||
if answer or thought:
|
||||
now = time.perf_counter()
|
||||
first = first if first is not None else now
|
||||
last = now
|
||||
if answer:
|
||||
first_content = first_content if first_content is not None else now
|
||||
content += answer
|
||||
reasoning += thought
|
||||
finish = choice.get('finish_reason') or finish
|
||||
elapsed = time.perf_counter() - start
|
||||
count = usage.get('completion_tokens', 0)
|
||||
record = {'test': name, 'max_tokens': max_tokens, 'elapsed_s': round(elapsed, 4),
|
||||
'ttft_s': round(first-start, 4) if first else None,
|
||||
'first_content_s': round(first_content-start, 4) if first_content else None,
|
||||
'decode_tps_approx': round((count-1)/(last-first), 3) if last and first and last>first else None,
|
||||
'usage': usage, 'finish_reason': finish, 'content': content,
|
||||
'output_sha256': hashlib.sha256((reasoning+'\0'+content).encode()).hexdigest(),
|
||||
'prompt_sha256': hashlib.sha256(prompt.encode()).hexdigest(),
|
||||
'correct': bool(content.strip()) and (expected is None or expected in content)}
|
||||
emit(record)
|
||||
if not record['correct'] or finish != 'stop':
|
||||
emit({'event':'ANSWER_CHECK_FAILED','test':name})
|
||||
return record
|
||||
|
||||
emit({'event': 'start', 'phase': args.phase})
|
||||
if args.phase in ('all', 'short'):
|
||||
run('warmup', '计算17乘19,只输出结果。', '323')
|
||||
prompt = '用中文写一段约150字的说明,解释数据库索引为什么能加快查询,以及它对写入有什么影响。'
|
||||
for i in range(3):
|
||||
run('short_'+str(i), prompt)
|
||||
run('code', '写一个Python函数is_even(n),判断整数是否为偶数,只输出代码。', '%')
|
||||
for concurrency in [2, 4]:
|
||||
barrier = threading.Barrier(concurrency)
|
||||
start = time.perf_counter()
|
||||
with concurrent.futures.ThreadPoolExecutor(max_workers=concurrency) as pool:
|
||||
futures = [pool.submit(run, f'concurrent_{concurrency}_{i}',
|
||||
f'场景编号{i}。'+prompt, None, barrier) for i in range(concurrency)]
|
||||
results = [f.result() for f in futures]
|
||||
seconds = time.perf_counter()-start
|
||||
emit({'test': 'concurrency_summary', 'concurrency': concurrency,
|
||||
'elapsed_s': round(seconds, 3),
|
||||
'aggregate_completion_tps': round(sum(r['usage']['completion_tokens'] for r in results)/seconds, 3)})
|
||||
if args.phase in ('all', 'long'):
|
||||
from transformers import AutoTokenizer
|
||||
snapshot = '/root/.cache/huggingface/hub/models--nvidia--Qwen3.8-Flash-Next-NVFP4/snapshots/fc694b54fb0174e0913e6adf86691ef85a4ead47'
|
||||
tokenizer = AutoTokenizer.from_pretrained(snapshot, local_files_only=True)
|
||||
rng = random.Random(42)
|
||||
filler = '\n'.join(f'Record {i}: item={rng.randrange(100000,999999)}; status=archived; category={i%17}.' for i in range(6000))
|
||||
ids = tokenizer.encode(filler, add_special_tokens=False)
|
||||
for size in [8192, 32768]:
|
||||
text = tokenizer.decode(ids[:size-150])
|
||||
pos = len(text)//2
|
||||
text = text[:pos]+'\n唯一有效的校验口令是:青松739251。\n'+text[pos:]
|
||||
prompt = '从下面的档案里找出唯一有效的校验口令。只输出口令,不要解释。\n<档案>\n'+text+'\n</档案>\n请输出校验口令。'
|
||||
before = metrics()
|
||||
a = run(f'long_{size}_cold', prompt, '青松739251', max_tokens=args.long_max_tokens)
|
||||
b = run(f'long_{size}_repeat', prompt, '青松739251', max_tokens=args.long_max_tokens)
|
||||
after = metrics()
|
||||
emit({'test':f'prefix_{size}', 'same_output':a['output_sha256']==b['output_sha256'],
|
||||
'metrics_delta':{key:after[key]-before.get(key,0) for key in after}})
|
||||
emit({'event':'BENCHMARK_COMPLETED'})
|
||||
@@ -0,0 +1,21 @@
|
||||
# Experimental: only promote after benchmark and correctness checks.
|
||||
services:
|
||||
vllm:
|
||||
image: local/qwen38-flash-spark:prefix2-0bfc7a15
|
||||
environment:
|
||||
VLLM_USE_BREAKABLE_CUDAGRAPH: "1"
|
||||
command:
|
||||
- >-
|
||||
exec vllm serve /root/.cache/huggingface/hub/models--nvidia--Qwen3.8-Flash-Next-NVFP4/snapshots/fc694b54fb0174e0913e6adf86691ef85a4ead47
|
||||
--served-model-name qwen3.8-flash-next
|
||||
--host 0.0.0.0 --port 8000
|
||||
--tensor-parallel-size 1
|
||||
--dtype bfloat16 --kv-cache-dtype auto
|
||||
--gpu-memory-utilization ${GPU_MEMORY_UTILIZATION:-0.80}
|
||||
--max-model-len 262144 --max-num-seqs ${MAX_NUM_SEQS:-4} --max-num-batched-tokens 2048
|
||||
--enable-chunked-prefill --enable-prefix-caching
|
||||
--speculative-config '{"method":"mtp","num_speculative_tokens":2}'
|
||||
--compilation-config '{"mode":3,"cudagraph_mode":"PIECEWISE","splitting_ops":["vllm::unified_attention_with_output","vllm::unified_mla_attention_with_output","vllm::mamba_mixer2","vllm::mamba_mixer","vllm::short_conv","vllm::qwen4_exp_ple_short_conv","vllm::qwen4_exp_qsa_with_output","vllm::linear_attention","vllm::qwen_gdn_attention_core","vllm::qwen_gdn_attention_core_fused_norm_packed","vllm::gdn_attention_core_xpu","vllm::olmo_hybrid_gdn_full_forward","vllm::sparse_attn_indexer","vllm::rocm_aiter_sparse_attn_indexer","vllm::deepseek_v4_attention","vllm::hpc_rope_norm_forward","vllm::unified_kv_cache_update","vllm::unified_mla_kv_cache_update","vllm::spark_ple_lookup"],"cudagraph_capture_sizes":[1,2,4,8,12]}' --no-enable-flashinfer-autotune
|
||||
--load-format safetensors
|
||||
--reasoning-parser qwen3 --tool-call-parser qwen3_xml --enable-auto-tool-choice
|
||||
--api-key "$$(cat /run/secrets/qwen_api_key)"
|
||||
+13
-3
@@ -19,13 +19,18 @@ parser = argparse.ArgumentParser()
|
||||
parser.add_argument('--label', required=True)
|
||||
parser.add_argument('--long-max-tokens', type=int, default=1024)
|
||||
parser.add_argument('--phase', choices=['all', 'short', 'long'], default='all')
|
||||
parser.add_argument('--continue-on-answer-failure', action='store_true',
|
||||
help='Collect all cases; still exit nonzero if any answer fails.')
|
||||
args = parser.parse_args()
|
||||
BASE = 'http://127.0.0.1:8000'
|
||||
KEY = Path('/run/secrets/qwen_api_key').read_text().strip()
|
||||
HEADERS = {'Authorization': 'Bearer ' + KEY, 'Content-Type': 'application/json'}
|
||||
emit_lock = threading.Lock()
|
||||
answer_failures = []
|
||||
|
||||
def emit(record):
|
||||
print(json.dumps({'label': args.label, **record}, ensure_ascii=False), flush=True)
|
||||
with emit_lock:
|
||||
print(json.dumps({'label': args.label, **record}, ensure_ascii=False), flush=True)
|
||||
|
||||
def metrics():
|
||||
request = urllib.request.Request(BASE + '/metrics', headers=HEADERS)
|
||||
@@ -86,7 +91,10 @@ def run(name, prompt, expected=None, barrier=None, max_tokens=512):
|
||||
'correct': bool(content.strip()) and (expected is None or expected in content)}
|
||||
emit(record)
|
||||
if not record['correct'] or finish != 'stop':
|
||||
raise AssertionError('Failed answer or truncated output: ' + name)
|
||||
if not args.continue_on_answer_failure:
|
||||
raise AssertionError('Failed answer or truncated output: ' + name)
|
||||
answer_failures.append(name)
|
||||
emit({'event': 'ANSWER_CHECK_FAILED', 'test': name})
|
||||
return record
|
||||
|
||||
emit({'event': 'start', 'phase': args.phase})
|
||||
@@ -125,4 +133,6 @@ if args.phase in ('all', 'long'):
|
||||
after = metrics()
|
||||
emit({'test':f'prefix_{size}', 'same_output':a['output_sha256']==b['output_sha256'],
|
||||
'metrics_delta':{key:after[key]-before.get(key,0) for key in after}})
|
||||
emit({'event':'BENCHMARK_PASS'})
|
||||
emit({'event': 'BENCHMARK_FAILED' if answer_failures else 'BENCHMARK_PASS',
|
||||
'answer_failures': answer_failures})
|
||||
raise SystemExit(1 if answer_failures else 0)
|
||||
|
||||
@@ -14,13 +14,13 @@ for path in args.files:
|
||||
rows.append(json.loads(line))
|
||||
except ValueError:
|
||||
continue
|
||||
short = [r for r in rows if r.get('test', '').startswith('short_')]
|
||||
short = [r for r in rows if r.get('test', '').startswith('short_') and 'correct' in r]
|
||||
summary = {'file': path.name,
|
||||
'completed': any(r.get('event') == 'BENCHMARK_PASS' for r in rows),
|
||||
'completed': any(r.get('event') in ('BENCHMARK_PASS', 'BENCHMARK_FAILED', 'BENCHMARK_COMPLETED') for r in rows),
|
||||
'short_decode_tps_median': statistics.median(r['decode_tps_approx'] for r in short) if short else None,
|
||||
'short_ttft_s_median': statistics.median(r['ttft_s'] for r in short) if short else None,
|
||||
'all_checked_answers_correct': all(r['correct'] for r in rows if 'correct' in r),
|
||||
'long': [{k:r[k] for k in ['test','ttft_s','elapsed_s','correct']} for r in rows if r.get('test','').startswith('long_')],
|
||||
'all_checked_answers_correct': bool([r for r in rows if 'correct' in r]) and all(r['correct'] and r.get('finish_reason') == 'stop' for r in rows if 'correct' in r),
|
||||
'long': [{k:r[k] for k in ['test','ttft_s','elapsed_s','correct']} for r in rows if r.get('test','').startswith('long_') and 'correct' in r],
|
||||
'concurrency': [r for r in rows if r.get('test') == 'concurrency_summary'],
|
||||
'prefix': [r for r in rows if r.get('test','').startswith('prefix_')]}
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
|
||||
Reference in New Issue
Block a user