This website requires JavaScript.
3059ff285a
Merge pull request 'Promote validated 128K Spark baseline and consolidate performance evidence' (#1 ) from perf/compile-prefix-cache into main
main
ruihan
2026-09-18 01:52:33 +08:00
7f3464fb05
Finalize migration acceptance summary
perf/compile-prefix-cache
ruihan
2026-09-18 01:51:17 +08:00
73a88a4e8c
Record Spark 128K migration acceptance and archive inventory
ruihan
2026-09-18 01:50:30 +08:00
e290793f2b
Set 128K eager baseline and organize deployment documentation
ruihan
2026-09-18 01:33:57 +08:00
d7e1e745c3
docs(perf): profile CUDA graph coverage and PLE costs on Spark
ruihan
2026-09-18 01:22:13 +08:00
f1a8964072
Document long-context limits and host OOM at near-262K input
ruihan
2026-09-18 00:03:19 +08:00
48a1c9d7c4
Record successful 2048-token CUDA Graph retest and preserve failure samples
ruihan
2026-09-17 23:04:40 +08:00
1e2f48d1a4
Enable validated prefix caching and record DGX Spark optimization benchmarks
ruihan
2026-09-17 22:09:58 +08:00
5f3030260e
Add reproducible DGX Spark deployment for Qwen3.8 Flash Next NVFP4
ruihan
2026-09-17 16:45:50 +08:00