ruihan
|
e290793f2b
|
Set 128K eager baseline and organize deployment documentation
|
2026-09-18 01:33:57 +08:00 |
|
ruihan
|
d7e1e745c3
|
docs(perf): profile CUDA graph coverage and PLE costs on Spark
|
2026-09-18 01:22:13 +08:00 |
|
ruihan
|
f1a8964072
|
Document long-context limits and host OOM at near-262K input
|
2026-09-18 00:03:19 +08:00 |
|
ruihan
|
48a1c9d7c4
|
Record successful 2048-token CUDA Graph retest and preserve failure samples
|
2026-09-17 23:04:40 +08:00 |
|
ruihan
|
1e2f48d1a4
|
Enable validated prefix caching and record DGX Spark optimization benchmarks
|
2026-09-17 22:09:58 +08:00 |
|
ruihan
|
5f3030260e
|
Add reproducible DGX Spark deployment for Qwen3.8 Flash Next NVFP4
|
2026-09-17 16:45:50 +08:00 |
|