Logo
Explore Help
Register Sign In
ruihan/qwen38-flash-next-dgx-spark
Watch 1
Star 0
Fork 0
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
9 Commits 2 Branches 0 Tags
3059ff285ac874204bedbb6eae1f964d961e095c
Commit Graph
9 Commits
This Branch
This Branch
All Branches
Author SHA1 Message Date
ruihan 3059ff285a Merge pull request 'Promote validated 128K Spark baseline and consolidate performance evidence' (#1) from perf/compile-prefix-cache into main 2026-09-18 01:52:33 +08:00
ruihan 7f3464fb05 Finalize migration acceptance summary 2026-09-18 01:51:17 +08:00
ruihan 73a88a4e8c Record Spark 128K migration acceptance and archive inventory 2026-09-18 01:50:30 +08:00
ruihan e290793f2b Set 128K eager baseline and organize deployment documentation 2026-09-18 01:33:57 +08:00
ruihan d7e1e745c3 docs(perf): profile CUDA graph coverage and PLE costs on Spark 2026-09-18 01:22:13 +08:00
ruihan f1a8964072 Document long-context limits and host OOM at near-262K input 2026-09-18 00:03:19 +08:00
ruihan 48a1c9d7c4 Record successful 2048-token CUDA Graph retest and preserve failure samples 2026-09-17 23:04:40 +08:00
ruihan 1e2f48d1a4 Enable validated prefix caching and record DGX Spark optimization benchmarks 2026-09-17 22:09:58 +08:00
ruihan 5f3030260e Add reproducible DGX Spark deployment for Qwen3.8 Flash Next NVFP4 2026-09-17 16:45:50 +08:00
Powered by Gitea Version: 1.27.3 Page: 8ms Template: 1ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API