Promote the validated Spark deployment to a 131072-token eager baseline with prefix caching and MTP=2. The previous 262K default triggered host OOM near full capacity; default, no-prefix fallback and optional graph configs now use 128K, while the conservative fallback remains 32K.
This merges the prefix-cache fixes, CUDA Graph comparisons, long-context limits and Nsight evidence. Documentation distinguishes current settings from historical experiments. Spark now uses a Git-managed layout; legacy files and raw traces are archived, local credentials remain untracked, and completed download containers plus obsolete experimental image tags were removed.
Validation on Spark: four Compose variants parsed; model reload, arithmetic/Chinese/tool-call smoke, prefix-cache check, and first/repeated retrieval on 127987 input tokens with a 2048 output budget all passed. Minimum sampled MemAvailable was 16.777 GiB, host OOM count stayed 25, and system/user failed units were zero. Service is healthy with unless-stopped. No CI workflows exist; these are recorded device checks, not CI claims. Full concurrency and long-duration stability remain unverified.
Promote the validated Spark deployment to a 131072-token eager baseline with prefix caching and MTP=2. The previous 262K default triggered host OOM near full capacity; default, no-prefix fallback and optional graph configs now use 128K, while the conservative fallback remains 32K.
This merges the prefix-cache fixes, CUDA Graph comparisons, long-context limits and Nsight evidence. Documentation distinguishes current settings from historical experiments. Spark now uses a Git-managed layout; legacy files and raw traces are archived, local credentials remain untracked, and completed download containers plus obsolete experimental image tags were removed.
Validation on Spark: four Compose variants parsed; model reload, arithmetic/Chinese/tool-call smoke, prefix-cache check, and first/repeated retrieval on 127987 input tokens with a 2048 output budget all passed. Minimum sampled MemAvailable was 16.777 GiB, host OOM count stayed 25, and system/user failed units were zero. Service is healthy with unless-stopped. No CI workflows exist; these are recorded device checks, not CI claims. Full concurrency and long-duration stability remain unverified.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Promote the validated Spark deployment to a 131072-token eager baseline with prefix caching and MTP=2. The previous 262K default triggered host OOM near full capacity; default, no-prefix fallback and optional graph configs now use 128K, while the conservative fallback remains 32K.
This merges the prefix-cache fixes, CUDA Graph comparisons, long-context limits and Nsight evidence. Documentation distinguishes current settings from historical experiments. Spark now uses a Git-managed layout; legacy files and raw traces are archived, local credentials remain untracked, and completed download containers plus obsolete experimental image tags were removed.
Validation on Spark: four Compose variants parsed; model reload, arithmetic/Chinese/tool-call smoke, prefix-cache check, and first/repeated retrieval on 127987 input tokens with a 2048 output budget all passed. Minimum sampled MemAvailable was 16.777 GiB, host OOM count stayed 25, and system/user failed units were zero. Service is healthy with unless-stopped. No CI workflows exist; these are recorded device checks, not CI claims. Full concurrency and long-duration stability remain unverified.