Compare MTP steps and prefill batches; retain qualified MTP 2 baseline
This commit is contained in:
@@ -0,0 +1,730 @@
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [api_utils.py:347]
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [api_utils.py:347] █ █ █▄ ▄█
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [api_utils.py:347] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.3.1.dev3+g0bfc7a15d
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [api_utils.py:347] █▄█▀ █ █ █ █ model /model
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [api_utils.py:347] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [api_utils.py:347]
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [api_utils.py:286] non-default args: {'model_tag': '/model', 'enable_auto_tool_choice': True, 'tool_call_parser': 'qwen3_xml', 'host': '0.0.0.0', 'api_key': '***', 'model': '/model', 'dtype': 'bfloat16', 'max_model_len': 131072, 'served_model_name': ['qwen3.8-flash-next'], 'load_format': 'safetensors', 'reasoning_parser': 'qwen3', 'gpu_memory_utilization': 0.985, 'kv_cache_dtype': 'fp8', 'enable_prefix_caching': True, 'max_num_batched_tokens': 2048, 'max_num_seqs': 1, 'enable_chunked_prefill': True, 'enable_flashinfer_autotune': False, 'speculative_config': {'method': 'mtp', 'num_speculative_tokens': 2}, 'compilation_config': {'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': [], 'ir_enable_torch_wrap': None, 'splitting_ops': None, 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': None, 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.FULL: 2>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [1, 3], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': None, 'pass_config': {}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': None, 'static_all_moe_layers': []}, 'engram_config': EngramConfig(cpu_offload=True, embedding_across_dp=False, dp_shared_memory=False)}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [model.py:691] Resolved architecture: Qwen4ExpForConditionalGeneration
|
||||
(APIServer pid=1) INFO 09-18 14:14:09 [model.py:2024] Using max model len 131072
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [cache.py:345] Using fp8 data type to store kv cache. It reduces the GPU memory footprint and boosts the performance. Meanwhile, it may cause accuracy drop without a proper scaling factor
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [model.py:691] Resolved architecture: Qwen4ExpMTP
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [model.py:2024] Using max model len 262144
|
||||
(APIServer pid=1) WARNING 09-18 14:14:13 [speculative.py:1360] Enabling num_speculative_tokens > 1 will run multiple times of forward on same MTP layer,which may result in lower acceptance rate
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [speculative.py:1653] Overriding draft model max model len from 262144 to 131072
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [config.py:625] Mamba cache mode is set to 'align' for Qwen4ExpForConditionalGeneration by default when prefix caching is enabled
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [vllm.py:1271] Resolved Engram configuration: EngramConfig(cpu_offload=True, embedding_across_dp=False, dp_shared_memory=False)
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [vllm.py:781] Auto-enabling VLLM_USE_BREAKABLE_CUDAGRAPH=1. Set VLLM_USE_BREAKABLE_CUDAGRAPH=0 to opt out.
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [kernel.py:408] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'], gelu_and_mul_sparse=['triton', 'native'])
|
||||
(APIServer pid=1) WARNING 09-18 14:14:13 [vllm.py:2152] max_num_scheduled_tokens is set to 2048 based on the speculative decoding settings. This may lead to suboptimal performance. Consider increasing max_num_batched_tokens to accommodate the additional draft token slots, or decrease num_speculative_tokens.
|
||||
(APIServer pid=1) INFO 09-18 14:14:13 [compilation.py:331] Enabled custom fusions: norm_quant, act_quant
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] The `use_fast` parameter is deprecated and will be removed in a future version. Use `backend="torchvision"` instead of `use_fast=True`, or `backend="pil"` instead of `use_fast=False`.
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(EngineCore pid=113) INFO 09-18 14:14:25 [core.py:123] Initializing a V1 LLM engine (v0.3.1.dev3+g0bfc7a15d) with config: model='/model', speculative_config=SpeculativeConfig(method='mtp', model='/model', num_spec_tokens=2), tokenizer='/model', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=131072, download_dir=None, load_format=safetensors, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=modelopt_mixed, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=qwen3.8-flash-next, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['+quant_fp8', 'all', '+quant_fp8'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.FULL: 2>, 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 3], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 3, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'], gelu_and_mul_sparse=['triton', 'native']), enable_flashinfer_autotune=False, enable_cutedsl_warmup=True, enable_jit_warmup=True, moe_backend='auto', sparse_indexer_topk_backend='auto', linear_backend='auto', linear_backend_per_quant=None)
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(EngineCore pid=113) INFO 09-18 14:14:26 [parallel_state.py:1825] world_size=1 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_ed11f4f25e124a16965c184de0adfebf backend=nccl
|
||||
(EngineCore pid=113) INFO 09-18 14:14:26 [parallel_state.py:2269] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, ETP rank 0, EP rank 0, EPLB rank N/A
|
||||
(EngineCore pid=113) INFO 09-18 14:14:26 [gpu_worker.py:441] Using V2 Model Runner
|
||||
(EngineCore pid=113) INFO 09-18 14:14:27 [model_runner.py:387] Loading model from scratch...
|
||||
(EngineCore pid=113) INFO 09-18 14:14:27 [cuda.py:595] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention
|
||||
(EngineCore pid=113) INFO 09-18 14:14:27 [mm_encoder_attention.py:372] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention.
|
||||
(EngineCore pid=113) INFO 09-18 14:14:27 [qwen_gdn_linear_attn.py:176] Using FlashInfer GDN prefill kernel (requested=auto, head_k_dim=128).
|
||||
(EngineCore pid=113) INFO 09-18 14:14:27 [qwen_gdn_linear_attn.py:528] GDN decode kernel: cuda
|
||||
(EngineCore pid=113) INFO 09-18 14:14:29 [nvfp4.py:302] Using 'FLASHINFER_CUTLASS' NvFp4 MoE backend out of potential backends: ['FLASHINFER_TRTLLM', 'FLASHINFER_CUTEDSL', 'FLASHINFER_CUTEDSL_BATCHED', 'FLASHINFER_CUTLASS', 'VLLM_CUTLASS', 'MARLIN', 'HUMMING', 'EMULATION'].
|
||||
(APIServer pid=1) [transformers] Qwen3VL video processing does not apply the per-frame pixel cap the reference implementation (qwen-vl-utils) applies, so some videos cost far more tokens than they would there. In v5.22 the capped behavior will become the default and `cap_pixels_per_frame` will be removed. Pass `cap_pixels_per_frame=True` to adopt the reference behavior now, or `False` to keep the current behavior and silence this warning.
|
||||
(APIServer pid=1) INFO 09-18 14:14:31 [base.py:261] Multi-modal warmup completed in 12.320s
|
||||
(APIServer pid=1) INFO 09-18 14:14:33 [base.py:261] Readonly multi-modal warmup completed in 1.302s
|
||||
(EngineCore pid=113) INFO 09-18 14:15:03 [ngram_embedding.py:720] Initialized PLE embedding language_model.model.layers.1.ple.ple_embedding.ngram_embedding: quantization_method=Qwen4ExpPLEFp8EmbeddingMethod, weight_dtype=torch.float8_e4m3fn, weight_device=cpu, pinned=True
|
||||
(EngineCore pid=113) INFO 09-18 14:15:03 [flash_attn.py:1115] Using FlashAttention version 2
|
||||
(EngineCore pid=113) WARNING 09-18 14:15:03 [compilation.py:1350] Op 'quant_fp8' not present in model, enabling with '+quant_fp8' has no effect
|
||||
(EngineCore pid=113) INFO 09-18 14:15:03 [weight_utils.py:900] Filesystem type for checkpoints: EXT4. Checkpoint size: 123.57 GiB. Available RAM: 142.31 GiB.
|
||||
(EngineCore pid=113) INFO 09-18 14:15:03 [weight_utils.py:923] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 0% Completed | 0/11 [00:00<?, ?it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 9% Completed | 1/11 [00:01<00:16, 1.63s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 18% Completed | 2/11 [00:06<00:33, 3.74s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 27% Completed | 3/11 [00:12<00:36, 4.52s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 36% Completed | 4/11 [00:17<00:34, 4.87s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 45% Completed | 5/11 [00:23<00:30, 5.11s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 55% Completed | 6/11 [00:28<00:26, 5.26s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 64% Completed | 7/11 [00:34<00:21, 5.36s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 73% Completed | 8/11 [00:40<00:16, 5.47s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 82% Completed | 9/11 [00:41<00:08, 4.20s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 100% Completed | 11/11 [00:42<00:00, 2.38s/it]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 100% Completed | 11/11 [00:42<00:00, 3.83s/it]
|
||||
(EngineCore pid=113)
|
||||
(EngineCore pid=113) INFO 09-18 14:15:45 [default_loader.py:430] Loading weights took 42.29 seconds
|
||||
(EngineCore pid=113) INFO 09-18 14:15:45 [nvfp4.py:611] Using MoEPrepareAndFinalizeNoDPEPModular
|
||||
(EngineCore pid=113) INFO 09-18 14:15:46 [vllm.py:1271] Resolved Engram configuration: EngramConfig(cpu_offload=True, embedding_across_dp=False, dp_shared_memory=False)
|
||||
(EngineCore pid=113) INFO 09-18 14:15:46 [kernel.py:408] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'], gelu_and_mul_sparse=['triton', 'native'])
|
||||
(EngineCore pid=113) WARNING 09-18 14:15:46 [vllm.py:2152] max_num_scheduled_tokens is set to 2048 based on the speculative decoding settings. This may lead to suboptimal performance. Consider increasing max_num_batched_tokens to accommodate the additional draft token slots, or decrease num_speculative_tokens.
|
||||
(EngineCore pid=113) INFO 09-18 14:15:46 [compilation.py:331] Enabled custom fusions: norm_quant, act_quant
|
||||
(EngineCore pid=113) INFO 09-18 14:15:46 [fp8.py:433] Using DEEPGEMM Fp8 MoE backend out of potential backends: ['AITER', 'FLASHINFER_TRTLLM', 'FLASHINFER_CUTLASS', 'DEEPGEMM', 'TRITON', 'MARLIN', 'HUMMING', 'BATCHED_DEEPGEMM', 'BATCHED_TRITON', 'XPU', 'CPU', 'HPC'].
|
||||
(EngineCore pid=113) INFO 09-18 14:15:47 [weight_utils.py:900] Filesystem type for checkpoints: EXT4. Checkpoint size: 123.57 GiB. Available RAM: 142.17 GiB.
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 0% Completed | 0/11 [00:00<?, ?it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 18% Completed | 2/11 [00:00<00:01, 5.13it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 27% Completed | 3/11 [00:00<00:01, 4.03it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 36% Completed | 4/11 [00:01<00:01, 3.63it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 45% Completed | 5/11 [00:01<00:01, 3.41it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 55% Completed | 6/11 [00:01<00:01, 3.30it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 64% Completed | 7/11 [00:02<00:01, 3.24it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 73% Completed | 8/11 [00:02<00:00, 3.19it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 82% Completed | 9/11 [00:02<00:00, 3.69it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 100% Completed | 11/11 [00:02<00:00, 3.89it/s]
|
||||
(EngineCore pid=113)
|
||||
Loading safetensors checkpoint shards: 100% Completed | 11/11 [00:02<00:00, 3.68it/s]
|
||||
(EngineCore pid=113)
|
||||
(EngineCore pid=113) INFO 09-18 14:15:50 [default_loader.py:430] Loading weights took 3.00 seconds
|
||||
(EngineCore pid=113) INFO 09-18 14:15:50 [deep_gemm.py:196] deep_gemm not found in site-packages, trying vendored vllm.third_party.deep_gemm
|
||||
(EngineCore pid=113) INFO 09-18 14:15:50 [deep_gemm.py:223] DeepGEMM PDL enabled on vllm.third_party.deep_gemm.
|
||||
(EngineCore pid=113) INFO 09-18 14:15:50 [deep_gemm.py:136] DeepGEMM E8M0 enabled on current platform.
|
||||
(EngineCore pid=113) INFO 09-18 14:15:52 [fp8.py:733] Using MoEPrepareAndFinalizeNoDPEPModular
|
||||
(EngineCore pid=113) WARNING 09-18 14:15:52 [speculator.py:235] Draft model Qwen4ExpMTP does not support external multimodal embeddings. Embeddings from the target model will not be passed to the drafter; using text-only draft inputs instead.
|
||||
(EngineCore pid=113) INFO 09-18 14:15:53 [model_runner.py:419] Model loading took 76.36 GiB memory and 86.461433 seconds
|
||||
(EngineCore pid=113) INFO 09-18 14:15:53 [topk_topp_sampler.py:78] Using FlashInfer for top-p & top-k sampling.
|
||||
(EngineCore pid=113) INFO 09-18 14:15:53 [interface.py:918] Setting attention block size to 3184 tokens to ensure that attention page size is >= mamba page size.
|
||||
(EngineCore pid=113) INFO 09-18 14:15:53 [interface.py:942] Padding mamba page size by 0.38% to ensure that mamba page size and attention page size are exactly equal.
|
||||
(EngineCore pid=113) INFO 09-18 14:15:53 [utils.py:320] Using BLNHC KV cache layout.
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(EngineCore pid=113) [transformers] The `use_fast` parameter is deprecated and will be removed in a future version. Use `backend="torchvision"` instead of `use_fast=True`, or `backend="pil"` instead of `use_fast=False`.
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(EngineCore pid=113) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_interleaved', 'mrope_section'}
|
||||
(EngineCore pid=113) INFO 09-18 14:15:58 [encoder_runner.py:131] Encoder cache will be initialized with a budget of 16384 tokens, and profiled with 1 image items of the maximum feature size.
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:20 [kv_cache_utils.py:2219] Speculative decoding (method=mtp) is enabled but no KV cache group could be identified as the draft model's.
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:20 [compilation.py:1415] CUDAGraphMode.FULL is not supported with GDNAttentionBackend backend (support: AttentionCGSupport.UNIFORM_BATCH); setting cudagraph_mode=FULL_DECODE_ONLY
|
||||
(EngineCore pid=113) INFO 09-18 14:16:20 [speculator.py:119] Fused multi-step draft decode is not supported by attention backend(s) QWEN4_EXP_EXP_QSA_STATE; falling back to rebuilding attention metadata between draft steps.
|
||||
(EngineCore pid=113)
|
||||
Capturing CUDA graphs (FULL): 0%| | 0/1 [00:00<?, ?it/s]
|
||||
Capturing CUDA graphs (FULL): 100%|██████████| 1/1 [00:01<00:00, 1.83s/it]
|
||||
Capturing CUDA graphs (FULL): 100%|██████████| 1/1 [00:01<00:00, 1.83s/it]
|
||||
(EngineCore pid=113) INFO 09-18 14:16:22 [speculator.py:150] Capturing model for speculator...
|
||||
(EngineCore pid=113)
|
||||
Capturing prefill CUDA graphs (FULL): 0%| | 0/1 [00:00<?, ?it/s]
|
||||
Capturing prefill CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 2.17it/s]
|
||||
Capturing prefill CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 2.17it/s]
|
||||
(EngineCore pid=113)
|
||||
Capturing decode CUDA graphs (FULL): 0%| | 0/1 [00:00<?, ?it/s]
|
||||
Capturing decode CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 5.69it/s]
|
||||
Capturing decode CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 5.69it/s]
|
||||
(EngineCore pid=113) INFO 09-18 14:16:23 [model_runner.py:1057] Graph capturing finished in 3 secs, took 0.14 GiB
|
||||
(EngineCore pid=113) INFO 09-18 14:16:24 [gpu_worker.py:641] Available KV cache memory: 2.78 GiB
|
||||
(EngineCore pid=113) INFO 09-18 14:16:24 [gpu_worker.py:656] CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.9850 is equivalent to --gpu-memory-utilization=0.9833 without CUDA graph memory profiling. To maintain the same effective KV cache size as before, increase --gpu-memory-utilization to 0.9867. To disable, set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0.
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:24 [kv_cache_utils.py:2219] Speculative decoding (method=mtp) is enabled but no KV cache group could be identified as the draft model's.
|
||||
(EngineCore pid=113) INFO 09-18 14:16:24 [kv_cache_utils.py:2404] GPU KV cache size: 146,622 tokens, Maximum concurrency for 131,072 tokens per request: 1.12x
|
||||
(EngineCore pid=113) INFO 09-18 14:16:24 [kernel_warmup.py:171] JIT kernel warmup starting.
|
||||
(EngineCore pid=113) INFO 09-18 14:16:24 [kernel_warmup.py:184] JIT kernel warmup finished in 0.00s.
|
||||
(EngineCore pid=113) INFO 09-18 14:16:24 [qwen_vl_triton_warmup.py:57] Warmed position embedding and vision rotary kernels on grids=[(1, 16, 16), (1, 16, 2), (1, 2, 16), (1, 2, 2)].
|
||||
(EngineCore pid=113) INFO 09-18 14:16:24 [qwen_vl_triton_warmup.py:98] Warmed M-RoPE Triton kernels.
|
||||
(EngineCore pid=113) INFO 09-18 14:16:24 [mamba_triton_warmup.py:42] Warmed Mamba batch_memcpy_kernel.
|
||||
(EngineCore pid=113) INFO 09-18 14:16:25 [qwen4_exp_qsa_warmup.py:71] Warmed up Qwen4Exp QSA decode kernels: ((1, 1), (2, 1), (3, 1)).
|
||||
(EngineCore pid=113) INFO 09-18 14:16:30 [qwen4_exp_qsa_warmup.py:85] Warmed up Qwen4Exp QSA sparse attention kernels: ((32, 2, 1), (32, 4, 4), (64, 1, 2), (64, 4, 4), (64, 8, 4), (64, 33, 8), (128, 4, 4), (128, 8, 4)).
|
||||
(EngineCore pid=113) INFO 09-18 14:16:30 [kernel_warmup.py:254] Skipping FlashInfer autotune because it is disabled.
|
||||
(EngineCore pid=113)
|
||||
Capturing CUDA graphs (FULL): 0%| | 0/1 [00:00<?, ?it/s]
|
||||
Capturing CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 9.91it/s]
|
||||
Capturing CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 9.89it/s]
|
||||
(EngineCore pid=113) INFO 09-18 14:16:39 [speculator.py:150] Capturing model for speculator...
|
||||
(EngineCore pid=113)
|
||||
Capturing prefill CUDA graphs (FULL): 0%| | 0/1 [00:00<?, ?it/s]
|
||||
Capturing prefill CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 118.66it/s]
|
||||
(EngineCore pid=113)
|
||||
Capturing decode CUDA graphs (FULL): 0%| | 0/1 [00:00<?, ?it/s]
|
||||
Capturing decode CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 176.98it/s]
|
||||
(EngineCore pid=113) INFO 09-18 14:16:39 [model_runner.py:1057] Graph capturing finished in 1 secs, took 0.13 GiB
|
||||
(EngineCore pid=113) INFO 09-18 14:16:39 [gpu_worker.py:824] CUDA graph pool memory: 0.13 GiB (actual), 0.14 GiB (estimated), difference: 0.01 GiB (5.8%).
|
||||
(EngineCore pid=113) INFO 09-18 14:16:39 [gpu_worker.py:887] Free memory on device (82.59/83.05 GiB) on startup. Desired GPU memory utilization is (0.985, 81.8 GiB). Actual usage is 77.85 GiB for consumed memory (weights + non-torch), 1.17 GiB for peak activation, and 0.13 GiB for CUDAGraph memory. Replace gpu_memory_utilization config with `--kv-cache-memory=2685284312` (2.5 GiB) to fit into requested memory, or `--kv-cache-memory=3531951104` (3.29 GiB) to fully utilize gpu memory. Current kv cache memory in use is 2.78 GiB.
|
||||
(EngineCore pid=113) INFO 09-18 14:16:40 [jit_monitor.py:84] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:41 [torch_utils.py:274] OMP_NUM_THREADS=8 is set; leaving Torch threads at 8 for serving. Multi-threaded torch CPU ops during serving can degrade performance through spin-wait contention and cgroup CPU-quota throttling.
|
||||
(EngineCore pid=113) INFO 09-18 14:16:41 [core.py:380] init engine (profile, create kv cache, warmup model) took 47.40 s
|
||||
(EngineCore pid=113) INFO 09-18 14:16:41 [kv_cache_utils.py:762] kv cache group sizes [3184, 3184, 3184, 3184, 8, 3184]
|
||||
(EngineCore pid=113) INFO 09-18 14:16:41 [kv_cache_utils.py:763] kv lcm block sizes 3184
|
||||
(EngineCore pid=113) INFO 09-18 14:16:41 [kernel.py:408] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'], gelu_and_mul_sparse=['triton', 'native'])
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [entry.py:132] Supported tasks: ['generate']
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [factories.py:76] Scale-out endpoints are disabled. Set --enable-scale-out to enable them.
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [parser_manager.py:34] "auto" tool choice has been enabled.
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'mrope_section', 'mrope_interleaved'}
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [hf.py:642] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this.
|
||||
(APIServer pid=1) WARNING 09-18 14:16:41 [model.py:1769] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 1.0, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [entry.py:136] Starting vLLM server on http://0.0.0.0:8000
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:60] Available routes are:
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /openapi.json, Methods: HEAD, GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /docs, Methods: HEAD, GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /redoc, Methods: HEAD, GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /load, Methods: GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /version, Methods: GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /health, Methods: GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /metrics, Methods: GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /tokenize, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /detokenize, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/models, Methods: GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /ping, Methods: GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /ping, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /invocations, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/chat/completions, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/chat/completions/batch, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/responses, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/responses/{response_id}, Methods: GET
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/completions, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/messages, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /v1/messages/count_tokens, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /generative_scoring, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /scale_elastic_ep, Methods: POST
|
||||
(APIServer pid=1) INFO 09-18 14:16:41 [launcher.py:69] Route: /is_scaling_elastic_ep, Methods: POST
|
||||
(APIServer pid=1) INFO: Started server process [1]
|
||||
(APIServer pid=1) INFO: Waiting for application startup.
|
||||
(APIServer pid=1) INFO: Application startup complete.
|
||||
(APIServer pid=1) INFO: 127.0.0.1:52136 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36256 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36262 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:46 [jit_monitor.py:140] Triton kernel JIT compilation during inference: layer_norm_fwd_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:46 [jit_monitor.py:140] Triton kernel JIT compilation during inference: _count_expert_num_tokens. This causes a latency spike; consider extending warmup to cover this shape/config.
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:46 [jit_monitor.py:140] Triton kernel JIT compilation during inference: _compute_local_logits_stats_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:46 [jit_monitor.py:140] Triton kernel JIT compilation during inference: _rejection_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
|
||||
(EngineCore pid=113) WARNING 09-18 14:16:47 [jit_monitor.py:140] Triton kernel JIT compilation during inference: _resample_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36278 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36282 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36290 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36304 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36306 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36310 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36316 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36332 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36342 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36352 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:16:52 [loggers.py:323] Engine 000: Avg prompt throughput: 53.5 tokens/s, Avg generation throughput: 60.9 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
|
||||
(APIServer pid=1) INFO 09-18 14:16:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.91, Accepted throughput: 39.90 tokens/s, Drafted throughput: 41.72 tokens/s, Accepted: 440 tokens, Drafted: 460 tokens, Per-position acceptance rate: 0.987, 0.926, Avg Draft acceptance rate: 95.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36364 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36378 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36380 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33586 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33600 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
[rank0]:[W918 14:16:56.003832297 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 377487360 bytes (free: 64946176, total: 89173131264).
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33614 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33630 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:17:02 [loggers.py:323] Engine 000: Avg prompt throughput: 2790.6 tokens/s, Avg generation throughput: 72.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:17:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.72, Accepted throughput: 45.70 tokens/s, Drafted throughput: 53.00 tokens/s, Accepted: 457 tokens, Drafted: 530 tokens, Per-position acceptance rate: 0.925, 0.800, Avg Draft acceptance rate: 86.2%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33642 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:17:12 [loggers.py:323] Engine 000: Avg prompt throughput: 7.7 tokens/s, Avg generation throughput: 124.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:17:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.27, Accepted throughput: 69.90 tokens/s, Drafted throughput: 110.19 tokens/s, Accepted: 699 tokens, Drafted: 1102 tokens, Per-position acceptance rate: 0.737, 0.532, Avg Draft acceptance rate: 63.4%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:35366 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43852 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:17:22 [loggers.py:323] Engine 000: Avg prompt throughput: 7.7 tokens/s, Avg generation throughput: 118.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:17:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.14, Accepted throughput: 62.90 tokens/s, Drafted throughput: 110.40 tokens/s, Accepted: 629 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.707, 0.433, Avg Draft acceptance rate: 57.0%
|
||||
(APIServer pid=1) INFO 09-18 14:17:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 128.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:17:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.30, Accepted throughput: 72.60 tokens/s, Drafted throughput: 111.60 tokens/s, Accepted: 726 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.758, 0.543, Avg Draft acceptance rate: 65.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:44946 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:17:42 [loggers.py:323] Engine 000: Avg prompt throughput: 7.7 tokens/s, Avg generation throughput: 119.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:17:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.16, Accepted throughput: 64.09 tokens/s, Drafted throughput: 110.38 tokens/s, Accepted: 641 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.707, 0.455, Avg Draft acceptance rate: 58.1%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:40726 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:17:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 123.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:17:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.22, Accepted throughput: 68.10 tokens/s, Drafted throughput: 111.40 tokens/s, Accepted: 681 tokens, Drafted: 1114 tokens, Per-position acceptance rate: 0.727, 0.496, Avg Draft acceptance rate: 61.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:41494 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:51114 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:18:02 [loggers.py:323] Engine 000: Avg prompt throughput: 15.8 tokens/s, Avg generation throughput: 124.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:18:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.27, Accepted throughput: 69.80 tokens/s, Drafted throughput: 109.80 tokens/s, Accepted: 698 tokens, Drafted: 1098 tokens, Per-position acceptance rate: 0.754, 0.517, Avg Draft acceptance rate: 63.6%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53922 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:18:12 [loggers.py:323] Engine 000: Avg prompt throughput: 7.9 tokens/s, Avg generation throughput: 129.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:18:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.32, Accepted throughput: 73.50 tokens/s, Drafted throughput: 111.00 tokens/s, Accepted: 735 tokens, Drafted: 1110 tokens, Per-position acceptance rate: 0.766, 0.559, Avg Draft acceptance rate: 66.2%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:57802 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:18:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 127.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:18:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.28, Accepted throughput: 71.50 tokens/s, Drafted throughput: 111.80 tokens/s, Accepted: 715 tokens, Drafted: 1118 tokens, Per-position acceptance rate: 0.758, 0.521, Avg Draft acceptance rate: 64.0%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:36444 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:18:32 [loggers.py:323] Engine 000: Avg prompt throughput: 7.9 tokens/s, Avg generation throughput: 128.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:18:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.31, Accepted throughput: 72.70 tokens/s, Drafted throughput: 110.80 tokens/s, Accepted: 727 tokens, Drafted: 1108 tokens, Per-position acceptance rate: 0.765, 0.547, Avg Draft acceptance rate: 65.6%
|
||||
(APIServer pid=1) INFO 09-18 14:18:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 129.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:18:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.32, Accepted throughput: 73.70 tokens/s, Drafted throughput: 111.79 tokens/s, Accepted: 737 tokens, Drafted: 1118 tokens, Per-position acceptance rate: 0.764, 0.555, Avg Draft acceptance rate: 65.9%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38818 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:34668 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:40728 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:18:52 [loggers.py:323] Engine 000: Avg prompt throughput: 15.8 tokens/s, Avg generation throughput: 120.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:18:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.20, Accepted throughput: 65.80 tokens/s, Drafted throughput: 109.99 tokens/s, Accepted: 658 tokens, Drafted: 1100 tokens, Per-position acceptance rate: 0.709, 0.487, Avg Draft acceptance rate: 59.8%
|
||||
(APIServer pid=1) INFO 09-18 14:19:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 122.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:19:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.20, Accepted throughput: 67.00 tokens/s, Drafted throughput: 111.40 tokens/s, Accepted: 670 tokens, Drafted: 1114 tokens, Per-position acceptance rate: 0.716, 0.487, Avg Draft acceptance rate: 60.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:50984 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:19:12 [loggers.py:323] Engine 000: Avg prompt throughput: 7.9 tokens/s, Avg generation throughput: 119.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:19:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.16, Accepted throughput: 64.20 tokens/s, Drafted throughput: 110.80 tokens/s, Accepted: 642 tokens, Drafted: 1108 tokens, Per-position acceptance rate: 0.702, 0.457, Avg Draft acceptance rate: 57.9%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:57380 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:37540 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:19:22 [loggers.py:323] Engine 000: Avg prompt throughput: 7.9 tokens/s, Avg generation throughput: 122.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:19:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.22, Accepted throughput: 67.59 tokens/s, Drafted throughput: 110.59 tokens/s, Accepted: 676 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.714, 0.508, Avg Draft acceptance rate: 61.1%
|
||||
(APIServer pid=1) INFO 09-18 14:19:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 121.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:19:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.17, Accepted throughput: 65.50 tokens/s, Drafted throughput: 111.59 tokens/s, Accepted: 655 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.710, 0.464, Avg Draft acceptance rate: 58.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43060 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:19:42 [loggers.py:323] Engine 000: Avg prompt throughput: 7.5 tokens/s, Avg generation throughput: 124.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:19:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.26, Accepted throughput: 69.60 tokens/s, Drafted throughput: 110.60 tokens/s, Accepted: 696 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.741, 0.517, Avg Draft acceptance rate: 62.9%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:45568 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:19:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 137.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:19:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.47, Accepted throughput: 81.90 tokens/s, Drafted throughput: 111.20 tokens/s, Accepted: 819 tokens, Drafted: 1112 tokens, Per-position acceptance rate: 0.831, 0.642, Avg Draft acceptance rate: 73.7%
|
||||
(APIServer pid=1) INFO 09-18 14:20:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 150.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.1%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:20:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.71, Accepted throughput: 95.00 tokens/s, Drafted throughput: 110.80 tokens/s, Accepted: 950 tokens, Drafted: 1108 tokens, Per-position acceptance rate: 0.919, 0.796, Avg Draft acceptance rate: 85.7%
|
||||
(APIServer pid=1) INFO 09-18 14:20:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 145.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.1%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:20:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.63, Accepted throughput: 90.20 tokens/s, Drafted throughput: 110.79 tokens/s, Accepted: 902 tokens, Drafted: 1108 tokens, Per-position acceptance rate: 0.888, 0.740, Avg Draft acceptance rate: 81.4%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:41446 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:20:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 138.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.6%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:20:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.51, Accepted throughput: 83.29 tokens/s, Drafted throughput: 110.59 tokens/s, Accepted: 833 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.837, 0.669, Avg Draft acceptance rate: 75.3%
|
||||
(APIServer pid=1) INFO 09-18 14:20:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 134.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.6%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:20:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.43, Accepted throughput: 79.10 tokens/s, Drafted throughput: 110.39 tokens/s, Accepted: 791 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.813, 0.620, Avg Draft acceptance rate: 71.6%
|
||||
(APIServer pid=1) INFO 09-18 14:20:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 134.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.6%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:20:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.45, Accepted throughput: 79.80 tokens/s, Drafted throughput: 110.19 tokens/s, Accepted: 798 tokens, Drafted: 1102 tokens, Per-position acceptance rate: 0.815, 0.633, Avg Draft acceptance rate: 72.4%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:38588 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:20:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 131.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 26.2%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:20:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.43, Accepted throughput: 77.00 tokens/s, Drafted throughput: 108.00 tokens/s, Accepted: 770 tokens, Drafted: 1080 tokens, Per-position acceptance rate: 0.743, 0.683, Avg Draft acceptance rate: 71.3%
|
||||
(APIServer pid=1) INFO 09-18 14:21:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 154.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 26.2%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:21:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.88, Accepted throughput: 100.70 tokens/s, Drafted throughput: 107.40 tokens/s, Accepted: 1007 tokens, Drafted: 1074 tokens, Per-position acceptance rate: 0.952, 0.924, Avg Draft acceptance rate: 93.8%
|
||||
(APIServer pid=1) INFO 09-18 14:21:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 143.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.7%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:21:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.66, Accepted throughput: 89.40 tokens/s, Drafted throughput: 108.00 tokens/s, Accepted: 894 tokens, Drafted: 1080 tokens, Per-position acceptance rate: 0.833, 0.822, Avg Draft acceptance rate: 82.8%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:42148 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:21:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 151.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.7%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:21:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.80, Accepted throughput: 97.59 tokens/s, Drafted throughput: 108.19 tokens/s, Accepted: 976 tokens, Drafted: 1082 tokens, Per-position acceptance rate: 0.906, 0.898, Avg Draft acceptance rate: 90.2%
|
||||
(APIServer pid=1) [transformers] Token indices sequence length is longer than the specified maximum sequence length for this model (268298 > 262144). Running this sequence through the model will result in indexing errors
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48854 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48862 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48874 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48886 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48894 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48902 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48914 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:21:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 140.5 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO 09-18 14:21:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.86, Accepted throughput: 91.39 tokens/s, Drafted throughput: 98.39 tokens/s, Accepted: 914 tokens, Drafted: 984 tokens, Per-position acceptance rate: 0.929, 0.929, Avg Draft acceptance rate: 92.9%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48926 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48928 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48934 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48948 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48950 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48964 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48968 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48976 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:21:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 67.7%, Prefix cache hit rate: 0.0%, MM cache hit rate: 66.7%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:49396 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38790 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38796 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38802 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:21:52 [loggers.py:323] Engine 000: Avg prompt throughput: 13599.9 tokens/s, Avg generation throughput: 37.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.3%, Prefix cache hit rate: 46.6%, MM cache hit rate: 71.4%
|
||||
(APIServer pid=1) INFO 09-18 14:21:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.80, Accepted throughput: 11.95 tokens/s, Drafted throughput: 13.30 tokens/s, Accepted: 239 tokens, Drafted: 266 tokens, Per-position acceptance rate: 0.940, 0.857, Avg Draft acceptance rate: 89.8%
|
||||
(APIServer pid=1) INFO 09-18 14:22:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 73.8%, Prefix cache hit rate: 46.6%, MM cache hit rate: 71.4%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43834 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43840 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43848 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43858 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43874 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43878 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:43890 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:22:12 [loggers.py:323] Engine 000: Avg prompt throughput: 13594.6 tokens/s, Avg generation throughput: 70.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.1%, Prefix cache hit rate: 52.5%, MM cache hit rate: 77.8%
|
||||
(APIServer pid=1) INFO 09-18 14:22:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.77, Accepted throughput: 22.55 tokens/s, Drafted throughput: 25.50 tokens/s, Accepted: 451 tokens, Drafted: 510 tokens, Per-position acceptance rate: 0.929, 0.839, Avg Draft acceptance rate: 88.4%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:51984 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:58388 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:22:22 [loggers.py:323] Engine 000: Avg prompt throughput: 1215.2 tokens/s, Avg generation throughput: 113.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 52.5%, MM cache hit rate: 77.8%
|
||||
(APIServer pid=1) INFO 09-18 14:22:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.29, Accepted throughput: 63.79 tokens/s, Drafted throughput: 99.19 tokens/s, Accepted: 638 tokens, Drafted: 992 tokens, Per-position acceptance rate: 0.754, 0.532, Avg Draft acceptance rate: 64.3%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:52418 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:22:32 [loggers.py:323] Engine 000: Avg prompt throughput: 8.6 tokens/s, Avg generation throughput: 139.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 52.5%, MM cache hit rate: 77.8%
|
||||
(APIServer pid=1) INFO 09-18 14:22:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.53, Accepted throughput: 84.50 tokens/s, Drafted throughput: 110.60 tokens/s, Accepted: 845 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.848, 0.680, Avg Draft acceptance rate: 76.4%
|
||||
(APIServer pid=1) INFO 09-18 14:22:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 134.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 52.5%, MM cache hit rate: 77.8%
|
||||
(APIServer pid=1) INFO 09-18 14:22:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.40, Accepted throughput: 78.30 tokens/s, Drafted throughput: 112.00 tokens/s, Accepted: 783 tokens, Drafted: 1120 tokens, Per-position acceptance rate: 0.811, 0.588, Avg Draft acceptance rate: 69.9%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:56790 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53910 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:22:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 33.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 58.5%, Prefix cache hit rate: 44.1%, MM cache hit rate: 77.8%
|
||||
(APIServer pid=1) INFO 09-18 14:22:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.42, Accepted throughput: 19.70 tokens/s, Drafted throughput: 27.80 tokens/s, Accepted: 197 tokens, Drafted: 278 tokens, Per-position acceptance rate: 0.799, 0.619, Avg Draft acceptance rate: 70.9%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47694 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47700 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:23:02 [loggers.py:323] Engine 000: Avg prompt throughput: 12912.6 tokens/s, Avg generation throughput: 24.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 30.8%, Prefix cache hit rate: 43.3%, MM cache hit rate: 81.8%
|
||||
(APIServer pid=1) INFO 09-18 14:23:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.74, Accepted throughput: 15.50 tokens/s, Drafted throughput: 17.80 tokens/s, Accepted: 155 tokens, Drafted: 178 tokens, Per-position acceptance rate: 0.921, 0.820, Avg Draft acceptance rate: 87.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47708 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:23:12 [loggers.py:323] Engine 000: Avg prompt throughput: 1215.4 tokens/s, Avg generation throughput: 118.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 43.3%, MM cache hit rate: 81.8%
|
||||
(APIServer pid=1) INFO 09-18 14:23:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.24, Accepted throughput: 65.70 tokens/s, Drafted throughput: 105.99 tokens/s, Accepted: 657 tokens, Drafted: 1060 tokens, Per-position acceptance rate: 0.743, 0.496, Avg Draft acceptance rate: 62.0%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:53052 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59822 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:23:22 [loggers.py:323] Engine 000: Avg prompt throughput: 8.4 tokens/s, Avg generation throughput: 120.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 43.3%, MM cache hit rate: 81.8%
|
||||
(APIServer pid=1) INFO 09-18 14:23:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.18, Accepted throughput: 65.40 tokens/s, Drafted throughput: 110.39 tokens/s, Accepted: 654 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.716, 0.469, Avg Draft acceptance rate: 59.2%
|
||||
(APIServer pid=1) INFO 09-18 14:23:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 121.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 43.3%, MM cache hit rate: 81.8%
|
||||
(APIServer pid=1) INFO 09-18 14:23:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.18, Accepted throughput: 66.00 tokens/s, Drafted throughput: 111.60 tokens/s, Accepted: 660 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.704, 0.478, Avg Draft acceptance rate: 59.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:34388 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:23:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 61.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 46.2%, Prefix cache hit rate: 37.4%, MM cache hit rate: 81.8%
|
||||
(APIServer pid=1) INFO 09-18 14:23:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.06, Accepted throughput: 31.80 tokens/s, Drafted throughput: 59.99 tokens/s, Accepted: 318 tokens, Drafted: 600 tokens, Per-position acceptance rate: 0.673, 0.387, Avg Draft acceptance rate: 53.0%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:50974 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:23:52 [loggers.py:323] Engine 000: Avg prompt throughput: 12598.6 tokens/s, Avg generation throughput: 2.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 81.5%, Prefix cache hit rate: 37.4%, MM cache hit rate: 81.8%
|
||||
(APIServer pid=1) INFO 09-18 14:23:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.80, Accepted throughput: 1.80 tokens/s, Drafted throughput: 2.00 tokens/s, Accepted: 18 tokens, Drafted: 20 tokens, Per-position acceptance rate: 0.900, 0.900, Avg Draft acceptance rate: 90.0%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:50664 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38410 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38412 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:24:02 [loggers.py:323] Engine 000: Avg prompt throughput: 1529.0 tokens/s, Avg generation throughput: 105.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 36.8%, MM cache hit rate: 84.6%
|
||||
(APIServer pid=1) INFO 09-18 14:24:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.35, Accepted throughput: 60.20 tokens/s, Drafted throughput: 89.20 tokens/s, Accepted: 602 tokens, Drafted: 892 tokens, Per-position acceptance rate: 0.791, 0.558, Avg Draft acceptance rate: 67.5%
|
||||
(APIServer pid=1) INFO 09-18 14:24:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 118.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 36.8%, MM cache hit rate: 84.6%
|
||||
(APIServer pid=1) INFO 09-18 14:24:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.13, Accepted throughput: 62.90 tokens/s, Drafted throughput: 111.40 tokens/s, Accepted: 629 tokens, Drafted: 1114 tokens, Per-position acceptance rate: 0.689, 0.440, Avg Draft acceptance rate: 56.5%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38508 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:37090 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:24:22 [loggers.py:323] Engine 000: Avg prompt throughput: 8.7 tokens/s, Avg generation throughput: 123.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 36.8%, MM cache hit rate: 84.6%
|
||||
(APIServer pid=1) INFO 09-18 14:24:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.23, Accepted throughput: 67.80 tokens/s, Drafted throughput: 110.60 tokens/s, Accepted: 678 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.732, 0.494, Avg Draft acceptance rate: 61.3%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:55060 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:24:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 100.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 36.9%, Prefix cache hit rate: 32.5%, MM cache hit rate: 84.6%
|
||||
(APIServer pid=1) INFO 09-18 14:24:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.33, Accepted throughput: 57.60 tokens/s, Drafted throughput: 86.60 tokens/s, Accepted: 576 tokens, Drafted: 866 tokens, Per-position acceptance rate: 0.762, 0.568, Avg Draft acceptance rate: 66.5%
|
||||
(APIServer pid=1) INFO 09-18 14:24:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 78.5%, Prefix cache hit rate: 32.5%, MM cache hit rate: 84.6%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:41864 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:44650 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:44654 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:44658 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:24:52 [loggers.py:323] Engine 000: Avg prompt throughput: 14128.6 tokens/s, Avg generation throughput: 77.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 32.0%, MM cache hit rate: 86.7%
|
||||
(APIServer pid=1) INFO 09-18 14:24:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.37, Accepted throughput: 22.20 tokens/s, Drafted throughput: 32.40 tokens/s, Accepted: 444 tokens, Drafted: 648 tokens, Per-position acceptance rate: 0.765, 0.605, Avg Draft acceptance rate: 68.5%
|
||||
(APIServer pid=1) INFO 09-18 14:25:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 130.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 32.0%, MM cache hit rate: 86.7%
|
||||
(APIServer pid=1) INFO 09-18 14:25:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.33, Accepted throughput: 74.49 tokens/s, Drafted throughput: 111.79 tokens/s, Accepted: 745 tokens, Drafted: 1118 tokens, Per-position acceptance rate: 0.767, 0.565, Avg Draft acceptance rate: 66.6%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:52324 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:25:12 [loggers.py:323] Engine 000: Avg prompt throughput: 8.7 tokens/s, Avg generation throughput: 124.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 32.0%, MM cache hit rate: 86.7%
|
||||
(APIServer pid=1) INFO 09-18 14:25:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.25, Accepted throughput: 69.30 tokens/s, Drafted throughput: 110.99 tokens/s, Accepted: 693 tokens, Drafted: 1110 tokens, Per-position acceptance rate: 0.737, 0.512, Avg Draft acceptance rate: 62.4%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:54114 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47722 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:25:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 112.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.3%, Prefix cache hit rate: 28.7%, MM cache hit rate: 86.7%
|
||||
(APIServer pid=1) INFO 09-18 14:25:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.32, Accepted throughput: 63.90 tokens/s, Drafted throughput: 96.79 tokens/s, Accepted: 639 tokens, Drafted: 968 tokens, Per-position acceptance rate: 0.756, 0.564, Avg Draft acceptance rate: 66.0%
|
||||
(APIServer pid=1) INFO 09-18 14:25:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 75.4%, Prefix cache hit rate: 28.7%, MM cache hit rate: 86.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56320 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56330 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56334 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:25:42 [loggers.py:323] Engine 000: Avg prompt throughput: 14128.9 tokens/s, Avg generation throughput: 67.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 28.3%, MM cache hit rate: 88.2%
|
||||
(APIServer pid=1) INFO 09-18 14:25:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.45, Accepted throughput: 19.80 tokens/s, Drafted throughput: 27.30 tokens/s, Accepted: 396 tokens, Drafted: 546 tokens, Per-position acceptance rate: 0.806, 0.645, Avg Draft acceptance rate: 72.5%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:34304 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:25:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 126.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 28.3%, MM cache hit rate: 88.2%
|
||||
(APIServer pid=1) INFO 09-18 14:25:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.27, Accepted throughput: 70.90 tokens/s, Drafted throughput: 111.80 tokens/s, Accepted: 709 tokens, Drafted: 1118 tokens, Per-position acceptance rate: 0.757, 0.512, Avg Draft acceptance rate: 63.4%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33198 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:26:02 [loggers.py:323] Engine 000: Avg prompt throughput: 8.5 tokens/s, Avg generation throughput: 117.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 28.3%, MM cache hit rate: 88.2%
|
||||
(APIServer pid=1) INFO 09-18 14:26:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.12, Accepted throughput: 62.09 tokens/s, Drafted throughput: 110.39 tokens/s, Accepted: 621 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.681, 0.444, Avg Draft acceptance rate: 56.2%
|
||||
(APIServer pid=1) INFO 09-18 14:26:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 130.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 28.3%, MM cache hit rate: 88.2%
|
||||
(APIServer pid=1) INFO 09-18 14:26:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.34, Accepted throughput: 74.39 tokens/s, Drafted throughput: 111.39 tokens/s, Accepted: 744 tokens, Drafted: 1114 tokens, Per-position acceptance rate: 0.772, 0.564, Avg Draft acceptance rate: 66.8%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:52376 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:51278 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:26:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 69.2%, Prefix cache hit rate: 25.7%, MM cache hit rate: 88.2%
|
||||
(APIServer pid=1) INFO 09-18 14:26:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.80, Accepted throughput: 0.40 tokens/s, Drafted throughput: 1.00 tokens/s, Accepted: 4 tokens, Drafted: 10 tokens, Per-position acceptance rate: 0.600, 0.200, Avg Draft acceptance rate: 40.0%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56900 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56902 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56906 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:26:32 [loggers.py:323] Engine 000: Avg prompt throughput: 14128.0 tokens/s, Avg generation throughput: 53.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 25.4%, MM cache hit rate: 89.5%
|
||||
(APIServer pid=1) INFO 09-18 14:26:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.74, Accepted throughput: 33.90 tokens/s, Drafted throughput: 39.00 tokens/s, Accepted: 339 tokens, Drafted: 390 tokens, Per-position acceptance rate: 0.923, 0.815, Avg Draft acceptance rate: 86.9%
|
||||
(APIServer pid=1) INFO 09-18 14:26:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 112.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 25.4%, MM cache hit rate: 89.5%
|
||||
(APIServer pid=1) INFO 09-18 14:26:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.01, Accepted throughput: 56.60 tokens/s, Drafted throughput: 111.59 tokens/s, Accepted: 566 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.658, 0.357, Avg Draft acceptance rate: 50.7%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:50340 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:37702 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:26:52 [loggers.py:323] Engine 000: Avg prompt throughput: 8.7 tokens/s, Avg generation throughput: 117.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 25.4%, MM cache hit rate: 89.5%
|
||||
(APIServer pid=1) INFO 09-18 14:26:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.12, Accepted throughput: 62.10 tokens/s, Drafted throughput: 110.80 tokens/s, Accepted: 621 tokens, Drafted: 1108 tokens, Per-position acceptance rate: 0.704, 0.417, Avg Draft acceptance rate: 56.0%
|
||||
(APIServer pid=1) INFO 09-18 14:27:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 120.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 25.4%, MM cache hit rate: 89.5%
|
||||
(APIServer pid=1) INFO 09-18 14:27:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.17, Accepted throughput: 65.20 tokens/s, Drafted throughput: 111.40 tokens/s, Accepted: 652 tokens, Drafted: 1114 tokens, Per-position acceptance rate: 0.707, 0.463, Avg Draft acceptance rate: 58.5%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:41528 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:27:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 45.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 53.8%, Prefix cache hit rate: 23.2%, MM cache hit rate: 89.5%
|
||||
(APIServer pid=1) INFO 09-18 14:27:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.25, Accepted throughput: 25.40 tokens/s, Drafted throughput: 40.80 tokens/s, Accepted: 254 tokens, Drafted: 408 tokens, Per-position acceptance rate: 0.745, 0.500, Avg Draft acceptance rate: 62.3%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:34370 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:60904 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:27:22 [loggers.py:323] Engine 000: Avg prompt throughput: 12913.2 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.1%, Prefix cache hit rate: 23.2%, MM cache hit rate: 90.0%
|
||||
(APIServer pid=1) INFO 09-18 14:27:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.66, Accepted throughput: 14.10 tokens/s, Drafted throughput: 17.00 tokens/s, Accepted: 141 tokens, Drafted: 170 tokens, Per-position acceptance rate: 0.894, 0.765, Avg Draft acceptance rate: 82.9%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:60920 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33700 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:27:32 [loggers.py:323] Engine 000: Avg prompt throughput: 1215.4 tokens/s, Avg generation throughput: 98.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 23.0%, MM cache hit rate: 90.5%
|
||||
(APIServer pid=1) INFO 09-18 14:27:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.09, Accepted throughput: 51.20 tokens/s, Drafted throughput: 93.60 tokens/s, Accepted: 512 tokens, Drafted: 936 tokens, Per-position acceptance rate: 0.660, 0.434, Avg Draft acceptance rate: 54.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:44936 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:27:42 [loggers.py:323] Engine 000: Avg prompt throughput: 8.7 tokens/s, Avg generation throughput: 134.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 23.0%, MM cache hit rate: 90.5%
|
||||
(APIServer pid=1) INFO 09-18 14:27:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.43, Accepted throughput: 79.30 tokens/s, Drafted throughput: 110.60 tokens/s, Accepted: 793 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.817, 0.617, Avg Draft acceptance rate: 71.7%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:45276 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:27:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 118.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 23.0%, MM cache hit rate: 90.5%
|
||||
(APIServer pid=1) INFO 09-18 14:27:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.11, Accepted throughput: 62.10 tokens/s, Drafted throughput: 111.80 tokens/s, Accepted: 621 tokens, Drafted: 1118 tokens, Per-position acceptance rate: 0.699, 0.411, Avg Draft acceptance rate: 55.5%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:34630 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:28:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 76.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 43.1%, Prefix cache hit rate: 21.2%, MM cache hit rate: 90.5%
|
||||
(APIServer pid=1) INFO 09-18 14:28:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.16, Accepted throughput: 41.10 tokens/s, Drafted throughput: 70.60 tokens/s, Accepted: 411 tokens, Drafted: 706 tokens, Per-position acceptance rate: 0.711, 0.453, Avg Draft acceptance rate: 58.2%
|
||||
(APIServer pid=1) INFO 09-18 14:28:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 84.6%, Prefix cache hit rate: 21.2%, MM cache hit rate: 90.5%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47082 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:37824 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:49298 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:37832 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:28:22 [loggers.py:323] Engine 000: Avg prompt throughput: 14128.0 tokens/s, Avg generation throughput: 98.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 21.0%, MM cache hit rate: 91.3%
|
||||
(APIServer pid=1) INFO 09-18 14:28:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.41, Accepted throughput: 28.70 tokens/s, Drafted throughput: 40.60 tokens/s, Accepted: 574 tokens, Drafted: 812 tokens, Per-position acceptance rate: 0.793, 0.621, Avg Draft acceptance rate: 70.7%
|
||||
(APIServer pid=1) INFO 09-18 14:28:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 124.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 21.0%, MM cache hit rate: 91.3%
|
||||
(APIServer pid=1) INFO 09-18 14:28:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.24, Accepted throughput: 69.10 tokens/s, Drafted throughput: 111.20 tokens/s, Accepted: 691 tokens, Drafted: 1112 tokens, Per-position acceptance rate: 0.745, 0.498, Avg Draft acceptance rate: 62.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53742 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:28:42 [loggers.py:323] Engine 000: Avg prompt throughput: 8.5 tokens/s, Avg generation throughput: 111.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 21.0%, MM cache hit rate: 91.3%
|
||||
(APIServer pid=1) INFO 09-18 14:28:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.02, Accepted throughput: 56.50 tokens/s, Drafted throughput: 110.40 tokens/s, Accepted: 565 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.650, 0.373, Avg Draft acceptance rate: 51.2%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:49254 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:44080 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:28:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 111.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.8%, Prefix cache hit rate: 19.5%, MM cache hit rate: 91.3%
|
||||
(APIServer pid=1) INFO 09-18 14:28:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.40, Accepted throughput: 64.89 tokens/s, Drafted throughput: 92.59 tokens/s, Accepted: 649 tokens, Drafted: 926 tokens, Per-position acceptance rate: 0.803, 0.598, Avg Draft acceptance rate: 70.1%
|
||||
(APIServer pid=1) INFO 09-18 14:29:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 76.9%, Prefix cache hit rate: 19.5%, MM cache hit rate: 91.3%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53800 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53804 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53806 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:29:12 [loggers.py:323] Engine 000: Avg prompt throughput: 14128.1 tokens/s, Avg generation throughput: 75.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 19.4%, MM cache hit rate: 92.0%
|
||||
(APIServer pid=1) INFO 09-18 14:29:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.57, Accepted throughput: 23.05 tokens/s, Drafted throughput: 29.40 tokens/s, Accepted: 461 tokens, Drafted: 588 tokens, Per-position acceptance rate: 0.864, 0.704, Avg Draft acceptance rate: 78.4%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:51408 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:29:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 118.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 19.4%, MM cache hit rate: 92.0%
|
||||
(APIServer pid=1) INFO 09-18 14:29:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.13, Accepted throughput: 62.89 tokens/s, Drafted throughput: 111.59 tokens/s, Accepted: 629 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.703, 0.425, Avg Draft acceptance rate: 56.4%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:60596 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:29:32 [loggers.py:323] Engine 000: Avg prompt throughput: 8.7 tokens/s, Avg generation throughput: 124.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 19.4%, MM cache hit rate: 92.0%
|
||||
(APIServer pid=1) INFO 09-18 14:29:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.25, Accepted throughput: 68.89 tokens/s, Drafted throughput: 110.59 tokens/s, Accepted: 689 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.732, 0.514, Avg Draft acceptance rate: 62.3%
|
||||
(APIServer pid=1) INFO 09-18 14:29:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 119.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 19.4%, MM cache hit rate: 92.0%
|
||||
(APIServer pid=1) INFO 09-18 14:29:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.15, Accepted throughput: 64.20 tokens/s, Drafted throughput: 111.20 tokens/s, Accepted: 642 tokens, Drafted: 1112 tokens, Per-position acceptance rate: 0.716, 0.439, Avg Draft acceptance rate: 57.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:60372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:41996 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:29:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 11.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 66.2%, Prefix cache hit rate: 18.1%, MM cache hit rate: 92.0%
|
||||
(APIServer pid=1) INFO 09-18 14:29:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.23, Accepted throughput: 6.40 tokens/s, Drafted throughput: 10.40 tokens/s, Accepted: 64 tokens, Drafted: 104 tokens, Per-position acceptance rate: 0.769, 0.462, Avg Draft acceptance rate: 61.5%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59394 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59402 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59410 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:30:02 [loggers.py:323] Engine 000: Avg prompt throughput: 14127.4 tokens/s, Avg generation throughput: 41.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 18.0%, MM cache hit rate: 92.6%
|
||||
(APIServer pid=1) INFO 09-18 14:30:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.76, Accepted throughput: 26.00 tokens/s, Drafted throughput: 29.60 tokens/s, Accepted: 260 tokens, Drafted: 296 tokens, Per-position acceptance rate: 0.912, 0.845, Avg Draft acceptance rate: 87.8%
|
||||
(APIServer pid=1) INFO 09-18 14:30:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 117.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 18.0%, MM cache hit rate: 92.6%
|
||||
(APIServer pid=1) INFO 09-18 14:30:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.10, Accepted throughput: 61.50 tokens/s, Drafted throughput: 111.39 tokens/s, Accepted: 615 tokens, Drafted: 1114 tokens, Per-position acceptance rate: 0.666, 0.438, Avg Draft acceptance rate: 55.2%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:43480 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:48952 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:30:22 [loggers.py:323] Engine 000: Avg prompt throughput: 8.7 tokens/s, Avg generation throughput: 119.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 18.0%, MM cache hit rate: 92.6%
|
||||
(APIServer pid=1) INFO 09-18 14:30:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.16, Accepted throughput: 63.90 tokens/s, Drafted throughput: 110.39 tokens/s, Accepted: 639 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.699, 0.458, Avg Draft acceptance rate: 57.9%
|
||||
(APIServer pid=1) INFO 09-18 14:30:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 120.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 18.0%, MM cache hit rate: 92.6%
|
||||
(APIServer pid=1) INFO 09-18 14:30:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.16, Accepted throughput: 65.00 tokens/s, Drafted throughput: 111.59 tokens/s, Accepted: 650 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.706, 0.459, Avg Draft acceptance rate: 58.2%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:50314 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:30:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 52.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 50.8%, Prefix cache hit rate: 16.9%, MM cache hit rate: 92.6%
|
||||
(APIServer pid=1) INFO 09-18 14:30:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.16, Accepted throughput: 28.10 tokens/s, Drafted throughput: 48.40 tokens/s, Accepted: 281 tokens, Drafted: 484 tokens, Per-position acceptance rate: 0.665, 0.496, Avg Draft acceptance rate: 58.1%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:50034 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:58630 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:30:52 [loggers.py:323] Engine 000: Avg prompt throughput: 12598.1 tokens/s, Avg generation throughput: 13.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 16.8%, MM cache hit rate: 92.9%
|
||||
(APIServer pid=1) INFO 09-18 14:30:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.71, Accepted throughput: 8.40 tokens/s, Drafted throughput: 9.80 tokens/s, Accepted: 84 tokens, Drafted: 98 tokens, Per-position acceptance rate: 0.878, 0.837, Avg Draft acceptance rate: 85.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:58644 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38342 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:31:02 [loggers.py:323] Engine 000: Avg prompt throughput: 1529.0 tokens/s, Avg generation throughput: 110.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 16.7%, MM cache hit rate: 93.1%
|
||||
(APIServer pid=1) INFO 09-18 14:31:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.35, Accepted throughput: 63.19 tokens/s, Drafted throughput: 93.59 tokens/s, Accepted: 632 tokens, Drafted: 936 tokens, Per-position acceptance rate: 0.778, 0.573, Avg Draft acceptance rate: 67.5%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:49568 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:31:12 [loggers.py:323] Engine 000: Avg prompt throughput: 8.5 tokens/s, Avg generation throughput: 126.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 16.7%, MM cache hit rate: 93.1%
|
||||
(APIServer pid=1) INFO 09-18 14:31:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.30, Accepted throughput: 71.59 tokens/s, Drafted throughput: 110.39 tokens/s, Accepted: 716 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.766, 0.531, Avg Draft acceptance rate: 64.9%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:53428 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:31:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 120.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 16.7%, MM cache hit rate: 93.1%
|
||||
(APIServer pid=1) INFO 09-18 14:31:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.16, Accepted throughput: 65.00 tokens/s, Drafted throughput: 111.59 tokens/s, Accepted: 650 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.703, 0.462, Avg Draft acceptance rate: 58.2%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53214 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:31:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 78.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.0%, Prefix cache hit rate: 15.8%, MM cache hit rate: 93.1%
|
||||
(APIServer pid=1) INFO 09-18 14:31:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.02, Accepted throughput: 39.70 tokens/s, Drafted throughput: 77.60 tokens/s, Accepted: 397 tokens, Drafted: 776 tokens, Per-position acceptance rate: 0.631, 0.392, Avg Draft acceptance rate: 51.2%
|
||||
(APIServer pid=1) INFO 09-18 14:31:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 81.5%, Prefix cache hit rate: 15.8%, MM cache hit rate: 93.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59288 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:56286 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59294 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59298 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:31:52 [loggers.py:323] Engine 000: Avg prompt throughput: 14129.1 tokens/s, Avg generation throughput: 88.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:31:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.40, Accepted throughput: 25.75 tokens/s, Drafted throughput: 36.80 tokens/s, Accepted: 515 tokens, Drafted: 736 tokens, Per-position acceptance rate: 0.788, 0.611, Avg Draft acceptance rate: 70.0%
|
||||
(APIServer pid=1) INFO 09-18 14:32:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 120.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:32:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.15, Accepted throughput: 64.30 tokens/s, Drafted throughput: 111.40 tokens/s, Accepted: 643 tokens, Drafted: 1114 tokens, Per-position acceptance rate: 0.713, 0.442, Avg Draft acceptance rate: 57.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:40376 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:32:12 [loggers.py:323] Engine 000: Avg prompt throughput: 8.7 tokens/s, Avg generation throughput: 129.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:32:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.35, Accepted throughput: 74.70 tokens/s, Drafted throughput: 110.40 tokens/s, Accepted: 747 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.775, 0.578, Avg Draft acceptance rate: 67.7%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:51024 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:54062 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:32:22 [loggers.py:323] Engine 000: Avg prompt throughput: 25.2 tokens/s, Avg generation throughput: 127.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:32:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.35, Accepted throughput: 73.19 tokens/s, Drafted throughput: 108.79 tokens/s, Accepted: 732 tokens, Drafted: 1088 tokens, Per-position acceptance rate: 0.774, 0.572, Avg Draft acceptance rate: 67.3%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:54076 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:54078 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53562 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:32:32 [loggers.py:323] Engine 000: Avg prompt throughput: 38.1 tokens/s, Avg generation throughput: 133.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:32:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.48, Accepted throughput: 79.80 tokens/s, Drafted throughput: 107.80 tokens/s, Accepted: 798 tokens, Drafted: 1078 tokens, Per-position acceptance rate: 0.831, 0.649, Avg Draft acceptance rate: 74.0%
|
||||
(APIServer pid=1) INFO 09-18 14:32:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 128.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:32:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.32, Accepted throughput: 73.10 tokens/s, Drafted throughput: 110.60 tokens/s, Accepted: 731 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.750, 0.571, Avg Draft acceptance rate: 66.1%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:36210 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53576 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:32:52 [loggers.py:323] Engine 000: Avg prompt throughput: 13.4 tokens/s, Avg generation throughput: 131.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:32:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.42, Accepted throughput: 77.20 tokens/s, Drafted throughput: 109.00 tokens/s, Accepted: 772 tokens, Drafted: 1090 tokens, Per-position acceptance rate: 0.796, 0.620, Avg Draft acceptance rate: 70.8%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:38118 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47594 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47610 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47622 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:33:02 [loggers.py:323] Engine 000: Avg prompt throughput: 59.9 tokens/s, Avg generation throughput: 137.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:33:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.59, Accepted throughput: 84.40 tokens/s, Drafted throughput: 106.20 tokens/s, Accepted: 844 tokens, Drafted: 1062 tokens, Per-position acceptance rate: 0.868, 0.721, Avg Draft acceptance rate: 79.5%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47632 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:47648 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:45364 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:33:12 [loggers.py:323] Engine 000: Avg prompt throughput: 27.9 tokens/s, Avg generation throughput: 142.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:33:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.64, Accepted throughput: 88.50 tokens/s, Drafted throughput: 108.00 tokens/s, Accepted: 885 tokens, Drafted: 1080 tokens, Per-position acceptance rate: 0.891, 0.748, Avg Draft acceptance rate: 81.9%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:45372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:45386 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:47374 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:33:22 [loggers.py:323] Engine 000: Avg prompt throughput: 32.5 tokens/s, Avg generation throughput: 28.2 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 15.6%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO 09-18 14:33:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.86, Accepted throughput: 18.40 tokens/s, Drafted throughput: 19.80 tokens/s, Accepted: 184 tokens, Drafted: 198 tokens, Per-position acceptance rate: 0.970, 0.889, Avg Draft acceptance rate: 92.9%
|
||||
(APIServer pid=1) INFO 09-18 14:33:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 15.6%, MM cache hit rate: 93.5%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:51446 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:47488 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:50100 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:54884 - "GET /metrics HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:49228 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:45014 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:45300 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:46998 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:51132 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.16.100.3:62103 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.16.100.3:62104 - "GET /v1/models HTTP/1.1" 401 Unauthorized
|
||||
(APIServer pid=1) INFO: 127.0.0.1:58480 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:46524 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:53330 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:55746 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:56396 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:55482 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:56702 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:36636 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:34060 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:44360 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:52262 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:60346 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:32864 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:47770 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:59988 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:47698 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:44138 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:52264 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:52764 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:54950 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:37836 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:50426 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:42744 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:36986 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:46082 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33154 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:33156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35076 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35092 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35102 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35110 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35124 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35134 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35136 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35150 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35160 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35174 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35188 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35192 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35200 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35204 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:50:02 [loggers.py:323] Engine 000: Avg prompt throughput: 399.3 tokens/s, Avg generation throughput: 105.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.1%, Prefix cache hit rate: 15.6%, MM cache hit rate: 93.8%
|
||||
(APIServer pid=1) INFO 09-18 14:50:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.84, Accepted throughput: 0.68 tokens/s, Drafted throughput: 0.74 tokens/s, Accepted: 679 tokens, Drafted: 740 tokens, Per-position acceptance rate: 0.957, 0.878, Avg Draft acceptance rate: 91.8%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:35214 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56506 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56510 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:56514 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:50:12 [loggers.py:323] Engine 000: Avg prompt throughput: 1821.0 tokens/s, Avg generation throughput: 96.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:50:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.32, Accepted throughput: 55.10 tokens/s, Drafted throughput: 83.40 tokens/s, Accepted: 551 tokens, Drafted: 834 tokens, Per-position acceptance rate: 0.763, 0.559, Avg Draft acceptance rate: 66.1%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:55476 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:50:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 127.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:50:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.29, Accepted throughput: 72.00 tokens/s, Drafted throughput: 111.60 tokens/s, Accepted: 720 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.758, 0.532, Avg Draft acceptance rate: 64.5%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:37334 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:50:32 [loggers.py:323] Engine 000: Avg prompt throughput: 7.7 tokens/s, Avg generation throughput: 122.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:50:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.21, Accepted throughput: 66.99 tokens/s, Drafted throughput: 110.39 tokens/s, Accepted: 670 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.725, 0.489, Avg Draft acceptance rate: 60.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:37138 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:50:42 [loggers.py:323] Engine 000: Avg prompt throughput: 7.7 tokens/s, Avg generation throughput: 121.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:50:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.20, Accepted throughput: 66.00 tokens/s, Drafted throughput: 110.39 tokens/s, Accepted: 660 tokens, Drafted: 1104 tokens, Per-position acceptance rate: 0.728, 0.467, Avg Draft acceptance rate: 59.8%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:53092 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:50:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 126.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:50:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.27, Accepted throughput: 70.60 tokens/s, Drafted throughput: 111.59 tokens/s, Accepted: 706 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.742, 0.523, Avg Draft acceptance rate: 63.3%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:39218 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:39220 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:51:02 [loggers.py:323] Engine 000: Avg prompt throughput: 15.8 tokens/s, Avg generation throughput: 121.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:51:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.22, Accepted throughput: 66.50 tokens/s, Drafted throughput: 109.40 tokens/s, Accepted: 665 tokens, Drafted: 1094 tokens, Per-position acceptance rate: 0.729, 0.486, Avg Draft acceptance rate: 60.8%
|
||||
(APIServer pid=1) INFO 09-18 14:51:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 127.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:51:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.27, Accepted throughput: 71.20 tokens/s, Drafted throughput: 112.00 tokens/s, Accepted: 712 tokens, Drafted: 1120 tokens, Per-position acceptance rate: 0.750, 0.521, Avg Draft acceptance rate: 63.6%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:41374 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:37916 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:51:22 [loggers.py:323] Engine 000: Avg prompt throughput: 7.9 tokens/s, Avg generation throughput: 128.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:51:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.32, Accepted throughput: 73.29 tokens/s, Drafted throughput: 110.79 tokens/s, Accepted: 733 tokens, Drafted: 1108 tokens, Per-position acceptance rate: 0.769, 0.554, Avg Draft acceptance rate: 66.2%
|
||||
(APIServer pid=1) INFO 09-18 14:51:32 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 128.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:51:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.30, Accepted throughput: 72.70 tokens/s, Drafted throughput: 111.60 tokens/s, Accepted: 727 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.758, 0.545, Avg Draft acceptance rate: 65.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:49184 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:51:42 [loggers.py:323] Engine 000: Avg prompt throughput: 7.9 tokens/s, Avg generation throughput: 127.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:51:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.30, Accepted throughput: 72.30 tokens/s, Drafted throughput: 111.00 tokens/s, Accepted: 723 tokens, Drafted: 1110 tokens, Per-position acceptance rate: 0.769, 0.533, Avg Draft acceptance rate: 65.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:39208 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:55292 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:39224 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:51:52 [loggers.py:323] Engine 000: Avg prompt throughput: 15.8 tokens/s, Avg generation throughput: 125.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:51:52 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.29, Accepted throughput: 71.00 tokens/s, Drafted throughput: 109.99 tokens/s, Accepted: 710 tokens, Drafted: 1100 tokens, Per-position acceptance rate: 0.744, 0.547, Avg Draft acceptance rate: 64.5%
|
||||
(APIServer pid=1) INFO 09-18 14:52:02 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 120.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:52:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.15, Accepted throughput: 64.39 tokens/s, Drafted throughput: 111.59 tokens/s, Accepted: 644 tokens, Drafted: 1116 tokens, Per-position acceptance rate: 0.697, 0.457, Avg Draft acceptance rate: 57.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53966 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:52:12 [loggers.py:323] Engine 000: Avg prompt throughput: 7.9 tokens/s, Avg generation throughput: 123.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:52:12 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.23, Accepted throughput: 67.80 tokens/s, Drafted throughput: 110.59 tokens/s, Accepted: 678 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.722, 0.505, Avg Draft acceptance rate: 61.3%
|
||||
(APIServer pid=1) INFO: 127.0.0.1:53434 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:52:22 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 121.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:52:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.18, Accepted throughput: 66.00 tokens/s, Drafted throughput: 111.39 tokens/s, Accepted: 660 tokens, Drafted: 1114 tokens, Per-position acceptance rate: 0.715, 0.470, Avg Draft acceptance rate: 59.2%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:58938 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:52:32 [loggers.py:323] Engine 000: Avg prompt throughput: 7.9 tokens/s, Avg generation throughput: 121.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:52:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.19, Accepted throughput: 65.60 tokens/s, Drafted throughput: 110.60 tokens/s, Accepted: 656 tokens, Drafted: 1106 tokens, Per-position acceptance rate: 0.712, 0.474, Avg Draft acceptance rate: 59.3%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59878 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59890 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59894 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59900 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59908 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59912 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59914 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59930 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59932 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59944 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59950 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:52:42 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 105.8 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 15.7%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO 09-18 14:52:42 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.20, Accepted throughput: 57.80 tokens/s, Drafted throughput: 96.20 tokens/s, Accepted: 578 tokens, Drafted: 962 tokens, Per-position acceptance rate: 0.711, 0.491, Avg Draft acceptance rate: 60.1%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59952 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59958 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59968 - "POST /tokenize HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:59980 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:55976 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:52:52 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 69.2%, Prefix cache hit rate: 14.9%, MM cache hit rate: 94.6%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:32818 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:32834 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:32836 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:53:02 [loggers.py:323] Engine 000: Avg prompt throughput: 13599.1 tokens/s, Avg generation throughput: 35.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.8%, Prefix cache hit rate: 21.8%, MM cache hit rate: 94.7%
|
||||
(APIServer pid=1) INFO 09-18 14:53:02 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.79, Accepted throughput: 11.35 tokens/s, Drafted throughput: 12.70 tokens/s, Accepted: 227 tokens, Drafted: 254 tokens, Per-position acceptance rate: 0.913, 0.874, Avg Draft acceptance rate: 89.4%
|
||||
(APIServer pid=1) INFO 09-18 14:53:12 [loggers.py:323] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 76.9%, Prefix cache hit rate: 21.8%, MM cache hit rate: 94.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53668 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53676 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53688 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53696 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53700 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 127.0.0.1:44356 - "GET /health HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53716 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53720 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:53:22 [loggers.py:323] Engine 000: Avg prompt throughput: 11222.0 tokens/s, Avg generation throughput: 125.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 21.8%, MM cache hit rate: 94.7%
|
||||
(APIServer pid=1) INFO 09-18 14:53:22 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.77, Accepted throughput: 40.25 tokens/s, Drafted throughput: 45.40 tokens/s, Accepted: 805 tokens, Drafted: 908 tokens, Per-position acceptance rate: 0.943, 0.830, Avg Draft acceptance rate: 88.7%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53732 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:53748 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO 09-18 14:53:32 [loggers.py:323] Engine 000: Avg prompt throughput: 17.3 tokens/s, Avg generation throughput: 144.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 21.8%, MM cache hit rate: 94.7%
|
||||
(APIServer pid=1) INFO 09-18 14:53:32 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 2.66, Accepted throughput: 90.39 tokens/s, Drafted throughput: 108.99 tokens/s, Accepted: 904 tokens, Drafted: 1090 tokens, Per-position acceptance rate: 0.899, 0.760, Avg Draft acceptance rate: 82.9%
|
||||
(APIServer pid=1) INFO: 172.21.0.1:40154 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
(APIServer pid=1) INFO: 172.21.0.1:40170 - "POST /v1/chat/completions HTTP/1.1" 200 OK
|
||||
Reference in New Issue
Block a user