Qualify and enable MTP 2 with paired workload evidence and fallback

This commit is contained in:
2026-09-18 22:37:54 +08:00
parent 06bcbcfd4e
commit 918b78a3b8
31 changed files with 3376 additions and 7 deletions
@@ -0,0 +1,7 @@
{"name": "invoice-extraction", "passed": true, "elapsed_s": 1.853186526999707, "response": {"id": "chatcmpl-9004d24804c21cdc", "object": "chat.completion", "created": 1789702200, "model": "qwen3.8-flash-next", "choices": [{"index": 0, "message": {"role": "assistant", "content": "\n\n{\n \"supplier\": \"North Star Ltd\",\n \"currency\": \"CNY\",\n \"total\": 476,\n \"due_date\": \"2026-10-03\"\n}", "refusal": null, "annotations": null, "audio": null, "function_call": null, "reasoning": "We need answer user's request. Need produce final only JSON. Need compute total: 3*125=375, 2*48=96, shipping 20, discount 15. total = 375+96+20-15=476. Keys supplier, currency, total, due_date. JSON only. Ensure integer total.\n"}, "logprobs": null, "finish_reason": "stop", "stop_reason": null, "token_ids": null, "routed_experts": null}], "service_tier": null, "system_fingerprint": "vllm-0.3.1.dev3+g0bfc7a15d-9c4a3436", "usage": {"prompt_tokens": 132, "total_tokens": 270, "completion_tokens": 138, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 85}}, "prompt_logprobs": null, "prompt_token_ids": null, "prompt_text": null, "kv_transfer_params": null, "ec_transfer_params": null, "metrics": null}}
{"name": "deduplicate", "passed": true, "elapsed_s": 3.305690247001621, "response": {"id": "chatcmpl-8c78159521c60876", "object": "chat.completion", "created": 1789702202, "model": "qwen3.8-flash-next", "choices": [{"index": 0, "message": {"role": "assistant", "content": "\n\n{\"ids\":[\"A\",\"B\",\"D\"],\"total\":48}", "refusal": null, "annotations": null, "audio": null, "function_call": null, "reasoning": "We need answer user's request. Need parse input orders: \"A pending 12; B paid 20; A paid 15; C cancelled 99; D paid 8; B paid 25. Keep the last occurrence of each ID, discard cancelled orders, and output ONLY JSON {\"ids\":[...],\"total\":...}. Sort IDs alphabetically.\"\n\nNeed produce final only JSON. Let's compute.\nOrders in sequence:\nA pending 12\nB paid 20\nA paid 15\nC cancelled 99\nD paid 8\nB paid 25\nKeep last occurrence of each ID: A last paid 15, B last paid 25, C last cancelled 99, D paid 8. Discard cancelled orders: remove C. Remaining A 15, B 25, D 8. Sort IDs alphabetically: A, B, D. Total = 15+25+8=48. Output {\"ids\":[\"A\",\"B\",\"D\"],\"total\":48}. Ensure ONLY JSON no markdown.\n"}, "logprobs": null, "finish_reason": "stop", "stop_reason": null, "token_ids": null, "routed_experts": null}], "service_tier": null, "system_fingerprint": "vllm-0.3.1.dev3+g0bfc7a15d-9c4a3436", "usage": {"prompt_tokens": 120, "total_tokens": 370, "completion_tokens": 250, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 233}}, "prompt_logprobs": null, "prompt_token_ids": null, "prompt_text": null, "kv_transfer_params": null, "ec_transfer_params": null, "metrics": null}}
{"name": "schedule", "passed": true, "elapsed_s": 3.0858910660008405, "response": {"id": "chatcmpl-b6ceeff440eaa531", "object": "chat.completion", "created": 1789702205, "model": "qwen3.8-flash-next", "choices": [{"index": 0, "message": {"role": "assistant", "content": "\n\n{\n \"earliest_finish\": 10,\n \"critical_path\": [\"A\", \"B\", \"D\"]\n}", "refusal": null, "annotations": null, "audio": null, "function_call": null, "reasoning": "We need answer user's request. Need parse problem. \"Jobs A takes 3 hours; B takes 5 hours after A; C takes 4 hours after A; D takes 2 hours after both B and C. Unlimited workers, start at hour 0. Output only JSON with earliest_finish and critical_path (array of job letters).\"\n\nWe need compute earliest finish times. A duration 3, starts 0 finishes 3. B after A duration 5 => starts 3 finishes 8. C after A duration 4 => starts 3 finishes 7. D after both B and C duration 2 => starts max(8,7)=8 finishes 10. Earliest finish overall 10. Critical path longest path: A->B->D (3+5+2=10). Output only JSON. Need ensure valid JSON. {\"earliest_finish\":10,\"critical_path\":[\"A\",\"B\",\"D\"]}. No extra text.\n"}, "logprobs": null, "finish_reason": "stop", "stop_reason": null, "token_ids": null, "routed_experts": null}], "service_tier": null, "system_fingerprint": "vllm-0.3.1.dev3+g0bfc7a15d-9c4a3436", "usage": {"prompt_tokens": 110, "total_tokens": 343, "completion_tokens": 233, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 201}}, "prompt_logprobs": null, "prompt_token_ids": null, "prompt_text": null, "kv_transfer_params": null, "ec_transfer_params": null, "metrics": null}}
{"name": "chinese-policy", "passed": true, "elapsed_s": 3.033975599999394, "response": {"id": "chatcmpl-9c030a36836653bb", "object": "chat.completion", "created": 1789702208, "model": "qwen3.8-flash-next", "choices": [{"index": 0, "message": {"role": "assistant", "content": "\n\n{\n \"甲\": \"主管\",\n \"乙\": \"经理\",\n \"丙\": \"经理\",\n \"丁\": \"总监\"\n}", "refusal": null, "annotations": null, "audio": null, "function_call": null, "reasoning": "We need answer user's request. Need produce final only JSON. Need parse Chinese:\nRule: amount <=5000 supervisor (主管); >5000 and <=20000 manager (经理); >20000 director (总监). Urgent doesn't change approver.\nApplications: 甲 5000 -> <=5000 => 主管. 乙 5001 加急 -> >5000 <=20000 => 经理. 丙 20000 -> <=20000 and >5000 => 经理. 丁 20001 -> >20000 => 总监.\nOutput only JSON keys 甲乙丙丁 values. Ensure valid JSON. Could be {\"甲\":\"主管\",\"乙\":\"经理\",\"丙\":\"经理\",\"丁\":\"总监\"}.\n"}, "logprobs": null, "finish_reason": "stop", "stop_reason": null, "token_ids": null, "routed_experts": null}], "service_tier": null, "system_fingerprint": "vllm-0.3.1.dev3+g0bfc7a15d-9c4a3436", "usage": {"prompt_tokens": 148, "total_tokens": 377, "completion_tokens": 229, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 191}}, "prompt_logprobs": null, "prompt_token_ids": null, "prompt_text": null, "kv_transfer_params": null, "ec_transfer_params": null, "metrics": null}}
{"name": "merge-intervals", "passed": false, "error": "the JSON object must be str, bytes or bytearray, not NoneType"}
{"name": "binary-search", "passed": true, "validation_stdout": "passed\n", "validation_stderr": "", "response": {"id": "chatcmpl-89efbe814cd5878e", "object": "chat.completion", "created": 1789702318, "model": "qwen3.8-flash-next", "choices": [{"index": 0, "message": {"role": "assistant", "content": "\n\n{\"code\":\"def lower_bound(nums, target):\\n lo = 0\\n hi = len(nums)\\n while lo < hi:\\n mid = (lo + hi) // 2\\n if nums[mid] < target:\\n lo = mid + 1\\n else:\\n hi = mid\\n return lo\"}", "refusal": null, "annotations": null, "audio": null, "function_call": null, "reasoning": "We need answer user's request. Need produce final ONLY JSON object with a code string. Need satisfy constraints: Write Python function lower_bound. Return index first element >= target in sorted list nums, or len(nums) if none. Parameters nums, target. Reply ONLY JSON object with a code string. Use no imports, type annotations, decorators, top-level statements or helper functions. Allowed builtins: sorted,len,range,min,max,list,enumerate,sum. Allowed methods: append,sort,copy.\n\nNeed output JSON object likely {\"code\": \"def lower_bound(nums, target):\\n ...\"}. No imports etc. Function only. Need maybe binary search using len, range? No type annotations. Could implement:\ndef lower_bound(nums, target):\n lo = 0\n hi = len(nums)\n while lo < hi:\n mid = (lo + hi) // 2\n if nums[mid] < target:\n lo = mid + 1\n else:\n hi = mid\n return lo\n\nThis uses no disallowed builtins. JSON string must escape newlines. Ensure only JSON object. Could include code string. Need no top-level statements. Good.\n\nDouble-check allowed builtins: len allowed. // operator okay. No helper. No annotations.\n\nNeed final only JSON. Ensure valid JSON. Could be:\n{\"code\":\"def lower_bound(nums, target):\\n lo = 0\\n hi = len(nums)\\n while lo < hi:\\n mid = (lo + hi) // 2\\n if nums[mid] < target:\\n lo = mid + 1\\n else:\\n hi = mid\\n return lo\"}\n\nNo trailing newline? fine.\n"}, "logprobs": null, "finish_reason": "stop", "stop_reason": null, "token_ids": null, "routed_experts": null}], "service_tier": null, "system_fingerprint": "vllm-0.3.1.dev3+g0bfc7a15d-9c4a3436", "usage": {"prompt_tokens": 134, "total_tokens": 572, "completion_tokens": 438, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 362}}, "prompt_logprobs": null, "prompt_token_ids": null, "prompt_text": null, "kv_transfer_params": null, "ec_transfer_params": null, "metrics": null}}
{"name": "tool-arguments", "passed": true, "response": {"id": "chatcmpl-817dda016167c20b", "object": "chat.completion", "created": 1789702324, "model": "qwen3.8-flash-next", "choices": [{"index": 0, "message": {"role": "assistant", "content": null, "refusal": null, "annotations": null, "audio": null, "function_call": null, "tool_calls": [{"id": "chatcmpl-tool-8dbfec4209335069", "type": "function", "function": {"name": "create_ticket", "arguments": "{\"title\": \"GPU cooling failure\", \"priority\": \"urgent\", \"asset\": \"GPU-092\"}"}}], "reasoning": "We need to respond to user: \"Call create_ticket with title exactly GPU cooling failure, priority urgent, and asset exactly GPU-092.\" Need call tool. Need final maybe brief. Ensure exact params.\n"}, "logprobs": null, "finish_reason": "tool_calls", "stop_reason": null, "token_ids": null, "routed_experts": null}], "service_tier": null, "system_fingerprint": "vllm-0.3.1.dev3+g0bfc7a15d-9c4a3436", "usage": {"prompt_tokens": 363, "total_tokens": 464, "completion_tokens": 101, "prompt_tokens_details": null, "completion_tokens_details": {"reasoning_tokens": 44}}, "prompt_logprobs": null, "prompt_token_ids": null, "prompt_text": null, "kv_transfer_params": null, "ec_transfer_params": null, "metrics": null}}