diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/identity.txt b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/identity.txt new file mode 100644 index 000000000..5467820c7 --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/identity.txt @@ -0,0 +1,20 @@ +date: 2026-09-21T15:47:10+09:00 +host: spark-101 +kernel: 7.0.0-1019-nvidia +arch: aarch64 +nvidia_driver: 580.178.04 +gpu: NVIDIA GB10, 12.1 +MLX_CUDA_ARCHITECTURES: 121 +cuda_toolkit: 13.0 +mlx_pin: 81ba1c6a0e50a9268b931579c2d4f1158b9aab5a +rustc_on_path: rustc 1.97.1 (8bab26f4f 2026-07-14) +rustc_used_for_build: rustc 1.97.1 (8bab26f4f 2026-07-14) +git_commit: 6f9a0982d7854b653c7b354fe0f1f0b26af41a7f +git_dirty: no +uptime: up 23 hours, 52 minutes +mem_total_gib: 121.7 +binary: /tmp/mlxcel1797.izrEDT/mlxcel1797-server +binary_sha256: f2ba16883e8aed7aecf57212f5a81f4516eaeddcc618252bc5a6f464f015248c +binary_mtime: 2026-09-21T15:47:10+09:00 +model /home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-4bit: bytes=3061131800 model_type=qwen3_5 quantization=affine +model /home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash: bytes=1074861667 model_type=qwen3 quantization=none diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.classic-close.log b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.classic-close.log new file mode 100644 index 000000000..0e49852b4 --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.classic-close.log @@ -0,0 +1,38 @@ +2026-09-21T06:54:44.139710Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-09-21T06:54:44.139753Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-09-21T06:54:44.139888Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=1 kv_unified=false n_parallel=4 prefill_chunk_size=512 max_kv_size=None +2026-09-21T06:54:44.139945Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-09-21T06:54:44.139948Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-09-21T06:54:44.139951Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-09-21T06:54:44.489887Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-09-21T06:54:44.489921Z INFO mlxcel::server::startup: Tokenizer recognizes a think marker pair; defaulting chat_template kwarg `enable_thinking=true` (upstream PR #1114) think_start=Some("") think_end=Some("") think_start_tokens_len=1 think_end_tokens_len=1 +2026-09-21T06:54:44.490492Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-09-21T06:54:44.839550Z INFO mlxcel::server::startup: Warming up model... +2026-09-21T06:54:44.840004Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-09-21T06:54:45.172015Z INFO mlxcel::server::model_provider::model_worker: Model qwen3.5-4b-4bit loaded in 0.332s (resident after load: 0.00 GB) worker_model_id=qwen3.5-4b-4bit load_seconds=0.331970264 active_bytes=0 peak_bytes=0 cache_bytes=0 limit_bytes=124128085401 +2026-09-21T06:54:45.172088Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=1, max_queue_depth=32, prefill_chunk_size=512, max_batch_prefill=4, decode_storage=auto) +2026-09-21T06:54:45.172492Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 3062487 blocks (32 layers, 32-token blocks) +2026-09-21T06:54:45.172506Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 256 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-09-21T06:54:45.172525Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=32 kv_cache_mode_total_layers=32 +2026-09-21T06:54:45.173662Z INFO prefill{seq_id=seq-0 prompt_len=1 cached=0 start=0}: mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-09-21T06:54:46.060687Z INFO prefill{seq_id=seq-0 prompt_len=1 cached=0 start=0}: mlxcel::server::batch::scheduler::prefill: prompt-cache: request completed during prefill: cached=0/1 prompt tokens, total 887ms prompt_tokens=1 cached_tokens=0 generation_time_ms=887 +2026-09-21T06:54:46.060744Z INFO mlxcel::server::startup: Warmup complete +2026-09-21T06:54:46.060838Z INFO mlxcel::server::startup: model_type=Qwen35VLM: enabling native video_url content block support +2026-09-21T06:54:46.061753Z WARN mlxcel::server::startup: The loaded model accepts video input, but `ffmpeg` and `ffprobe` are not both on PATH, so every video request will be refused. The check is cached for the life of the process: restart the server after installing ffmpeg. +2026-09-21T06:54:46.062528Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18797 +2026-09-21T06:54:46.062545Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-09-21T06:54:46.062548Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-09-21T06:54:46.062550Z INFO mlxcel::server::startup: Endpoints: +2026-09-21T06:54:46.062551Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-09-21T06:54:46.062552Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-09-21T06:54:46.062554Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-09-21T06:54:46.062555Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-09-21T06:54:46.062556Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-09-21T06:54:46.062557Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-09-21T06:54:46.062559Z INFO mlxcel::server::startup: GET /props - Server properties +2026-09-21T06:54:46.062560Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-09-21T06:54:46.062561Z INFO mlxcel::server::startup: GET /health - Health check +2026-09-21T06:54:50.818671Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 4012ms prompt_tokens=158 cached_tokens=0 generation_time_ms=4012 +2026-09-21T06:54:54.345526Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3524ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3524 +2026-09-21T06:54:57.869074Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3521ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3521 +2026-09-21T06:55:01.395604Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3524ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3524 diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.classic-open.log b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.classic-open.log new file mode 100644 index 000000000..6dca9cbdd --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.classic-open.log @@ -0,0 +1,38 @@ +2026-09-21T06:48:06.437539Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-09-21T06:48:06.437652Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-09-21T06:48:06.438105Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=1 kv_unified=false n_parallel=4 prefill_chunk_size=512 max_kv_size=None +2026-09-21T06:48:06.438178Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-09-21T06:48:06.438181Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-09-21T06:48:06.438183Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-09-21T06:48:06.763072Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-09-21T06:48:06.763199Z INFO mlxcel::server::startup: Tokenizer recognizes a think marker pair; defaulting chat_template kwarg `enable_thinking=true` (upstream PR #1114) think_start=Some("") think_end=Some("") think_start_tokens_len=1 think_end_tokens_len=1 +2026-09-21T06:48:06.764223Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-09-21T06:48:07.088497Z INFO mlxcel::server::startup: Warming up model... +2026-09-21T06:48:07.089259Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-09-21T06:48:07.462941Z INFO mlxcel::server::model_provider::model_worker: Model qwen3.5-4b-4bit loaded in 0.374s (resident after load: 0.00 GB) worker_model_id=qwen3.5-4b-4bit load_seconds=0.373617994 active_bytes=0 peak_bytes=0 cache_bytes=0 limit_bytes=124128085401 +2026-09-21T06:48:07.463048Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=1, max_queue_depth=32, prefill_chunk_size=512, max_batch_prefill=4, decode_storage=auto) +2026-09-21T06:48:07.463513Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 3062487 blocks (32 layers, 32-token blocks) +2026-09-21T06:48:07.463527Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 256 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-09-21T06:48:07.463641Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=32 kv_cache_mode_total_layers=32 +2026-09-21T06:48:07.465664Z INFO prefill{seq_id=seq-0 prompt_len=1 cached=0 start=0}: mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-09-21T06:48:08.392504Z INFO prefill{seq_id=seq-0 prompt_len=1 cached=0 start=0}: mlxcel::server::batch::scheduler::prefill: prompt-cache: request completed during prefill: cached=0/1 prompt tokens, total 928ms prompt_tokens=1 cached_tokens=0 generation_time_ms=928 +2026-09-21T06:48:08.393130Z INFO mlxcel::server::startup: Warmup complete +2026-09-21T06:48:08.393303Z INFO mlxcel::server::startup: model_type=Qwen35VLM: enabling native video_url content block support +2026-09-21T06:48:08.394254Z WARN mlxcel::server::startup: The loaded model accepts video input, but `ffmpeg` and `ffprobe` are not both on PATH, so every video request will be refused. The check is cached for the life of the process: restart the server after installing ffmpeg. +2026-09-21T06:48:08.395806Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18797 +2026-09-21T06:48:08.395823Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-09-21T06:48:08.395827Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-09-21T06:48:08.395828Z INFO mlxcel::server::startup: Endpoints: +2026-09-21T06:48:08.395829Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-09-21T06:48:08.395831Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-09-21T06:48:08.395832Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-09-21T06:48:08.395833Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-09-21T06:48:08.395834Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-09-21T06:48:08.395836Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-09-21T06:48:08.395837Z INFO mlxcel::server::startup: GET /props - Server properties +2026-09-21T06:48:08.395838Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-09-21T06:48:08.395839Z INFO mlxcel::server::startup: GET /health - Health check +2026-09-21T06:48:13.105736Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3975ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3975 +2026-09-21T06:48:16.608087Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3499ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3499 +2026-09-21T06:48:20.120235Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3510ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3510 +2026-09-21T06:48:23.633141Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3510ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3510 diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w2.log b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w2.log new file mode 100644 index 000000000..4a85f125d --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w2.log @@ -0,0 +1,51 @@ +2026-09-21T06:49:24.992872Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-09-21T06:49:24.992919Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-09-21T06:49:24.993058Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=1 kv_unified=false n_parallel=4 prefill_chunk_size=512 max_kv_size=None +2026-09-21T06:49:24.993112Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-09-21T06:49:24.993115Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-09-21T06:49:24.993117Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-09-21T06:49:25.342234Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-09-21T06:49:25.342270Z INFO mlxcel::server::startup: Tokenizer recognizes a think marker pair; defaulting chat_template kwarg `enable_thinking=true` (upstream PR #1114) think_start=Some("") think_end=Some("") think_start_tokens_len=1 think_end_tokens_len=1 +2026-09-21T06:49:25.342842Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-09-21T06:49:25.698486Z INFO mlxcel::server::startup: Warming up model... +2026-09-21T06:49:25.698527Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-09-21T06:49:26.048383Z INFO mlxcel::server::model_provider::model_worker: Model qwen3.5-4b-4bit loaded in 0.350s (resident after load: 0.00 GB) worker_model_id=qwen3.5-4b-4bit load_seconds=0.349810107 active_bytes=0 peak_bytes=0 cache_bytes=0 limit_bytes=124128085401 +2026-09-21T06:49:26.048457Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=1, max_queue_depth=32, prefill_chunk_size=512, max_batch_prefill=4, decode_storage=auto, speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=2, block_size_from=operator override, explicit_kind=true)) +2026-09-21T06:49:26.048863Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 3062487 blocks (32 layers, 32-token blocks) +2026-09-21T06:49:26.048876Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 256 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-09-21T06:49:26.048896Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=32 kv_cache_mode_total_layers=32 +2026-09-21T06:49:27.517156Z INFO mlxcel::models::speculative_exactness: MTP exactness probe passed: verify block is byte-identical to the single-token chain block_size=2 +2026-09-21T06:49:27.517184Z INFO mlxcel::server::batch::speculative_burst: Lazy-loading drafter from /home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash (kind=Some(Dflash)) +2026-09-21T06:49:27.517560Z INFO mlxcel::server::batch::speculative_burst: Drafter loaded (kind=dflash, 0 ms) +2026-09-21T06:49:27.518467Z INFO mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-09-21T06:49:27.539910Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=2 rounds=0 proposed_tokens=0 accepted_tokens=0 acceptance_rate=0.0 emitted_per_verify=0.0 zero_accept_rounds=0 partial_accept_rounds=0 full_accept_rounds=0 prefill_verify_ms=0.866993 first_bonus_ms=21.36098 first_hidden_ms=0.00696 bind_reset_ms=0.087342 draft_ms=0.0 verify_ms=0.0 target_argmax_sync_ms=0.0 logprobs_ms=0.0 walk_ms=0.0 hidden_concat_ms=0.0 rollback_ms=0.0 decode_ms=0.0 +2026-09-21T06:49:27.540144Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=1 generated_tokens=1 burst_ms=1490 +2026-09-21T06:49:27.540151Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-0 burst_wall_ms=1491.0195509999999 burst_active_ms=1491.0195509999999 slices=1 tokens_generated=1 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:49:27.540280Z INFO mlxcel::server::startup: Warmup complete +2026-09-21T06:49:27.540385Z INFO mlxcel::server::startup: model_type=Qwen35VLM: enabling native video_url content block support +2026-09-21T06:49:27.541656Z WARN mlxcel::server::startup: The loaded model accepts video input, but `ffmpeg` and `ffprobe` are not both on PATH, so every video request will be refused. The check is cached for the life of the process: restart the server after installing ffmpeg. +2026-09-21T06:49:27.542426Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18797 +2026-09-21T06:49:27.542435Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-09-21T06:49:27.542438Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-09-21T06:49:27.542440Z INFO mlxcel::server::startup: Endpoints: +2026-09-21T06:49:27.542441Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-09-21T06:49:27.542443Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-09-21T06:49:27.542444Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-09-21T06:49:27.542445Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-09-21T06:49:27.542447Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-09-21T06:49:27.542448Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-09-21T06:49:27.542449Z INFO mlxcel::server::startup: GET /props - Server properties +2026-09-21T06:49:27.542451Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-09-21T06:49:27.542452Z INFO mlxcel::server::startup: GET /health - Health check +2026-09-21T06:49:33.706149Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=2 rounds=114 proposed_tokens=114 accepted_tokens=86 acceptance_rate=0.7543859649122807 emitted_per_verify=1.7456140350877194 zero_accept_rounds=28 partial_accept_rounds=0 full_accept_rounds=86 prefill_verify_ms=5.039418 first_bonus_ms=584.160803 first_hidden_ms=0.010064 bind_reset_ms=0.031344 draft_ms=2817.84715 verify_ms=243.60834000000003 target_argmax_sync_ms=2372.036596 logprobs_ms=0.02332800000000002 walk_ms=0.07868400000000003 hidden_concat_ms=0.21332500000000004 rollback_ms=13.764792999999997 decode_ms=5465.96001 +2026-09-21T06:49:33.706752Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=6055 +2026-09-21T06:49:33.706788Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-1 burst_wall_ms=6055.963497 burst_active_ms=6055.963497 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:49:36.688396Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=2 rounds=114 proposed_tokens=114 accepted_tokens=86 acceptance_rate=0.7543859649122807 emitted_per_verify=1.7456140350877194 zero_accept_rounds=28 partial_accept_rounds=0 full_accept_rounds=86 prefill_verify_ms=2.789697 first_bonus_ms=120.44864700000001 first_hidden_ms=0.028735 bind_reset_ms=0.085807 draft_ms=366.5048819999998 verify_ms=230.36967999999996 target_argmax_sync_ms=2223.9670730000003 logprobs_ms=0.01899199999999999 walk_ms=0.061472000000000006 hidden_concat_ms=0.219642 rollback_ms=12.891338000000003 decode_ms=2853.268224 +2026-09-21T06:49:36.688732Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=2977 +2026-09-21T06:49:36.688776Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-2 burst_wall_ms=2977.156654 burst_active_ms=2977.156654 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:49:39.692914Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=2 rounds=114 proposed_tokens=114 accepted_tokens=86 acceptance_rate=0.7543859649122807 emitted_per_verify=1.7456140350877194 zero_accept_rounds=28 partial_accept_rounds=0 full_accept_rounds=86 prefill_verify_ms=2.66602 first_bonus_ms=122.498706 first_hidden_ms=0.015296 bind_reset_ms=0.082703 draft_ms=367.421775 verify_ms=241.678898 target_argmax_sync_ms=2232.3261150000003 logprobs_ms=0.019839000000000013 walk_ms=0.062046999999999984 hidden_concat_ms=0.22460700000000003 rollback_ms=12.655679 decode_ms=2875.0160830000004 +2026-09-21T06:49:39.693244Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=3000 +2026-09-21T06:49:39.693287Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-3 burst_wall_ms=3000.659458 burst_active_ms=3000.659458 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:49:42.697288Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=2 rounds=114 proposed_tokens=114 accepted_tokens=86 acceptance_rate=0.7543859649122807 emitted_per_verify=1.7456140350877194 zero_accept_rounds=28 partial_accept_rounds=0 full_accept_rounds=86 prefill_verify_ms=2.536246 first_bonus_ms=120.437795 first_hidden_ms=0.015216 bind_reset_ms=0.077967 draft_ms=375.9427449999999 verify_ms=243.70482399999995 target_argmax_sync_ms=2223.5196660000006 logprobs_ms=0.021328000000000003 walk_ms=0.05954900000000002 hidden_concat_ms=0.22059199999999998 rollback_ms=13.086343999999999 decode_ms=2877.03246 +2026-09-21T06:49:42.697559Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=3000 +2026-09-21T06:49:42.697601Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-4 burst_wall_ms=3000.426782 burst_active_ms=3000.426782 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w3.log b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w3.log new file mode 100644 index 000000000..39898a781 --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w3.log @@ -0,0 +1,51 @@ +2026-09-21T06:50:44.017525Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-09-21T06:50:44.017567Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-09-21T06:50:44.017698Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=1 kv_unified=false n_parallel=4 prefill_chunk_size=512 max_kv_size=None +2026-09-21T06:50:44.017760Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-09-21T06:50:44.017762Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-09-21T06:50:44.017765Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-09-21T06:50:44.340620Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-09-21T06:50:44.340656Z INFO mlxcel::server::startup: Tokenizer recognizes a think marker pair; defaulting chat_template kwarg `enable_thinking=true` (upstream PR #1114) think_start=Some("") think_end=Some("") think_start_tokens_len=1 think_end_tokens_len=1 +2026-09-21T06:50:44.341210Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-09-21T06:50:44.667858Z INFO mlxcel::server::startup: Warming up model... +2026-09-21T06:50:44.668331Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-09-21T06:50:45.028685Z INFO mlxcel::server::model_provider::model_worker: Model qwen3.5-4b-4bit loaded in 0.360s (resident after load: 0.00 GB) worker_model_id=qwen3.5-4b-4bit load_seconds=0.360325624 active_bytes=0 peak_bytes=0 cache_bytes=0 limit_bytes=124128085401 +2026-09-21T06:50:45.028763Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=1, max_queue_depth=32, prefill_chunk_size=512, max_batch_prefill=4, decode_storage=auto, speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=3, block_size_from=operator override, explicit_kind=true)) +2026-09-21T06:50:45.029181Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 3062487 blocks (32 layers, 32-token blocks) +2026-09-21T06:50:45.029196Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 256 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-09-21T06:50:45.029218Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=32 kv_cache_mode_total_layers=32 +2026-09-21T06:50:46.660713Z INFO mlxcel::models::speculative_exactness: MTP exactness probe passed: verify block is byte-identical to the single-token chain block_size=3 +2026-09-21T06:50:46.660747Z INFO mlxcel::server::batch::speculative_burst: Lazy-loading drafter from /home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash (kind=Some(Dflash)) +2026-09-21T06:50:46.661150Z INFO mlxcel::server::batch::speculative_burst: Drafter loaded (kind=dflash, 0 ms) +2026-09-21T06:50:46.662104Z INFO mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-09-21T06:50:46.684798Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=3 rounds=0 proposed_tokens=0 accepted_tokens=0 acceptance_rate=0.0 emitted_per_verify=0.0 zero_accept_rounds=0 partial_accept_rounds=0 full_accept_rounds=0 prefill_verify_ms=0.907668 first_bonus_ms=22.669054000000003 first_hidden_ms=0.008336 bind_reset_ms=0.030576 draft_ms=0.0 verify_ms=0.0 target_argmax_sync_ms=0.0 logprobs_ms=0.0 walk_ms=0.0 hidden_concat_ms=0.0 rollback_ms=0.0 decode_ms=0.0 +2026-09-21T06:50:46.684967Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=1 generated_tokens=1 burst_ms=1655 +2026-09-21T06:50:46.684976Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-0 burst_wall_ms=1655.5268839999999 burst_active_ms=1655.5268839999999 slices=1 tokens_generated=1 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:50:46.685063Z INFO mlxcel::server::startup: Warmup complete +2026-09-21T06:50:46.685156Z INFO mlxcel::server::startup: model_type=Qwen35VLM: enabling native video_url content block support +2026-09-21T06:50:46.685711Z WARN mlxcel::server::startup: The loaded model accepts video input, but `ffmpeg` and `ffprobe` are not both on PATH, so every video request will be refused. The check is cached for the life of the process: restart the server after installing ffmpeg. +2026-09-21T06:50:46.686524Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18797 +2026-09-21T06:50:46.686541Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-09-21T06:50:46.686545Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-09-21T06:50:46.686547Z INFO mlxcel::server::startup: Endpoints: +2026-09-21T06:50:46.686548Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-09-21T06:50:46.686550Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-09-21T06:50:46.686551Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-09-21T06:50:46.686553Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-09-21T06:50:46.686554Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-09-21T06:50:46.686556Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-09-21T06:50:46.686557Z INFO mlxcel::server::startup: GET /props - Server properties +2026-09-21T06:50:46.686559Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-09-21T06:50:46.686560Z INFO mlxcel::server::startup: GET /health - Health check +2026-09-21T06:50:53.206519Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=3 rounds=91 proposed_tokens=182 accepted_tokens=109 acceptance_rate=0.5989010989010989 emitted_per_verify=2.1868131868131866 zero_accept_rounds=27 partial_accept_rounds=19 full_accept_rounds=45 prefill_verify_ms=5.046424 first_bonus_ms=599.3535880000001 first_hidden_ms=0.0084 bind_reset_ms=0.027695 draft_ms=2381.873432 verify_ms=206.602133 target_argmax_sync_ms=2291.583825999999 logprobs_ms=0.012304 walk_ms=0.06073599999999999 hidden_concat_ms=0.14649299999999998 rollback_ms=26.686075 decode_ms=4916.782923999999 +2026-09-21T06:50:53.207081Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=5521 +2026-09-21T06:50:53.207120Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-1 burst_wall_ms=5521.951357 burst_active_ms=5521.951357 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:50:56.152906Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=3 rounds=91 proposed_tokens=182 accepted_tokens=109 acceptance_rate=0.5989010989010989 emitted_per_verify=2.1868131868131866 zero_accept_rounds=27 partial_accept_rounds=19 full_accept_rounds=45 prefill_verify_ms=2.935063 first_bonus_ms=119.746444 first_hidden_ms=0.008 bind_reset_ms=0.044848 draft_ms=317.46032899999994 verify_ms=218.80341699999997 target_argmax_sync_ms=2238.6714439999996 logprobs_ms=0.008304 walk_ms=0.06215999999999999 hidden_concat_ms=0.16542 rollback_ms=31.804025 decode_ms=2820.244314 +2026-09-21T06:50:56.153216Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=2943 +2026-09-21T06:50:56.153259Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-2 burst_wall_ms=2943.43799 burst_active_ms=2943.43799 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:50:59.103353Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=3 rounds=91 proposed_tokens=182 accepted_tokens=109 acceptance_rate=0.5989010989010989 emitted_per_verify=2.1868131868131866 zero_accept_rounds=27 partial_accept_rounds=19 full_accept_rounds=45 prefill_verify_ms=2.787385 first_bonus_ms=116.640179 first_hidden_ms=0.015088 bind_reset_ms=0.086239 draft_ms=318.965685 verify_ms=215.09529099999992 target_argmax_sync_ms=2250.130131 logprobs_ms=0.008736000000000006 walk_ms=0.06255900000000003 hidden_concat_ms=0.15551900000000007 rollback_ms=31.881363 decode_ms=2827.595127 +2026-09-21T06:50:59.103648Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=2947 +2026-09-21T06:50:59.103690Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-3 burst_wall_ms=2947.498701 burst_active_ms=2947.498701 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:51:02.055005Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=3 rounds=91 proposed_tokens=182 accepted_tokens=109 acceptance_rate=0.5989010989010989 emitted_per_verify=2.1868131868131866 zero_accept_rounds=27 partial_accept_rounds=19 full_accept_rounds=45 prefill_verify_ms=2.920808 first_bonus_ms=118.201001 first_hidden_ms=0.0152 bind_reset_ms=0.08767899999999999 draft_ms=320.41396199999997 verify_ms=222.00846599999994 target_argmax_sync_ms=2241.846096 logprobs_ms=0.010144000000000005 walk_ms=0.06505399999999999 hidden_concat_ms=0.162831 rollback_ms=31.886941000000004 decode_ms=2827.3205580000003 +2026-09-21T06:51:02.055469Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=2949 +2026-09-21T06:51:02.055539Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-4 burst_wall_ms=2949.135253 burst_active_ms=2949.135253 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w4.log b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w4.log new file mode 100644 index 000000000..107ae877f --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w4.log @@ -0,0 +1,51 @@ +2026-09-21T06:52:03.411715Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-09-21T06:52:03.411757Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-09-21T06:52:03.411891Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=1 kv_unified=false n_parallel=4 prefill_chunk_size=512 max_kv_size=None +2026-09-21T06:52:03.411972Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-09-21T06:52:03.411976Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-09-21T06:52:03.411978Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-09-21T06:52:03.732661Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-09-21T06:52:03.732697Z INFO mlxcel::server::startup: Tokenizer recognizes a think marker pair; defaulting chat_template kwarg `enable_thinking=true` (upstream PR #1114) think_start=Some("") think_end=Some("") think_start_tokens_len=1 think_end_tokens_len=1 +2026-09-21T06:52:03.733249Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-09-21T06:52:04.058740Z INFO mlxcel::server::startup: Warming up model... +2026-09-21T06:52:04.059259Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-09-21T06:52:04.423155Z INFO mlxcel::server::model_provider::model_worker: Model qwen3.5-4b-4bit loaded in 0.364s (resident after load: 0.00 GB) worker_model_id=qwen3.5-4b-4bit load_seconds=0.363838214 active_bytes=0 peak_bytes=0 cache_bytes=0 limit_bytes=124128085401 +2026-09-21T06:52:04.423236Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=1, max_queue_depth=32, prefill_chunk_size=512, max_batch_prefill=4, decode_storage=auto, speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=4, block_size_from=operator override, explicit_kind=true)) +2026-09-21T06:52:04.423669Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 3062487 blocks (32 layers, 32-token blocks) +2026-09-21T06:52:04.423683Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 256 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-09-21T06:52:04.423705Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=32 kv_cache_mode_total_layers=32 +2026-09-21T06:52:06.135971Z INFO mlxcel::models::speculative_exactness: MTP exactness probe passed: verify block is byte-identical to the single-token chain block_size=4 +2026-09-21T06:52:06.136004Z INFO mlxcel::server::batch::speculative_burst: Lazy-loading drafter from /home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash (kind=Some(Dflash)) +2026-09-21T06:52:06.136419Z INFO mlxcel::server::batch::speculative_burst: Drafter loaded (kind=dflash, 0 ms) +2026-09-21T06:52:06.137419Z INFO mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-09-21T06:52:06.161483Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=4 rounds=0 proposed_tokens=0 accepted_tokens=0 acceptance_rate=0.0 emitted_per_verify=0.0 zero_accept_rounds=0 partial_accept_rounds=0 full_accept_rounds=0 prefill_verify_ms=0.955125 first_bonus_ms=24.02868 first_hidden_ms=0.010192 bind_reset_ms=0.037919 draft_ms=0.0 verify_ms=0.0 target_argmax_sync_ms=0.0 logprobs_ms=0.0 walk_ms=0.0 hidden_concat_ms=0.0 rollback_ms=0.0 decode_ms=0.0 +2026-09-21T06:52:06.161654Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=1 generated_tokens=1 burst_ms=1737 +2026-09-21T06:52:06.161661Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-0 burst_wall_ms=1737.72296 burst_active_ms=1737.72296 slices=1 tokens_generated=1 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:52:06.161742Z INFO mlxcel::server::startup: Warmup complete +2026-09-21T06:52:06.161838Z INFO mlxcel::server::startup: model_type=Qwen35VLM: enabling native video_url content block support +2026-09-21T06:52:06.163303Z WARN mlxcel::server::startup: The loaded model accepts video input, but `ffmpeg` and `ffprobe` are not both on PATH, so every video request will be refused. The check is cached for the life of the process: restart the server after installing ffmpeg. +2026-09-21T06:52:06.164027Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18797 +2026-09-21T06:52:06.164034Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-09-21T06:52:06.164037Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-09-21T06:52:06.164038Z INFO mlxcel::server::startup: Endpoints: +2026-09-21T06:52:06.164039Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-09-21T06:52:06.164041Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-09-21T06:52:06.164042Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-09-21T06:52:06.164043Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-09-21T06:52:06.164044Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-09-21T06:52:06.164046Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-09-21T06:52:06.164047Z INFO mlxcel::server::startup: GET /props - Server properties +2026-09-21T06:52:06.164048Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-09-21T06:52:06.164050Z INFO mlxcel::server::startup: GET /health - Health check +2026-09-21T06:52:12.474036Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=4 rounds=79 proposed_tokens=237 accepted_tokens=121 acceptance_rate=0.510548523206751 emitted_per_verify=2.518987341772152 zero_accept_rounds=22 partial_accept_rounds=34 full_accept_rounds=23 prefill_verify_ms=5.130228 first_bonus_ms=594.76559 first_hidden_ms=0.011312 bind_reset_ms=0.036704 draft_ms=2124.3539459999993 verify_ms=189.45603799999998 target_argmax_sync_ms=2443.49924 logprobs_ms=0.020736 walk_ms=0.060672 hidden_concat_ms=0.13118199999999997 rollback_ms=40.66935600000001 decode_ms=4805.575708 +2026-09-21T06:52:12.474648Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=5406 +2026-09-21T06:52:12.474687Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-1 burst_wall_ms=5406.38198 burst_active_ms=5406.38198 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:52:15.462256Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=4 rounds=79 proposed_tokens=237 accepted_tokens=121 acceptance_rate=0.510548523206751 emitted_per_verify=2.518987341772152 zero_accept_rounds=22 partial_accept_rounds=34 full_accept_rounds=23 prefill_verify_ms=2.842143 first_bonus_ms=117.253052 first_hidden_ms=0.013231999999999999 bind_reset_ms=0.06385500000000001 draft_ms=254.212343 verify_ms=182.86321799999996 target_argmax_sync_ms=2377.8071260000006 logprobs_ms=0.016448 walk_ms=0.04862399999999999 hidden_concat_ms=0.129262 rollback_ms=39.6562 decode_ms=2862.547173 +2026-09-21T06:52:15.462577Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=2983 +2026-09-21T06:52:15.462631Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-2 burst_wall_ms=2983.195706 burst_active_ms=2983.195706 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:52:18.454578Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=4 rounds=79 proposed_tokens=237 accepted_tokens=121 acceptance_rate=0.510548523206751 emitted_per_verify=2.518987341772152 zero_accept_rounds=22 partial_accept_rounds=34 full_accept_rounds=23 prefill_verify_ms=2.689537 first_bonus_ms=119.110945 first_hidden_ms=0.009184000000000001 bind_reset_ms=0.04952 draft_ms=256.80983499999996 verify_ms=186.41631499999997 target_argmax_sync_ms=2376.9477769999994 logprobs_ms=0.014991999999999997 walk_ms=0.04660800000000001 hidden_concat_ms=0.13274899999999998 rollback_ms=39.49270500000001 decode_ms=2866.255389 +2026-09-21T06:52:18.454886Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=2988 +2026-09-21T06:52:18.454949Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-3 burst_wall_ms=2988.501258 burst_active_ms=2988.501258 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:52:21.498071Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=4 rounds=79 proposed_tokens=237 accepted_tokens=121 acceptance_rate=0.510548523206751 emitted_per_verify=2.518987341772152 zero_accept_rounds=22 partial_accept_rounds=34 full_accept_rounds=23 prefill_verify_ms=2.83792 first_bonus_ms=118.712384 first_hidden_ms=0.016062999999999997 bind_reset_ms=0.09215899999999999 draft_ms=284.26618 verify_ms=199.6296499999999 target_argmax_sync_ms=2369.039821000001 logprobs_ms=0.01864000000000001 walk_ms=0.05428800000000001 hidden_concat_ms=0.1850069999999999 rollback_ms=51.50579000000002 decode_ms=2911.791648 +2026-09-21T06:52:21.498576Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=3034 +2026-09-21T06:52:21.498647Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-4 burst_wall_ms=3034.152918 burst_active_ms=3034.152918 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w6.log b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w6.log new file mode 100644 index 000000000..b43a063b1 --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w6.log @@ -0,0 +1,51 @@ +2026-09-21T06:53:22.863529Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-09-21T06:53:22.863572Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-09-21T06:53:22.863706Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=1 kv_unified=false n_parallel=4 prefill_chunk_size=512 max_kv_size=None +2026-09-21T06:53:22.863782Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-09-21T06:53:22.863785Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-09-21T06:53:22.863787Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-09-21T06:53:23.204404Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-09-21T06:53:23.204438Z INFO mlxcel::server::startup: Tokenizer recognizes a think marker pair; defaulting chat_template kwarg `enable_thinking=true` (upstream PR #1114) think_start=Some("") think_end=Some("") think_start_tokens_len=1 think_end_tokens_len=1 +2026-09-21T06:53:23.205006Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-09-21T06:53:23.552049Z INFO mlxcel::server::startup: Warming up model... +2026-09-21T06:53:23.552349Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-09-21T06:53:23.887427Z INFO mlxcel::server::model_provider::model_worker: Model qwen3.5-4b-4bit loaded in 0.335s (resident after load: 0.00 GB) worker_model_id=qwen3.5-4b-4bit load_seconds=0.33504395 active_bytes=0 peak_bytes=0 cache_bytes=0 limit_bytes=124128085401 +2026-09-21T06:53:23.887501Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=1, max_queue_depth=32, prefill_chunk_size=512, max_batch_prefill=4, decode_storage=auto, speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=6, block_size_from=operator override, explicit_kind=true)) +2026-09-21T06:53:23.887896Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 3062487 blocks (32 layers, 32-token blocks) +2026-09-21T06:53:23.887909Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 256 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-09-21T06:53:23.887932Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=32 kv_cache_mode_total_layers=32 +2026-09-21T06:53:25.733561Z INFO mlxcel::models::speculative_exactness: MTP exactness probe passed: verify block is byte-identical to the single-token chain block_size=6 +2026-09-21T06:53:25.733592Z INFO mlxcel::server::batch::speculative_burst: Lazy-loading drafter from /home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash (kind=Some(Dflash)) +2026-09-21T06:53:25.734003Z INFO mlxcel::server::batch::speculative_burst: Drafter loaded (kind=dflash, 0 ms) +2026-09-21T06:53:25.735087Z INFO mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-09-21T06:53:25.759353Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=6 rounds=0 proposed_tokens=0 accepted_tokens=0 acceptance_rate=0.0 emitted_per_verify=0.0 zero_accept_rounds=0 partial_accept_rounds=0 full_accept_rounds=0 prefill_verify_ms=1.038182 first_bonus_ms=24.216985 first_hidden_ms=0.008544 bind_reset_ms=0.051711 draft_ms=0.0 verify_ms=0.0 target_argmax_sync_ms=0.0 logprobs_ms=0.0 walk_ms=0.0 hidden_concat_ms=0.0 rollback_ms=0.0 decode_ms=0.0 +2026-09-21T06:53:25.759514Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=1 generated_tokens=1 burst_ms=1870 +2026-09-21T06:53:25.759520Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-0 burst_wall_ms=1871.377592 burst_active_ms=1871.377592 slices=1 tokens_generated=1 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:53:25.759595Z INFO mlxcel::server::startup: Warmup complete +2026-09-21T06:53:25.759690Z INFO mlxcel::server::startup: model_type=Qwen35VLM: enabling native video_url content block support +2026-09-21T06:53:25.760605Z WARN mlxcel::server::startup: The loaded model accepts video input, but `ffmpeg` and `ffprobe` are not both on PATH, so every video request will be refused. The check is cached for the life of the process: restart the server after installing ffmpeg. +2026-09-21T06:53:25.761379Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18797 +2026-09-21T06:53:25.761386Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-09-21T06:53:25.761389Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-09-21T06:53:25.761391Z INFO mlxcel::server::startup: Endpoints: +2026-09-21T06:53:25.761392Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-09-21T06:53:25.761393Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-09-21T06:53:25.761395Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-09-21T06:53:25.761396Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-09-21T06:53:25.761397Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-09-21T06:53:25.761399Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-09-21T06:53:25.761400Z INFO mlxcel::server::startup: GET /props - Server properties +2026-09-21T06:53:25.761401Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-09-21T06:53:25.761403Z INFO mlxcel::server::startup: GET /health - Health check +2026-09-21T06:53:32.108541Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=6 rounds=63 proposed_tokens=315 accepted_tokens=136 acceptance_rate=0.43174603174603177 emitted_per_verify=3.1587301587301586 zero_accept_rounds=16 partial_accept_rounds=36 full_accept_rounds=11 prefill_verify_ms=5.180510999999999 first_bonus_ms=579.306353 first_hidden_ms=0.010688 bind_reset_ms=0.031936 draft_ms=1730.3702990000002 verify_ms=166.91050499999994 target_argmax_sync_ms=3054.5889279999997 logprobs_ms=0.007711999999999999 walk_ms=0.06580799999999999 hidden_concat_ms=0.09932699999999997 rollback_ms=44.36057299999999 decode_ms=5001.170961 +2026-09-21T06:53:32.109124Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=5586 +2026-09-21T06:53:32.109164Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-1 burst_wall_ms=5586.429642 burst_active_ms=5586.429642 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:53:35.650365Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=6 rounds=63 proposed_tokens=315 accepted_tokens=136 acceptance_rate=0.43174603174603177 emitted_per_verify=3.1587301587301586 zero_accept_rounds=16 partial_accept_rounds=36 full_accept_rounds=11 prefill_verify_ms=3.0737949999999996 first_bonus_ms=121.402173 first_hidden_ms=0.014096000000000001 bind_reset_ms=0.10275100000000001 draft_ms=213.79313399999998 verify_ms=179.78492499999996 target_argmax_sync_ms=2961.4058910000003 logprobs_ms=0.010416000000000002 walk_ms=0.063184 hidden_concat_ms=0.11918399999999998 rollback_ms=53.31381000000001 decode_ms=3412.952307 +2026-09-21T06:53:35.650864Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=3538 +2026-09-21T06:53:35.650947Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-2 burst_wall_ms=3538.215132 burst_active_ms=3538.215132 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:53:39.221287Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=6 rounds=63 proposed_tokens=315 accepted_tokens=136 acceptance_rate=0.43174603174603177 emitted_per_verify=3.1587301587301586 zero_accept_rounds=16 partial_accept_rounds=36 full_accept_rounds=11 prefill_verify_ms=4.414535 first_bonus_ms=118.5672 first_hidden_ms=0.016784 bind_reset_ms=0.140318 draft_ms=232.5790429999999 verify_ms=194.944277 target_argmax_sync_ms=2945.068726 logprobs_ms=0.017904000000000003 walk_ms=0.066784 hidden_concat_ms=0.161774 rollback_ms=67.628989 decode_ms=3444.917807 +2026-09-21T06:53:39.221604Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=3568 +2026-09-21T06:53:39.221651Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-3 burst_wall_ms=3568.454576 burst_active_ms=3568.454576 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 +2026-09-21T06:53:42.778747Z INFO mlxcel::server::batch::dflash_target: DFlash diagnostics block_size=6 rounds=63 proposed_tokens=315 accepted_tokens=136 acceptance_rate=0.43174603174603177 emitted_per_verify=3.1587301587301586 zero_accept_rounds=16 partial_accept_rounds=36 full_accept_rounds=11 prefill_verify_ms=3.52928 first_bonus_ms=120.23686000000001 first_hidden_ms=0.010304 bind_reset_ms=0.055071 draft_ms=211.58493600000003 verify_ms=180.489591 target_argmax_sync_ms=2974.4406919999997 logprobs_ms=0.013600000000000006 walk_ms=0.059439999999999986 hidden_concat_ms=0.13692499999999994 rollback_ms=60.137314 decode_ms=3430.7773519999996 +2026-09-21T06:53:42.779188Z INFO mlxcel::server::batch::speculative_burst: Speculative burst completed prompt_tokens=158 generated_tokens=200 burst_ms=3555 +2026-09-21T06:53:42.779258Z INFO mlxcel::server::batch::scheduler::speculative_finalize: speculative B=1 burst finalized (burst_wall_ms is the max single-tick HOL stall on concurrent rows) seq_id=seq-4 burst_wall_ms=3555.141806 burst_active_ms=3555.141806 slices=1 tokens_generated=200 rounds=0 accepted_draft_tokens=0 hol_waiters=0 diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/summary.txt b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/summary.txt new file mode 100644 index 000000000..3e43901aa --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/summary.txt @@ -0,0 +1,25 @@ +## affine (qwen3.5-4b-4bit + qwen3.5-4b-dflash), main after #1939 + +Classic bracket: opening 56.93 to 57.11, closing 56.71 to 56.75 tok/s. They DO NOT OVERLAP: the session drifted, treat every middle arm as suspect. + +| arm | n | e2e tok/s mean (min to max) | vs classic | separates from classic | acceptance | emitted per verify | round device sync ms | NVRM delta | resolved block_size | +| --- | ---: | ---: | ---: | --- | ---: | ---: | ---: | ---: | --- | +| classic-open | 3 | 57.00 (56.93 to 57.11) | 1.00x | | | | | +0 | classic | +| w2 | 3 | 66.74 (66.56 to 67.09) | 1.17x | yes, above every classic run | 0.754 | 1.75 | 19.9 | +0 | 2 | +| w3 | 3 | 67.81 (67.76 to 67.89) | 1.19x | yes, above every classic run | 0.599 | 2.19 | 24.8 | +0 | 3 | +| w4 | 3 | 66.50 (65.71 to 66.96) | 1.17x | yes, above every classic run | 0.511 | 2.52 | 30.3 | +0 | 4 | +| w6 | 3 | 56.23 (56.01 to 56.47) | 0.99x | yes, below every classic run | 0.432 | 3.16 | 47.4 | +0 | 6 | +| classic-close | 3 | 56.73 (56.71 to 56.75) | 1.00x | | | | | +0 | classic | + +Greedy identity: the two classic arms are byte-identical to each other. +Greedy identity: classic-close == classic +Greedy identity: classic-open == classic +Greedy identity: w2 != classic +Greedy identity: w3 != classic +Greedy identity: w4 != classic +Greedy identity: w6 != classic + +w2: speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=2, block_size_from=operator override, explicit_kind=true) +w3: speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=3, block_size_from=operator override, explicit_kind=true) +w4: speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=4, block_size_from=operator override, explicit_kind=true) +w6: speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=6, block_size_from=operator override, explicit_kind=true) diff --git a/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/widths.jsonl b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/widths.jsonl new file mode 100644 index 000000000..f6674d8e4 --- /dev/null +++ b/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/widths.jsonl @@ -0,0 +1,6 @@ +{"arm": "classic-open", "width": "classic", "cmd": ["/tmp/mlxcel1797.izrEDT/mlxcel1797-server", "-m", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-4bit", "--port", "18797", "--ignore-eos", "--max-batch-size", "1"], "arm_env": {}, "load1_before": 0.3740234375, "ci_job_running": false, "gate_wait_s": 0, "driver_wait_s": 0, "mem_wait_s": 0, "mem_available_gib_before": 116.2, "nvrm_window_before": 0, "nvrm_total_before": 0, "log": "/home/inureyes/Development/mlxcel-wt-main/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.classic-open.log", "startup_s": 3.0, "model_id": "qwen3.5-4b-4bit", "prompt_tokens": 158, "warmup_s": 3.978, "warmup_chars": 823, "runs": [{"wall_s": 3.502, "e2e_tok_s": 57.113, "chunks": 200, "chars": 823, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n *,\n logger: logging.Logger = None,\n max_attempts: int = None,\n max_delay: float = None,\n base_delay: float = None,\n jitter: bool = None,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n This decorator is useful for retrying HTTP calls that might fail due to\n transient network issues. It uses exponential backoff with jitter to\n avoid thundering"}, {"wall_s": 3.512, "e2e_tok_s": 56.941, "chunks": 200, "chars": 823, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n *,\n logger: logging.Logger = None,\n max_attempts: int = None,\n max_delay: float = None,\n base_delay: float = None,\n jitter: bool = None,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n This decorator is useful for retrying HTTP calls that might fail due to\n transient network issues. It uses exponential backoff with jitter to\n avoid thundering"}, {"wall_s": 3.513, "e2e_tok_s": 56.934, "chunks": 200, "chars": 823, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n *,\n logger: logging.Logger = None,\n max_attempts: int = None,\n max_delay: float = None,\n base_delay: float = None,\n jitter: bool = None,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n This decorator is useful for retrying HTTP calls that might fail due to\n transient network issues. It uses exponential backoff with jitter to\n avoid thundering"}], "speculative_line": null, "diagnostics": [], "diagnostics_block_sizes": [], "diagnostics_rounds": [], "declined_to_classic": false, "speculative_ran": false, "nvrm_total_after": 0, "nvrm_delta": 0, "nvrm_window_after": 0} +{"arm": "w2", "width": "2", "cmd": ["/tmp/mlxcel1797.izrEDT/mlxcel1797-server", "-m", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-4bit", "--port", "18797", "--ignore-eos", "--max-batch-size", "1", "--draft-model", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash", "--draft-kind", "dflash", "--draft-block-size", "2"], "arm_env": {}, "load1_before": 0.2333984375, "ci_job_running": false, "gate_wait_s": 0, "driver_wait_s": 0, "mem_wait_s": 0, "mem_available_gib_before": 116.2, "nvrm_window_before": 0, "nvrm_total_before": 0, "log": "/home/inureyes/Development/mlxcel-wt-main/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w2.log", "startup_s": 3.0, "model_id": "qwen3.5-4b-4bit", "prompt_tokens": 158, "warmup_s": 6.061, "warmup_chars": 820, "runs": [{"wall_s": 2.981, "e2e_tok_s": 67.088, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}, {"wall_s": 3.004, "e2e_tok_s": 66.57, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}, {"wall_s": 3.005, "e2e_tok_s": 66.56, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}], "speculative_line": "speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=2, block_size_from=operator override, explicit_kind=true)", "diagnostics": [{"block_size": 2, "rounds": 0, "proposed_tokens": 0, "accepted_tokens": 0, "acceptance_rate": 0.0, "emitted_per_verify": 0.0, "draft_ms": 0.0, "verify_ms": 0.0, "target_argmax_sync_ms": 0.0, "decode_ms": 0.0}, {"block_size": 2, "rounds": 114, "proposed_tokens": 114, "accepted_tokens": 86, "acceptance_rate": 0.7543859649122807, "emitted_per_verify": 1.7456140350877194, "draft_ms": 2817.84715, "verify_ms": 243.60834000000003, "target_argmax_sync_ms": 2372.036596, "decode_ms": 5465.96001}, {"block_size": 2, "rounds": 114, "proposed_tokens": 114, "accepted_tokens": 86, "acceptance_rate": 0.7543859649122807, "emitted_per_verify": 1.7456140350877194, "draft_ms": 366.5048819999998, "verify_ms": 230.36967999999996, "target_argmax_sync_ms": 2223.9670730000003, "decode_ms": 2853.268224}, {"block_size": 2, "rounds": 114, "proposed_tokens": 114, "accepted_tokens": 86, "acceptance_rate": 0.7543859649122807, "emitted_per_verify": 1.7456140350877194, "draft_ms": 367.421775, "verify_ms": 241.678898, "target_argmax_sync_ms": 2232.3261150000003, "decode_ms": 2875.0160830000004}, {"block_size": 2, "rounds": 114, "proposed_tokens": 114, "accepted_tokens": 86, "acceptance_rate": 0.7543859649122807, "emitted_per_verify": 1.7456140350877194, "draft_ms": 375.9427449999999, "verify_ms": 243.70482399999995, "target_argmax_sync_ms": 2223.5196660000006, "decode_ms": 2877.03246}], "diagnostics_block_sizes": [2], "diagnostics_rounds": [0, 114], "declined_to_classic": false, "speculative_ran": true, "nvrm_total_after": 0, "nvrm_delta": 0, "nvrm_window_after": 0} +{"arm": "w3", "width": "3", "cmd": ["/tmp/mlxcel1797.izrEDT/mlxcel1797-server", "-m", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-4bit", "--port", "18797", "--ignore-eos", "--max-batch-size", "1", "--draft-model", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash", "--draft-kind", "dflash", "--draft-block-size", "3"], "arm_env": {}, "load1_before": 0.19873046875, "ci_job_running": false, "gate_wait_s": 0, "driver_wait_s": 0, "mem_wait_s": 0, "mem_available_gib_before": 116.2, "nvrm_window_before": 0, "nvrm_total_before": 0, "log": "/home/inureyes/Development/mlxcel-wt-main/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w3.log", "startup_s": 4.0, "model_id": "qwen3.5-4b-4bit", "prompt_tokens": 158, "warmup_s": 5.524, "warmup_chars": 820, "runs": [{"wall_s": 2.946, "e2e_tok_s": 67.89, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}, {"wall_s": 2.95, "e2e_tok_s": 67.792, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}, {"wall_s": 2.952, "e2e_tok_s": 67.759, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}], "speculative_line": "speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=3, block_size_from=operator override, explicit_kind=true)", "diagnostics": [{"block_size": 3, "rounds": 0, "proposed_tokens": 0, "accepted_tokens": 0, "acceptance_rate": 0.0, "emitted_per_verify": 0.0, "draft_ms": 0.0, "verify_ms": 0.0, "target_argmax_sync_ms": 0.0, "decode_ms": 0.0}, {"block_size": 3, "rounds": 91, "proposed_tokens": 182, "accepted_tokens": 109, "acceptance_rate": 0.5989010989010989, "emitted_per_verify": 2.1868131868131866, "draft_ms": 2381.873432, "verify_ms": 206.602133, "target_argmax_sync_ms": 2291.583825999999, "decode_ms": 4916.782923999999}, {"block_size": 3, "rounds": 91, "proposed_tokens": 182, "accepted_tokens": 109, "acceptance_rate": 0.5989010989010989, "emitted_per_verify": 2.1868131868131866, "draft_ms": 317.46032899999994, "verify_ms": 218.80341699999997, "target_argmax_sync_ms": 2238.6714439999996, "decode_ms": 2820.244314}, {"block_size": 3, "rounds": 91, "proposed_tokens": 182, "accepted_tokens": 109, "acceptance_rate": 0.5989010989010989, "emitted_per_verify": 2.1868131868131866, "draft_ms": 318.965685, "verify_ms": 215.09529099999992, "target_argmax_sync_ms": 2250.130131, "decode_ms": 2827.595127}, {"block_size": 3, "rounds": 91, "proposed_tokens": 182, "accepted_tokens": 109, "acceptance_rate": 0.5989010989010989, "emitted_per_verify": 2.1868131868131866, "draft_ms": 320.41396199999997, "verify_ms": 222.00846599999994, "target_argmax_sync_ms": 2241.846096, "decode_ms": 2827.3205580000003}], "diagnostics_block_sizes": [3], "diagnostics_rounds": [0, 91], "declined_to_classic": false, "speculative_ran": true, "nvrm_total_after": 0, "nvrm_delta": 0, "nvrm_window_after": 0} +{"arm": "w4", "width": "4", "cmd": ["/tmp/mlxcel1797.izrEDT/mlxcel1797-server", "-m", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-4bit", "--port", "18797", "--ignore-eos", "--max-batch-size", "1", "--draft-model", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash", "--draft-kind", "dflash", "--draft-block-size", "4"], "arm_env": {}, "load1_before": 0.21142578125, "ci_job_running": false, "gate_wait_s": 0, "driver_wait_s": 0, "mem_wait_s": 0, "mem_available_gib_before": 116.2, "nvrm_window_before": 0, "nvrm_total_before": 0, "log": "/home/inureyes/Development/mlxcel-wt-main/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w4.log", "startup_s": 4.0, "model_id": "qwen3.5-4b-4bit", "prompt_tokens": 158, "warmup_s": 5.411, "warmup_chars": 820, "runs": [{"wall_s": 2.987, "e2e_tok_s": 66.958, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}, {"wall_s": 2.992, "e2e_tok_s": 66.847, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}, {"wall_s": 3.044, "e2e_tok_s": 65.708, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}], "speculative_line": "speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=4, block_size_from=operator override, explicit_kind=true)", "diagnostics": [{"block_size": 4, "rounds": 0, "proposed_tokens": 0, "accepted_tokens": 0, "acceptance_rate": 0.0, "emitted_per_verify": 0.0, "draft_ms": 0.0, "verify_ms": 0.0, "target_argmax_sync_ms": 0.0, "decode_ms": 0.0}, {"block_size": 4, "rounds": 79, "proposed_tokens": 237, "accepted_tokens": 121, "acceptance_rate": 0.510548523206751, "emitted_per_verify": 2.518987341772152, "draft_ms": 2124.3539459999993, "verify_ms": 189.45603799999998, "target_argmax_sync_ms": 2443.49924, "decode_ms": 4805.575708}, {"block_size": 4, "rounds": 79, "proposed_tokens": 237, "accepted_tokens": 121, "acceptance_rate": 0.510548523206751, "emitted_per_verify": 2.518987341772152, "draft_ms": 254.212343, "verify_ms": 182.86321799999996, "target_argmax_sync_ms": 2377.8071260000006, "decode_ms": 2862.547173}, {"block_size": 4, "rounds": 79, "proposed_tokens": 237, "accepted_tokens": 121, "acceptance_rate": 0.510548523206751, "emitted_per_verify": 2.518987341772152, "draft_ms": 256.80983499999996, "verify_ms": 186.41631499999997, "target_argmax_sync_ms": 2376.9477769999994, "decode_ms": 2866.255389}, {"block_size": 4, "rounds": 79, "proposed_tokens": 237, "accepted_tokens": 121, "acceptance_rate": 0.510548523206751, "emitted_per_verify": 2.518987341772152, "draft_ms": 284.26618, "verify_ms": 199.6296499999999, "target_argmax_sync_ms": 2369.039821000001, "decode_ms": 2911.791648}], "diagnostics_block_sizes": [4], "diagnostics_rounds": [0, 79], "declined_to_classic": false, "speculative_ran": true, "nvrm_total_after": 0, "nvrm_delta": 0, "nvrm_window_after": 0} +{"arm": "w6", "width": "6", "cmd": ["/tmp/mlxcel1797.izrEDT/mlxcel1797-server", "-m", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-4bit", "--port", "18797", "--ignore-eos", "--max-batch-size", "1", "--draft-model", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash", "--draft-kind", "dflash", "--draft-block-size", "6"], "arm_env": {}, "load1_before": 0.189453125, "ci_job_running": false, "gate_wait_s": 0, "driver_wait_s": 0, "mem_wait_s": 0, "mem_available_gib_before": 116.2, "nvrm_window_before": 0, "nvrm_total_before": 0, "log": "/home/inureyes/Development/mlxcel-wt-main/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.w6.log", "startup_s": 4.0, "model_id": "qwen3.5-4b-4bit", "prompt_tokens": 158, "warmup_s": 5.589, "warmup_chars": 820, "runs": [{"wall_s": 3.542, "e2e_tok_s": 56.47, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}, {"wall_s": 3.571, "e2e_tok_s": 56.011, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}, {"wall_s": 3.558, "e2e_tok_s": 56.213, "chunks": 200, "chars": 820, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n logger: logging.Logger = None,\n **kwargs,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n Args:\n func: The function to retry.\n backoff_policy: The backoff policy to use.\n logger: The logger to use for logging.\n **kwargs: Additional arguments to pass to the function.\n\n Returns:\n The result of the function.\n \"\"\"\n if"}], "speculative_line": "speculative=dflash (drafter=/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash, block_size=6, block_size_from=operator override, explicit_kind=true)", "diagnostics": [{"block_size": 6, "rounds": 0, "proposed_tokens": 0, "accepted_tokens": 0, "acceptance_rate": 0.0, "emitted_per_verify": 0.0, "draft_ms": 0.0, "verify_ms": 0.0, "target_argmax_sync_ms": 0.0, "decode_ms": 0.0}, {"block_size": 6, "rounds": 63, "proposed_tokens": 315, "accepted_tokens": 136, "acceptance_rate": 0.43174603174603177, "emitted_per_verify": 3.1587301587301586, "draft_ms": 1730.3702990000002, "verify_ms": 166.91050499999994, "target_argmax_sync_ms": 3054.5889279999997, "decode_ms": 5001.170961}, {"block_size": 6, "rounds": 63, "proposed_tokens": 315, "accepted_tokens": 136, "acceptance_rate": 0.43174603174603177, "emitted_per_verify": 3.1587301587301586, "draft_ms": 213.79313399999998, "verify_ms": 179.78492499999996, "target_argmax_sync_ms": 2961.4058910000003, "decode_ms": 3412.952307}, {"block_size": 6, "rounds": 63, "proposed_tokens": 315, "accepted_tokens": 136, "acceptance_rate": 0.43174603174603177, "emitted_per_verify": 3.1587301587301586, "draft_ms": 232.5790429999999, "verify_ms": 194.944277, "target_argmax_sync_ms": 2945.068726, "decode_ms": 3444.917807}, {"block_size": 6, "rounds": 63, "proposed_tokens": 315, "accepted_tokens": 136, "acceptance_rate": 0.43174603174603177, "emitted_per_verify": 3.1587301587301586, "draft_ms": 211.58493600000003, "verify_ms": 180.489591, "target_argmax_sync_ms": 2974.4406919999997, "decode_ms": 3430.7773519999996}], "diagnostics_block_sizes": [6], "diagnostics_rounds": [0, 63], "declined_to_classic": false, "speculative_ran": true, "nvrm_total_after": 0, "nvrm_delta": 0, "nvrm_window_after": 0} +{"arm": "classic-close", "width": "classic", "cmd": ["/tmp/mlxcel1797.izrEDT/mlxcel1797-server", "-m", "/home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-4bit", "--port", "18797", "--ignore-eos", "--max-batch-size", "1"], "arm_env": {}, "load1_before": 0.1796875, "ci_job_running": false, "gate_wait_s": 0, "driver_wait_s": 0, "mem_wait_s": 0, "mem_available_gib_before": 116.2, "nvrm_window_before": 0, "nvrm_total_before": 0, "log": "/home/inureyes/Development/mlxcel-wt-main/docs/benchmark_results/data/draft-block-width-post-1939-gb10-2026-09-21/logs/server.classic-close.log", "startup_s": 3.0, "model_id": "qwen3.5-4b-4bit", "prompt_tokens": 158, "warmup_s": 4.014, "warmup_chars": 823, "runs": [{"wall_s": 3.527, "e2e_tok_s": 56.711, "chunks": 200, "chars": 823, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n *,\n logger: logging.Logger = None,\n max_attempts: int = None,\n max_delay: float = None,\n base_delay: float = None,\n jitter: bool = None,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n This decorator is useful for retrying HTTP calls that might fail due to\n transient network issues. It uses exponential backoff with jitter to\n avoid thundering"}, {"wall_s": 3.524, "e2e_tok_s": 56.754, "chunks": 200, "chars": 823, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n *,\n logger: logging.Logger = None,\n max_attempts: int = None,\n max_delay: float = None,\n base_delay: float = None,\n jitter: bool = None,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n This decorator is useful for retrying HTTP calls that might fail due to\n transient network issues. It uses exponential backoff with jitter to\n avoid thundering"}, {"wall_s": 3.526, "e2e_tok_s": 56.722, "chunks": 200, "chars": 823, "completion_tokens": 200, "text": " \"\"\"Generate a sequence of delays for retrying.\"\"\"\n delays = []\n for i in range(self.max_attempts):\n delay = self.base_delay * (2 ** i)\n delay = min(delay, self.max_delay)\n if self.jitter:\n delay = delay + random.uniform(0, delay)\n delays.append(delay)\n return delays\n\n\ndef retry_with_backoff(\n func,\n backoff_policy: BackoffPolicy = None,\n *,\n logger: logging.Logger = None,\n max_attempts: int = None,\n max_delay: float = None,\n base_delay: float = None,\n jitter: bool = None,\n):\n \"\"\"Retry a function with exponential backoff and jitter.\n\n This decorator is useful for retrying HTTP calls that might fail due to\n transient network issues. It uses exponential backoff with jitter to\n avoid thundering"}], "speculative_line": null, "diagnostics": [], "diagnostics_block_sizes": [], "diagnostics_rounds": [], "declined_to_classic": false, "speculative_ran": false, "nvrm_total_after": 0, "nvrm_delta": 0, "nvrm_window_after": 0} diff --git a/docs/benchmark_results/draft-block-width-post-1939-gb10-2026-09-21.md b/docs/benchmark_results/draft-block-width-post-1939-gb10-2026-09-21.md new file mode 100644 index 000000000..a05ea2570 --- /dev/null +++ b/docs/benchmark_results/draft-block-width-post-1939-gb10-2026-09-21.md @@ -0,0 +1,72 @@ +# Draft block width after the #1939 prefill fix (GB10, 2026-09-21) + +Issue #1797 seeded a measured default of 4 for `(12, 1, Affine)` on a DFlash drafter, from a sweep taken before PR #1939 fixed the DFlash burst's prompt prefill. That fix changes the KV and gated-delta state every round starts from, which changes acceptance, which changes throughput per width. This re-measures the curve on current main so the shipped default rests on the tree that ships it. + +Nothing here is differenced against the #1797 table. That sweep ran on a different binary and a different tree; it is the prior conclusion to confirm or overturn, not a set of rows to subtract from. The reference is this session's own classic brackets. + +**The ordering moved. Width 3 is faster than the shipped default of 4, by more than this session's drift.** + +## Host + +``` +host: spark-101 +kernel: 7.0.0-1019-nvidia +arch: aarch64 +nvidia_driver: 580.178.04 +gpu: NVIDIA GB10, 12.1 +MLX_CUDA_ARCHITECTURES: 121 +cuda_toolkit: 13.0 +mlx_pin: 81ba1c6a0e50a9268b931579c2d4f1158b9aab5a +rustc_used_for_build: rustc 1.97.1 (8bab26f4f 2026-07-14) +git_commit: 6f9a0982d7854b653c7b354fe0f1f0b26af41a7f +git_dirty: no +mem_total_gib: 121.7 +binary_sha256: f2ba16883e8aed7aecf57212f5a81f4516eaeddcc618252bc5a6f464f015248c +``` + +The binary is built from a detached worktree at `origin/main`, not from the issue #1935 branch whose exactness gate declines this burst. That distinction is not theoretical: the first attempt at this sweep copied a stale artifact out of the shared target directory and measured the gated branch, which the harness caught and reported as `DECLINED TO CLASSIC` on the first width rather than as a number. The binary used here was verified to carry none of that branch's gate before it ran. + +## Method + +`sweep_server_widths.py` from the #1797 harness, unchanged, on `models/mlx/qwen3.5-4b-4bit` with `models/mlx/qwen3.5-4b-dflash`. Widths 2, 3, 4 and 6, n = 3 per arm after a discarded warm-up, with the harness's own classic arm first and last as the drift control. One server per arm, `--ignore-eos --max-batch-size 1`, the harness's fixed 158-token prompt, 200 tokens per request, `MLX_ENABLE_TF32=1`. + +The host gate is the #1820 one the harness imports: a sustained-quiet CPU predicate matching on `/proc//comm` with stopped processes dropped, a foreign-model check, a memory floor, and the cumulative `NV_ERR_NO_MEMORY` trip wire. It held this sweep for several minutes while CI ran on the same host and released it when the runner's worker finished; no arm was timed against a compiler or a CI job. The driver count was 0 before and 0 after, per arm. + +## Result + +Classic bracket: opening 56.93 to 57.11, closing 56.71 to 56.75 tok/s. They DO NOT OVERLAP: the session drifted, treat every middle arm as suspect. + +| arm | n | e2e tok/s mean (min to max) | vs classic | separates from classic | acceptance | emitted per verify | round device sync ms | NVRM delta | resolved block_size | +| --- | ---: | ---: | ---: | --- | ---: | ---: | ---: | ---: | --- | +| classic-open | 3 | 57.00 (56.93 to 57.11) | 1.00x | | | | | +0 | classic | +| w2 | 3 | 66.74 (66.56 to 67.09) | 1.17x | yes, above every classic run | 0.754 | 1.75 | 19.9 | +0 | 2 | +| w3 | 3 | 67.81 (67.76 to 67.89) | 1.19x | yes, above every classic run | 0.599 | 2.19 | 24.8 | +0 | 3 | +| w4 | 3 | 66.50 (65.71 to 66.96) | 1.17x | yes, above every classic run | 0.511 | 2.52 | 30.3 | +0 | 4 | +| w6 | 3 | 56.23 (56.01 to 56.47) | 0.99x | yes, below every classic run | 0.432 | 3.16 | 47.4 | +0 | 6 | +| classic-close | 3 | 56.73 (56.71 to 56.75) | 1.00x | | | | | +0 | classic | + +## Reading + +**The brackets do not overlap, and the finding survives it anyway.** The opening classic arm ran at 56.93 to 57.11 and the closing one at 56.71 to 56.75, so the session drifted downward by about 0.3 tok/s, half a percent. That is the resolution floor for everything between them, and the harness says so rather than printing a clean table over a dirty run. Two numbers describe that drift and they should not be added together: the brackets' full spread, 56.71 to 57.11, is 0.40 tok/s, and the shift between their means, 57.00 to 56.73, is 0.27. Width 3's slowest run (67.76) is 0.80 tok/s above width 4's fastest (66.96), which is twice the full spread. Width 4 ran after width 3, so the drift works against width 3 rather than for it, and correcting width 4 upward by the whole mean shift still leaves the two ranges disjoint. + +**Width 3 is the best of the four, width 2 and width 4 do not separate from each other, and width 6 is a loss.** Width 3 at 1.19x classic separates from both its neighbours: its range clears width 2's fastest run (67.09) and width 4's (66.96). Width 2 and width 4 overlap (66.56 to 67.09 against 65.71 to 66.96) and this sweep does not order them. Width 6 at 0.99x is below every classic run: the verify block has stopped paying for itself there. + +**Acceptance is what moves, and it moves monotonically while the block widens.** 0.754 at width 2, 0.599 at 3, 0.511 at 4, 0.432 at 6, against emitted-per-verify of 1.75, 2.19, 2.52 and 3.16 and a per-round device-sync cost of 19.9, 24.8, 30.3 and 47.4 ms. The product of those is what a width is worth, and on this tree it peaks at 3 rather than at 4. The shipped default sits one step past the peak. + +**Every speculative arm's greedy text differs from classic**, at every width including 2 and 3, which is issue #1935 reproduced independently here: this is a `main` binary with no gate, and the harness's own identity check reports it. + +``` +Greedy identity: the two classic arms are byte-identical to each other. +Greedy identity: classic-close == classic +Greedy identity: classic-open == classic +Greedy identity: w2 != classic +Greedy identity: w3 != classic +Greedy identity: w4 != classic +Greedy identity: w6 != classic +``` + +## What follows + +The measured default for `(12, 1, Affine)` on a DFlash drafter should be re-examined against 3. This record does not change it: the sweep covers one pairing on one host. The width is not moot for it either, which is worth saying because PR #1944 makes the issue #1935 gate decline this pairing on CUDA: `MLXCEL_MTP_ALLOW_INEXACT=1` is a documented escape, and an operator who sets it gets exactly this curve, so the default still decides what they run at. On the evidence here 4 is not the right number for this pairing on this host and 3 is the better one. The finding is filed so that rests on a measurement of the tree that ships rather than on a pre-#1939 one. + +Data: `data/draft-block-width-post-1939-gb10-2026-09-21/`.