Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
date: 2026-09-21T15:47:10+09:00
host: spark-101
kernel: 7.0.0-1019-nvidia
arch: aarch64
nvidia_driver: 580.178.04
gpu: NVIDIA GB10, 12.1
MLX_CUDA_ARCHITECTURES: 121
cuda_toolkit: 13.0
mlx_pin: 81ba1c6a0e50a9268b931579c2d4f1158b9aab5a
rustc_on_path: rustc 1.97.1 (8bab26f4f 2026-07-14)
rustc_used_for_build: rustc 1.97.1 (8bab26f4f 2026-07-14)
git_commit: 6f9a0982d7854b653c7b354fe0f1f0b26af41a7f
git_dirty: no
uptime: up 23 hours, 52 minutes
mem_total_gib: 121.7
binary: /tmp/mlxcel1797.izrEDT/mlxcel1797-server
binary_sha256: f2ba16883e8aed7aecf57212f5a81f4516eaeddcc618252bc5a6f464f015248c
binary_mtime: 2026-09-21T15:47:10+09:00
model /home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-4bit: bytes=3061131800 model_type=qwen3_5 quantization=affine
model /home/inureyes/Development/mlxcel-wt-main/models/mlx/qwen3.5-4b-dflash: bytes=1074861667 model_type=qwen3 quantization=none
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
2026-09-21T06:54:44.139710Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins
2026-09-21T06:54:44.139753Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0
2026-09-21T06:54:44.139888Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=1 kv_unified=false n_parallel=4 prefill_chunk_size=512 max_kv_size=None
2026-09-21T06:54:44.139945Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA)
2026-09-21T06:54:44.139948Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set)
2026-09-21T06:54:44.139951Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB
2026-09-21T06:54:44.489887Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"]
2026-09-21T06:54:44.489921Z INFO mlxcel::server::startup: Tokenizer recognizes a think marker pair; defaulting chat_template kwarg `enable_thinking=true` (upstream PR #1114) think_start=Some("<think>") think_end=Some("</think>") think_start_tokens_len=1 think_end_tokens_len=1
2026-09-21T06:54:44.490492Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256
2026-09-21T06:54:44.839550Z INFO mlxcel::server::startup: Warming up model...
2026-09-21T06:54:44.840004Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model...
2026-09-21T06:54:45.172015Z INFO mlxcel::server::model_provider::model_worker: Model qwen3.5-4b-4bit loaded in 0.332s (resident after load: 0.00 GB) worker_model_id=qwen3.5-4b-4bit load_seconds=0.331970264 active_bytes=0 peak_bytes=0 cache_bytes=0 limit_bytes=124128085401
2026-09-21T06:54:45.172088Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=1, max_queue_depth=32, prefill_chunk_size=512, max_batch_prefill=4, decode_storage=auto)
2026-09-21T06:54:45.172492Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 3062487 blocks (32 layers, 32-token blocks)
2026-09-21T06:54:45.172506Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 256 blocks per layer (fused decode serves a layer only while its rows fit one slab)
2026-09-21T06:54:45.172525Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=32 kv_cache_mode_total_layers=32
2026-09-21T06:54:45.173662Z INFO prefill{seq_id=seq-0 prompt_len=1 cached=0 start=0}: mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved
2026-09-21T06:54:46.060687Z INFO prefill{seq_id=seq-0 prompt_len=1 cached=0 start=0}: mlxcel::server::batch::scheduler::prefill: prompt-cache: request completed during prefill: cached=0/1 prompt tokens, total 887ms prompt_tokens=1 cached_tokens=0 generation_time_ms=887
2026-09-21T06:54:46.060744Z INFO mlxcel::server::startup: Warmup complete
2026-09-21T06:54:46.060838Z INFO mlxcel::server::startup: model_type=Qwen35VLM: enabling native video_url content block support
2026-09-21T06:54:46.061753Z WARN mlxcel::server::startup: The loaded model accepts video input, but `ffmpeg` and `ffprobe` are not both on PATH, so every video request will be refused. The check is cached for the life of the process: restart the server after installing ffmpeg.
2026-09-21T06:54:46.062528Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18797
2026-09-21T06:54:46.062545Z INFO mlxcel::server::startup: Detected 1 GPU(s)
2026-09-21T06:54:46.062548Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin)
2026-09-21T06:54:46.062550Z INFO mlxcel::server::startup: Endpoints:
2026-09-21T06:54:46.062551Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions
2026-09-21T06:54:46.062552Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions
2026-09-21T06:54:46.062554Z INFO mlxcel::server::startup: GET /v1/models - List models
2026-09-21T06:54:46.062555Z INFO mlxcel::server::startup: POST /completion - llama-server native completion
2026-09-21T06:54:46.062556Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text
2026-09-21T06:54:46.062557Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens
2026-09-21T06:54:46.062559Z INFO mlxcel::server::startup: GET /props - Server properties
2026-09-21T06:54:46.062560Z INFO mlxcel::server::startup: GET /slots - Slot status
2026-09-21T06:54:46.062561Z INFO mlxcel::server::startup: GET /health - Health check
2026-09-21T06:54:50.818671Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 4012ms prompt_tokens=158 cached_tokens=0 generation_time_ms=4012
2026-09-21T06:54:54.345526Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3524ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3524
2026-09-21T06:54:57.869074Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3521ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3521
2026-09-21T06:55:01.395604Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3524ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3524
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
2026-09-21T06:48:06.437539Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins
2026-09-21T06:48:06.437652Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0
2026-09-21T06:48:06.438105Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=1 kv_unified=false n_parallel=4 prefill_chunk_size=512 max_kv_size=None
2026-09-21T06:48:06.438178Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA)
2026-09-21T06:48:06.438181Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set)
2026-09-21T06:48:06.438183Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB
2026-09-21T06:48:06.763072Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"]
2026-09-21T06:48:06.763199Z INFO mlxcel::server::startup: Tokenizer recognizes a think marker pair; defaulting chat_template kwarg `enable_thinking=true` (upstream PR #1114) think_start=Some("<think>") think_end=Some("</think>") think_start_tokens_len=1 think_end_tokens_len=1
2026-09-21T06:48:06.764223Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256
2026-09-21T06:48:07.088497Z INFO mlxcel::server::startup: Warming up model...
2026-09-21T06:48:07.089259Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model...
2026-09-21T06:48:07.462941Z INFO mlxcel::server::model_provider::model_worker: Model qwen3.5-4b-4bit loaded in 0.374s (resident after load: 0.00 GB) worker_model_id=qwen3.5-4b-4bit load_seconds=0.373617994 active_bytes=0 peak_bytes=0 cache_bytes=0 limit_bytes=124128085401
2026-09-21T06:48:07.463048Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=1, max_queue_depth=32, prefill_chunk_size=512, max_batch_prefill=4, decode_storage=auto)
2026-09-21T06:48:07.463513Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 3062487 blocks (32 layers, 32-token blocks)
2026-09-21T06:48:07.463527Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 256 blocks per layer (fused decode serves a layer only while its rows fit one slab)
2026-09-21T06:48:07.463641Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=32 kv_cache_mode_total_layers=32
2026-09-21T06:48:07.465664Z INFO prefill{seq_id=seq-0 prompt_len=1 cached=0 start=0}: mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved
2026-09-21T06:48:08.392504Z INFO prefill{seq_id=seq-0 prompt_len=1 cached=0 start=0}: mlxcel::server::batch::scheduler::prefill: prompt-cache: request completed during prefill: cached=0/1 prompt tokens, total 928ms prompt_tokens=1 cached_tokens=0 generation_time_ms=928
2026-09-21T06:48:08.393130Z INFO mlxcel::server::startup: Warmup complete
2026-09-21T06:48:08.393303Z INFO mlxcel::server::startup: model_type=Qwen35VLM: enabling native video_url content block support
2026-09-21T06:48:08.394254Z WARN mlxcel::server::startup: The loaded model accepts video input, but `ffmpeg` and `ffprobe` are not both on PATH, so every video request will be refused. The check is cached for the life of the process: restart the server after installing ffmpeg.
2026-09-21T06:48:08.395806Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18797
2026-09-21T06:48:08.395823Z INFO mlxcel::server::startup: Detected 1 GPU(s)
2026-09-21T06:48:08.395827Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin)
2026-09-21T06:48:08.395828Z INFO mlxcel::server::startup: Endpoints:
2026-09-21T06:48:08.395829Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions
2026-09-21T06:48:08.395831Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions
2026-09-21T06:48:08.395832Z INFO mlxcel::server::startup: GET /v1/models - List models
2026-09-21T06:48:08.395833Z INFO mlxcel::server::startup: POST /completion - llama-server native completion
2026-09-21T06:48:08.395834Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text
2026-09-21T06:48:08.395836Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens
2026-09-21T06:48:08.395837Z INFO mlxcel::server::startup: GET /props - Server properties
2026-09-21T06:48:08.395838Z INFO mlxcel::server::startup: GET /slots - Slot status
2026-09-21T06:48:08.395839Z INFO mlxcel::server::startup: GET /health - Health check
2026-09-21T06:48:13.105736Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3975ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3975
2026-09-21T06:48:16.608087Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3499ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3499
2026-09-21T06:48:20.120235Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3510ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3510
2026-09-21T06:48:23.633141Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/158 prompt tokens, total 3510ms prompt_tokens=158 cached_tokens=0 generation_time_ms=3510
Loading
Loading