How can I improve the prefilling speed for this model?

#49
by BipedalBit - opened

When running Qwen3.8-27b-IQ4_XS, I can get up to 1000 t/s on prefilling speed on my NVIDIA 4070Ti Super 16G, while for Qwen3.8-Flash-Next-IQ4_XS, I see that even with 64G RAM, the maximum prefilling speed is only around 50 t/s. How can I improve the prefilling speed for this model?
If the prefilling speed is high enough and TTFT is small enough, I think a generation speed of 10+ t/s would also be acceptable.

Set the batch and ub values to high numbers, and you'll see a several-fold improvement with large prompts. Keep in mind that this consumes additional video memory. Try the following:
1)
-b 2048
-ub 2048
2)
-b 4096
-ub 4096

I saw at https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/discussions/35#6a928b54e2a71b3463183b92 that someone has already tried to refine the currently globally applied mmap, making only ngram use mmap while keeping other compute-intensive parts in VRAM or RAM as much as possible.
I've already tried this set of configurations myself:

--no-mmproj-offload
--kv-unified
--kv-offload
--ctx-size ${ctx_size_224k}
--parallel 1
--threads 14
--threads-batch 20
--flash-attn on
--cpu-moe
--no-warmup
--tensor-read-lazy on
--override-kv "qwen4exp.attention.indexer.top_k=int:4096"
--jinja
--reasoning on
--reasoning-preserve
--reasoning-effort medium
--reasoning-format deepseek
--temp 1
--top-k 20

On a 16G VRAM + 32G RAM + NVMe setup, I can get a maximum prefilling speed of 20 t/s and a maximum generation speed of 10 t/s, but it's unstable.
Apart from ordering more RAM, I don't have a better configuration plan to optimize the prefilling speed.

image

Set the batch and ub values to high numbers, and you'll see a several-fold improvement with large prompts. Keep in mind that this consumes additional video memory. Try the following:
1)
-b 2048
-ub 2048
2)
-b 4096
-ub 4096

impressive effect

--no-mmproj-offload
--kv-unified
--kv-offload
--ctx-size ${ctx_size_128k}
--parallel 1
--threads 14
--threads-batch 20
--batch-size 2048
--ubatch-size 2048
--flash-attn on
--cpu-moe
--no-warmup
--tensor-read-lazy on
--override-kv "qwen4exp.attention.indexer.top_k=int:4096"
--jinja
--reasoning on
--reasoning-preserve
--reasoning-effort medium
--reasoning-format deepseek
--temp 1
--top-k 20
0.17.298.487 I srv    load_model: loaded multimodal model, '/ssd-models/unsloth/Qwen3.8-Flash-Next-GGUF/mmproj-BF16.gguf'
0.18.491.956 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 131072, kv_unified = 'true'
0.18.506.937 I srv  llama_server: model loaded
0.18.506.987 I srv  llama_server: listening on http://0.0.0.0:5813
0.19.650.222 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
0.19.652.900 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
0.45.816.057 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   2048, progress = 0.33, t =  26.15 s / 78.31 tokens per second
1.06.627.123 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   4096, progress = 0.66, t =  46.96 s / 87.22 tokens per second
1.30.390.595 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   4202, progress = 0.67, t =  70.74 s / 59.40 tokens per second
1.52.746.595 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6238, progress = 1.00, t =  93.09 s / 67.01 tokens per second
1.54.865.296 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6246, progress = 1.00, t =  95.21 s / 65.60 tokens per second
2.09.424.337 I slot print_timing: id  0 | task 0 | n_gen =    100, tg =   7.07 t/s, tg_3s =   7.14 t/s
2.12.522.129 I slot print_timing: id  0 | task 0 | n_gen =    129, tg =   7.48 t/s, tg_3s =   9.36 t/s
2.15.552.494 I slot print_timing: id  0 | task 0 | n_gen =    154, tg =   7.60 t/s, tg_3s =   8.25 t/s
2.18.677.214 I slot print_timing: id  0 | task 0 | n_gen =    179, tg =   7.65 t/s, tg_3s =   8.00 t/s
2.21.688.769 I slot print_timing: id  0 | task 0 | n_gen =    198, tg =   7.50 t/s, tg_3s =   6.31 t/s
2.24.790.513 I slot print_timing: id  0 | task 0 | n_gen =    229, tg =   7.76 t/s, tg_3s =   9.99 t/s
2.27.801.878 I slot print_timing: id  0 | task 0 | n_gen =    262, tg =   8.06 t/s, tg_3s =  10.96 t/s
2.30.812.962 I slot print_timing: id  0 | task 0 | n_gen =    299, tg =   8.42 t/s, tg_3s =  12.29 t/s
2.33.888.412 I slot print_timing: id  0 | task 0 | n_gen =    335, tg =   8.68 t/s, tg_3s =  11.71 t/s
2.37.121.745 I slot print_timing: id  0 | task 0 | prompt eval time =   95765.52 ms /  6250 tokens (   15.32 ms per token,    65.26 tokens per second)
2.37.121.754 I slot print_timing: id  0 | task 0 |        eval time =   41701.73 ms /   358 tokens (  116.81 ms per token,     8.56 tokens per second)
2.37.121.755 I slot print_timing: id  0 | task 0 |       total time =  137467.25 ms /  6608 tokens
2.37.121.758 I slot print_timing: id  0 | task 0 |    graphs reused =        355
2.37.182.870 I slot      release: id  0 | task 0 | stop processing: n_tokens = 6607, truncated = 0
2.37.809.287 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = 108992968159

I think this is about the limit I can accept.

--no-mmproj-offload
--kv-unified
--kv-offload
--ctx-size ${ctx_size_128k}
--parallel 1
--threads 14
--threads-batch 20
--batch-size 2560
--ubatch-size 2560
--flash-attn on
--cpu-moe
--no-warmup
--tensor-read-lazy on
--override-kv "qwen4exp.attention.indexer.top_k=int:4096"
--jinja
--reasoning on
--reasoning-preserve
--reasoning-effort medium
--reasoning-format deepseek
--temp 1
--top-k 20
0.19.819.857 I srv    load_model: loaded multimodal model, '/ssd-models/unsloth/Qwen3.8-Flash-Next-GGUF/mmproj-BF16.gguf'
0.24.070.589 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 131072, kv_unified = 'true'
0.24.115.807 I srv  llama_server: model loaded
0.24.116.280 I srv  llama_server: listening on http://0.0.0.0:5813
0.24.770.874 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
0.24.773.679 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
0.48.878.580 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   2560, progress = 0.41, t =  24.09 s / 106.26 tokens per second
1.14.623.936 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   3690, progress = 0.59, t =  49.84 s / 74.04 tokens per second
1.36.393.012 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6238, progress = 1.00, t =  71.61 s / 87.11 tokens per second
1.38.535.448 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   6246, progress = 1.00, t =  73.76 s / 84.68 tokens per second
1.53.902.563 I slot print_timing: id  0 | task 0 | n_gen =    100, tg =   6.68 t/s, tg_3s =   6.74 t/s
1.56.916.809 I slot print_timing: id  0 | task 0 | n_gen =    130, tg =   7.23 t/s, tg_3s =   9.95 t/s
1.59.929.698 I slot print_timing: id  0 | task 0 | n_gen =    157, tg =   7.48 t/s, tg_3s =   8.96 t/s
2.02.999.748 I slot print_timing: id  0 | task 0 | n_gen =    188, tg =   7.82 t/s, tg_3s =  10.10 t/s
2.06.016.592 I slot print_timing: id  0 | task 0 | n_gen =    223, tg =   8.24 t/s, tg_3s =  11.60 t/s
2.09.023.448 I slot print_timing: id  0 | task 0 | n_gen =    249, tg =   8.28 t/s, tg_3s =   8.65 t/s
2.12.027.587 I slot print_timing: id  0 | task 0 | n_gen =    277, tg =   8.38 t/s, tg_3s =   9.32 t/s
2.15.082.746 I slot print_timing: id  0 | task 0 | n_gen =    310, tg =   8.58 t/s, tg_3s =  10.80 t/s
2.18.087.882 I slot print_timing: id  0 | task 0 | n_gen =    334, tg =   8.54 t/s, tg_3s =   7.99 t/s
2.19.419.222 I slot print_timing: id  0 | task 0 | prompt eval time =   74301.39 ms /  6250 tokens (   11.89 ms per token,    84.12 tokens per second)
2.19.419.224 I slot print_timing: id  0 | task 0 |        eval time =   40343.97 ms /   347 tokens (  116.60 ms per token,     8.58 tokens per second)
2.19.419.225 I slot print_timing: id  0 | task 0 |       total time =  114645.36 ms /  6597 tokens
2.19.419.226 I slot print_timing: id  0 | task 0 |    graphs reused =        344
2.19.465.734 I slot      release: id  0 | task 0 | stop processing: n_tokens = 6596, truncated = 0
2.20.035.786 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = 109502656355

image

I think this is about the limit I can accept.

Taking in account that typical harness (omp, hermess, dsh) takes roughly 25-32k tokens for system prompt, instructions, skills and tools descriptions to load into context with 100 t/s prefill speed you'll need to wait about 5 minuts each time after starting a project. With context growing up to 100k + tokens in case of compacting (when a model does one single prefill from beginning to the end of the context) the amount of time to wait will be equal to 15 minutes or more. And those numbers are exaggerated since prefill speed tends to drop the larger amount of tokens being processed. So at the moment it's barely usable for large contexts.

I think this is about the limit I can accept.

Taking in account that typical harness (omp, hermess, dsh) takes roughly 25-32k tokens for system prompt, instructions, skills and tools descriptions to load into context with 100 t/s prefill speed you'll need to wait about 5 minuts each time after starting a project. With context growing up to 100k + tokens in case of compacting (when a model does one single prefill from beginning to the end of the context) the amount of time to wait will be equal to 15 minutes or more. And those numbers are exaggerated since prefill speed tends to drop the larger amount of tokens being processed. So at the moment it's barely usable for large contexts.

Makes sense. Maybe it's only suitable for those well-planned overnight /goal runs.

Well I am using same model as well. My hardware is rtx3090ti+rtx5060ti+128gb ddr4 3000 mhz system ram. With bunch of tinkering testing and pushing the homelab to limit, at most I could reach at 192k context window kvq8 is at lowest 11 t/s at highest 18 t/s decode speeds and average 480-490 t/s prompt processing speed. But it started as low as 180 t/s prompt speed. Real change was what I disabled mmap and forced rest of the model to load at ram. Than my prompt processing speeds spiked to average 400ish. Than I started to squeeze more layers into GPUs, as more layers goes into it both decode and prompt processing speeds are increased and at one point it started to sit comfortably sit at average 480-490 t/s which I use right now.

But for you there is an issue, remember whole model is 93.7gb in size, so your vram+ram combination is a big bottleneck. You cant disable mmap like me and load whole model along side with your context window. What I might try to do if I were you to switch to lower quant model or maybe make you context window smaller so it takes less space in your gpu, that way you can put more layers in gpu which will boost your decode and prompt processing speeds. you cant increase Ubatch more since as you increase Ubatch, llama.cpp reserve more gpu vram for processing, so less vram left for model itself. you need to tinker the knobs a bit basically.

For example drop your ubatch size to open up more vram, than drop your context window a bit as well and possible do kvq8 to squezee a bit more to grab more vram too. Than try to put 1 or 2 more layers into gpu and check if the speeds are increased or not.

Another choice is switching to lower quant along side what I suggested about. IQ3_K_XL is 3,7 gb is smaller. thats like putting in 1-2 more layer into gpu too and according to unsloth difference between IQ4_X_S and IQ3_K_XL isnt much as well.

You are working at really tight margins. Each GB counts for you.

Today I'm picking up an extra 32G of RAM, and I've also been digging deeper into balancing -ub and --n-cpu-moe. I'll post an update with the results later.

Well I am using same model as well. My hardware is rtx3090ti+rtx5060ti+128gb ddr4 3000 mhz system ram. With bunch of tinkering testing and pushing the homelab to limit, at most I could reach at 192k context window kvq8 is at lowest 11 t/s at highest 18 t/s decode speeds and average 480-490 t/s prompt processing speed. But it started as low as 180 t/s prompt speed. Real change was what I disabled mmap and forced rest of the model to load at ram. Than my prompt processing speeds spiked to average 400ish. Than I started to squeeze more layers into GPUs, as more layers goes into it both decode and prompt processing speeds are increased and at one point it started to sit comfortably sit at average 480-490 t/s which I use right now.

But for you there is an issue, remember whole model is 93.7gb in size, so your vram+ram combination is a big bottleneck. You cant disable mmap like me and load whole model along side with your context window. What I might try to do if I were you to switch to lower quant model or maybe make you context window smaller so it takes less space in your gpu, that way you can put more layers in gpu which will boost your decode and prompt processing speeds. you cant increase Ubatch more since as you increase Ubatch, llama.cpp reserve more gpu vram for processing, so less vram left for model itself. you need to tinker the knobs a bit basically.

For example drop your ubatch size to open up more vram, than drop your context window a bit as well and possible do kvq8 to squezee a bit more to grab more vram too. Than try to put 1 or 2 more layers into gpu and check if the speeds are increased or not.

Another choice is switching to lower quant along side what I suggested about. IQ3_K_XL is 3,7 gb is smaller. thats like putting in 1-2 more layer into gpu too and according to unsloth difference between IQ4_X_S and IQ3_K_XL isnt much as well.

You are working at really tight margins. Each GB counts for you.

For RAM-rich environments, have you tried using GGML_CUDA_REGISTER_HOST=1 GGML_SCHED_PREFETCH_EXPERTS=1 (locked memory + expert prefetching)? I've learned that this can further noticeably improve prefilling speed.

Soo, when I was replying yesterday I was at work and didnt had specific configs I use etc and couldnt run tests, but after your reply I tested everything in more structured way so it will be maybe a benchmark for someone.
First hardware; rtx3090ti+rtx5060ti+128gb ddr4 3000 mhz system ram+ i7 14700k cpu. First biggest limitation for me is I am using consumer grade montherboard, so I only get dual channel ram bandwith which is theorotically is 48 GB/s but in practice you get less out of it. So it slow downs the things a bit when model like this one offloads to ram a lot. Second big limitation for me is bandwith of rtx 5060ti. While rtx 3090ti have 1008 gb/s bandwith, rtx 5060ti only have 448 gb/s. It creates a bottleneck there too.

Second focus is my llama.cpp settings and the build. Before I tested with your suggestions I was using llama.cpp build b10684 with these settings ;
--cache-ram 8192 \ --parallel 1 \ --n-gpu-layers 99 \ --n-cpu-moe 33 \ --split-mode layer \ --tensor-split 85,15 \ --fit off \ --flash-attn on \
--ctx-size 196608 \ --cache-type-k q8_0 \ --cache-type-v q8_0
--batch-size 8192 \ --ubatch-size 2048 \ --threads 16 \ --threads-batch 16
--load-mode none \ --lazy-mode off \ --jinja \ --reasoning-effort medium \ --reasoning-preserve
--temp 1.0 \ --top-p 0.95 \ --top-k 20 \ --min-p 0.0 \ --presence-penalty 0.0 \ --repeat-penalty 1.0
I did all my testings at 70k to 85k context windows, warmed up.
So base setting results; ~76K context window, Decode Speed; 13.92 t/s , Prompt Processing Speed; 476.57 t/s
Than I did same test with GGML_CUDA_REGISTER_HOST=1 added to it; ~73K context window, Decode Speed; 14.77 t/s , Prompt Processing Speed; 479.30 t/s
I sadly could do 3rd test with GGML_SCHED_PREFETCH_EXPERTS=1 or use both of them because looks like prefetch expert stuff is a fork of a llama.cpp. Its not in the official released build yet.

Than I decided to update my build to latest llama.cpp one and it was b10715. and did same tests from scratch
base setting results; ~79K context window, Decode Speed;16.93 t/s , Prompt Processing Speed; 480.80 t/s
Than I did same test with GGML_CUDA_REGISTER_HOST=1 added to it; ~83K context window, Decode Speed; 16.80 t/s , Prompt Processing Speed; 474.82 t/s
I sadly could do 3rd test with GGML_SCHED_PREFETCH_EXPERTS=1 or use both of them because looks like prefetch expert stuff is a fork of a llama.cpp. Its not in the official released build yet. So instead I tried this CUDA_SCALE_LAUNCH_QUEUES=4x; ~86K context window, Decode Speed; 16.44 t/s , Prompt Processing Speed; 468.95 t/s
I honestly new llama.cpp build showed 2-2.5% increase at my Prompt processing speed, bug biggest gain I got was from Decoding speed which whopping got 11,5% base increase decode speed and higher decode speeds at longer context windows. So even though those settings didnt do anything for me. It was nice to test and see all of the results.

If anyone have any other settings which may improve things, I am happy to be a test subject. Just shoot your ideas and I can give it a try.

In the bench mark picture, I got the all check point speeds from context window at 0 token to the end of it, and compared them. So Graph is more accurate than these average values I wrote in the post, since with growing context window of each msg, decode and prompt processing speed degrade. So comparing check points is much more accurate.

ChatGPT Image Aug 31, 2026, 02_49_31 PM

on a 16 gb vram and 64 gb ddr5 ram on IQ4_XS i am getting 200 tps prompt prefill... get latest llama.cpp from master...
here is the recipe try on ur end

n-gpu-layers = 49
ncmoe = 44 # ur choice
ctx-size = 124000 # ur choice
ctk=q4_0 #hadamard rotation for help
ctv=q4_0
temp = 0.8
top-p = 0.95
top-k = 20
min-p = 0.0
reasoning-budget = 2048
chat-template-kwargs = {"preserve_thinking": true,"reasoning_effort": "medium"}
b=2048
ub=1024
np=1
flash-attn = on
t=6 # play with core numbers
tb=8 # play with logical threads
jinja=on
lm=mmap

Give some queries/time, model will move quickly to a higher prompt prefill [using amd 9600x]

Today I'm picking up an extra 32G of RAM, and I've also been digging deeper into balancing -ub and --n-cpu-moe. I'll post an update with the results later.

After adding the new RAM, I got held up for a while by a SecureBoot issue. Now let's take a look at how the optimization performs on the Nvidia 4070Ti Super 16G VRAM + 64G RAM setup.
Just as I mentioned earlier, I'll run a controlled comparison of different balancing schemes between --n-cpu-moe and -ub to see how they affect performance.

llama.cpp version: b10712

common configs:

--model /ssd-models/unsloth/Qwen3.8-Flash-Next-GGUF/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf
--mmproj /ssd-models/unsloth/Qwen3.8-Flash-Next-GGUF/mmproj-BF16.gguf
--no-mmproj-offload
-ctk q4_0
-ctv q4_0
--kv-unified
--kv-offload
--keep -1
--ctx-size ${ctx_size_128k}
--parallel 1
--threads 14
--threads-batch 20
--flash-attn on
--n-gpu-layers ${n_gpu_layers_max}
--tensor-read-lazy on
--override-kv "qwen4exp.attention.indexer.top_k=int:2048"
--jinja
--reasoning on
--reasoning-preserve
--reasoning-effort medium
--reasoning-format deepseek
--temp 1
--top-k 20
request body from Open WebUI ```json { "stream": true, "model": "qwen3.8-flash-next:q4-128k", "messages": [ { "role": "user", "content": "θ‡ͺζˆ‘δ»‹η»δΈ€δΈ‹" } ], "tools": [ { "type": "function", "function": { "name": "get_current_timestamp", "description": "Get the current Unix timestamp in seconds.", "parameters": { "properties": {}, "type": "object" } } }, { "type": "function", "function": { "name": "calculate_timestamp", "description": "Get the current Unix timestamp, optionally adjusted by days, weeks, months, or years.\nUse this to calculate timestamps for date filtering in search functions.\nExamples: \"last week\" = weeks_ago=1, \"3 days ago\" = days_ago=3, \"a year ago\" = years_ago=1", "parameters": { "properties": { "days_ago": { "default": 0, "description": "Number of days to subtract from current time (default: 0)", "type": "integer" }, "weeks_ago": { "default": 0, "description": "Number of weeks to subtract from current time (default: 0)", "type": "integer" }, "months_ago": { "default": 0, "description": "Number of months to subtract from current time (default: 0)", "type": "integer" }, "years_ago": { "default": 0, "description": "Number of years to subtract from current time (default: 0)", "type": "integer" } }, "type": "object" } } }, { "type": "function", "function": { "name": "ask_user", "description": "Ask the user clarifying questions before continuing.\nUse this when the next step depends on user intent, preference, or a tradeoff that cannot be inferred safely.", "parameters": { "properties": { "questions": { "description": "1-3 question objects, each with id, header, question, and 2-3 options. Each option needs label and description.", "items": { "additionalProperties": true, "type": "object" }, "type": "array" }, "allow_other": { "default": true, "description": "Whether users may enter a free-form answer instead of choosing one of the options", "type": "boolean" }, "timeout_ms": { "default": 120000, "description": "How long the browser should keep the prompt open before cancelling it", "type": "integer" } }, "required": [ "questions" ], "type": "object" } } }, { "type": "function", "function": { "name": "list_knowledge_bases", "description": "List the user's accessible knowledge bases so a relevant internal source\ncan be chosen.", "parameters": { "properties": { "count": { "default": 10, "description": "Maximum number of KBs to return (default: 10)", "type": "integer" }, "skip": { "default": 0, "description": "Number of results to skip for pagination (default: 0)", "type": "integer" } }, "type": "object" } } }, { "type": "function", "function": { "name": "search_knowledge_bases", "description": "Search the user's accessible knowledge bases by name and description to find\na relevant internal source.", "parameters": { "properties": { "query": { "description": "The search query to find matching knowledge bases", "type": "string" }, "count": { "default": 5, "description": "Maximum number of results to return (default: 5)", "type": "integer" }, "skip": { "default": 0, "description": "Number of results to skip for pagination (default: 0)", "type": "integer" } }, "required": [ "query" ], "type": "object" } } }, { "type": "function", "function": { "name": "query_knowledge_bases", "description": "Search knowledge bases by semantic similarity to query.\nFinds KBs whose name/description match the meaning of your query.\nHelpful for discovering which knowledge base to query next.", "parameters": { "properties": { "query": { "description": "Natural language query describing what you're looking for", "type": "string" }, "count": { "default": 5, "description": "Maximum results (default: 5)", "type": "integer" } }, "required": [ "query" ], "type": "object" } } }, { "type": "function", "function": { "name": "grep_knowledge_files", "description": "Search for exact text across knowledge files. Returns matching lines with line numbers.\nUnlike query_knowledge_files (semantic/vector search), this performs exact string matching.\nAutomatically detects regex patterns (e.g. \"error|warn\", \"version \\d+\").\nHelpful for literal strings, identifiers, error messages, or regex-style searches.", "parameters": { "properties": { "pattern": { "description": "The text pattern to search for (regex auto-detected)", "type": "string" }, "file_id": { "description": "Optional file ID to search within a single file only", "type": "string" }, "case_insensitive": { "default": false, "description": "If true, ignore case when matching (default: false)", "type": "boolean" }, "count_only": { "default": false, "description": "If true, return only match counts per file (default: false)", "type": "boolean" } }, "required": [ "pattern" ], "type": "object" } } }, { "type": "function", "function": { "name": "search_knowledge_files", "description": "Search files by filename across knowledge bases the user has access to.\nWhen the model has attached knowledge, searches only within attached KBs and files.\nHelpful when looking for a specific document or file name.", "parameters": { "properties": { "query": { "description": "The search query to find matching files by filename", "type": "string" }, "knowledge_id": { "description": "Optional KB id to limit search to a specific knowledge base", "type": "string" }, "count": { "default": 5, "description": "Maximum number of results to return (default: 5)", "type": "integer" }, "skip": { "default": 0, "description": "Number of results to skip for pagination (default: 0)", "type": "integer" } }, "required": [ "query" ], "type": "object" } } }, { "type": "function", "function": { "name": "query_knowledge_files", "description": "Search knowledge base files using semantic/vector search. Searches across collections (KBs),\nindividual files, and notes that the user has access to.\nHelpful for internal documentation, uploaded knowledge, and attached model knowledge.", "parameters": { "properties": { "query": { "description": "The search query to find semantically relevant content", "type": "string" }, "knowledge_ids": { "description": "Optional list of KB ids to limit search to specific knowledge bases", "items": { "type": "string" }, "type": "array" }, "count": { "default": 5, "description": "Maximum number of results to return (default: 5)", "type": "integer" } }, "required": [ "query" ], "type": "object" } } }, { "type": "function", "function": { "name": "view_knowledge_file", "description": "Get the content of a file from a knowledge base. Supports pagination for large files.", "parameters": { "properties": { "file_id": { "description": "The ID of the file to retrieve", "type": "string" }, "offset": { "default": 0, "description": "Character offset to start reading from (default: 0)", "type": "integer" }, "max_chars": { "default": 10000, "description": "Maximum characters to return (a server-side hard cap applies)", "type": "integer" }, "line_numbers": { "default": false, "description": "If true, prefix each line with its 1-indexed line number", "type": "boolean" }, "start_line": { "description": "Optional 1-indexed start line (overrides offset/max_chars when set)", "type": "integer" }, "end_line": { "description": "Optional 1-indexed end line (inclusive)", "type": "integer" } }, "required": [ "file_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "search_chats", "description": "Search the user's previous chat conversations by title and message content,\nexcluding the current chat. Helpful for finding details from earlier\nconversations when they are not already visible in the current context.\nExact phrase matches are preferred, and descriptive keyword queries are\nsupported.", "parameters": { "properties": { "query": { "description": "Exact phrase or descriptive keyword query to find matching previous chats", "type": "string" }, "count": { "default": 5, "description": "Maximum number of results to return (default: 5)", "type": "integer" }, "start_timestamp": { "description": "Only include chats updated after this Unix timestamp (seconds)", "type": "integer" }, "end_timestamp": { "description": "Only include chats updated before this Unix timestamp (seconds)", "type": "integer" } }, "required": [ "query" ], "type": "object" } } }, { "type": "function", "function": { "name": "view_chat", "description": "Get the full conversation history of a chat by its ID after a relevant\nprevious chat has been identified.", "parameters": { "properties": { "chat_id": { "description": "The ID of the chat to retrieve", "type": "string" } }, "required": [ "chat_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "search_memories", "description": "Search or browse saved memories by content, path, type, or memory ID.", "parameters": { "properties": { "query": { "default": "", "description": "Optional query to search memory content and path", "type": "string" }, "count": { "default": 5, "description": "Number of memories to return (default 5)", "type": "integer" }, "type": { "default": "all", "description": "\"user\", \"context\", or \"all\"", "type": "string" }, "path": { "description": "Optional memory path to search around", "type": "string" }, "memory_id": { "description": "Optional exact memory ID to read", "type": "string" } }, "type": "object" } } }, { "type": "function", "function": { "name": "list_memory_paths", "description": "List saved memory paths to find existing memory groups before writing or moving memories.", "parameters": { "properties": { "query": { "default": "", "description": "Optional query to filter memory paths or contents", "type": "string" }, "count": { "default": 100, "description": "Maximum number of paths to return", "type": "integer" }, "type": { "default": "all", "description": "\"user\", \"context\", or \"all\"", "type": "string" } }, "type": "object" } } }, { "type": "function", "function": { "name": "read_memory_path", "description": "Read saved memories at a memory path, including nearby parent and child paths.", "parameters": { "properties": { "path": { "description": "Memory path to read", "type": "string" }, "count": { "default": 50, "description": "Maximum number of memories to return", "type": "integer" }, "type": { "default": "all", "description": "\"user\", \"context\", or \"all\"", "type": "string" }, "include_children": { "default": true, "description": "Include memories under child paths", "type": "boolean" } }, "required": [ "path" ], "type": "object" } } }, { "type": "function", "function": { "name": "list_memories", "description": "List all stored memories for the user, including IDs and timestamps.", "parameters": { "properties": {}, "type": "object" } } }, { "type": "function", "function": { "name": "update_memory", "description": "Apply a batch of memory changes after learning enduring information.\n\nUse type \"user\" for facts, preferences, or instructions about the user.\nUse type \"context\" for other durable context that may help future chats.\nDo not save one-off activity, meals, routine daily events, temporary mood, or other short-lived details\nunless the user explicitly asks you to remember them.\nPath is optional. Use it as a stable memory address to group related memories.\nPrefer an existing path from list_memory_paths when one fits.\nLeave path empty when no useful grouping is clear.\n\nOperation shapes:\n- {\"action\": \"add\", \"content\": \"...\", \"type\": \"user\"|\"context\", \"path\": \"...\"}\n- {\"action\": \"replace\", \"id\": \"...\", \"content\": \"...\", \"type\": \"user\"|\"context\", \"path\": \"...\"}\n- {\"action\": \"move\", \"id\": \"...\", \"path\": \"...\"}\n- {\"action\": \"remove\", \"id\": \"...\"}", "parameters": { "properties": { "operations": { "description": "Memory operations to apply in one request", "items": { "additionalProperties": true, "type": "object" }, "type": "array" } }, "required": [ "operations" ], "type": "object" } } }, { "type": "function", "function": { "name": "add_memory", "description": "Save enduring information that can improve future chats.\n\nSave stable preferences, goals, projects, relationships, habits, and standing instructions.\nDo not save one-off activity, meals, routine daily events, temporary mood, or other short-lived details\nunless the user explicitly asks you to remember them.", "parameters": { "properties": { "content": { "description": "The memory content to store", "type": "string" }, "type": { "default": "user", "description": "Use \"user\" for facts/preferences about the user, or \"context\" for other durable context", "type": "string" }, "path": { "description": "Optional stable memory address for grouping related memories", "type": "string" } }, "required": [ "content" ], "type": "object" } } }, { "type": "function", "function": { "name": "replace_memory_content", "description": "Update an existing saved memory by its ID when its content needs correction.", "parameters": { "properties": { "memory_id": { "description": "The ID of the memory to update", "type": "string" }, "content": { "description": "The new content for the memory", "type": "string" }, "type": { "description": "Optional \"user\" or \"context\" type for the updated memory", "type": "string" }, "path": { "description": "Optional stable memory address for grouping related memories", "type": "string" } }, "required": [ "memory_id", "content" ], "type": "object" } } }, { "type": "function", "function": { "name": "delete_memory", "description": "Delete a saved memory by its ID.", "parameters": { "properties": { "memory_id": { "description": "The ID of the memory to delete", "type": "string" } }, "required": [ "memory_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "search_notes", "description": "Search the user's saved notes by title and content.", "parameters": { "properties": { "query": { "description": "The search query to find matching notes", "type": "string" }, "count": { "default": 5, "description": "Maximum number of results to return (default: 5)", "type": "integer" }, "start_timestamp": { "description": "Only include notes updated after this Unix timestamp (seconds)", "type": "integer" }, "end_timestamp": { "description": "Only include notes updated before this Unix timestamp (seconds)", "type": "integer" } }, "required": [ "query" ], "type": "object" } } }, { "type": "function", "function": { "name": "view_note", "description": "Get the full content of a note by its ID.", "parameters": { "properties": { "note_id": { "description": "The ID of the note to retrieve", "type": "string" } }, "required": [ "note_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "write_note", "description": "Create a new note with the given title and content.", "parameters": { "properties": { "title": { "description": "The title of the new note", "type": "string" }, "content": { "description": "The markdown content for the note", "type": "string" } }, "required": [ "title", "content" ], "type": "object" } } }, { "type": "function", "function": { "name": "replace_note_content", "description": "Update an existing note by replacing the whole markdown content or applying range operations.", "parameters": { "properties": { "note_id": { "description": "The ID of the note to update", "type": "string" }, "content": { "description": "The new markdown content for a whole-note update", "type": "string" }, "operations": { "description": "Optional note operations:\n- {\"action\": \"replace\", \"content\": \"...\"}\n- {\"action\": \"replace_range\", \"start\": 0, \"end\": 10, \"content\": \"...\", \"expected\": \"...\"}", "items": { "additionalProperties": true, "type": "object" }, "type": "array" }, "title": { "description": "Optional new title for the note", "type": "string" } }, "required": [ "note_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "create_tasks", "description": "Create a visible task checklist for multi-step work so progress can be shown in chat.", "parameters": { "properties": { "tasks": { "description": "List of task items. Each item: content (string, required), status (pending|in_progress|completed|cancelled, default pending), id (optional, auto-generated).", "items": { "properties": { "id": { "description": "Unique identifier for the task. Auto-generated if omitted.", "type": "string" }, "content": { "description": "Task description.", "type": "string" }, "status": { "default": "pending", "description": "Task status.", "enum": [ "pending", "in_progress", "completed", "cancelled" ], "type": "string" } }, "required": [ "content" ], "type": "object" }, "type": "array" } }, "required": [ "tasks" ], "type": "object" } } }, { "type": "function", "function": { "name": "update_task", "description": "Mark a single visible task item as completed, in_progress, pending, or cancelled.", "parameters": { "properties": { "id": { "description": "The task ID to update", "type": "string" }, "status": { "default": "completed", "description": "New status: completed, in_progress, pending, or cancelled (default: completed)", "type": "string" } }, "required": [ "id" ], "type": "object" } } }, { "type": "function", "function": { "name": "create_automation", "description": "Create a scheduled automation that runs a prompt on a recurring or one-time schedule.\nUse this when the user wants to schedule a task to run automatically.\nThe automation will use the current chat model.\n\nThe rrule parameter must be a valid iCalendar RRULE string. Common examples:\n- Every day at 9am: \"DTSTART:20250101T090000\\nRRULE:FREQ=DAILY\"\n- Every Monday at 8am: \"DTSTART:20250106T080000\\nRRULE:FREQ=WEEKLY;BYDAY=MO\"\n- Every hour: \"RRULE:FREQ=HOURLY;INTERVAL=1\"\n- Every 30 minutes: \"RRULE:FREQ=MINUTELY;INTERVAL=30\"\n- Once at a specific time: \"DTSTART:20250415T140000\\nRRULE:FREQ=DAILY;COUNT=1\"\n- First day of every month: \"DTSTART:20250101T090000\\nRRULE:FREQ=MONTHLY;BYMONTHDAY=1\"\n\nThe DTSTART time should reflect the desired execution time. Use COUNT=1 for one-time automations.", "parameters": { "properties": { "name": { "description": "A short descriptive name for the automation", "type": "string" }, "prompt": { "description": "The prompt/instructions to execute on each run", "type": "string" }, "rrule": { "description": "An iCalendar RRULE string defining the schedule", "type": "string" }, "folder_id": { "description": "Optional owner-owned folder ID for generated chats", "type": "string" } }, "required": [ "name", "prompt", "rrule" ], "type": "object" } } }, { "type": "function", "function": { "name": "update_automation", "description": "Update an existing automation. Only the provided fields are changed; omitted fields stay the same.", "parameters": { "properties": { "automation_id": { "description": "The ID of the automation to update", "type": "string" }, "name": { "description": "New name for the automation (optional)", "type": "string" }, "prompt": { "description": "New prompt/instructions (optional)", "type": "string" }, "rrule": { "description": "New iCalendar RRULE schedule string (optional). See create_automation for format examples.", "type": "string" }, "model_id": { "description": "New model ID to use (optional); blank values are ignored", "type": "string" }, "folder_id": { "default": "", "description": "New owner-owned folder ID (optional); omit or pass blank to keep unchanged, pass null to clear", "type": "string" } }, "required": [ "automation_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "list_automations", "description": "List the user's scheduled automations.", "parameters": { "properties": { "status": { "description": "Filter by status: \"active\", \"paused\", or omit for all", "type": "string" }, "folder_id": { "description": "Optional owner-owned folder ID filter; pass an empty string to clear the folder filter", "type": "string" }, "count": { "default": 10, "description": "Maximum number of automations to return (default: 10)", "type": "integer" } }, "type": "object" } } }, { "type": "function", "function": { "name": "toggle_automation", "description": "Pause or resume a scheduled automation. If active, it will be paused. If paused, it will be resumed.", "parameters": { "properties": { "automation_id": { "description": "The ID of the automation to toggle", "type": "string" } }, "required": [ "automation_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "delete_automation", "description": "Delete a scheduled automation and all its run history.", "parameters": { "properties": { "automation_id": { "description": "The ID of the automation to delete", "type": "string" } }, "required": [ "automation_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "search_calendar_events", "description": "Search calendar events, reminders, and scheduled items by text and/or date range.\nHelpful for finding upcoming events, reminders, or schedule items.", "parameters": { "properties": { "query": { "description": "Search text to match against event title, description, or location (optional)", "type": "string" }, "start": { "description": "Only return events starting at or after this datetime, e.g. \"2026-04-20 00:00\" (optional)", "type": "string" }, "end": { "description": "Only return events starting before this datetime, e.g. \"2026-04-27 00:00\" (optional)", "type": "string" }, "count": { "default": 10, "description": "Maximum number of events to return (default: 10)", "type": "integer" } }, "type": "object" } } }, { "type": "function", "function": { "name": "create_calendar_event", "description": "Create a calendar event, reminder, or alarm. Use this when the user wants to\nschedule an event, set a reminder, create an alarm, or says things like\n\"remind me\", \"don't let me forget\", \"notify me at\", or \"add to my calendar\".\nFor simple reminders, omit end/location/all_day and set reminder_minutes to 0.", "parameters": { "properties": { "title": { "description": "Event or reminder title (e.g. \"Team standup\", \"Take medicine\", \"Call mom\")", "type": "string" }, "start": { "description": "Start datetime in the user's local time (e.g. \"2026-04-20 09:00\")", "type": "string" }, "end": { "description": "End datetime in the user's local time (optional β€” omit for reminders or point-in-time events)", "type": "string" }, "description": { "description": "Event description or notes (optional)", "type": "string" }, "calendar_id": { "description": "Target calendar ID (optional, uses default calendar if omitted)", "type": "string" }, "all_day": { "default": false, "description": "Whether this is an all-day event (default: false)", "type": "boolean" }, "location": { "description": "Event location (optional)", "type": "string" }, "reminder_minutes": { "description": "Minutes before the event to send a notification (optional, default: 10). Use 0 for \"at time of event\", -1 for no notification.", "type": "integer" } }, "required": [ "title", "start" ], "type": "object" } } }, { "type": "function", "function": { "name": "update_calendar_event", "description": "Update an existing calendar event. Only provided fields are changed;\nomitted fields stay the same.", "parameters": { "properties": { "event_id": { "description": "The ID of the event to update", "type": "string" }, "title": { "description": "New event title (optional)", "type": "string" }, "description": { "description": "New event description (optional)", "type": "string" }, "start": { "description": "New start datetime string in your local time, e.g. \"2026-04-20 09:00\" (optional)", "type": "string" }, "end": { "description": "New end datetime string in your local time (optional)", "type": "string" }, "all_day": { "description": "Whether this is an all-day event (optional)", "type": "boolean" }, "location": { "description": "New event location (optional)", "type": "string" }, "is_cancelled": { "description": "Set to true to cancel the event (optional)", "type": "boolean" }, "reminder_minutes": { "description": "Minutes before the event to send a reminder notification (optional). Use 0 for \"at time of event\", -1 for no reminder. Accepts any positive integer for custom timing (e.g. 120 for 2 hours before).", "type": "integer" } }, "required": [ "event_id" ], "type": "object" } } }, { "type": "function", "function": { "name": "delete_calendar_event", "description": "Delete a calendar event permanently.", "parameters": { "properties": { "event_id": { "description": "The ID of the event to delete", "type": "string" } }, "required": [ "event_id" ], "type": "object" } } } ] } ```
This request body contains a message like 'introduce yourself' along with a bunch of tool injections.
Metric Config 1 (--batch-size 4096 --ubatch-size 2048 --n-cpu-moe 45) Config 2 (--batch-size 8192 --ubatch-size 4096 --n-cpu-moe 48)
Prompt eval total time 35,358.14 ms 26,352.95 ms
Prompt eval tokens 6,250 6,250
Prompt eval ms/token 5.66 ms 4.22 ms
Prompt eval speed (pp) 176.76 tokens/s 237.17 tokens/s
Peak prompt processing speed ~179.29 t/s ~241.91 t/s
Generation eval total time 18,605.04 ms 17,303.79 ms
Generation eval tokens 339 270
Generation eval ms/token 55.04 ms 64.33 ms
Generation speed (tg) 18.17 tokens/s 15.55 tokens/s
Peak generation speed ~18.36 t/s ~16.58 t/s
Total time 53,963.17 ms 43,656.74 ms
Total tokens processed 6,589 6,520
Graphs reused 576 655

Summary

  • Prefilling (pp) speed: Config 2 is faster (237 t/s) than Config 1 (177 t/s).
  • Generation (tg) speed: Config 1 is faster(18.2 t/s) than Config 2 (15.6 t/s).
  • Total time: Config 2 completes the task the fastest overall, despite having the slowest generation, due to its significantly faster prompt processing.
  • Trade-off: Larger batch/ubatch sizes boost prefilling but tend to hurt generation throughput.

I think this is about the limit I can accept.

Taking in account that typical harness (omp, hermess, dsh) takes roughly 25-32k tokens for system prompt, instructions, skills and tools descriptions to load into context with 100 t/s prefill speed you'll need to wait about 5 minuts each time after starting a project. With context growing up to 100k + tokens in case of compacting (when a model does one single prefill from beginning to the end of the context) the amount of time to wait will be equal to 15 minutes or more. And those numbers are exaggerated since prefill speed tends to drop the larger amount of tokens being processed. So at the moment it's barely usable for large contexts.

For the hermes-agent scenario, most requests look like this β€” it doesn't seem too bad because the prompts almost always hit the cache. I think it's similar in dsh and CC as well.

image

To be fair those numbers are actually great. You are really optimizing it well. @ 39k context window usage having prefill at 270.26 is quite a good number. What I realized when you first run the model server, until it warms up prefill speeds starts low and climb until up to a sweet spot, than after a sweet spot is passed it starts to gradually drop as context window is getting bigger and bigger.
From my own tests when I warm up the model and give it a 80k context , it starts at 640 t/s PP speed and as it get processed and context is filling up slowly drops. lowest I saw was 450 t/s @80k context. Probably if I push it at higher context it will drop to 350 t/s values maybe 300 t/s. Honestly I still think its usable. Which leads to next part of my reply.

  • Total time: Config 2 completes the task the fastest overall, despite having the slowest generation, due to its significantly faster prompt processing.

After reading this I started to think about why I am so stuck up to PP speeds and valuing it over decode? Of course I inject a lot of documents to it or it have access to my project files and it does lots of reading etc. So I always prioritized PP over decode. Than a thought come to me and I went into hermes and checked last 3 day's token usage and it was as follows;
Input: 4,292,555 tokens
Output: 601,322 tokens
Cached: 54,870,667 tokens
Total: 59,764,544 tokens
So with my speeds as it follows;
PP: ~480.8 t/s @80k context
Decode: ~16.93 t/s @80kcontext
Time to churn all these tokens was;
New input:
4.29M / 480.8 β‰ˆ 2.48 hours
Output:
601K / 16.93 β‰ˆ 9.87 hours

80% of my time is spend on decoding while 20% of it used to prompt processing. So in theory if I can sacrifice some prompt processing speed to gain worthwile decode speed, the time I spend to accomplish a task will be shorter. What matters most is how much time you spend to complete the task after all. not how much time you spend at input or output?

So my plan is to play around with ubatch etc to lower prompt processing speeds a bit to gain some decode speed. I made gpt to do some calculations for me to aim for some numbers as a goal so I guess its something like this;
PP 460 + TG 17.3 β†’ not really worth it
PP 460 + TG 17.6 β†’ worthwhile
PP 450 + TG 18.0 β†’ definitely worthwhile
PP 430 + TG 18.0 β†’ still worthwhile
PP 420 + TG 18.5 β†’ very worthwhile
PP 400 + TG 19.0 β†’ still a good trade

Lets see how its gonna turn out.

Sign up or log in to comment