Qwen3.8 Flash-Next Uncensored for Gufo โ€” Q4 Mix

A Gufo-compatible GGUF conversion of orcarouter/Qwen3.8-Flash-Next-Uncensored. This is a quantization of that checkpoint, not a new fine-tune. It was built and tested on AMD Strix Halo (gfx1151) with Gufo commit eb91584.

Q4 Mix is this release's shorthand for its mixed tensor recipe. It is not a stock Q4_K_M preset: routed expert gate/up weights are Q4_K, expert down weights are Q5_1, the large per-layer embedding table is IQ4_NL, and the token embeddings, output head, and most dense projections are Q8_0. The quality-sensitive output head remains at higher precision. The exact recipe and build notes are included.

Files

File Purpose Size SHA-256
Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix.gguf Text target with embedded PLE weights 109,776,678,624 bytes (102.23 GiB) c7f2331227aaefd1947b15fd6a31adeb29a4609a4dec5713670743b61ec44b7a
mmproj-BF16.gguf Matching BF16 vision projector 907,543,360 bytes 0c2760d6f3e6686500989ad6ccc28aab69abfb272a3419ac33d4e26e7708ce8a

The optional Q8_0 MTP predictor is not included. The tested sidecar is mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf from jcbtc's CIRU Orca release. Gufo also runs this target without MTP.

Run with Gufo

Build or install Gufo for Strix Halo, then download this repository's GGUF files. This release is public and ungated. No Hugging Face account or access request is required to download these files.

hf download vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF \
  --include '*.gguf' --local-dir ./orca-gufo

gufo serve --host 127.0.0.1 --port 8080 --sessions 2 llm \
  --model ./orca-gufo/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix.gguf \
  --mmproj ./orca-gufo/mmproj-BF16.gguf \
  --served-model-name orca-gufo --context 262144

For speculative MTP, download the tested sidecar and add the following options to the gufo serve ... llm command:

hf download jcbtc/Qwen3.8-Flash-CIRU-STRIX-Orca \
  --include 'mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf' \
  --local-dir ./ciru-sidecar

# Add to gufo serve ... llm:
--speculative mtp \
--mtp-model ./ciru-sidecar/mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf \
--draft-tokens 6

The API is OpenAI-compatible at http://127.0.0.1:8080/v1; send model ID orca-gufo. Gufo's Qwen formatter is compiled into the runtime and does not load an external Jinja template. Thinking and effort can be controlled per request with reasoning_effort; --think off is a server default if desired. See Gufo's Flash-Next guide and server API documentation.

The repository contains a BF16 vision projector, but image input was tested only as a simple controlled smoke check. Generic GGUF runners and other GPUs have not been qualified for this particular package.

What was tested

The target loaded with Gufo on a Ryzen AI MAX+ 395 / Radeon 8060S (gfx1151). Text, a required function tool call, and a simple image question passed. The target also loaded with the linked MTP sidecar. A 262,144-token context configuration was tested with one session and short text/tool/image requests. The operator subsequently ran two concurrent 262,144-token text sessions on bare-bones Linux and reported roughly 40 tokens/s during long-context development-harness use. A live last-request /metrics gauge read 40.31 tokens/s, but it did not record prompt depth or aggregate two-user throughput. These are informal observations, not a controlled two-user benchmark or a test of two fully filled contexts.

For a controlled single-session comparison, the same prompts were served by Gufo and the existing jcbtc CIRU Orca package on the same host at 32K context, MTP enabled, greedy sampling, and thinking off. Each case had one warm-up and three measured streaming requests with cache_n=0:

Case Prompt / output tokens Gufo median wall CIRU median wall Gufo speedup
Mixed short 371 / 128 5.65 s 6.82 s 1.21ร—
Mixed medium 2,678 / 128 6.52 s 13.48 s 2.07ร—
Mixed long 10,583 / 128 12.20 s 35.28 s 2.89ร—
Repetitive short 369 / 109 Gufo (110 CIRU) 2.49 s 3.62 s 1.45ร—

Prompt token counts matched across runtimes. The repetitive visible output matched, but CIRU counted it as 110 completion tokens. Gufo MTP reproduced Gufo autoregressive visible output on all 12 measured prompts. These results compare complete serving setups: GGUF quantization and runtime kernels both differ, so they do not isolate a format effect. They are not a broad quality evaluation. Raw JSON and the benchmark client are included.

Provenance and license

The source and original base repositories contain identical copies of the Qwen Community License 1.0, reproduced here with the weights. The source model card currently labels its metadata apache-2.0; this release follows the license text distributed with the source weights and original base. Review LICENSE for its terms. This conversion is unofficial and is not endorsed by Qwen, OrcaRouter, jcbtc, or the Gufo project.

Downloads last month
41,273
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF

Quantized
(33)
this model

Space using vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF 1