Instructions to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Use Docker
docker model run hf.co/vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
- LM Studio
- Jan
- vLLM
How to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
- Ollama
How to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with Ollama:
ollama run hf.co/vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
- Unsloth Desktop
- Pi
How to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with Docker Model Runner:
docker model run hf.co/vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
- Lemonade
How to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF-BF16
List all available models
lemonade list
- Hermes Agent
How to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8 Flash-Next Uncensored for Gufo โ Q4 Mix
A Gufo-compatible GGUF conversion of orcarouter/Qwen3.8-Flash-Next-Uncensored. This is a quantization of that checkpoint, not a new fine-tune. It was built and tested on AMD Strix Halo (gfx1151) with Gufo commit eb91584.
Q4 Mix is this release's shorthand for its mixed tensor recipe. It is not a stock Q4_K_M preset: routed expert gate/up weights are Q4_K, expert down weights are Q5_1, the large per-layer embedding table is IQ4_NL, and the token embeddings, output head, and most dense projections are Q8_0. The quality-sensitive output head remains at higher precision. The exact recipe and build notes are included.
Files
| File | Purpose | Size | SHA-256 |
|---|---|---|---|
Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix.gguf |
Text target with embedded PLE weights | 109,776,678,624 bytes (102.23 GiB) | c7f2331227aaefd1947b15fd6a31adeb29a4609a4dec5713670743b61ec44b7a |
mmproj-BF16.gguf |
Matching BF16 vision projector | 907,543,360 bytes | 0c2760d6f3e6686500989ad6ccc28aab69abfb272a3419ac33d4e26e7708ce8a |
The optional Q8_0 MTP predictor is not included. The tested sidecar is mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf from jcbtc's CIRU Orca release. Gufo also runs this target without MTP.
Run with Gufo
Build or install Gufo for Strix Halo, then download this repository's GGUF files. This release is public and ungated. No Hugging Face account or access request is required to download these files.
hf download vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF \
--include '*.gguf' --local-dir ./orca-gufo
gufo serve --host 127.0.0.1 --port 8080 --sessions 2 llm \
--model ./orca-gufo/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix.gguf \
--mmproj ./orca-gufo/mmproj-BF16.gguf \
--served-model-name orca-gufo --context 262144
For speculative MTP, download the tested sidecar and add the following options to the gufo serve ... llm command:
hf download jcbtc/Qwen3.8-Flash-CIRU-STRIX-Orca \
--include 'mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf' \
--local-dir ./ciru-sidecar
# Add to gufo serve ... llm:
--speculative mtp \
--mtp-model ./ciru-sidecar/mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf \
--draft-tokens 6
The API is OpenAI-compatible at http://127.0.0.1:8080/v1; send model ID orca-gufo. Gufo's Qwen formatter is compiled into the runtime and does not load an external Jinja template. Thinking and effort can be controlled per request with reasoning_effort; --think off is a server default if desired. See Gufo's Flash-Next guide and server API documentation.
The repository contains a BF16 vision projector, but image input was tested only as a simple controlled smoke check. Generic GGUF runners and other GPUs have not been qualified for this particular package.
What was tested
The target loaded with Gufo on a Ryzen AI MAX+ 395 / Radeon 8060S (gfx1151). Text, a required function tool call, and a simple image question passed. The target also loaded with the linked MTP sidecar. A 262,144-token context configuration was tested with one session and short text/tool/image requests. The operator subsequently ran two concurrent 262,144-token text sessions on bare-bones Linux and reported roughly 40 tokens/s during long-context development-harness use. A live last-request /metrics gauge read 40.31 tokens/s, but it did not record prompt depth or aggregate two-user throughput. These are informal observations, not a controlled two-user benchmark or a test of two fully filled contexts.
For a controlled single-session comparison, the same prompts were served by Gufo and the existing jcbtc CIRU Orca package on the same host at 32K context, MTP enabled, greedy sampling, and thinking off. Each case had one warm-up and three measured streaming requests with cache_n=0:
| Case | Prompt / output tokens | Gufo median wall | CIRU median wall | Gufo speedup |
|---|---|---|---|---|
| Mixed short | 371 / 128 | 5.65 s | 6.82 s | 1.21ร |
| Mixed medium | 2,678 / 128 | 6.52 s | 13.48 s | 2.07ร |
| Mixed long | 10,583 / 128 | 12.20 s | 35.28 s | 2.89ร |
| Repetitive short | 369 / 109 Gufo (110 CIRU) | 2.49 s | 3.62 s | 1.45ร |
Prompt token counts matched across runtimes. The repetitive visible output matched, but CIRU counted it as 110 completion tokens. Gufo MTP reproduced Gufo autoregressive visible output on all 12 measured prompts. These results compare complete serving setups: GGUF quantization and runtime kernels both differ, so they do not isolate a format effect. They are not a broad quality evaluation. Raw JSON and the benchmark client are included.
Provenance and license
- BF16 source: orcarouter/Qwen3.8-Flash-Next-Uncensored, revision
8336e613. - Original base: Qwen/Qwen3.8-Flash-Next.
- Converter and quantizer: llama.cpp, local build at commit
d3e63dbfcfdac54bc7925a8ec5f44057035ac127. - Inference engine: Gufo, tested at commit
eb915840ffb62a8ec4b5c1adb41b04b5c1c75892. - Optional MTP sidecar: jcbtc/Qwen3.8-Flash-CIRU-STRIX-Orca; credit and download remain with its publisher.
The source and original base repositories contain identical copies of the Qwen Community License 1.0, reproduced here with the weights. The source model card currently labels its metadata apache-2.0; this release follows the license text distributed with the source weights and original base. Review LICENSE for its terms. This conversion is unofficial and is not endorsed by Qwen, OrcaRouter, jcbtc, or the Gufo project.
- Downloads last month
- 41,273
We're not able to determine the quantization variants.
Model tree for vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF
Base model
Qwen/Qwen3.8-Flash-Next