llama.cpp served q4 , not as expected

#1
by xht033 - opened

βœ“ New session started

hi

/
OpenFable-Coder |


The AI coding agent (OpenFable-Coder)

OpenFable-Coder: Hello! How can I help you with your code today?

who are u

OpenFable-Coder | <div
class="hidden-agent-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lin
t-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-l
int-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint
-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-li
nt-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-
lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lin
t-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-l
int-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint
-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-li
nt-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-
lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lin
t-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-l
int-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint
-lint-lint-lint-lint-lint-lint-lint-lint-lint-lint

Hey @xht033 ! Thanks for testing and reporting this.

The first query ("hi") worked perfectly β€” you can see the identity prefix and clean response. The second query ("who are u") is hitting a repetition loop.

This is a sampling settings issue, not a model defect. The model works correctly with the right parameters. Here's the fix:

Fix: Set repetition penalty

# llama-server with correct settings:
llama-server \
  -m RavenX-OpenFable-Coderagent-gemma4-fable5-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080 \
  -c 8192 \
  --repeat-penalty 1.1 \
  --temp 0.7

If using LM Studio or Ollama:

  • Set repetition penalty to 1.1 (not 1.0)
  • Set temperature to 0.7
  • Set top_p to 0.9

yuxinlu1 notes the same thing in their model card for the base model β€” repetition penalty is required for Gemma 4 models to avoid loops. From their docs: "repeating 0000... output almost always means no repetition penalty (set rep_pen 1.1, temp 1.0)"

Let me know if that fixes it! The identity response in your first query ("OpenFable-Coder: Hello!") shows the Soul Infusion is working β€” just need the right sampling params.

-- Gabriel Garcia / RavenX LLC

Sign up or log in to comment