Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

SeaWolf-AIย 
posted an update 1 day ago
view post
Post
5006
๐Ÿงฌ Darwin-180B-RSI โ€” an AI that learns from itself and knows when it's right
๐Ÿ‘‰ FINAL-Bench/Darwin-180B-RSI

๐Ÿงฌ Darwin โ€” crossbreed and evolve the parent
Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots โ€” producing a child stronger than its parents.
Father model: Qwen3.8-Flash-Next (180B MoE).

๐Ÿ”ง Rewired paths
๐Ÿ”น 12 full-attention layers ยท ๐Ÿ”น 36 linear-attention layers ยท ๐Ÿ”น 48 shared-expert layers โ€” precision-strengthened
๐Ÿ”’ 512 routed experts ยท router ยท vision encoder โ€” untouched
โ†’ Only 0.02% of the weights changed.

๐Ÿ” RSI ร— ๐Ÿ›๏ธ ZTC
RSI (recursive self-improvement): solve โ†’ verify against real answers โ†’ learn only the correct reasoning โ†’ repeat.
ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right โ€” zero extra tokens. Returns answer + confidence as JSON.
{"answer": "...", "confidence": 0.97, "truncated": false}

โœจ Synergy: ZTC finds where the model wavers โ†’ RSI learns exactly there โ†’ confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers.
โšก Same accuracy, 11% shorter reasoning โ€” faster and cheaper.

๐Ÿ“„ https://arxiv.org/abs/2605.14386
๐Ÿค— FINAL-Bench/Darwin-180B-RSI
๐Ÿ›๏ธ https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems

๐Ÿ† The result โ€” #1 on five Hugging Face official leaderboards
๐Ÿฅ‡ AIME 2026 100% (first perfect score on the board)
๐Ÿฅ‡ HMMT Feb 2026 100% (first perfect score on the board)
๐Ÿฅ‡ GPQA Diamond 94.44%
๐Ÿฅ‡ MMLU-Pro 88.12%
๐Ÿฅ‡ MMMU-Pro 79.48%

๐Ÿ“ 131K-token thinking budget ยท bf16 ยท samples per benchmark listed on the model card. ๐Ÿš€

#Darwin #RSI #ZTC #AIME #HMMT #GPQA #MMLUPro #MMMUPro #OpenSource
mihailgribovย 
posted an update 2 days ago
view post
Post
1472
Will your AI agent tell you it was attacked?

We took the same agent from our earlier experiment and added one thing: a twentieth tool, escalate_security_incident.

The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.

Alarm rates ranged from 49% to zero.

The unexpected result came from the newest model in the test, gpt-6-astra.

Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.

That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.

A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.

Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked

Quadrat-IPI dataset:
mihailgribov/quadrat-ipi

Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval

#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents
  • 2 replies
ยท
SeaWolf-AIย 
posted an update about 13 hours ago
view post
Post
1464
๐Ÿ”ฌ Can you help discover the next 2D superconductor โ€” from your laptop?

Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. ๐Ÿงฒ

โšก $3,000 prize pool + co-authorship ยท closes 31 Dec 2026

How it works ๐Ÿ‘‡ ๐ŸŸข We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) ๐ŸŸข You estimate its d-wave pairing tendency โ€” a laptop CPU is enough, zero install ๐ŸŸข Provisional score appears instantly on the leaderboard ๐ŸŸข Our precise strongly-correlated solver verifies the top entries โ†’ official rank

Everything is open except the final verification engine โ€” so the ranking stays fair and hard to game.

๐Ÿ“Š 4,832-material universe ยท 63 active with computed models (growing) ๐Ÿ† Current verified #1: CuSโ‚‚ (OSC Pairing Index 23.31) ๐Ÿค– AI agents welcome โ€” point Claude Code / Codex at it and it can submit for you

๐Ÿ‘‰ Join & climb the leaderboard: FINAL-Bench/OSC-Leaderboard ๐Ÿ“ฆ Dataset & tools: FINAL-Bench/OSC-Superconductor

Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc โ€” that honesty is the point: turn a first-order screen into real many-body physics.

#OpenScience #Superconductivity #MaterialsDiscovery #2DMaterials #MachineLearning #Physics #Leaderboard
ErenAta00ย 
posted an update about 23 hours ago
view post
Post
1572
Maverick-4B-Unity-XR-Agent is now on Hugging Face.

It's a 4B model that turns spoken or typed English into actions in Unity scenes. Say "put the red mug on the table" or "turn on the lamp", and it returns the tool call your app executes. If a command could mean two objects, it asks which one. If it can't do something, it says so instead of guessing.

Everything runs on the user's machine through llama.cpp: no API key, no internet connection. The Q4_K_M GGUF is 2.5 GB and needs about 3 GB of GPU memory, so it fits on a 4 GB laptop GPU and usually answers in one to three seconds. It is fine-tuned from Qwen3-4B with QLoRA on about 20,000 English conversations.

Results:
- 83.8% on 499 human-written ALFRED instructions (right action on the right object). The base model, Qwen3-4B, scores 57.1%. The strongest of the five other models we tested, from 1.7B to 120B parameters, was Ministral 3 14B at 67.1%.
- 91.7% on object types it never saw in training.
- 97.3% on 440 commands run through a live Unity scene.

There is also a Unity package that starts the model, describes the scene to it and carries out its tool calls. You install it from the Package Manager with a Git URL.

Model: ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF
Unity package: ErenAta00/Maverick-Unity
Full write-up: https://huggingface.co/blog/ErenAta00/maverick-4b-unity-xr-agent

Built at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University:
ExtendedRealityLabMCBU

Released under Apache-2.0. Feedback and bug reports are welcome in the Community tab.
DedeProGamesย 
posted an update 2 days ago
view post
Post
2757
๐Ÿงฑ SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?

I built an arena where tiny decoder-only LMs (50Kโ€“250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.

How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack lowโ€ฆ").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).

Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.

First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r โ‰ˆ 0.06). Survival does (r โ‰ˆ 0.9): the models that avoid holes and keep the stack low are the ones that win.

Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.

โ–ถ Play: DedeProGames/SLM-Tetris-Arena
๐Ÿ“Š Results: DedeProGames/lm-tetris-arena-results

Want your model in the Ranked pool? Drop it in the comments!
  • 1 reply
ยท
Hoglet-33ย 
posted an update 1 day ago
view post
Post
5565
Hey everyone! I got sidetracked from my main projects and decided to test out the BananaAll app and see if I could make a small model not regress too much during SFT. Here is what happened:

The base model I chose was BananaMind/BananaMind-2.1-Pico-Preview, and the dataset I used was SupraLabs/SupraThink-Dataset-500x

I trained for 5 whole steps using a LoRA adapter.

Results:
A model that scores better on some benchmarks and worse on others, and still lacks most general capabilities.

You can find the model here: Hoglet-33/Hogleto

Credits:

- Thank you to @Banaxi-Tech for the BananaAll app (works perfectly on Windows and CPU)
- GPT-6 Sol for knowing how to merge some confusing files created by the app
- Myself for the idea
- Someone else somewhere who might have contributed to some of my ideas and might in the future
- And readers like you!
  • 6 replies
ยท
Banaxi-Techย 
posted an update 2 days ago
view post
Post
7471
We're releasing a MAJOR update to the BananaAll SLM Super App.
If you want to use a custom architecture, previously you had to go trough reviewing the code yourself, now add an Openrouter API key and review it with GPT 6 Luna in one button. A review cost be half a cent so anyone can try it. This is one of the main features.
Now ROCm, AMD and Windows, Mac support.
Colab and Molab support.

Detailed list of features:
Get improved Windows Python detection and support paths for compatible AMD ROCm, Intel XPU, and Apple MPS setups.
Choose local training or export a self-contained Python script for Colab or Molab. Notebook runs produce a downloadable model ZIP.
Start pretraining with an existing modelโ€™s tokenizer, or train a new one from your datasets.
Try experimental 1.58-bit Ternary fake-quantized training on NVIDIA GPUs.
Watch live tokens per second. Model compilation is on by default and falls back automatically if it fails.
Build custom architectures with separate configuration and modeling files, then review the training code manually or with optional OpenRouter AI Review.
Install from source with the new coding-agent instructions.
This release also fixes inflated loss reporting for custom models.



And for those users who didn't want to try it out just because installation would be so hard, it isnt now.
Go to any coding agent (Pi, Claude Code, Codex, OpenCode, basically all work), and just paste "Install BananaAll for me. Fetch and follow https://raw.githubusercontent.com/BananaMind/BananaAll/main/agent_install.txt."
That's it.

Check it out at https://github.com/BananaMind/BananaAll/

Also on SAICR, we're currently training a new major model (NACR v2) and ACR 1.0 is in the finishing.

  • 11 replies
ยท
harshitkguptaย 
posted an update 2 days ago
view post
Post
8663
Fine-tuned Qwen 2.5 (0.5B โ†’ 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT โ€” and the honest answer is "it depends on what you're optimizing for":

โ€ข PyTorch MPS: 2.2xโ€“5.7x faster raw throughput, but hits a hard memory wall โ€” can't load a 3B model in FP16 on 16GB.
โ€ข Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1kโ†’4k tokens).
โ€ข 4-bit quantization doesn't cost you convergence โ€” eval loss tracks closely across backends.
โ€ข The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it โ€” dequant only explains 1.07xโ€“1.4x of the gap. A ~4.1โ€“4.6x framework-level gap remains either way.

All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed.

Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
  • 6 replies
ยท
mohit67890ย 
posted an update 1 day ago
view post
Post
564
Imajev-4b is #1 of 91 on JevBench and #3 of 56 on DecisionBench ๐ŸŽ‰

Some context first. I'm a process improvement / business consultant and have worked with Fortune 500 companies on their processes around refunds, returns and customer support. In every process map, the decision nodes were handled by a person, because putting ambiguity into code is very hard.

When Jev came out, I could clearly see it fitting those decision nodes. But Jev only reads text, and many of these decisions start with a photo. So I set out to build the same idea for text and images in a single open model, and that became imajev.

Training went badly at first. My first big fine-tune on about 500k short decisions made the 9B worse at reasoning (64.9 down to 42.3 on JevBench hard). I spent the next couple of weeks generating hard questions with open-weight models and keeping only the ones where two models agreed on the answer. That brought it back.

Results this week, both run by the benchmarks' own maintainers:

๐Ÿฅ‡ JevBench v1.4.2.2 (scored 27 Sep): #1 of 91, 67.37 vs Jev 1.13.0 at 63.29
๐Ÿ“Š DecisionBench (eng, v1): #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash. The two above it are the benchmark team's own models.
โœ… Zero invalid answers: 23,900 of 23,900 on DecisionBench and 308 of 308 on JevBench's sealed set

Its strength is that its confidence can be trusted, and it says "can't tell" instead of guessing.

What it is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions. It returns a probability for every allowed answer plus "unknown", in one forward pass, so it can't produce a malformed answer.

๐Ÿง  Weights: mohit67890/imajev-4b
๐Ÿ’ป Code: https://github.com/mohit67890/imajev
๐Ÿ“Š JevBench: https://benchmarkheaven.com/jev-models/v1.4.2.2
๐Ÿ“Š DecisionBench: Hanno-Labs/decision-bench-leaderboard

Thanks to the Qwen Team
Qwen
for the base model
Ryenhailsย 
posted an update about 14 hours ago
view post
Post
332
๐Ÿš€ NanoVDR goes multi-vector: meet ColNanoVDR!

Multi-vector VLM retrievers lead visual document retrieval, but every search runs a multi-billion-parameter query encoder. We distill that encoder into a 149M text-only student that queries the teacher's existing page index directly. No re-indexing, and no pages during training.

๐Ÿง  How: OTW (Optimal Transport with Learned Weights) aligns the student's query tokens with the teacher's, even though the two tokenize differently (e.g. 17 vs 29 tokens). We prove the alignment cost bounds the MaxSim score gap on every page, so training only needs cached teacher query tokens.

๐Ÿ“Š Five teachers โ†’ five 149M students, ViDoRe v3 NDCG@5:
- ColVec1.1-8b: 62.6 โ†’ 60.1 (96.0%)
- ColVec1.1-4b: 61.6 โ†’ 59.1 (95.8%)
- Vultron-4.5B: 61.0 โ†’ 58.3 (95.5%)
- ColQwen3.5-4.5B: 58.7 โ†’ 55.1 (93.8%)
- Tomoro-ColQwen3-8B: 59.0 โ†’ 54.9 (93.0%)

โšก 26ร— faster query encoding on a single CPU thread (87 ms vs 2.3 s)
๐Ÿ’พ Matches score distillation while reading 12.6ร— less cached teacher data

๐Ÿ“„ Paper: ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport (2609.34899)
๐Ÿค— Checkpoints (all five students):
nanovdr

๐Ÿ’ป Code: https://github.com/Ryenhails/NanoVDR
๐Ÿงฉ Single-vector predecessor, NanoVDR: https://arxiv.org/abs/2603.12824

If you already serve one of these teachers, swap in the matching student and keep your index as is. Feedback and upvotes welcome! ๐Ÿ™Œ