Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

SeaWolf-AIΒ 
posted an update about 24 hours ago
view post
Post
2776
πŸ’» Data-center AI, now on a laptop: POCKET-Darwin-180B

We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.

πŸ“¦ 360 GB β†’ 111 GB (4-bit GGUF, 4 files)
πŸ–₯️ No GPU: one server CPU (16 threads) at 18.4–21.0 tokens/s
πŸ’» RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s
🧊 128 GB mini PC: whole model in memory, no GPU needed
🎯 MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%

How?
Β· Only ~3B of 180B parameters are active per token (10 of 512 experts)
Β· llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough
Β· Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)

Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.

Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.

πŸ“ Article: https://huggingface.co/blog/FINAL-Bench/data-center-ai-now-on-a-laptop-pocket-darwin-180b
πŸ€— Model: FINAL-Bench/POCKET-Darwin-180B-GGUF
🧬 Original: FINAL-Bench/Darwin-180B-RSI

#Darwin #RSI #GGUF #llamacpp #OnDevice #MoE
  • 19 replies
Β·
cafkafkΒ 
posted an update 2 days ago
view post
Post
4337
I'm working on a local model that can be run on stuff like an an RTX 3080 at 60 t/s. I'm focused on making it good at Rust and Nix. It's... not done yet, but doing great already. Had a 28/100 on livebench v6 earlier, which obviously isn't amazing.

But for my first real model, very exciting. Can't wait to share more as I get closer to the finish line in the coming weeks.
  • 6 replies
Β·
pollixΒ 
posted an update 3 days ago
view post
Post
6350
stuntd 0.1.2 is out πŸŽ‰

stuntd sits in front of your LLM, learns its typed decisions and answers the confident ones locally with a small head on the Laya encoder by @convaiinnovations . About 20ms on GPU and 60ms on CPU, and anything it isn't sure about still goes to the big model.

New in 0.1.2:
- decisions with several fields, like category + urgency + needs_human in one call, answered locally only when every field is sure
- the Anthropic Messages API learns too, not only OpenAI
- auto_retrain: the daemon retrains a site in the background once enough new traffic comes in, so collect, train, shadow and live run on their own
- serve --lazy loads the checkpoint on the first request

Try it in the browser: pollix/stuntd
Code: https://github.com/bladedevoff/stuntd

pip install -U stuntd
  • 1 reply
Β·
mohit67890Β 
posted an update 1 day ago
view post
Post
4060
πŸ₯‡ Imajev-4b is #1 of 50 on Image JevBench (v0.1.4, 29 Sep 2026), the leaderboard for AI models that make decisions from images.

A small demo built on imajev-4b: a closet stylist πŸ‘—
Request you to star it here so we can make it better - https://github.com/mohit67890/imajev.

Tap a piece and it reads the photo (red 75%, checked 99%), then the app picks bottoms, shoes and a bag from your own closet in the colours you like. Change your colours and the outfit changes.

Under the hood it's one request with one photo and 4 typed questions. Every option gets a probability, so the app applies its rules (one pattern per outfit) and ranks what's left. About 1.1 s per outfit on a Mac (MLX). Every % in the video is the model's real answer.

Also, thanks to @zenmagnets for the FP8 version for Blackwell GPUs: same calibration, 0.1 pt less accuracy on all 23,900 DecisionBench rows, ~1.6Γ— faster πŸ™
zenmagnets/Imajev-4B-FP8-SM120

🧠 Weights: mohit67890/imajev-4b
πŸš€ Demo: mohit67890/imajev
πŸ“Š Leaderboard: https://benchmarkheaven.com/image-jev-bench
  • 1 reply
Β·
Banaxi-TechΒ 
posted an update 1 day ago
view post
Post
4026
ACR 1.0 launch is being prepared and researched now!
Also I'm going to vacation tomorrow but it should still be released!

saicr
  • 1 reply
Β·
pollixΒ 
posted an update 1 day ago
view post
Post
3858
First stuntd model is on the Hub :)

pollix/stuntd-support-triage is three small heads on the Laya encoder that triage a support ticket in one request: category, urgency and needs_human. About 50 MB each, all three answers come back at a p50 of 71ms through the daemon.

On 1,000 tickets they never saw, each head answers on its own when it's sure: category 99.9%, needs_human 92%, urgency 76%. A ticket only skips the big model when all three are sure, that's 72.7% of them, and all three are right on 97.1% of those.

It's the support demo from the repo, so the tickets are generated and the teacher is a rule. The point is to show what a head looks like and how fast it is, then you train the same thing on your own traffic with your own LLM as the teacher.

hf download pollix/stuntd-support-triage --local-dir support-heads


Model: pollix/stuntd-support-triage
Everything in one place: pollix/stuntd-6abe0a33303828e10c72ab41
Code: https://github.com/bladedevoff/stuntd
  • 3 replies
Β·
DedeProGamesΒ 
posted an update 3 days ago
view post
Post
6226
🧱 SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?

I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.

How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).

Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.

First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r β‰ˆ 0.06). Survival does (r β‰ˆ 0.9): the models that avoid holes and keep the stack low are the ones that win.

Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.

β–Ά Play: DedeProGames/SLM-Tetris-Arena
πŸ“Š Results: DedeProGames/lm-tetris-arena-results

Want your model in the Ranked pool? Drop it in the comments!
  • 1 reply
Β·
DedeProGamesΒ 
posted an update about 11 hours ago
view post
Post
192
how is this possible
  • 5 replies
Β·
SeaWolf-AIΒ 
posted an update 4 days ago
view post
Post
5464
πŸ”¬ Can you help discover the next 2D superconductor β€” from your laptop?

Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. 🧲

⚑ $3,000 prize pool + co-authorship · closes 31 Dec 2026

How it works πŸ‘‡ 🟒 We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) 🟒 You estimate its d-wave pairing tendency β€” a laptop CPU is enough, zero install 🟒 Provisional score appears instantly on the leaderboard 🟒 Our precise strongly-correlated solver verifies the top entries β†’ official rank

Everything is open except the final verification engine β€” so the ranking stays fair and hard to game.

πŸ“Š 4,832-material universe Β· 63 active with computed models (growing) πŸ† Current verified #1: CuSβ‚‚ (OSC Pairing Index 23.31) πŸ€– AI agents welcome β€” point Claude Code / Codex at it and it can submit for you

πŸ‘‰ Join & climb the leaderboard: FINAL-Bench/OSC-Leaderboard πŸ“¦ Dataset & tools: FINAL-Bench/OSC-Superconductor

Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc β€” that honesty is the point: turn a first-order screen into real many-body physics.

#OpenScience #Superconductivity #MaterialsDiscovery #2DMaterials #MachineLearning #Physics #Leaderboard
DavidAUΒ 
posted an update 16 days ago
view post
Post
15686
Qwen 3.5 9B - The Defiant, 27B power ; now with Qwen 3.8 Reasoning modes.

640 ARC-C for both 8bit and 4bit. Model exceeds 7 of 7 benchmarks for Qwen 3.5 9B, Qwen3.5 27B, Qwen3.6 35B-A3B, and meets Qwen 3.6 27B in some cases... and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads)

NEW - Qwen 3.8 Reasoning Modes: 2 MTP quants (Q6/Q8) Now with 5 reasoning modes (2 new - Spoon / Einstein), and 5 instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model control at the chat/message level). Model name has "plusIQ" in the name.

(there is also a extra robust "tools" version too.)

DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

PS: I have posted a 22k output example using "spoon" mode too.

UPDATE:
This model (and a few more) will soon have 12 reasoning and 12 instruct modes plus interactive help, model embedded system to select the best reasoning/instruct mode[s] for all use cases.

Final testing is in progress...
  • 6 replies
Β·