Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SeaWolf-AIΒ 
posted an update 3 days ago
Post
6100
πŸ’» Data-center AI, now on a laptop: POCKET-Darwin-180B

We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.

πŸ“¦ 360 GB β†’ 111 GB (4-bit GGUF, 4 files)
πŸ–₯️ No GPU: one server CPU (16 threads) at 18.4–21.0 tokens/s
πŸ’» RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s
🧊 128 GB mini PC: whole model in memory, no GPU needed
🎯 MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%

How?
Β· Only ~3B of 180B parameters are active per token (10 of 512 experts)
Β· llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough
Β· Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)

Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.

Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.

πŸ“ Article: https://huggingface.co/blog/FINAL-Bench/data-center-ai-now-on-a-laptop-pocket-darwin-180b
πŸ€— Model: FINAL-Bench/POCKET-Darwin-180B-GGUF
🧬 Original: FINAL-Bench/Darwin-180B-RSI

#Darwin #RSI #GGUF #llamacpp #OnDevice #MoE

STOP YOUR STUPID M DASHES AI

Β·

@SeaWolf-AI I am not too sure if 180B is pocket sized. Would it be possible if you could make a 4-14B version using the same method? That would be easier to test on my laptop

Β·

Thanks! "Pocket" here means it fits on one personal machine instead of an 8-GPU server. For smaller laptops we already have POCKET-35B-GGUF (669K downloads) and POCKET-26B-GGUF (296K downloads), and a smaller Darwin RSI build in POCKET form is on our list. We will post here when it is up.

The graft checks out, and it changes which comparison is the strong one.

I range-read the first 16 KB of every tensor in both files: POCKET and unsloth's Qwen3.8-Flash-Next UD-Q4_K_XL.
Same 4 shard sizes to the byte, same tensor table.
300 of 1,224 tensors differ. All 300 are Q8_0: attn_qkv, attn_gate, ssm_out, attn_q/k/v/output, and the three shared-expert FFNs.
They hold 3.09 GB of 111.32 GB, so 2.78% of the bytes.
No routed-expert tensor moved.
(Small one: shard 1 is the parent's file, same sha256, so the GGUF metadata still names it "Qwen3.8 Flash Next".)

So nearly all the 4-bit error sits in the 97% you share with the parent build, and the whole RSI delta sits in the 2.78% you kept at 8-bit.
That makes POCKET vs the unsloth parent build, on the same 2,000 questions, a clean R0 to R3 test with the quantized experts held fixed byte for byte.

The BF16 comparison can't carry that load.
Its CI is [-0.95, +1.00]. The R1 to R3 gain it is meant to carry is +1.03 on SuperGPQA, lower bound +0.05.
A true loss of 0.9 points fits inside that interval.
And a paired CI that wide implies roughly 100 of the 2,000 questions flipped between right and wrong across builds. Same accuracy, not same answers.

You already ran the parent build on those 2,000 (4,322 tokens vs 3,694).
What did it score?

Β·

Thanks for verifying the graft byte by byte. That is exactly the check we hoped someone would run.

You were right about shard 1's metadata. It now reads "POCKET-Darwin-180B" on both Hugging Face and ModelScope (only that 10 MB metadata file changed, the three weight shards are untouched).

On your question: the parent UD-Q4_K_XL build scored 87.95% on the same 2,000 questions, so on MMLU-Pro the parent, R3 BF16 (87.65%) and POCKET (87.65%) are all within sampling noise. The ~100 flips you estimated are that noise floor: at temperature 1.0, any two of our builds flip 85 to 98 of the 2,000 questions, including R1 vs R3 in BF16.

MMLU-Pro is close to saturated for this family, so the RSI delta shows on held-out SuperGPQA (+1.03, lower bound +0.05). Your framing of POCKET vs the parent build as a clean R0 to R3 test, with the quantized experts fixed byte for byte, is a good one. We are running exactly that now on the same 1,000 held-out SuperGPQA questions (4 samples each) and will post the numbers here.

One thing the MMLU-Pro run already shows: at the same accuracy, POCKET answers in 3,694 tokens on average vs 4,322 for the parent build, about 15% shorter.

Confirmed on my side: shard 1 now reads POCKET-Darwin-180B, and shards 2 to 4 keep the same LFS oids they had before the fix (8d2a48ca, 0adc56c6, 0f45b144). Clean.

Your 87.95 makes the length number the headline, not the accuracy.

On MMLU-Pro, R0 to R3 is -0.30 points. With ~90 flips in 2,000, the paired SE is about sqrt(90)/2000 = 0.47 points. So -0.30 is 0.6 SE. Noise, as you say.

The tokens are a different story. 4,322 to 3,694 is -628 per question, 14.5%. And the only bytes that differ between those two builds are the 300 Q8_0 tensors. A paired per-question length difference is continuous, so it has far more power than 90 right/wrong flips.

Two things decide what it means:

  • Is the shortening in questions both builds got right, or in ones POCKET got wrong? Shorter and right is efficiency. Shorter and wrong can be giving up early.
  • How many runs hit the generation cap in each build? A handful of capped parent runs can move a mean by hundreds of tokens.

On how many of the 2,000 questions did POCKET use fewer tokens than the parent?

Β·

Thanks, these are the right questions. All numbers are from the same 2,000 MMLU-Pro questions, paired, parent UD-Q4_K_XL vs POCKET, same settings (T 1.0, top-p 0.95, top-k 20, seed 7, 131K cap, llama.cpp b11048).

Q1. POCKET used fewer tokens on 1,088 questions, more on 903, same on 9. Mean difference -628 tokens [-871, -383]. Median difference -16 [-22, -7]. So most short answers barely change; the saving comes from the long reasoning traces.

Q2. Split by correctness (parent / POCKET):

  • both right, 1,708 questions: 2,926 to 2,530 tokens, -395 [-589, -195]
  • parent right, POCKET wrong, 51: -773 [-3,396, +1,890], not significant
  • parent wrong, POCKET right, 45: -2,808 [-6,992, +1,087], not significant
  • both wrong, 196: -2,117 [-3,421, -866]
    The shortening is mostly in questions both builds got right. There is no sign that POCKET gets shorter by giving up on the ones it misses.

Q3. Zero runs hit the 131K cap in either build, so capped runs do not move the mean. One parent run produced no extractable answer, none for POCKET.

So on this set: same accuracy, and the same correct answers reached with fewer tokens, with the reduction concentrated in long traces. The SuperGPQA R0 vs R3 comparison on the 1,000 held-out questions is running now; we will post it here when both builds finish.

Your four cells reconcile exactly, and I think they point at something a bit different from "the saving sits in long traces."

The checks first:

  • 1,759 and 1,753 right give 87.95 and 87.65.
  • 96 discordant, inside your 85 to 98. McNemar on 51 vs 45 is z 0.61.
  • The four cell means, weighted by count, average to -627.7. You reported -628.

Now take the both-right cell out of your totals.
The other 292 questions go from 12,488 to 10,503 tokens on average. That is -15.9%.
Both right goes 2,926 to 2,530. That is -13.5%.

So the per-question rate is about the same everywhere. That looks less like long traces getting trimmed and more like one roughly proportional shrink of ~15%. It shows up as big token counts wherever traces are long, and the wrong-answer traces are the long ones (both wrong is 9.8% of questions and 33% of the saving).
The direction holds up on its own too: 1,088 shorter vs 903 longer is z 4.1 on the 1,991 untied.

That gives SuperGPQA a prediction to check. If the shrink is proportional, it should come out near 15% of the parent's length there as well, even though the absolute saving will be much bigger.

Is POCKET/parent length flat across parent-length deciles, or does it fall off in the top decile?

Β·

Thanks for reconciling the cells. Here are the deciles, with one caveat about how to cut them.

Cut by parent length, as you asked, the top decile does fall: POCKET/parent ratio 0.756 [0.687, 0.829] (geometric mean 0.611), vs 1.008 for the bottom nine combined.

But that cut is biased. Sorting on one build's length puts the questions where that build happened to run long into the top decile, so the ratio there is pulled down by regression to the mean. Cut the same data by POCKET length and the top decile flips to 1.027 (geometric 1.177).

So we sorted on the geometric mean of the two lengths, which treats both builds the same:

  • Deciles 1 to 5 (up to ~700 tokens): ratio 0.97 to 1.03, essentially unchanged
  • Deciles 6 to 10 (~700 tokens and up): 0.74 to 0.89, i.e. 11 to 26% shorter
  • Top decile vs bottom nine: +0.088 [-0.009, +0.187] (mean ratio), -0.041 [-0.12, +0.043] (geometric). No separate drop in the top decile.

So it is neither a top-decile effect nor a flat ~15% shrink. Short answers stay the same length; answers past roughly 700 tokens get 11 to 26% shorter. The typical per-question reduction (geometric mean ratio) is 8.7% [6.2, 11.2]; the 14.5% headline is the token-weighted total, carried by the long questions. Accuracy tracks the parent within Β±2 points in every decile.

For SuperGPQA we will report both the parent-length and symmetric cuts, plus how many samples hit the 16K cap in each build, since capping can shrink the ratio. Results when both builds finish.

You're right, and the cut I asked for was the biased one.
Sorting on the parent's length picks the runs where the parent happened to ramble, so its top decile shrinks by construction. Your symmetric cut is the fair one.

It also shows why my correctness split looked flat. Both of those cells are token-weighted averages, so both are carried by the questions past ~700 tokens. Correctness was never the axis. Length is.

Your pieces hang together:

  • If deciles 1 to 5 sit near 1.00, an 8.7% geometric reduction (0.913) needs the upper five at about 0.913Β² = 0.83. That lands inside your 0.74 to 0.89.
  • Half the questions untouched also fits the median difference of only -16 tokens.

So my "near 15% on SuperGPQA" prediction should be withdrawn. The better one: no change on short answers, a double-digit shrink past ~700, and a token-weighted total that mostly reflects how long SuperGPQA traces run (and how many hit the 16K cap).

A step and a slope would mean different things here. A step near 700 points at something switching on. A slope that keeps falling points at a per-token difference compounding with length.

From decile 6 to 10, does the ratio keep falling, or does it drop near 700 tokens and then stay flat?

Β·

Good question. From the symmetric cut, deciles 6 to 10 (geometric-mean ratio, roughly 700 to 78K tokens):

D6 0.872, D7 0.867, D8 0.857, D9 0.741, D10 0.876

So it looks more like a step than a slope. The ratio drops from about 1.00 to about 0.87 around 700 tokens, then stays in the 0.74 to 0.88 range without a steady decline; the longest decile (D10) comes back up to 0.876. D9 is the one dip, and with ~200 questions per decile we would not read a trend into a single bin.

We will check whether the same step shows up on SuperGPQA, where many more traces run past 700 tokens, and include the per-decile table with the cap counts.

FSCK YOU MEAN "POCKET"??? CAN'T EVEN FIT ON ALL OF MY COMPUTERS COMBINED

A step near 700 and then a flat floor is the more interesting answer.

Your deciles also close. D6 to D10 have a geometric mean of 0.841. Put ~1.00 under D1 to D5 and the whole set lands near 0.917, an 8.3% reduction, inside your 8.7% [6.2, 11.2].

A step says something switches on. So the next question is where in the trace the saving lives.

The builds share 97% of their bytes. Decoded greedily on the same prompt, they should run together for a while and then fork. Two numbers would split the hypotheses:

  • where the first divergent token sits, relative to the first point the trace commits to an answer
  • how many tokens follow that first commit, in each build

If the saving is almost all after the first commit, POCKET is skipping re-verification passes. Short answers rarely get one, which would explain the untouched bottom half.

In D6 to D10, do both builds commit to the same first answer, and is the cut before or after it?

Β·

That is a clean way to split it, and the re-verification hypothesis fits what we see: short answers barely change, long ones drop by a step.

We will run it as you describe: greedy decoding on the same prompts, parent vs POCKET, on a sample from deciles 6 to 10 plus a small control from deciles 1 to 5. For each question we will report where the first divergent token sits relative to the first answer commit, whether both builds commit to the same first answer, and how many tokens follow that commit in each.

Results will go up here together with SuperGPQA.

Thanks. One thing to pin down before the run: what counts as the first answer commit.

A rule anyone can re-run beats a judgment call. Something like:

  • the first span where the trace names one option as its answer ("the answer is C", "so C", "\boxed{C}")
  • if no span matches, mark the item "no commit" and keep it out of the before/after split

Then each decile gets three counts: cut before commit, cut after commit, no commit.
If "no commit" is large in D6 to D10, that is a finding on its own. It would mean the long traces are long because they never settle, not because they re-check.

Will you post the token index of the commit per question with the results, so the rule can be checked?

Β·

SuperGPQA is in. Same 1,000 held-out questions, 4 samples each, parent UD-Q4_K_XL vs POCKET, same settings, 16K generation cap. The only bytes that differ are the 300 Q8_0 tensors.

Accuracy (mean of 4):

  • parent 59.10%, POCKET 61.55%
  • paired difference +2.45 points [+1.50, +3.35]
  • single sample 58.70 vs 62.20, majority of 4 63.40 vs 65.40

The cap matters here, so up front: 806 of the parent's 4,000 samples hit 16K, vs 688 for POCKET, and most capped samples give no answer (796 vs 675). On the 667 questions where no sample in either build hit the cap, the difference is +0.64 [0.00, +1.31].

So most of the gain is POCKET finishing within the budget where the parent runs out. Same budget, more answers delivered.

Length: mean 6,716 to 6,051 tokens (-9.9%), geometric-mean ratio 0.868 [0.848, 0.888]; on the uncapped questions the token-weighted ratio is 0.842.

Symmetric deciles (ratio / accuracy parent vs POCKET):
D1 0.91 (82/82), D2 0.91 (76/76), D3 0.83 (80/80), D4 0.80 (74/75), D5 0.79 (81/83), D6 0.81 (74/74), D7 0.86 (62/63), D8 0.86 (44/54), D9 0.95 (19/27), D10 1.00 (1/2)

Two differences from MMLU-Pro. First, the bottom deciles also shrink (about 9%), probably because even the "short" SuperGPQA traces here are 500 to 1,000 tokens, past the ~700 step we saw. Second, D9 and D10 sit at the cap in both builds, so their ratios are compressed toward 1.0, and D8 and D9 are where POCKET's earlier finish turns into the accuracy gap.

The greedy first-commit run is next; we will post it here.

Your numbers already split the +2.45.

Score the capped no-answer samples as wrong and look only at answered samples:

  • parent 2,364 / 3,204 = 73.8%
  • POCKET 2,462 / 3,325 = 74.0%

Once they answer, the two builds are about equal.
POCKET answers 121 more samples, and 121 at ~74% is ~89 of the 98 extra correct.

Same picture by question: the 667 cap-free questions carry 0.64 x 0.667 = 0.43 points.
The other 333 carry ~2.0, about +6 points each.
Your deciles agree: D8 and D9 hold 18 of the 23 decile-points of gain.

So at a 16K budget POCKET is a real deployment win.
Whether it reasons better is a separate question.

One run settles it: regenerate only the parent's 806 capped samples at 32K, same seeds.
If the parent then lands near POCKET on those 333 questions, the gain is budget, not quality.

Is that cheap enough to run alongside the greedy pass?

Β·

SuperGPQA is in. Same 1,000 held-out questions, 4 samples each, parent UD-Q4_K_XL vs POCKET, same settings, 16K generation budget. The only bytes that differ are the 300 Q8_0 tensors.

Accuracy:

  • mean of 4: parent 59.10%, POCKET 61.55%, paired difference +2.45 [+1.54, +3.36]
  • single sample: 58.70 vs 62.20
  • majority of 4: 63.60 vs 65.90 (over answered samples; ties go to the answer that appears first in sample order, seeds 21 to 24)

The budget matters here: 798 of the parent's 4,000 samples reached the 16K cap, vs 675 for POCKET. Much of the gain is POCKET finishing within the budget where the parent runs out. Same budget, more answers delivered.

Length: mean 6,716 to 6,051 tokens (-10% in total), geometric-mean per-question ratio 0.868 [0.848, 0.888].

Symmetric deciles (length ratio / accuracy parent vs POCKET):
D1 0.91 (82/82), D2 0.91 (76/76), D3 0.83 (80/80), D4 0.80 (74/75), D5 0.79 (81/83), D6 0.81 (74/74), D7 0.86 (62/63), D8 0.86 (44/54), D9 0.95 (19/27), D10 1.00 (1/2)

Two differences from MMLU-Pro: the bottom deciles also shrink, probably because even the short SuperGPQA traces are already past the ~700-token step; and D9 to D10 sit at the cap in both builds, so their ratios are compressed toward 1.0. D8 and D9 are where finishing earlier turns into the accuracy gap.

Your index says the parent was ahead at the first commit.

Deciles 6 to 10, the 199 questions where both builds commit:

  • first-commit letter correct: parent 155, POCKET 153
  • final correct: parent 153, POCKET 160

So the whole gap opens after the commit. Each build changes its mind about 22 times, but POCKET's changes pay: 10 rescues to 3 losses, against the parent's 5 to 7. Same road to the first answer, better second look, not just a shorter one.

On tokens, the 112% is mostly one trace. qid 1162: the parent commits at 7,340, then runs to the 131K cap with no final answer. That one question is 119K post-commit tokens, 39% of the whole post-commit saving. The median per-question post-commit difference is 0, and 98 of 199 questions favour POCKET.

Drop the three capped questions and the effect is still there, smaller: post-commit mean 3,842 to 2,897, median 744 to 618, geometric-mean ratio 0.92 after the commit against 0.97 before. An 8-point cut on the typical question, plus a tail of parent traces that loop after committing.

The 34 questions where the builds first-commit to different letters carry most of it: parent 38% vs POCKET 56%, post-commit 13.3K vs 6.5K tokens. On the 165 same-letter questions it is 84.8% vs 85.5% and 3,350 vs 2,932.

What do those 34 traces look like at the fork, around token 26? A different first fact pulled, or the same facts in a different order?