English | Español
C99 LLM runtime for low-RAM machines — GGUF → G2BX (own format) → inference. Weights are memory-mapped: your real RAM budget is KV cache + activations + tokenizer, not the model file.
Runs a 3B model on 145 MB of RAM and a large model on a 2 GB machine (
--swap: KV backed on disk, footprint ≈ 37 MB).
Measured on Intel i5-6200U · 2C/4T · DDR3L dual-channel (~9.4 GB/s bus ceiling).
Sustained decode with --fast (high priority + OpenMP + quantized KV).
| Model | Weights (mmap) | Runtime RAM | decode | prefill |
|---|---|---|---|---|
| Qwen2.5-3B Q4_0 | 1992 MB | 145 MB | 4.3 | 7.8 |
| Qwen3-0.6B Q4_0 | 319 MB | 511 MB | 25.0 | 47.9 |
| LFM2.5-1.2B Q4_0S | 567 MB | 631 MB | 15.7 | 17.9 |
| Llama-3.2-1B F16 | 804 MB | 644 MB | 13.8 | — |
| SmolLM2-135M Q4_0 | 72 MB | 40 MB | 59.5 | — |
Measured 2026-08-31 on i5-6200U with bench -n 32 (min3). At ~25 tok/s decode
pinned to DDR3L ~9 GB/s; prefill hits compute ~27 GMAC/s. (--mv speedups
removed in v5.1.0: the skip destroyed ppl ×1400 — speed of a broken model.)
make # Linux / MinGW · OpenMP + AVX2 recommended
make test # smoke test on a synthetic model
# Pack a GGUF once, run forever
./gguf2bin2 pack Qwen3-0.6B-Q8_0.gguf qwen.g2bx --q4
./gguf2bin2 chat qwen.g2bx --no-think --fast--threads N picks OpenMP threads; physical cores is the sweet spot.
| Knob | Effect |
|---|---|
--q8-kv |
KV cache F32→Q8_0: ~3.8× less RAM (94 % greedy argmax agreement) |
-c N |
Context sizes to your session, not the model's 262k max |
--max-ram MB |
Auto: enables Q8 KV first, then halves context until it fits |
--swap [PATH] |
KV cache backed on disk → 37 MB heap even for big models |
$ gguf2bin2 run qwen.g2bx "Hello" --max-ram 2048 # ideal for 2 GB machines| Component | Counts? | Fix |
|---|---|---|
| Weights (mmap page cache) | ❌ evictable | — |
| KV cache | ✅ | --q8-kv, -c N, --swap |
| Buffers / activations | ✅ small | — |
| Tokenizer (~250k vocab) | ✅ ~30–60 MB | — |
gguf2bin2 pack model.gguf out.g2bx [--q4] # GGUF → G2BX (--q4: half the bytes)
gguf2bin2 info model.g2bx # slots, geometry, types
gguf2bin2 verify model.g2bx # header + CRC + geometry, no weights load
gguf2bin2 run m.g2bx "prompt" [-n N] [-t T] [--bos] [--gpu]
gguf2bin2 chat m.g2bx [--no-think] [--fast] [--swap]
gguf2bin2 bench m.g2bx [-n 32] [--prefill 256]
gguf2bin2 ppl m.g2bx -f text.txt # quality harness (perplexity)
gguf2bin2 vkinfo # Vulkan probeSampling: quickselect top-k O(n) + Gumbel-max + xorshift64* reproducible via --seed.
The head GEMV (vocab×dim, the heaviest layer) gets split between CPU and GPU with automatic calibration:
[gpu] worker ready
[gpu] dual band: cpu=[0..44855) gpu=[44855..65536) tc=6.1ms tg=13.1ms
- Vulkan worker in a child process — if the driver crashes or hangs, the runtime falls back to CPU-only without interrupting generation.
- Vulkan loader bypass: loads the ICD directly from DriverStore (useful on systems with a broken
Khronos\Vulkan\Driversregistry). - Automatic optimal split:
gpu = vocab·tc/(tc+tg)measured on the first token; if the GPU is >4× slower than the CPU it shuts itself off. - Supports Q4_0 and Q4_0S heads. Bit-identical output vs the CPU path (greedy).
⚠️ On hardware where the iGPU shares the RAM bus with the CPU (HD 520 + DDR3L) there is no net gain — calibration detects it and disables itself. The real payoff comes with a dGPU with dedicated VRAM.
Perplexity (internal corpus, SmolLM2-135M-Instruct): uniform Q4_0 = 73.7 (broken), native Q4_K_M = 48.7, Q6_K = 48.0, Q8_0 = 48.0. Native K-quants keep Q8_0-level quality within 1.4 % while running faster than uniform Q4_0.
Measured 2026-09-05 in separate CLI processes, using the pre-edit README+README.es+G2BX_SPEC corpus. These are short local checks, not a standard benchmark; model provenance and quantization paths need further validation.
| Model / format | ppl | vs base |
|---|---|---|
| Qwen3-0.6B Q4_0 | 25.4 | — |
| Qwen3-0.6B Q8_0 | 20.8 | −18 % |
| LFM2.5-1.2B q4max | 28.1 | — |
| LFM2.5-1.2B q4s | 34.9 | +24 % |
| Llama-3.2-1B baseline (256 tokens) | 83.966 | — |
| Llama-3.2-1B Q4_VVC (256 tokens) | 1177.176 | ×14 |
| Llama-3.2-1B baseline (128 tokens) | 82.325 | — |
| Llama-3.2-1B Q4_0S_PSY (128 tokens) | 2380555.838 | ×28916 |
Qwen and LFM2 rows use 1024 tokens. LFM2 Q4_0S increased ppl by about 24%. The tested PSY/VVC artifacts showed severe degradation. For PSY a root cause was found in v5.1.1: both AVX2 kernels read the nibbles 2 bytes off, so the PSY row above measures a broken kernel, not the format; it needs re-measuring. VVC is a plain 3-bit uniform quant with one scale per 256 (no inter-row prediction is implemented). Both remain not recommended until re-validated. The types at 0x80+ are custom G2BX types, not standard GGUF types supported by other runtimes.
Included numerical validation: make kvtest (F32 vs Q8 KV), tools/prefilltest
(bit-exact batched prefill), tools/qkcheck (46/46 K-quant kernels).
Removed in v5.1.0 (measured, see docs/ROADMAP_PERF.md Phase 7): --mv and
--bvh (large ppl increases at the tested ratios), cyber-* commands (reported accuracy and particle-loop loss were
not real — experiment archived in experimental/cyber-mrna/), dead
OrderBook/HDR/ZRAM/FM-index paths. --cyber <lora.bin> still loads real
LoRA adapters (v1/v2).
The removal of mv_ratio/bvh_keep changes the public g2b_config layout:
recompile API clients against the new header; do not mix old headers and new libraries.
📦 G2BX format & supported types
G2BX | ver:u16 | arch:u8 | flags:u8 | ModelCfg | n_slots:u32 | Slot[] | data[] 64B-aligned | tokenizer | [v3: crc32 + "G2BX"]
Slot: role:u8 layer:u16 type:u8 nbytes:u32 off:u64
Full spec: docs/G2BX_SPEC.md (v1/v2/v3 compat matrix; own types live at 0x80+ since v3).
| Type | Load | Fused matmul |
|---|---|---|
| F32 / F16 | yes (pack→Q4_0 if weight) | F32 path / dequant |
| Q4_0 / Q4_1 / Q5_0 | yes | AVX2 fused |
| Q8_0 | yes | AVX2 fused |
| Q4_0S (own, fp16 shared scale /256) | yes | AVX2 batched |
| Q2_K … Q6_K | yes (dequant) | AVX2 fused (Q4_K/Q6_K), scalar rest |
| IQ* / Q8_K | partial | — |
🗂 Project layout
include/gguf2bin.h public API (sessions) · src/internal/ shared internals
src/model.c G2BX load/free, geometry, RAM budget, synth
src/kv.c F32/Q8 KV cache, disk swap, runtime alloc, TLS
src/forward_*.c dense (+dispatch) / lfm2 / qwen35-hybrid / batched prefill
src/l1_gguf.c GGUF parser (mmap) · src/l4_gbin.c G2BX packer
src/l2_codec.c fused dequant + matmul (AVX2) · src/l3_math.c norms/rope/softmax
src/l6_token.c BPE tokenizer · src/l7_vulkan.c dual-band GPU (child process)
src/g2b_api.c API implementation · src/os_mm.c mmap/files · src/sampler.c
src/opts.c CLI flags · src/g2bx_io.c G2BX reader/writer/CRC/verify
src/main.c CLI · shaders/ Q4_0/Q4_0S GEMV · tools/ harnesses + fuzz/
docs/G2BX_SPEC.md format spec · docs/RESEARCH.md notes + roadmap
📜 Changelog
- CI green again (red since Phase 3):
strdup/fseeko/ftello/clock_gettimewere implicitly declared under-std=c99— on Linux x86-64 the truncatedstrduppointer crashedmake test, andftellotruncated offsets >2 GB. Also fixed: fuzz job YAML (>folded the clang lines apart),fmemopenin the harnesses, MinGWcopyunder MSYS2sh, CMake include dirs, AVX2 flags in the sanitizer job. - Tokenizer:
u2b[289]overflowed (68 remapped bytes → indices up to 323); bytes 0x7F–0xA0/0xAD decoded wrong (€, à, emojis).tok_read_sectionleaked on early errors. - Q4_0S_PSY kernels (decode + batched) read nibbles 2 bytes off.
- qwen35: attention
wousedn=diminstead ofn_heads*head_dim. LFM2/qwen35 recurrent state was never reset atpos 0(ppl windows,chat_reset, compaction and the Android app inherited the previous sequence). - LoRA: batched prefill skipped the adapter (now falls back to sequential); loader validates rank and reads.
- pack --prune:
ffn_downcopy assumed one block per group (broken for F16/F32/Q8_0 down); OOM mid-prune now aborts. Tensors >4 GB are rejected instead of truncatingSlot.nbytes. - Untrusted .g2bx hardening (the Android app downloads from any URL): slot
nbytesvalidated against geometry, geometry caps against i32 overflow, blob-past-EOF and offset-overflow checks, non-mmap fallback read from the right offset. - Android JNI:
freeModelcould free the model whilegeneratestill ran (use-after-free); tokens are emitted as whole UTF-8 characters via UTF-16 (NewStringUTFbroke on split characters and emojis). - Chat (LFM2): the empty
<think></think>block (Qwen3enable_thinking=Falseconvention) was injected into LFM2.5 too, whose template has no such block; the model opened every answer "correcting itself". Found and verified on the releaselfm25-1.2b-q4s.g2bx(now answers "Paris" / "Madrid"). - ppl: windows after the first now restart with BOS, like llama.cpp (LFM2 at
-c 128: 485 → 109; single-window results unchanged). g2b_packno longer leaks the Q4_0S/PSY/VVC mode into later calls; chat prompts are no longer truncated at 4/9 KB; default--swapfile is per-process and opened withO_NOFOLLOW.
- CRC32 footer: every new
.g2bxends with[crc32 of everything before][magic]; the loader verifies on open (warming the page cache as a side effect) and rejects truncated/corrupt files with a clear message. Newverifycommand (header + slots + types + geometry + CRC without loading weights). - Internal types at 0x80+: Q4_0S/PSY/VVC leave IDs 25/26/27 (I16/I32/I64 in ggml today — a real collision). v1/v2 files are normalized on load; the packer already rejected native I16/I32/I64.
- Written spec:
docs/G2BX_SPEC.md(layout, field-by-field LE, v1/v2/v3 compat matrix). Header serialization is now explicit LE ing2bx_io(byte-identical on x86). - Harness-guarded refactors: l5 split, public API,
os_mm/sampler/opts/g2bx_iomodules, CI + CMake.
- IQ1_S integer dot (
madd+SAD, act Q8, 1 hsum/escala por 32): the 264 IQ1_S slots (~3.4 GB) ran scalar fallback; the old fused prototype existed but was never dispatched. Wired intomatmul_q/matmul_q_b, validated by newtools/iq1check(vs exact-Q8 math: maxrel 5e-4). - Q3_K integer dot (values −4..3, per-16 scales, bias −32): covers the 248k head (521 MB) + dense Q3_K models. Validated by new
tools/q3kcheckincl. a 256-position one-hot sweep (caught a half-vectorcvtepi8bug pre-ship: high 8 elems silently dropped). - 27B hybrid (Qwen3.8, hybrid → always sequential decode): stock 132.7 s → 79.6 s (−40 %, 1.67×) for prompt+2 tokens, warm page cache, i5-6200U. Greedy output differs in argmax (Q8 approximation on 1.5-bit weights — both outputs are IQ1_S-grade mojibake); math bounded by the harnesses above.
- Weight traffic ÷4 in
matmul_q4_0_b/matmul_q4_0s_b: each weight row is unpacked once and reused for 4 tokens (was: re-streamed per token, 16× per batch). Qwen3-0.6B Q4_0 prefill 38.9 → 53.7 tok/s (+38 %) and Qwen2.5-3B Q4_0 7.0 → 9.6 tok/s (+37 %) (interleaved A/B on i5-6200U). Bit-exact (prefilltestdiff 0 incl. 3B GQA,q4bcheck5/5, ppl identical 58.709). Decode untouched (3B: 5.4 = 5.4); 27B IQ1_S hybrid output byte-identical to stock.
- Q4_0S_PSY: 2 escalas fp16 por 256 (132B vs 130B). La mejora de calidad anunciada queda retirada: la prueba local de 128 tokens dio ppl 2380555.838 frente a 82.325 base. El soporte permanece, desaconsejado hasta validar la causa. Fallback IQ usa TLS para evitar malloc por fila.
- --mv 0.0..1.0: tunable skip of FFN (dense) / SSM delta (hybrid) via hash + 2-bit predictor.
25.0 → 40.1 tok/s (+60%)on Qwen3-0.6B Q4_0. [RETIRADO v5.1: ppl 25.4 → 35 629 (×1400) a ratio 0.1, 352 904 a 0.5. La velocidad era la de un modelo roto. Ver Phase 7.]
- Batched (prefill) kernel with deferred accumulation: same treatment as the decode kernel. Qwen2.5-3B prefill 4.3 → 7.8 tok/s (+81 %), bit-exact (
tools/prefilltest). Sets the stage for speculative verification. - Dual band CPU+GPU head GEMV: Vulkan worker in a child process (crash-proof), loader bypass loading the ICD straight from DriverStore, automatic split calibration with self-shutdown when the GPU doesn't help. Heads Q4_0/Q4_0S, bit-identical output.
- --drop N (ShortGPT): measures per-block Block Influence during a quick calibration and skips the N least influential blocks. On LFM2.5-1.2B it doesn't pay off (min BI 0.106).
- Decode Q4_0 kernel with deferred accumulation: one hsum per row instead of one per block. Qwen2.5-3B 3.1 → 4.3 tok/s (2.9× vs v3.5); Qwen3-0.6B +10 %.
- Q5_0 end-to-end (fused AVX2 kernel with high-bits LUT).
- Measured quality table (
pplcommand).
- Fused AVX2 kernels for Q4_K and Q6_K (
maddubs+ m·Σx correction term). Before, any K-quant pack fell to the 2–5× slower fallback. Validated byte-by-byte withtools/qkcheck. pplcommand; min-of-3 bench (thermal throttling lies).
- K-quants fixed against official ggml (
deq_q3_K/q4_K/q5_Kbroken since v3.4: half the tensor unwritten + wrong scale interleave). - Batched prefill (B=8) bit-exact; geometry validation at load; persistent Q8 activation scratch (−210 malloc/free per token).
- Prefill without logits (only the last token computes vocab×dim): prefill 1.32×.
- GQA-major attention: each K/V row dequantized once per head group.
- softmax/silu AVX2 with fast exp (rel err < 2e-7); rmsnorm fix for non-multiple-of-32 tails.
- New sampling: O(n) quickselect top-k, Gumbel-max, xorshift64*, reproducible
--seed. - GGUF via mmap; long-chat context compaction.
- Q4_0 AVX2 (2 blocks/iter, ILP), AVX2 attention, full K-quant dequant.
- Q8_0 KV cache (
--q8-kv), effective context, RAM budget (--max-ram), disk swap. - LLaMA RoPE fix (−2.0/head_dim step) and NEOX vs LLaMA: the historical root of corrupt output.