Skip to content
AnonymoDGHPublic

About

Custom C99 runtime: GGUF → G2BX → inference (Qwen3 / Qwen2 / Llama).

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

gguf2bin2

English | Español

C99 LLM runtime for low-RAM machines — GGUF → G2BX (own format) → inference. Weights are memory-mapped: your real RAM budget is KV cache + activations + tokenizer, not the model file.

C99 AVX2 Vulkan RAM

Runs a 3B model on 145 MB of RAM and a large model on a 2 GB machine (--swap: KV backed on disk, footprint ≈ 37 MB).


⚡ Speed

Measured on Intel i5-6200U · 2C/4T · DDR3L dual-channel (~9.4 GB/s bus ceiling). Sustained decode with --fast (high priority + OpenMP + quantized KV).

Model Weights (mmap) Runtime RAM decode prefill
Qwen2.5-3B Q4_0 1992 MB 145 MB 4.3 7.8
Qwen3-0.6B Q4_0 319 MB 511 MB 25.0 47.9
LFM2.5-1.2B Q4_0S 567 MB 631 MB 15.7 17.9
Llama-3.2-1B F16 804 MB 644 MB 13.8 —
SmolLM2-135M Q4_0 72 MB 40 MB 59.5 —

Measured 2026-08-31 on i5-6200U with bench -n 32 (min3). At ~25 tok/s decode pinned to DDR3L ~9 GB/s; prefill hits compute ~27 GMAC/s. (--mv speedups removed in v5.1.0: the skip destroyed ppl ×1400 — speed of a broken model.)

🚀 Quick start

make            # Linux / MinGW · OpenMP + AVX2 recommended
make test       # smoke test on a synthetic model

# Pack a GGUF once, run forever
./gguf2bin2 pack Qwen3-0.6B-Q8_0.gguf qwen.g2bx --q4
./gguf2bin2 chat qwen.g2bx --no-think --fast

--threads N picks OpenMP threads; physical cores is the sweet spot.

🧠 RAM knobs

Knob Effect
--q8-kv KV cache F32→Q8_0: ~3.8× less RAM (94 % greedy argmax agreement)
-c N Context sizes to your session, not the model's 262k max
--max-ram MB Auto: enables Q8 KV first, then halves context until it fits
--swap [PATH] KV cache backed on disk → 37 MB heap even for big models
$ gguf2bin2 run qwen.g2bx "Hello" --max-ram 2048   # ideal for 2 GB machines

What counts against a 2 GB budget?

Component Counts? Fix
Weights (mmap page cache) ❌ evictable —
KV cache ✅ --q8-kv, -c N, --swap
Buffers / activations ✅ small —
Tokenizer (~250k vocab) ✅ ~30–60 MB —

🎛 CLI

gguf2bin2 pack model.gguf out.g2bx [--q4]     # GGUF → G2BX (--q4: half the bytes)
gguf2bin2 info model.g2bx                     # slots, geometry, types
gguf2bin2 verify model.g2bx                   # header + CRC + geometry, no weights load
gguf2bin2 run m.g2bx "prompt" [-n N] [-t T] [--bos] [--gpu]
gguf2bin2 chat m.g2bx [--no-think] [--fast] [--swap]
gguf2bin2 bench m.g2bx [-n 32] [--prefill 256]
gguf2bin2 ppl m.g2bx -f text.txt              # quality harness (perplexity)
gguf2bin2 vkinfo                              # Vulkan probe

Sampling: quickselect top-k O(n) + Gumbel-max + xorshift64* reproducible via --seed.

🎮 Dual band CPU+GPU (--gpu)

The head GEMV (vocab×dim, the heaviest layer) gets split between CPU and GPU with automatic calibration:

[gpu] worker ready
[gpu] dual band: cpu=[0..44855) gpu=[44855..65536)  tc=6.1ms tg=13.1ms
  • Vulkan worker in a child process — if the driver crashes or hangs, the runtime falls back to CPU-only without interrupting generation.
  • Vulkan loader bypass: loads the ICD directly from DriverStore (useful on systems with a broken Khronos\Vulkan\Drivers registry).
  • Automatic optimal split: gpu = vocab·tc/(tc+tg) measured on the first token; if the GPU is >4× slower than the CPU it shuts itself off.
  • Supports Q4_0 and Q4_0S heads. Bit-identical output vs the CPU path (greedy).

⚠️ On hardware where the iGPU shares the RAM bus with the CPU (HD 520 + DDR3L) there is no net gain — calibration detects it and disables itself. The real payoff comes with a dGPU with dedicated VRAM.

📐 Quality

Perplexity (internal corpus, SmolLM2-135M-Instruct): uniform Q4_0 = 73.7 (broken), native Q4_K_M = 48.7, Q6_K = 48.0, Q8_0 = 48.0. Native K-quants keep Q8_0-level quality within 1.4 % while running faster than uniform Q4_0.

Measured 2026-09-05 in separate CLI processes, using the pre-edit README+README.es+G2BX_SPEC corpus. These are short local checks, not a standard benchmark; model provenance and quantization paths need further validation.

Model / format ppl vs base
Qwen3-0.6B Q4_0 25.4 —
Qwen3-0.6B Q8_0 20.8 −18 %
LFM2.5-1.2B q4max 28.1 —
LFM2.5-1.2B q4s 34.9 +24 %
Llama-3.2-1B baseline (256 tokens) 83.966 —
Llama-3.2-1B Q4_VVC (256 tokens) 1177.176 ×14
Llama-3.2-1B baseline (128 tokens) 82.325 —
Llama-3.2-1B Q4_0S_PSY (128 tokens) 2380555.838 ×28916

Qwen and LFM2 rows use 1024 tokens. LFM2 Q4_0S increased ppl by about 24%. The tested PSY/VVC artifacts showed severe degradation. For PSY a root cause was found in v5.1.1: both AVX2 kernels read the nibbles 2 bytes off, so the PSY row above measures a broken kernel, not the format; it needs re-measuring. VVC is a plain 3-bit uniform quant with one scale per 256 (no inter-row prediction is implemented). Both remain not recommended until re-validated. The types at 0x80+ are custom G2BX types, not standard GGUF types supported by other runtimes.

Included numerical validation: make kvtest (F32 vs Q8 KV), tools/prefilltest (bit-exact batched prefill), tools/qkcheck (46/46 K-quant kernels).

Removed in v5.1.0 (measured, see docs/ROADMAP_PERF.md Phase 7): --mv and --bvh (large ppl increases at the tested ratios), cyber-* commands (reported accuracy and particle-loop loss were not real — experiment archived in experimental/cyber-mrna/), dead OrderBook/HDR/ZRAM/FM-index paths. --cyber <lora.bin> still loads real LoRA adapters (v1/v2). The removal of mv_ratio/bvh_keep changes the public g2b_config layout: recompile API clients against the new header; do not mix old headers and new libraries.

📦 G2BX format & supported types
G2BX | ver:u16 | arch:u8 | flags:u8 | ModelCfg | n_slots:u32 | Slot[] | data[] 64B-aligned | tokenizer | [v3: crc32 + "G2BX"]
Slot: role:u8 layer:u16 type:u8 nbytes:u32 off:u64

Full spec: docs/G2BX_SPEC.md (v1/v2/v3 compat matrix; own types live at 0x80+ since v3).

Type Load Fused matmul
F32 / F16 yes (pack→Q4_0 if weight) F32 path / dequant
Q4_0 / Q4_1 / Q5_0 yes AVX2 fused
Q8_0 yes AVX2 fused
Q4_0S (own, fp16 shared scale /256) yes AVX2 batched
Q2_K … Q6_K yes (dequant) AVX2 fused (Q4_K/Q6_K), scalar rest
IQ* / Q8_K partial —
🗂 Project layout
include/gguf2bin.h   public API (sessions) · src/internal/ shared internals
src/model.c          G2BX load/free, geometry, RAM budget, synth
src/kv.c             F32/Q8 KV cache, disk swap, runtime alloc, TLS
src/forward_*.c      dense (+dispatch) / lfm2 / qwen35-hybrid / batched prefill
src/l1_gguf.c        GGUF parser (mmap) · src/l4_gbin.c  G2BX packer
src/l2_codec.c       fused dequant + matmul (AVX2) · src/l3_math.c  norms/rope/softmax
src/l6_token.c       BPE tokenizer · src/l7_vulkan.c  dual-band GPU (child process)
src/g2b_api.c        API implementation · src/os_mm.c  mmap/files · src/sampler.c
src/opts.c           CLI flags · src/g2bx_io.c  G2BX reader/writer/CRC/verify
src/main.c           CLI · shaders/  Q4_0/Q4_0S GEMV · tools/  harnesses + fuzz/
docs/G2BX_SPEC.md    format spec · docs/RESEARCH.md  notes + roadmap
📜 Changelog

v5.1.1 — review fixes

  • CI green again (red since Phase 3): strdup/fseeko/ftello/clock_gettime were implicitly declared under -std=c99 — on Linux x86-64 the truncated strdup pointer crashed make test, and ftello truncated offsets >2 GB. Also fixed: fuzz job YAML (> folded the clang lines apart), fmemopen in the harnesses, MinGW copy under MSYS2 sh, CMake include dirs, AVX2 flags in the sanitizer job.
  • Tokenizer: u2b[289] overflowed (68 remapped bytes → indices up to 323); bytes 0x7F–0xA0/0xAD decoded wrong (€, à, emojis). tok_read_section leaked on early errors.
  • Q4_0S_PSY kernels (decode + batched) read nibbles 2 bytes off.
  • qwen35: attention wo used n=dim instead of n_heads*head_dim. LFM2/qwen35 recurrent state was never reset at pos 0 (ppl windows, chat_reset, compaction and the Android app inherited the previous sequence).
  • LoRA: batched prefill skipped the adapter (now falls back to sequential); loader validates rank and reads.
  • pack --prune: ffn_down copy assumed one block per group (broken for F16/F32/Q8_0 down); OOM mid-prune now aborts. Tensors >4 GB are rejected instead of truncating Slot.nbytes.
  • Untrusted .g2bx hardening (the Android app downloads from any URL): slot nbytes validated against geometry, geometry caps against i32 overflow, blob-past-EOF and offset-overflow checks, non-mmap fallback read from the right offset.
  • Android JNI: freeModel could free the model while generate still ran (use-after-free); tokens are emitted as whole UTF-8 characters via UTF-16 (NewStringUTF broke on split characters and emojis).
  • Chat (LFM2): the empty <think></think> block (Qwen3 enable_thinking=False convention) was injected into LFM2.5 too, whose template has no such block; the model opened every answer "correcting itself". Found and verified on the release lfm25-1.2b-q4s.g2bx (now answers "Paris" / "Madrid").
  • ppl: windows after the first now restart with BOS, like llama.cpp (LFM2 at -c 128: 485 → 109; single-window results unchanged).
  • g2b_pack no longer leaks the Q4_0S/PSY/VVC mode into later calls; chat prompts are no longer truncated at 4/9 KB; default --swap file is per-process and opened with O_NOFOLLOW.

v5.0 — G2BX v3: CRC'd format + own type namespace

  • CRC32 footer: every new .g2bx ends with [crc32 of everything before][magic]; the loader verifies on open (warming the page cache as a side effect) and rejects truncated/corrupt files with a clear message. New verify command (header + slots + types + geometry + CRC without loading weights).
  • Internal types at 0x80+: Q4_0S/PSY/VVC leave IDs 25/26/27 (I16/I32/I64 in ggml today — a real collision). v1/v2 files are normalized on load; the packer already rejected native I16/I32/I64.
  • Written spec: docs/G2BX_SPEC.md (layout, field-by-field LE, v1/v2/v3 compat matrix). Header serialization is now explicit LE in g2bx_io (byte-identical on x86).
  • Harness-guarded refactors: l5 split, public API, os_mm/sampler/opts/g2bx_io modules, CI + CMake.

v4.9 — IQ1_S + Q3_K fused (27B desatascado)

  • IQ1_S integer dot (madd+SAD, act Q8, 1 hsum/escala por 32): the 264 IQ1_S slots (~3.4 GB) ran scalar fallback; the old fused prototype existed but was never dispatched. Wired into matmul_q/matmul_q_b, validated by new tools/iq1check (vs exact-Q8 math: maxrel 5e-4).
  • Q3_K integer dot (values −4..3, per-16 scales, bias −32): covers the 248k head (521 MB) + dense Q3_K models. Validated by new tools/q3kcheck incl. a 256-position one-hot sweep (caught a half-vector cvtepi8 bug pre-ship: high 8 elems silently dropped).
  • 27B hybrid (Qwen3.8, hybrid → always sequential decode): stock 132.7 s → 79.6 s (−40 %, 1.67×) for prompt+2 tokens, warm page cache, i5-6200U. Greedy output differs in argmax (Q8 approximation on 1.5-bit weights — both outputs are IQ1_S-grade mojibake); math bounded by the harnesses above.

v4.8 — Blocked prefill (G=4 token blocking)

  • Weight traffic ÷4 in matmul_q4_0_b / matmul_q4_0s_b: each weight row is unpacked once and reused for 4 tokens (was: re-streamed per token, 16× per batch). Qwen3-0.6B Q4_0 prefill 38.9 → 53.7 tok/s (+38 %) and Qwen2.5-3B Q4_0 7.0 → 9.6 tok/s (+37 %) (interleaved A/B on i5-6200U). Bit-exact (prefilltest diff 0 incl. 3B GQA, q4bcheck 5/5, ppl identical 58.709). Decode untouched (3B: 5.4 = 5.4); 27B IQ1_S hybrid output byte-identical to stock.

v4.7 — Q4_0S_PSY (psicoacústico) + fallback TLS

  • Q4_0S_PSY: 2 escalas fp16 por 256 (132B vs 130B). La mejora de calidad anunciada queda retirada: la prueba local de 128 tokens dio ppl 2380555.838 frente a 82.325 base. El soporte permanece, desaconsejado hasta validar la causa. Fallback IQ usa TLS para evitar malloc por fila.

v4.6 — Swapeculative MV Triple Band

  • --mv 0.0..1.0: tunable skip of FFN (dense) / SSM delta (hybrid) via hash + 2-bit predictor. 25.0 → 40.1 tok/s (+60%) on Qwen3-0.6B Q4_0. [RETIRADO v5.1: ppl 25.4 → 35 629 (×1400) a ratio 0.1, 352 904 a 0.5. La velocidad era la de un modelo roto. Ver Phase 7.]

v4.5

  • Batched (prefill) kernel with deferred accumulation: same treatment as the decode kernel. Qwen2.5-3B prefill 4.3 → 7.8 tok/s (+81 %), bit-exact (tools/prefilltest). Sets the stage for speculative verification.
  • Dual band CPU+GPU head GEMV: Vulkan worker in a child process (crash-proof), loader bypass loading the ICD straight from DriverStore, automatic split calibration with self-shutdown when the GPU doesn't help. Heads Q4_0/Q4_0S, bit-identical output.

v4.4

  • --drop N (ShortGPT): measures per-block Block Influence during a quick calibration and skips the N least influential blocks. On LFM2.5-1.2B it doesn't pay off (min BI 0.106).

v4.3

  • Decode Q4_0 kernel with deferred accumulation: one hsum per row instead of one per block. Qwen2.5-3B 3.1 → 4.3 tok/s (2.9× vs v3.5); Qwen3-0.6B +10 %.
  • Q5_0 end-to-end (fused AVX2 kernel with high-bits LUT).
  • Measured quality table (ppl command).

v4.2

  • Fused AVX2 kernels for Q4_K and Q6_K (maddubs + m·Σx correction term). Before, any K-quant pack fell to the 2–5× slower fallback. Validated byte-by-byte with tools/qkcheck.
  • ppl command; min-of-3 bench (thermal throttling lies).

v4.1

  • K-quants fixed against official ggml (deq_q3_K/q4_K/q5_K broken since v3.4: half the tensor unwritten + wrong scale interleave).
  • Batched prefill (B=8) bit-exact; geometry validation at load; persistent Q8 activation scratch (−210 malloc/free per token).

v4.0

  • Prefill without logits (only the last token computes vocab×dim): prefill 1.32×.
  • GQA-major attention: each K/V row dequantized once per head group.
  • softmax/silu AVX2 with fast exp (rel err < 2e-7); rmsnorm fix for non-multiple-of-32 tails.
  • New sampling: O(n) quickselect top-k, Gumbel-max, xorshift64*, reproducible --seed.
  • GGUF via mmap; long-chat context compaction.

v3.x

  • Q4_0 AVX2 (2 blocks/iter, ILP), AVX2 attention, full K-quant dequant.
  • Q8_0 KV cache (--q8-kv), effective context, RAM budget (--max-ram), disk swap.
  • LLaMA RoPE fix (−2.0/head_dim step) and NEOX vs LLaMA: the historical root of corrupt output.

About

Custom C99 runtime: GGUF → G2BX → inference (Qwen3 / Qwen2 / Llama).

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages