Skip to content
geistenPublic

About

Tiny dependency-free C23 inference engine for small LLMs — ternary BitNet, CPU-first, runs on a Raspberry Pi

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

2,726 Commits

Folders and files

geistlib

geistlib 👻

Run Microsoft BitNet 2B locally on a 4 GB Raspberry Pi 5. One binary and one model file. No Python, no cloud.

curl -L -o geist https://geisten.net/download/geistlib/latest/geist-linux-arm64
curl -L -o bitnet.gguf https://huggingface.co/microsoft/bitnet-b1.58-2B-4T-gguf/resolve/main/ggml-model-i2_s.gguf
chmod +x geist
./geist bitnet.gguf "The capital of France is"

Run it with only the model for a minimal REPL (each line completes independently; the model stays loaded). Cooling, cold-start times, errors and model limits: docs/PI5_BITNET.md.

On a Raspberry Pi 5: real-time BitNet b1.58 2B-4T text generation from a single dependency-free binary

Real time on a Raspberry Pi 5: ternary BitNet b1.58 2B-4T from a single dependency-free binary. No GPU, no driver stack.

  • Tested on a 4 GB Pi 5 (Raspberry Pi 5 Model B, 64-bit Raspberry Pi OS) — see the reference runs.
  • 15–18 decode tokens/s depending on context: 17.9 t/s at a short prompt, 15.0 t/s at a 512-token prompt (frozen protocol, 10 repeats, methodology).
  • ~1.2 GB download, model included. Offline after download: nothing leaves the device.

Prebuilt binaries ship for Linux arm64 and x86_64 (swap -linux-arm64 for -linux-x86_64). On macOS, build from source or use the prebuilt SDK.

CI License C Standard Platform Latest release Status Discussions Good first issues

Questions? → Discussions · Bug? → open an issue · Want to help? → good first issues


What it is

geistlib is a C23 inference engine shipped as a static library (libgeist.a) with no runtime dependencies. It loads GGUF models and produces tokens.

  • Ternary kernels, first-class. BitNet b1.58 stores every weight as −1/0/+1; geistlib runs it with integer-only dot products (ARM SDOT on the Pi, AVX-512 VNNI on x86). The 2B model is 1.1 GiB, about a third of a comparable 4-bit model.
  • Zero-copy weights. The GGUF is mmapped (or, when embedded in an executable, aliased out of its read-only data) and demand-paged. Some CPU kernels keep a repacked copy of a dtype for speed; docs/BACKENDS.md lists them and the switch for each.
  • An engine, not an application. No chat templates, tool use or system prompts — those belong to whatever links the library. The boundary is in docs/README.md.

Architecture (layers, load-time kernel binding, why C): docs/ARCHITECTURE.md.

Bring your own model

The slim CLI (release asset geist-linux-arm64 or geist-linux-x86_64, ~2 MB) runs any supported GGUF that carries its own tokenizer:

curl -L -o geist https://geisten.net/download/geistlib/latest/geist-linux-arm64
chmod +x geist
./geist model.gguf "your prompt" [max_new_tokens]

Model families: Gemma 4 (text, vision, audio), Llama, Qwen3, Qwen3.5/3.6/3.8 dense hybrids (incl. Ternary-Bonsai-2-27B) and BitNet b1.58. Downloads, sizes and RAM needs: docs/MODELS.md.

Build from source

macOS and Linux (arm64, x86-64). Needs gcc ≥ 14 or Apple clang ≥ 16, and make; on macOS also Homebrew libomp for multi-threading.

git clone https://github.com/geisten/geistlib && cd geistlib
make lib                     # auto-detects target; or TARGET=mac-omp | mac | pi5 | linux
make fetch-bench-model       # BitNet b1.58 2B-4T, ternary, ~1.2 GB
make run ARGS='gguf_artifacts/bitnet-2b4t-i2_s.gguf "The capital of France is"'
  • Ubuntu 24.04 ships gcc 13: apt install gcc-14 and pass CC=gcc-14.
  • Linux links OpenBLAS by default (libopenblas-dev); GEMM_PROVIDER=native builds without it.
  • make help lists every target and option.

make run builds and runs examples/simple_generate.c: load, prefill, decode against the STABLE API only — the program to copy when you embed the library.

Embed the library

Every release ships libgeist-<platform>.tar.gz (libgeist.a, include/*.h, LICENSE) for macos-arm64, linux-arm64 and linux-x86_64. Verify it against SHA256SUMS (sha256sum -c SHA256SUMS), then link:

cc -std=c23 -I libgeist-linux-arm64/include my_app.c \
   libgeist-linux-arm64/libgeist.a -fopenmp -lm -o my_app
# macOS: replace -fopenmp with -framework Accelerate "$(brew --prefix libomp)/lib/libomp.a"

The text path is geist_backend_create → geist_model_load → geist_session_create → geist_session_set_prompt → loop geist_session_decode_step. The header is the ABI: any language can call it through its FFI without bindings — examples/ffi/ has the complete integration in Python, Rust, Go and JavaScript (~30–40 lines each).

Backends

Backend Hardware Default
cpu_neon ARM64 with dot product (Pi 5, Apple Silicon, armv8.2+) arm64
cpu_x86 x86-64-v3 (AVX2) with runtime AVX-512/VNNI dispatch x86-64
cpu_scalar portable C, numerical reference always built
metal Apple GPU, experimental opt-in
vulkan Linux GPU, experimental opt-in

Select at build time with BACKENDS="..." (make clean when you change it). Details, knobs and memory use per backend: docs/BACKENDS.md.

How fast?

Same GGUF, greedy decode, both engines measured in the same run on the same box, thermally gated. geistlib beats Microsoft's bitnet.cpp on ternary BitNet on a Pi 5 (~2× decode) and an AMD 9950X, and matches-to-beats llama.cpp on the CPU paths:

Side-by-side terminal recording on one Raspberry Pi 5: geistlib finishes 110 tokens in 7.1 s (15.5 tok/s) while bitnet.cpp needs 11.7 s (9.3 tok/s), same GGUF, both greedy, thermally gated

One board, one GGUF, recorded sequentially with a thermal gate and shown side by side — how it was made, and more demos.

Decode-throughput scoreboard: geistlib divided by its baseline engine, decode tokens/s, grouped by system. Raspberry Pi 5 (Linux): BitNet decode 1.96x bitnet.cpp, BitNet prefill 0.99x bitnet.cpp, Gemma decode 1.1x llama.cpp. AMD Ryzen 9 9950X (Linux): BitNet decode 1.9x bitnet.cpp, Gemma decode 1.1x llama.cpp, Llama 3.2 decode 1.0x llama.cpp. Sub-parity rows are shown too.

Run them yourself: make bench measures geistlib and any llama.cpp / bitnet.cpp binary it finds on your machine, on the byte-identical GGUF, and prints the spread. Greedy output is checked bit-identical to the scalar reference before a speedup is quoted. Protocol, raw data and per-system tables: benchmark/. GPU numbers (Metal decodes Qwen3.8-27B at 1.41× llama.cpp Metal on an M1 Max): docs/BACKENDS.md.


Status

main is the experimental development branch; for binaries and a citable version use the latest release. The STABLE core (load → session → decode → tokenize) is the part to build on. EXPERIMENTAL surfaces (KV-cache modes, speculative decode, multimodal attach, GPU backends) may change between minor versions.

Direction — ternary and binary quantization as first-class citizens, hardware people own, one-step install, models that adapt — is laid out track by track in ROADMAP.md.

Contributing

make lib && make test      # builds libgeist.a, runs the C suite

Open fronts: NEON / AVX-512 microkernels, low-bit quantization (TQ2_0, IQ variants, below 1.58 bits), portability (Windows, x86-64 quant coverage, Vulkan prefill). Workflow and review bar: CONTRIBUTING.md; coding rules that source comments cite: AGENT.md. Unsure where to start? Ask in Discussions.

Documentation

All documents, with what each covers: docs/README.md.

Citation

Cite the version you used: each release tag carries its matching CITATION.cff, which GitHub's "Cite this repository" action reads.

License

Apache License 2.0 — permissive, with an explicit patent grant. See also NOTICE.

About

Tiny dependency-free C23 inference engine for small LLMs — ternary BitNet, CPU-first, runs on a Raspberry Pi

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages