Δ deltaverse · cryptoAGI

bankML

Verified 1-bit and ternary language models on the CPU you already have: exactly the same answers as llama.cpp, checked bit for bit, then faster.

Written 6 October 2026, while bankML 0.3.6 is the latest release, 0.3.7 is in its release gate, and 0.3.8 and 0.3.9 are being built on the way to the 0.4.0 milestone.

bankML in one paragraph

bankML runs language models on the computer you already have: a laptop, a small server, no graphics card needed. It specialises in 1-bit and ternary models, which store each weight in one or two bits instead of sixteen. An 8-billion-parameter model then fits in 1.2 GB (1-bit) or 2.3 GB (ternary). bankML gives exactly the same answers as llama.cpp, the reference engine, and checks that it does, bit for bit, before it is allowed to be faster. Every answer carries a receipt that says which model file produced it.

9.4–10×faster per ternary matrix than llama.cpp
≈ 8×whole ternary answers vs llama-server
0external dependencies
bit-exactagainst llama.cpp b11192

the full story: README · the technical report: TECHNICAL.md

The advantages

1. It is fast where it matters most

llama.cpp has no fast x86 code for ternary weights; it falls back to plain C. bankML wrote the missing kernel. Each matrix is 9.4–10× faster, and whole answers on the ternary model come out at about 8× llama-server's speed (2.3–2.4 tokens/s against 0.30 on the same laptop). The ternary model gives the better answers of the two, so this is the speed that counts.

PERFORMANCE.md · the kernel: modules/q2_0.md, bankML/q2_0.rs

2. It is exact, and proves it

Speed only counts if the numbers are the same. Every kernel, the tokenizer, the chat templates, the samplers and whole conversations are checked against llama.cpp b11192's own compiled code. These checks are called oracles, and a release ships only when every oracle passes in its release gate.

oracles.md · the gate: testing/release_gate.sh and its records in testing/results/

3. It refuses to answer from a model it cannot verify

Before a model answers, three gates run. A guard reads the file's header and refuses known-bad files. A pin checks the file's sha256 against the record of where it came from. The oracle, run ahead of time, has proven the arithmetic. Each answer then carries a receipt with the model's sha256 and the hashes of the request and the answer.

TECHNICAL.md §III.1 · receipts (usage.md §9) · the guard: bankML/gguf.rs · the pin: bankML/sha256.rs

4. It is small and has nothing to install around it

bankML is one Rust crate with zero external dependencies. The GGUF reader, sha256, f16 maths, the thread pool, the HTTP server and every kernel are written in the crate. You get one binary and nothing else to trust.

one page per source file · Cargo.toml

5. It speaks the languages your tools already use

6. It measures itself

bankML reports its own time to first token, tokens per second and, where the machine allows, the energy each token costs. It also limits how much of a graphics card it may use.

metrics.md and console.md (with 0.3.7) · the GPU: gpu.md

What people use it for

A private assistant on your own machineSavante, a chat page on 127.0.0.1:7873, answered by bankML — usage.md §4, playback.md
The engine behind mindXmindX's default model server on its VPS since 4 October 2026, in place of Ollama — OLLAMA.md
A drop-in Ollama or OpenAI serverbankml serve --native --registry — usage.md §6a, install.md
Your own model, verifiedbankml convert (safetensors → GGUF, byte-identical to llama.cpp's) and bankml create (a persona layer from a Modelfile) — convert.md, create.md
Structured answers for programsJSON mode, JSON schemas and grammars — usage.md
Inference inside another programlibbankml and bankml.h — CAPI.md
Agents with a provable historyreceipts, Merkle commitments of the conversation, THOT bundles and iNFTs — usage.md §8a, §8e
Meaning search over notesbge-m3 embeddings fused with keyword search — embedding.md

new here? start with ./install.sh: usage.md §1

How Rust helps bankML go fast and stay exact

Rust lets bankML write code as close to the hardware as C, while the compiler checks much of what C leaves to the programmer.

Where it is going

The road is laid out release by release in TODO.md, and each step counts only once its oracle passes.

Our hope for bankML

We hope bankML shows that a capable, private AI does not need a data centre or a graphics card: that an eight-billion-parameter model can run on the laptop in front of you and give answers you can check. We want every speed claim to come with proof that the numbers are the same, so that "faster" always means faster and correct. We hope the ternary kernel goes back upstream to llama.cpp, so everyone benefits from it. We want bankML to become the dependable, verifiable engine under Savante, mindX and the agents built on them. And we want it to stay small enough that one person can read all of it.

the authors' own words: the Thesis · the argument in full, with the contemporary field: the thesis · among other engines and papers: research.md · how it was built: BUILD_HISTORY.md