llama.cpp's 42x prompt lookup speedup lives in a fork

A fork of llama.cpp made prompt lookup drafting up to 42x faster. The cache-side numbers matter more, and none of it is merged upstream yet.

llama.cpp's 42x prompt lookup speedup lives in a fork

After a year of watching inference bundles grow, a fork of llama.cpp made prompt lookup drafting up to 42x faster using up to 2.6x less memory. It sat at the top of r/LocalLLaMA this week, and I traced every number in the writeup back to its benchmark before trusting any of them: per-drafted-token latency on the largest test corpus fell from 165.48 µs to 3.98 µs, and to 1.18 µs once Daniel Lemire added a precheck of his own, roughly 140x end to end. Acceptance rate stayed identical through every change. One catch sits above all of it. None of the four optimizations are merged upstream, so the speedup only exists if you run that fork.

Prompt lookup decoding is the cheap cousin of speculative drafting. Instead of training a small model to guess the next tokens, the engine looks for the current text in n-gram caches and proposes what followed last time. It needs no extra model and no training run. It works best when output repeats the input, which describes most coding and agent workloads: you paste a function, the model rewrites half of it back.

Where the 42x came from

llama.cpp keeps three n-gram caches behind prompt lookup: a context cache for the current session, a dynamic cache across runs, and a static cache you build from a corpus with llama-lookup-create. All three were built as a map of maps, an outer std::unordered_map over n-grams with inner maps over follower tokens.

That shape hid two costs. First, the inner maps were being copied on every drafting step; reading them by reference alone was worth 4.5x to 25.6x depending on corpus size. Second, std::unordered_map chains its buckets, which is slow to walk and rough on CPU caches. Swapping the outer map for a flat hash map, then for a constmap backed by sorted vectors, took the per-token cost at the 541 MB corpus down to 3.98 µs and cut peak memory from 3.47 GB to 1.31 GB. Lemire's contribution skips scoring candidate tokens that cannot pass the acceptance thresholds anyway, which is what pushes drafting to 1.18 µs per token. (The thresholds are hardcoded inside llama.cpp as of release b11182, so the precheck leans on behavior that could change upstream.)

The cache-side numbers nobody is quoting

Drafting gets the headlines. The cache work is what I'd steal. Building the static cache from the 541 MB corpus dropped from 5.49 s to 0.30 s, and the constmap loads in about 0.04 to 0.30 s across corpus sizes, against seconds before. The cache also stopped costing more than its file: 463 MB in RAM for a 467 MB corpus on disk.

Cache building runs once per corpus, so a five-second savings sounds like a rounding error until you count restarts. CLI tools and short-lived server sessions rebuild or reload state on every launch, and half a gigabyte of headroom matters on the 16 GB machine running a 30 B model alongside it. If you run local models on consumer hardware, this is the half of the PR series built for you.

The catch

All of this lives in jadidbourbaki's fork. Each optimization sits as an open pull request on that fork: the by-reference fix, the constmap, and Lemire's precheck among them. None has been sent to upstream llama.cpp, so running the fast path means tracking the fork, and a fork of llama.cpp drifts from upstream on a schedule nobody controls. The benchmarks are also the author's own: an Apple M4 Pro, a 4096-token context, WikiText-103 as the corpus, medians of 3 runs, with the measurement code published in a separate repo you can rerun. And the win is workload-shaped. If your output rarely repeats your prompt, lookup has less to find.

If you want to try it

The honest test takes three numbers. Draft acceptance first, which should match upstream in your workload because none of these changes touch which tokens get drafted. Then drafts per second on a prompt with real repetition, where the 42x-to-140x band either shows up or does not. Then static cache load time from a cold start, where the 5.49 s to 0.30 s difference is easiest to see. Run the author's bench repo before your own numbers if you want a baseline.

Where this lands

The broader lesson survives any merge decision. The fastest recent wins in local inference came from data-structure work, not model work: stop copying maps and flatten the lookup structure. The other half of the lesson is giving caches the same measurement rigor as the model itself. We reached the same conclusion from the product side with diff-hash caching in cora-code, which skips repeat code reviews by content address instead of re-running them, and that cache is now the part I watch most closely. Prompt lookup may or may not land upstream this quarter. The habit of treating caches as measured components is already the part worth copying.