Open Source · C
A tiny language model, run from scratch in C. A complete transformer inference engine in a few hundred lines, with no libraries doing the work: no PyTorch, no NumPy, no BLAS, no framework of any kind. It loads a pretrained model and generates text, and the whole point is to shrink it until it runs somewhere it has no business running, a chip instead of a computer.
Most people run a language model by importing a giant library and calling one function. mote is the opposite: it is the function, opened up. It takes a model that has already been trained, a set of numbers, and does the actual arithmetic that turns a prompt into the next word, then the next, one token at a time. Every part of that is written out by hand so you can read it end to end in an afternoon.
It does not train anything and it does not try to be the fastest engine in the world. It is built to be legible, and to be small, small enough that the same code can eventually run on a microcontroller with a few megabytes of memory instead of a laptop.
Point it at a small pretrained model, give it the start of a sentence, and it writes the rest. This is the real output, generated by the from-scratch engine at about a hundred tokens a second on a laptop, single-threaded.
$ ./mote stories15M.bin -i "Once upon a time, there was a little robot"
Once upon a time, there was a little robot. The robot liked to crawl. One day, the robot went to the park. He saw a big tree. The robot wanted to crawl to the tree. He went up, up, up the tree. He saw a bird singing, and stopped to listen. Then the robot saw a big hill. He was scared, but he was brave, and he crawled to the top. There was a party waiting for him. The robot and the bird were friends. They played and had fun.
[158 tokens, 109.4 tok/s]
This is the smallest model, a 15-million-parameter storyteller, chosen because you can watch it work. Further down, the same engine runs a real conversational model hundreds of times larger, still entirely on-device.
The same engine runs inside a native iOS app, with the model living in the app itself. No server, no network, no API. Put the phone in airplane mode and it still answers, because there is nothing to phone home to.
It started as a playground that continued your sentences; the current build holds a conversation with a real half-billion-parameter instruct model, quantized to 4-bit so it ships inside the app, with the true tokens-per-second printed under every reply. Same C engine as everywhere else, called from Swift.

A word comes in as a token, which is looked up as a row of numbers, its embedding. That vector then passes through a stack of identical blocks. Each block does two things: it lets the token look back at everything before it and pull in what is relevant (attention), and then it thinks on its own for a moment (a small feed-forward network). After the last block, the engine reads out a score for every possible next word and picks one. Feed that word back in and the loop repeats.
The details are the parts people usually let a library hide: RMSNorm to keep the numbers well-behaved, rotary embeddings so the model knows the order of the words, grouped-query attention with a cache so each new token reuses its history instead of recomputing it, and a SwiGLU feed-forward. In mote they are all just a few dozen lines of C each. There is exactly one loop that costs anything, the matrix multiply, and it is the only place the code reaches for extra threads.
The weights are memory-mapped straight off disk rather than loaded into a buffer, so the model appears instantly and the operating system pages in only what each step touches. It is the kind of detail that stops mattering on a laptop and starts mattering a lot on a chip.
That one hot loop, the matrix multiply, is where all the speed lives, so it is the one place the code gets clever: hand-written SIMD and a spread across cores. On the same laptop that made the from-scratch engine several times faster, and because it uses the instructions a phone also has, the gain carries onto the device.
Tokens per second
Model size on disk, MB
Measured on one laptop, so the on-device numbers differ, but the shape holds: int8 is a quarter of the size, int4 goes below a sixth, and the vectorized, threaded matmul multiplies the speed. int4 decodes a touch slower than int8 here because unpacking two weights from every byte costs a little arithmetic; on disk it wins outright.
"Quality held" is easy to claim, so the repo ships a perplexity harness instead: feed the model seventeen thousand held-out tokens it has never seen and score how well it predicts each next one. Lower is better, and the gap between rows is exactly what quantization costs.
| Weights | Perplexity | Cost |
|---|---|---|
| fp32 | 2.376 | — |
| int8 | 2.379 | +0.1% |
| int4 | 2.642 | +11% |
Getting the int4 number down was a lesson in trusting measurement over intuition. The obvious improvement, searching for a slightly clipped scale per group of weights, made the mathematical error smaller and the actual model worse: the outlier weights it clipped turned out to be the ones that matter. The scheme that won is almost insultingly simple, plain rounding with the scale signed so the largest weight in each group is stored exactly.
The subtler lesson came from fine-tuning: a gently-trained behavior can be smaller than 4-bit rounding noise, and simply vanish when the model is quantized. The fix is training through the exact same rounding the deployed engine applies, bit for bit. Every stage has to round the same way; a converter that rounds "slightly better" than training is rounding differently, and different is broken.
The story models are the teaching material; the engine also runs the real thing. It loads Qwen2.5, an open half-billion-parameter instruct model, through its own converter and a byte-level BPE tokenizer written from scratch in C, verified token-for-token identical against the reference implementation. Quantized to int4 the model drops from 525 MB to 309 MB and holds a correct conversation entirely offline.
And because the engine has no dependencies, compiling it to WebAssembly is just one more target: the entire engine is under 90 KB of wasm. The model downloads once, then chat runs inside the tab, with nothing sent anywhere. With hand-written wasm SIMD and a small worker pool it measures 48 tokens a second in desktop Chrome and 20 in iPhone Safari, up from 13.6 and 6.5 single-threaded. Still shy of the native app, but a working conversation from a URL with nothing installed.
It is live to try: mote-alpha.vercel.app — the model downloads once, then it runs offline and even installs to your home screen. The browser build also remembers, and the mechanism is worth its own section.
Long-term memory usually means shipping a second model, an embedding model, alongside the chat model. mote's trick is that a transformer already computes a rich internal description of whatever it reads. One small function exposes it: run the text through the model, average the hidden state at every token, and normalize. The chat model doubles as its own embedding model, at the cost of one extra C function.
every exchange
embed(what you said + what it replied) → store the pair with its vector
every new question
embed(question) → rank stored exchanges by similarity → hand the closest few to the model as notes
The honest part is what happens with the scores. Similarities from a small model's hidden states rank correctly but bunch into a narrow band, so there is no clean threshold between "relevant" and "not". Rather than pretending otherwise, the closest notes are always offered and the model itself was trained on three behaviors: answer from a note when it holds the answer, ignore it when it is beside the point, and say plainly when the answer was never mentioned.
Tell it your name and your dog's name, wipe the conversation entirely, ask again, and it answers from retrieval alone. Everything lives in the browser's local storage; nothing is uploaded, because there is nowhere to upload to.
In the same spirit as the perplexity table: honesty about the ceiling. The model that fits on a phone is half a billion parameters, hundreds of times smaller than the ones that run in data centers. It answers everyday questions, explains things, does small math, drafts short messages, and remembers what you tell it. It is genuinely useful and genuinely small.
It is not an oracle. It has real knowledge gaps, it can be talked out of a right answer, and it makes the kinds of mistakes a small model makes. The project treats that as a feature to be honest about, not a flaw to hide: a capable little assistant that lives in your pocket and keeps to itself by default, not a stand-in for the big models. Knowing exactly where that line falls is part of the work, and it points to the honest way past it, below.
A half-billion-parameter model does not know today's news, and no amount of on-device cleverness changes that. So it does not pretend to. The browser and phone builds add one explicit button, "web". Everything else still runs locally; only a question you choose to search ever leaves the device, and it is clearly the query you typed, nothing else.
When you tap it, the query goes to a small web-research service: it searches, then a larger model reads the results and writes one grounded answer with its sources. The reply comes back in its own card, tagged "from the web", so it is never mistaken for something the little local model knew on its own. Ask it who the second US president was and the on-device model will guess and get it wrong; tap web and it comes back with John Adams and the pages it read.
The split is the point. The tiny model handles conversation, and it does so privately and offline. Current facts are fetched, out loud, only when asked. Local by default, the web when you want it, and honest about which is which.
mote is being built in three steps, from the general to the almost absurd.
Step one · done
Inference in pure C
The full forward pass with no dependencies, generating coherent text. This is what runs above.
Step two · done
Quantize it
The weights pack down to 8-bit and then 4-bit integers, two weights to a byte, with the compressed arithmetic written by hand. Over six times smaller on disk, faster to run, and the quality cost of every step measured rather than assumed.
Step three · underway
Onto a microcontroller
The engine is already chip-shaped: it loads its weights from a plain array with no filesystem, and a small model fits a $6 microcontroller's memory. What is left is the board itself, and the real tokens per second and memory measured on it, no hand-waving.
A mote is the smallest thing you can still see, a speck of dust caught in a beam of light. That is the target: a language model shrunk until it is barely there, and still works.
loading commits…