A complete transformer inference engine in C with nothing underneath it. No PyTorch, no NumPy, no BLAS, so RMSNorm, rotary embeddings, grouped-query attention with a KV cache, SwiGLU and sampling are all written out by hand and every shortcut is closed. Quantization is where it got interesting. The obvious improvement, searching for a slightly clipped scale for each group of weights, made the mathematical error smaller and the model worse, because the outliers it clipped were exactly the ones that mattered; plain rounding won. int4 turns out to decode slower than int8, since unpacking two weights from a byte costs arithmetic you do not get back, and it only wins on disk. And a gently trained behaviour can be smaller than 4-bit rounding noise and simply vanish, so training has to round bit for bit the way the engine does. A converter that rounds slightly better than training is rounding differently, and different is broken. Under 90 KB of WebAssembly, a real chat model offline on a phone, a $6 microcontroller next.
A pocket 60.5 GHz radar that measures in microns, by reading the phase of the returning wave rather than its timing: one degree is 6.9 micrometres of movement, about a tenth of the width of a hair. Two things come out of that. Count a heartbeat across still air, and recover sound from a surface the device never touches. The whole architecture then hinges on one number nobody can calculate honestly in advance, which is how fast the link to the pre-certified module really sustains, because buying the module rather than the bare die costs three to four orders of magnitude of bandwidth and the audio you can recover is only ever half the sweep rate. So the sensor is made to answer for itself: every frame carries a flag meaning I could not deliver that one on time, and the first program written for the project does nothing but push the rate up until that flag starts firing. The board passes its electrical checks with zero errors and its traces are deliberately not laid yet, because the answer decides whether they get routed for the easy job or the hard one. My first hardware project, and the line between proven and assumed is drawn on purpose.
An autonomous quant platform arranged around a refusal: the language model is never allowed to emit code. Claude researches and writes each strategy as structured data, a content-hashed artifact, and a deterministic engine trades only the frozen artifact. So you validate the artifact rather than the model, a non-deterministic author yields perfectly reproducible behaviour, and if you removed Claude entirely the thing would keep trading safely on its frozen rules. It would just stop getting smarter. Nothing reaches the market without clearing an out-of-sample gate that raises its own bar as more ideas are tried, because an in-sample Sharpe of 3.72 that becomes minus 6.99 out of sample is the whole reason the gate exists, and a strategy that passes at 1.39 still gets rejected for a 65% drawdown. In July the scoreboard was caught flattering itself, with after-hours benchmark prices up to 3% stale quietly overstating the book for the entire live window. It was found by auditing a number that looked too good instead of enjoying it, and the smaller honest figure replaced it the same day. It trades paper, it has no validated alpha, and every simple strategy tested so far has failed honest validation, which on efficient large-cap markets is the correct result.
A voice that has to sound unhurried has a problem: every text-to-speech speed control slows the words themselves, which sounds like a tape running slow. So leave the speed alone. Render the sentence once at its natural rate, then find the silences the model already produced and lengthen only those, in proportion to how long each one already was. Every register in the ladder runs at speed 1.00 and still lands on its target words per second, 2.33 down to 1.25, because the silence does all the work. Getting there meant learning that a stop consonant contains a real 40 to 120 millisecond gap of nothing, which a naive pause detector cannot tell from a breath, and on one line it stretched a 60 millisecond consonant closure by 3.1 seconds, inside the word. Around it sits the loop: the model streamed token by token and cut into sentences as they complete, so audio starts after the first one rather than the whole reply; a pre-screened holding line if that first sentence misses its budget; and barge-in that stops the audio but keeps reading the stream, because letting it go marks the wrong line as the interrupted one. Rendering is capped at two slots that queue rather than shed, on the principle that a late sentence is still spoken and a refused one never is.
A ghostwriting studio for ministry authors, where the interesting part is how an agent is stopped from inventing. Every margin note the editor writes has to carry an anchor quote, and the anchor is checked as a verbatim substring of the manuscript before the note is allowed to exist. What makes it work is where the rejection goes: rather than discarding a failed note, the checker hands the model back the string anchor_quote does not appear verbatim in the chapter, copy the exact text and retry as the tool result, so it corrects itself inside the same loop with no orchestration code at all. The author's voice is a 45-field record built by an interview that merges monotonically: identity fields are write-once and the weights only ratchet up, so contradicting yourself at question seven cannot erase question one. Fifteen end-to-end evals gate the pipeline on real tokens. The verbatim gate proves a note is anchored to real text; it does not prove the model's claim about which sermon that text came from, and the drafting pass itself is governed by prompt rules rather than a verifier.
A schema and a discipline, given a renderer. It lays an emotional history out as a typed graph, and the rigour is entirely in what the types refuse to let you write down rather than in anything the machine concludes. Thirteen node kinds, six levels of evidence from observed through reported to hypothesis, and edges that either support or contradict, so a claim cannot be recorded without declaring how it is known. A hypothesis is never promoted to a fact, competing explanations sit side by side with no winner, every tension lists unknown among its valid readings, and softer provenance is drawn with a dashed border so you can see the uncertainty before you click. There is no inference here and it does not pretend to have any: nothing is scored, nothing propagates, and every judgement is authored by a person. The only real computation is the layout, a force simulation run to convergence once and then frozen, with the repulsion capped so 200 nodes stay framable on a phone.
A Claude Code skill that hands the assistant a whole persona, an accent, an attitude and an energy that all correspond, and holds it in character for an entire session while the engineering stays senior grade. The code, the comments and the commits stay clean; only the conversation takes on the voice. Five deeply rooted voices written as love letters rather than caricatures, an invent-your-own mode, and a rule set for staying human instead of sounding like an AI in a costume. One Markdown file, no dependencies, open source.
Things built to find something out rather than to ship.
An intelligence that lives in the building rather than a browser tab, and the machine specified to hold it. The want is continuity: one system with several bodies that has been indexing my work rather than being handed a summary of it, so that what we decided last week is retrieved instead of reinvented. Most of it never reaches a large model. Around 80% of requests are a nearest-neighbour match against an embedded command registry, another 15% are a single tool call whose reliability comes from masking generation to a grammar rather than from trusting the model, and only the rest escalates. But every tier has to be resident at once or the whole thing is just an app with extra steps, which is a problem when everything I run today time-shares one 12 GB card. So: four 96 GB Blackwell cards in one enclosure, specified from part numbers, in two purchase phases, with the power ledger and the places the platform genuinely cannot match a datacentre machine written down rather than glossed.
The single most common operation inside mote, the from-scratch C inference engine, is multiply an activation by a weight and add it to a running total. A whole neural network is mostly that, done billions of times. mak is that operation built as actual hardware, a signed int8 multiply-accumulate lane taken from Verilog all the way down to a SkyWater 130nm silicon layout on an open flow. Eight int8 lanes multiplied at once, summed through an adder tree, and accumulated across cycles into a 32-bit register. Verified against a software reference across forty dot products, synthesized to roughly 4,500 cells, then placed and routed to a GDSII, the same file a foundry is handed. Designed and laid out, not fabbed, one tile rather than an accelerator, and honest about being exactly that: the atomic operation of something I built, carried down to the layer where it becomes a physical object.
Shadowbox in front of a laptop; it finds your habit and says it out loud a beat before you do it. Not “you threw a jab” but “when you finish on a right, your left hand comes down”. Timing is the whole problem: a conventional pipeline recognises an action 250 ms after it ends, and the next begins 263 ms after that. The first reader measured a palm growing on its way to the lens, the right physics at arm’s length and the wrong physics standing back, where it caught 1 punch in 1,448 on real footage. The finished drill reads punches from the arm and the guard from its height after each combo, scored on fifteen public rounds held out by person: 47% of punches, and a phantom habit in 3% of shuffled rounds. The habit and the call are measured in simulation at those error rates, and it says so.
How do you measure the error rate of a change detector without labelling anything? Point it at a desert. Two cloud-free dates over ground that has not moved in a century, and whatever area it reports as changed is, by construction, its own false-positive rate. That test started at 7% on the Aswan desert and 38.9% on Lagos harbour, and it found four real bugs on the way down to 0.23% and 1.99%. The best one: Otsu's method has no absolute floor, so a relative threshold always splits something, and handed pure sensor noise it will confidently split that too. The next best: the engine reported the sensor's 10 metres per pixel while quietly resampling every area into a fixed 256 pixel tile, so a 6 km box was really 26 metres a pixel and every square-metre figure downstream of it was wrong. Run the same harness in reverse, painting a structure of known size into real imagery, and it finds 60 metres and misses 30, which is physics rather than a bug. On radar a whole-scene brightness threshold found zero vessels, because land backscatter swamps the statistics; taking them from the water class alone found 42. Free Sentinel optical and SAR, no API keys, no GPU, and a detector that is still a brightness baseline which has never seen a human label. It says so itself, before you have to ask.
Applied LLMs
Modeling & rigor
Services & data
Ship & operate
The doppelgänger of folklore terrifies because it is a copy of you that is not you. But the body is a pattern rebuilt from new matter, memory is re-saved each time it is opened, and everyone who knows you carries a version of you you will never meet. From Goethe's consoling double and the brain's misfiled selves to Yoruba twin figures and a chatbot built from a dead man's messages, an essay on why there was never an original, and what the copies owe each other.
Terence Tao predicted that AI-generated proofs would accumulate faster than they could be verified. This measures it. Every public proof claim on erdosproblems.com from July to September 2026, 251 of them and 97% made with AI, was traced to see whether anyone ever checked it. Of the 224 on problems still open, 48% drew any comment, 19% reached any verdict and 4% were accepted by the site's moderators. Verdicts arrive within a day or not at all: after 60 days, 82% of claims still have none. Attaching a Lean proof made no difference, 19 people delivered every verdict, and one complete, formally verified proof was dismissed with a mistaken "already solved" that nobody corrected.
You cannot label the earth, so nobody measures how often a satellite change detector invents change. This runs the pipeline over ground another community already certified as invariant, where every detection is a false alarm by construction. The field has not done this because it cannot: on an empty ground truth, F1 and IoU are undefined for a perfect result and identically zero for every imperfect one, so they cannot tell one false positive from five hundred thousand. On the six CEOS-endorsed desert sites, a threshold chosen from the data reports a median 21.41% of certified-stable ground as changed, against 0.04% with an absolute floor. Otsu turns out not to degrade at all: it holds near 27% and falls 26.97 points across a single 0.25-sigma step, so behaviour on scenes containing change predicts nothing about scenes without it. Seven defects surfaced, five invisible in the outputs, including a pipeline that reports thirty times less change once a large real structure appears.
A synthesized netlist is pure logic, but a layout a foundry could build is logic plus the machinery that lets it survive fabrication: antenna diodes, well taps, timing buffers, fill. Taking one signed int8 multiply-accumulate through the open Sky130 flow at lane counts from 1 to 64, I decompose every placed layout and watch the arithmetic share fall from 59 percent of the cells to 48, crossing below half, because the overhead outgrows the logic and its pieces scale against wiring and die area rather than the gates.
An attempt needs two separate efforts: one to make it work and one to put it in front of anyone. The optimal number of simultaneous attempts depends on their sum, but the heuristic builders actually use depends on build cost alone. Agent assistance collapsed build cost and left shipping cost untouched, so the two prescriptions have diverged. Simulation plus a 48-repository archive where every stalled project is stalled at a step building cannot resolve.
When a language model proposes trading strategies, the trial count that selection-bias corrections depend on is not the number of proposals but the effective number of independent ones. Borrowing the effective-number-of-tests estimator from statistical genetics gives a principled M_eff, and on the live ÀYÒ pipeline the size of the correction turns out to be a property of the prompting protocol: diverse angles yield near-independent proposals that need no correction, while drawing repeatedly from a single narrow angle collapses fourteen proposals to three effective trials and inflates the acceptance bar enough to reject genuine strategies.
A validation gate built to be un-foolable by construction: an LLM proposes strategies as frozen, content-hashed data artifacts, and a deterministic verifier trades only what survives walk-forward out-of-sample testing, a trials-adjusted acceptance bar, and robustness sweeps. On real markets it turned an in-sample Sharpe of 3.07 into a −4.90 rejection and unmasked a +0.13 out-of-sample "edge" as a 0-of-12 fluke — deploying nothing, which is the correct result on efficient markets.
A transparent, self-pruning, MCP-native memory engine that unifies consolidation, salience-scaled adaptive forgetting, contradiction supersession, per-tenant isolation, and an exportable audit trail — evaluated on the governance properties that recall benchmarks ignore.