Work

Everything I've built

The full archive — shipped products, experiments, and the small things in between, newest work first.

Featured

mote

CASE STUDY →Aug 2026

A complete transformer inference engine in C with nothing underneath it. No PyTorch, no NumPy, no BLAS, so RMSNorm, rotary embeddings, grouped-query attention with a KV cache, SwiGLU and sampling are all written out by hand and every shortcut is closed. Quantization is where it got interesting. The obvious improvement, searching for a slightly clipped scale for each group of weights, made the mathematical error smaller and the model worse, because the outliers it clipped were exactly the ones that mattered; plain rounding won. int4 turns out to decode slower than int8, since unpacking two weights from a byte costs arithmetic you do not get back, and it only wins on disk. And a gently trained behaviour can be smaller than 4-bit rounding noise and simply vanish, so training has to round bit for bit the way the engine does. A converter that rounds slightly better than training is rounding differently, and different is broken. Under 90 KB of WebAssembly, a real chat model offline on a phone, a $6 microcontroller next.

Embedded C
Transformer
Quantization
Open Source

ràdá

CASE STUDY →Aug 2026

A pocket 60.5 GHz radar that measures in microns, by reading the phase of the returning wave rather than its timing: one degree is 6.9 micrometres of movement, about a tenth of the width of a hair. Two things come out of that. Count a heartbeat across still air, and recover sound from a surface the device never touches. The whole architecture then hinges on one number nobody can calculate honestly in advance, which is how fast the link to the pre-certified module really sustains, because buying the module rather than the bare die costs three to four orders of magnitude of bandwidth and the audio you can recover is only ever half the sweep rate. So the sensor is made to answer for itself: every frame carries a flag meaning I could not deliver that one on time, and the first program written for the project does nothing but push the rate up until that flag starts firing. The board passes its electrical checks with zero errors and its traces are deliberately not laid yet, because the answer decides whether they get routed for the easy job or the hard one. My first hardware project, and the line between proven and assumed is drawn on purpose.

60.5 GHz radar
nRF52840
Embedded C
KiCad

àyò

CASE STUDY →Jul 2026

An autonomous quant platform arranged around a refusal: the language model is never allowed to emit code. Claude researches and writes each strategy as structured data, a content-hashed artifact, and a deterministic engine trades only the frozen artifact. So you validate the artifact rather than the model, a non-deterministic author yields perfectly reproducible behaviour, and if you removed Claude entirely the thing would keep trading safely on its frozen rules. It would just stop getting smarter. Nothing reaches the market without clearing an out-of-sample gate that raises its own bar as more ideas are tried, because an in-sample Sharpe of 3.72 that becomes minus 6.99 out of sample is the whole reason the gate exists, and a strategy that passes at 1.39 still gets rejected for a 65% drawdown. In July the scoreboard was caught flattering itself, with after-hours benchmark prices up to 3% stale quietly overstating the book for the entire live window. It was found by auditing a number that looked too good instead of enjoying it, and the smaller honest figure replaced it the same day. It trades paper, it has no validated alpha, and every simple strategy tested so far has failed honest validation, which on efficient large-cap markets is the correct result.

Python
Claude
Alpaca
pandas

sotto

CASE STUDY →Sep 2026

A voice that has to sound unhurried has a problem: every text-to-speech speed control slows the words themselves, which sounds like a tape running slow. So leave the speed alone. Render the sentence once at its natural rate, then find the silences the model already produced and lengthen only those, in proportion to how long each one already was. Every register in the ladder runs at speed 1.00 and still lands on its target words per second, 2.33 down to 1.25, because the silence does all the work. Getting there meant learning that a stop consonant contains a real 40 to 120 millisecond gap of nothing, which a naive pause detector cannot tell from a breath, and on one line it stretched a 60 millisecond consonant closure by 3.1 seconds, inside the word. Around it sits the loop: the model streamed token by token and cut into sentences as they complete, so audio starts after the first one rather than the whole reply; a pre-screened holding line if that first sentence misses its budget; and barge-in that stops the audio but keeps reading the stream, because letting it go marks the wrong line as the interrupted one. Rendering is capped at two slots that queue rather than shed, on the principle that a late sentence is still spoken and a refused one never is.

Python
FastAPI
Kokoro TTS
Self-hosted

the scribe

LIVE SITE ↗Jun 2026

A ghostwriting studio for ministry authors, where the interesting part is how an agent is stopped from inventing. Every margin note the editor writes has to carry an anchor quote, and the anchor is checked as a verbatim substring of the manuscript before the note is allowed to exist. What makes it work is where the rejection goes: rather than discarding a failed note, the checker hands the model back the string anchor_quote does not appear verbatim in the chapter, copy the exact text and retry as the tool result, so it corrects itself inside the same loop with no orchestration code at all. The author's voice is a 45-field record built by an interview that merges monotonically: identity fields are write-once and the weights only ratchet up, so contradicting yourself at question seven cannot erase question one. Fifteen end-to-end evals gate the pipeline on real tokens. The verbatim gate proves a note is anchored to real text; it does not prove the model's claim about which sermon that text came from, and the drafting pass itself is governed by prompt rules rather than a verifier.

TypeScript
Next.js
OpenAI
Postgres
Prisma

dimeji

CASE STUDY →Jul 2026

A schema and a discipline, given a renderer. It lays an emotional history out as a typed graph, and the rigour is entirely in what the types refuse to let you write down rather than in anything the machine concludes. Thirteen node kinds, six levels of evidence from observed through reported to hypothesis, and edges that either support or contradict, so a claim cannot be recorded without declaring how it is known. A hypothesis is never promoted to a fact, competing explanations sit side by side with no winner, every tension lists unknown among its valid readings, and softer provenance is drawn with a dashed border so you can see the uncertainty before you click. There is no inference here and it does not pretend to have any: nothing is scored, nothing propagates, and every judgement is authored by a person. The only real computation is the layout, a force simulation run to convergence once and then frozen, with the repulsion capped so 200 nodes stay framable on a phone.

Framework
React Flow
TypeScript
Epistemics

vibe

CASE STUDY →Jul 2026

A Claude Code skill that hands the assistant a whole persona, an accent, an attitude and an energy that all correspond, and holds it in character for an entire session while the engineering stays senior grade. The code, the comments and the commits stay clean; only the conversation takes on the voice. Five deeply rooted voices written as love letters rather than caricatures, an invent-your-own mode, and a rule set for staying human instead of sounding like an AI in a costume. One Markdown file, no dependencies, open source.

Claude Code
Skill
Open Source
Markdown

Experiments

node

CASE STUDY →Aug 2026

An intelligence that lives in the building rather than a browser tab, and the machine specified to hold it. The want is continuity: one system with several bodies that has been indexing my work rather than being handed a summary of it, so that what we decided last week is retrieved instead of reinvented. Most of it never reaches a large model. Around 80% of requests are a nearest-neighbour match against an embedded command registry, another 15% are a single tool call whose reliability comes from masking generation to a grammar rather than from trusting the model, and only the rest escalates. But every tier has to be resident at once or the whole thing is just an app with extra steps, which is a problem when everything I run today time-shares one 12 GB card. So: four 96 GB Blackwell cards in one enclosure, specified from part numbers, in two purchase phases, with the power ledger and the places the platform genuinely cannot match a datacentre machine written down rather than glossed.

Threadripper PRO
384 GB VRAM
Local inference
Hardware

mak

CASE STUDY →Aug 2026

The single most common operation inside mote, the from-scratch C inference engine, is multiply an activation by a weight and add it to a running total. A whole neural network is mostly that, done billions of times. mak is that operation built as actual hardware, a signed int8 multiply-accumulate lane taken from Verilog all the way down to a SkyWater 130nm silicon layout on an open flow. Eight int8 lanes multiplied at once, summed through an adder tree, and accumulated across cycles into a 32-bit register. Verified against a software reference across forty dot products, synthesized to roughly 4,500 cells, then placed and routed to a GDSII, the same file a foundry is handed. Designed and laid out, not fabbed, one tile rather than an accelerator, and honest about being exactly that: the atomic operation of something I built, carried down to the layer where it becomes a physical object.

Verilog · RTL
Sky130
OpenLane
Silicon

ika

CASE STUDY →Sep 2026

Shadowbox in front of a laptop; it finds your habit and says it out loud a beat before you do it. Not “you threw a jab” but “when you finish on a right, your left hand comes down”. Timing is the whole problem: a conventional pipeline recognises an action 250 ms after it ends, and the next begins 263 ms after that. The first reader measured a palm growing on its way to the lens, the right physics at arm’s length and the wrong physics standing back, where it caught 1 punch in 1,448 on real footage. The finished drill reads punches from the arm and the guard from its height after each combo, scored on fifteen public rounds held out by person: 47% of punches, and a phantom habit in 3% of shuffled rounds. The habit and the call are measured in simulation at those error rates, and it says so.

MediaPipe
PyTorch
Python
On-device

jùwọ̀n

CASE STUDY →Aug 2026

How do you measure the error rate of a change detector without labelling anything? Point it at a desert. Two cloud-free dates over ground that has not moved in a century, and whatever area it reports as changed is, by construction, its own false-positive rate. That test started at 7% on the Aswan desert and 38.9% on Lagos harbour, and it found four real bugs on the way down to 0.23% and 1.99%. The best one: Otsu's method has no absolute floor, so a relative threshold always splits something, and handed pure sensor noise it will confidently split that too. The next best: the engine reported the sensor's 10 metres per pixel while quietly resampling every area into a fixed 256 pixel tile, so a 6 km box was really 26 metres a pixel and every square-metre figure downstream of it was wrong. Run the same harness in reverse, painting a structure of known size into real imagery, and it finds 60 metres and misses 30, which is physics rather than a bug. On radar a whole-scene brightness threshold found zero vessels, because land backscatter swamps the statistics; taking them from the water class alone found 42. Free Sentinel optical and SAR, no API keys, no GPU, and a detector that is still a brightness baseline which has never seen a human label. It says so itself, before you have to ask.

Sentinel-2 · SAR
NumPy
Python
Calibration

Archive

folio

LIVE SITE ↗Apr 2026

A decision layer that sits on top of your product analytics. Connect PostHog or the Folio SDK, describe a growth goal in plain English, and Folio drafts the in-app nudges — banner, modal, and dashboard copy with CTAs — that you approve before they ship, then serves them through a headless decisions API with serve→click→convert attribution on every recommendation.

TypeScript
Next.js
Supabase
React

nūr

May 2026

A shipped, premium AI Qur'an companion. A SwiftUI iOS app on a FastAPI/Postgres backend with Claude-powered reflections behind a safety + eval layer, geolocated prayer times, multi-reciter recitation audio, subscriptions, home-screen widgets, and an automated, personalized retention/notifications engine.

SwiftUI
FastAPI
Postgres
Claude

atom

LIVE SITE ↗Jun 2026

A self-hosted LLM pretraining pipeline built from scratch — data prep, training, and evaluation for small language models, owned end to end rather than calling someone else's API. The deep end of the stack.

Python
PyTorch
CUDA

orator

Jun 2026

An AI public-speaking coach with real-time conversational voice agents (ElevenLabs) and an LLM feedback pass that scores delivery and coaches each rehearsal. Credit-based billing, Google auth, the works.

Next.js
FastAPI
ElevenLabs

veil

A competitive-intelligence OSINT platform: large-scale, resilient web scraping on Bright Data feeding structured, queryable intelligence — proxies, layout-drift tolerance, dedup, and freshness handled at scale.

Python
Bright Data
Postgres

greenlight

Jun 2026

Agentic prior-authorization for healthcare payers: LLM agents verify medical necessity with paragraph-level citations, gated by a deterministic citation verifier that blocks fabricated quotes outright — agents you can actually trust.

Python
LangGraph
Anthropic

glean

Jun 2026

An AI editing studio that turns a long recording into a cut master plus captioned vertical clips — Claude tool-use proposes the cut list, ffmpeg renders it. Background jobs, transcription, and storage wired end to end.

FastAPI
Claude
ffmpeg

whisk

May 2026

An "aesthetic" AI recipe app where a single codebase ships as six niche App Store apps via an env switch. SwiftUI (WidgetKit, Live Activities, StoreKit) on a FastAPI/Postgres backend, with Claude recipe generation and fridge-photo vision.

SwiftUI
FastAPI
Claude

medula

May 2026

A multi-agent AI platform — a Next.js studio over a distributed microservice backend (API gateway, agent engine, context engine), with Postgres and Qdrant powering long-term vector memory.

Next.js
FastAPI
Qdrant

nuj

Mar 2026

A relationship-wellness app for couples — AI daily moments, quizzes, check-ins, and shared journaling. A native Swift app (widget + share extensions) on a FastAPI/Supabase backend with pgvector and push.

SwiftUI
FastAPI
Supabase

lumi

Oct 2025

An AI voice-companion iOS app — real-time conversational agents (ElevenLabs) with a lock-screen Live Activity, on a FastAPI backend.

SwiftUI
ElevenLabs
FastAPI

netforge

GITHUB2024

A lightweight, cross-platform networking abstraction layer designed to simplify HTTP networking operations across iOS, macOS, and Linux.

Swift
SwiftNIO
SPM

hydrate

A health-focused iOS and watchOS application that helps users track their daily water intake in correlation with their health metrics.

Swift
HealthKit
CoreML
WatchKit

notely

2023

A podcasting app in SwiftUI: record, and move the audio between Apple devices without losing it. Early work, and plainly that rather than dressed up as more.

Swift
SwiftUI
AVFoundation

pixels

Advanced image processing iOS application leveraging CoreImage and Metal for high-fidelity editing capabilities. Features real-time photo synchronization across Apple devices using CloudKit, delivering exceptional performance for complex visual tasks.

Swift
Metal
CoreImage

itype

Smart keyboard utilizing transformer-based language models and federated learning for privacy-preserving text predictions. Implemented efficient on-device inference with model quantization.

Swift
Python
CoreML

kept journal

AI-powered journaling app featuring sentiment analysis, topic modeling, and personalized content recommendations using fine-tuned BERT models optimized for iOS.

Swift
Python
TensorFlow

kagaku

Content recommendation engine using collaborative filtering and neural network-based deep learning models, with custom loss functions for improved accuracy

Python
PyTorch
FastAPI

docgen

Command-line tool that automates comprehensive documentation generation through static code analysis. Features intelligent source code parsing, custom template engine for flexible documentation formats, and support for multiple programming languages with extensible architecture.

Python
AST
Jinja2

annotate

Advanced iPad note-taking platform leveraging machine learning for real-time handwriting analysis and semantic content organization. Implements on-device natural language processing for intelligent note categorization, featuring contextual learning algorithms for personalized content suggestions and study pattern optimization.

Swift
CoreML
PencilKit
NLP

lab controller

Sophisticated laboratory control system utilizing Bluetooth Low Energy (BLE) for precise Arduino microcontroller management. Features real-time bidirectional communication protocol, custom GATT service implementation, and hardware abstraction layer for seamless iOS-Arduino integration in laboratory environments.

Swift
CoreBluetooth
Arduino
C++

menudancer

macOS menu bar application featuring real-time audio visualization through an animated character. Implements low-latency audio processing using AVFoundation, with custom Core Animation choreography synced to frequency spectrum analysis of microphone input.

Swift
AVFoundation
CoreAnimation