Case study · Sep 2026
ika
Shadowbox in front of your laptop. It finds your habit, then says it out loud a beat before you do it. Not “you threw a jab”, which you know, but “after a right, your left hand comes down”, which you cannot see yourself.
The problem
Recognising a movement is solved, and it is already too late to be useful. A conventional pipeline names an action about 250 ms after it ends, and in a fight the next one starts 263 ms after that. So a coach that only recognises can only tell you what you already did.
The useful question is the other one. People repeat themselves under pressure. Can the repetition be found from a laptop camera, and can the next move be called early enough that hearing it still changes something? That is what ika was built to answer, as a terminal tool you run with one command: ika drill.
How it works
Everything runs on the laptop, in real time. The first round learns your resting guard in three seconds; ika history shows each habit session by session, so you can watch one fade.
Key decisions
Commit early, not certain
Holding the habit miner fixed and varying only how long recognition takes separates two things people usually blur. The call is right 58% of the time at every delay. The share of calls that arrive while they can still be used falls from 58% to 6%. The limit is latency, not intelligence, so ika decides on the start of a movement instead of waiting for its end.
Waiting for more certainty is tempting, and it works: against an opponent who feints a third of the time, accuracy goes from 78% to 97%. It costs the call going from 100 ms to 233 ms, inside a 263 ms window. Against someone who never feints, the cheap early call is already right 94% of the time, so the bar should depend on who is in front of you.
Geometry, not pixels
The easy assumption is that a fine-tuned image model beats hand-built geometry. On the same person-grouped split it does not, by a wide margin, and the same held for bodies: describing a strike as a wrist travelling toward another torso took punch recall on UT-Interaction from 21% to 84%. A strike is a relationship, not a shape.
Throw away the first reader
The first punch reader was elegant: a punch thrown at the lens barely moves in the image, it just gets bigger, so measure the palm growing, which is a depth-free closing rate.
It was the right physics at the wrong distance. Nobody shadowboxes 60 cm from a screen, and two metres back, over real labelled footage, it caught 1 punch in 1,448. The body was visible almost every frame, so punches are now read from the arm. The palm reader survives as ika drill --close.
Read the guard the way a coach does
A dropped guard was first an event: a wrist crossing a line. On real footage that found 30% of real drops, and a strong planted habit was named for 12% of simulated fighters. Now ika asks what a coach asks: how high was that hand, on average, in the second after this combo, compared with after your other combos. Averaged over every repetition, the noise mostly cancels.
Results
Measured on fifteen public follow-along rounds, fourteen people, three minutes each. Every punch was labelled twice, by two separate Claude instances that never saw each other’s labels, so their agreement is measured rather than assumed. The punch reader is always scored on a person it was not trained on.
| Measured | Result | Against |
|---|---|---|
| Punches caught, unseen people | 47% at 5.6 false/min | labellers: 90% at 4.9 |
| Front-facing people only | 51% at 3.8 false/min | |
| Palm reader, standing back | 1 in 1,448 | |
| Phantom guard habit, shuffled real rounds | 3% | under 5% |
| Strong habit named, 12 sessions (simulated) | 40% of fighters | 93% with perfect sensors |
| Call lead when the habit is real (simulated) | 0.8 s |
The loop works end to end, from a camera to a spoken call, and it rarely invents a habit. It is slow to be sure, because the punch reader misses about half of what is thrown and a two-punch setup needs both caught. For a strong habit that means a couple of weeks of short rounds.
What I learned
- Meet real people early. Every number was clean on synthetic fighters. Real footage broke the punch reader, the guard detector and two statistical tests, and each break changed the design.
- A null that should be clean finds bugs nothing else does. Shuffled real sessions exposed a miner reporting a habit 25% of the time, then 14%. A rank test on the guard heights brought it to 3%.
- Timing beats accuracy. A right answer that arrives late is worth nothing, and no better classifier fixes that.
- Say what the numbers are. The habit and the call are measured in simulation at real error rates, and the labels are a model’s. The page says so wherever they are used.
Limits
It needs you from head to hips in shot, so footwork and weight shift are not read at all. Side-on fighters read worse, and which arm threw a punch is unreliable when the shoulders overlap. The bench is 360p video; a webcam at two metres is unmeasured. The real test of the whole idea is someone who knows their own habit standing in front of it.
The full research, every command behind these numbers and the frozen side lanes (cursor control, fight footage, glasses) are in the README. 623 tests, CI across Python 3.10 to 3.12.
ìka, fingers in Yoruba. It started as a hand project and kept the name after it grew a body.