Case study · Sep 2026

ika

Shadowbox in front of your laptop. It finds your habit, then says it out loud a beat before you do it. Not “you threw a jab”, which you know, but “after a right, your left hand comes down”, which you cannot see yourself.

MediaPipe
PyTorch
Python
On-device
source →
Results · a real round
Of eight calls in twenty seconds, one warned 1.2 s before the drop
what ika sees
ika drill
108.3s punch_right
108.5s punch_left
108.9s punch_right
109.2s punch_left
109.4s punch_right
109.6s punch_left
110.0s punch_right
110.0s >> LEFT HAND (after punch_right)
110.2s punch_left
110.5s punch_right
111.2s guard_down_left
“left hand”
left hand dropped 1.2 s later
111.3s
Twenty seconds of a real round from a public follow-along workout, drawn from the body landmarks ika read rather than the video. The log and the calls are what ika drill printed replaying it, at the times it printed them. Of the eight calls here, one comes 1.2 s before the drop, one lands with it, and six are followed by nothing. Ticks on the bar mark each call; the solid black one is the early hit.
263 ms
between one move ending and the next beginning. The whole budget for a warning.
3%
of shuffled real rounds report a habit that is not there. It rarely makes things up.
0
servers, uploads or API keys. The camera, the models and the voice all stay on the laptop.

The problem

Recognising a movement is solved, and it is already too late to be useful. A conventional pipeline names an action about 250 ms after it ends, and in a fight the next one starts 263 ms after that. So a coach that only recognises can only tell you what you already did.

The useful question is the other one. People repeat themselves under pressure. Can the repetition be found from a laptop camera, and can the next move be called early enough that hearing it still changes something? That is what ika was built to answer, as a terminal tool you run with one command: ika drill.

How it works

01 · see
Find the body
MediaPipe pose, 33 points a frame, from two metres back. Found in 95 to 100% of frames.
02 · read
Read the punch
A small network over how far each wrist and elbow has left its own guard, and how fast.
03 · remember
Find the habit
What your guard does after each combo, tested against chance, kept across sessions.
04 · call
Say it early
When the setup starts, it speaks the habit before the hand moves.

Everything runs on the laptop, in real time. The first round learns your resting guard in three seconds; ika history shows each habit session by session, so you can watch one fade.

Key decisions

Commit early, not certain

Holding the habit miner fixed and varying only how long recognition takes separates two things people usually blur. The call is right 58% of the time at every delay. The share of calls that arrive while they can still be used falls from 58% to 6%. The limit is latency, not intelligence, so ika decides on the start of a movement instead of waiting for its end.

Key decision · latency
A right call that arrives late is a wrong call
call is right
arrives in time to be used
0%25%50%75%100%lost to delaytypical pipeline0100250400recognition delay, ms58% rightat every delay6% usefuldown from 58%
1,000 unseen actions, replayed against habits mined from 2,000 past ones. The window being shot into is 263 ms. At a 400 ms delay the median warning arrives 119 ms after the action it was warning about has already started.

Waiting for more certainty is tempting, and it works: against an opponent who feints a third of the time, accuracy goes from 78% to 97%. It costs the call going from 100 ms to 233 ms, inside a 263 ms window. Against someone who never feints, the cheap early call is already right 94% of the time, so the bar should depend on who is in front of you.

Key decision · certainty
Against a feinting opponent, being sure costs 133 ms
opponent never feints
opponent feints 35% of the time
60%70%80%90%100%waiting costs 133 ms+0.50100 ms0.70100 ms0.85100 ms0.95233 ms0.99267 mscommit threshold · resulting latencyno feints35% feints70% up to 99%
19,001 prefix samples, held-out accuracy 90.9% across all prefix lengths. Accuracy is counted over the calls actually committed to, and the classifier declines to commit at all on 3% to 7% of movements once the bar is above 0.95.

Geometry, not pixels

The easy assumption is that a fine-tuned image model beats hand-built geometry. On the same person-grouped split it does not, by a wide margin, and the same held for bodies: describing a strike as a wrist travelling toward another torso took punch recall on UT-Interaction from 21% to 84%. A strike is a relationship, not a shape.

Results · geometry vs pixels
Hand pose beats a fine-tuned image model on accuracy, size and speed
hand pose + small MLP
fine-tuned MobileNetV3
Accuracyhigher is better
hand pose
96.1%
MobileNetV3
80.8%
Parameterslower is better
hand pose
60,578
MobileNetV3
1,552,706
Time per framelower is better
hand pose
0.06 ms
MobileNetV3
5.52 ms
Fifteen points more accurate, twenty-six times smaller and ninety-two times faster, on a split where no person appears on both sides. The backbone is not being starved: it is fine-tuned on the same photos. Landmarks simply throw away everything about the image except the geometry, and the geometry is what the gesture is.

Throw away the first reader

The first punch reader was elegant: a punch thrown at the lens barely moves in the image, it just gets bigger, so measure the palm growing, which is a depth-free closing rate.

Key decision · geometry
A punch at the lens doesn't move. It grows.
tt + 1t + 2same spot, the span s gets widerthe cameralensstraight down the optical axisfrom the sided(ln s)/dt = 1 / time to contact
Measure the span between two points on the hand. Under a pinhole camera it grows as one over distance, so its fractional rate of growth is a depth-free closing rate: one over the time to contact.

It was the right physics at the wrong distance. Nobody shadowboxes 60 cm from a screen, and two metres back, over real labelled footage, it caught 1 punch in 1,448. The body was visible almost every frame, so punches are now read from the arm. The palm reader survives as ika drill --close.

Read the guard the way a coach does

A dropped guard was first an event: a wrist crossing a line. On real footage that found 30% of real drops, and a strong planted habit was named for 12% of simulated fighters. Now ika asks what a coach asks: how high was that hand, on average, in the second after this combo, compared with after your other combos. Averaged over every repetition, the noise mostly cancels.

Results

Measured on fifteen public follow-along rounds, fourteen people, three minutes each. Every punch was labelled twice, by two separate Claude instances that never saw each other’s labels, so their agreement is measured rather than assumed. The punch reader is always scored on a person it was not trained on.

MeasuredResultAgainst
Punches caught, unseen people47% at 5.6 false/minlabellers: 90% at 4.9
Front-facing people only51% at 3.8 false/min
Palm reader, standing back1 in 1,448
Phantom guard habit, shuffled real rounds3%under 5%
Strong habit named, 12 sessions (simulated)40% of fighters93% with perfect sensors
Call lead when the habit is real (simulated)0.8 s

The loop works end to end, from a camera to a spoken call, and it rarely invents a habit. It is slow to be sure, because the punch reader misses about half of what is thrown and a two-punch setup needs both caught. For a strong habit that means a couple of weeks of short rounds.

What I learned

  • Meet real people early. Every number was clean on synthetic fighters. Real footage broke the punch reader, the guard detector and two statistical tests, and each break changed the design.
  • A null that should be clean finds bugs nothing else does. Shuffled real sessions exposed a miner reporting a habit 25% of the time, then 14%. A rank test on the guard heights brought it to 3%.
  • Timing beats accuracy. A right answer that arrives late is worth nothing, and no better classifier fixes that.
  • Say what the numbers are. The habit and the call are measured in simulation at real error rates, and the labels are a model’s. The page says so wherever they are used.

Limits

It needs you from head to hips in shot, so footwork and weight shift are not read at all. Side-on fighters read worse, and which arm threw a punch is unreliable when the shoulders overlap. The bench is 360p video; a webcam at two metres is unmeasured. The real test of the whole idea is someone who knows their own habit standing in front of it.

The full research, every command behind these numbers and the frozen side lanes (cursor control, fight footage, glasses) are in the README. 623 tests, CI across Python 3.10 to 3.12.

ìka, fingers in Yoruba. It started as a hand project and kept the name after it grew a body.