AbstractTwo objects, one string, and a detector that keeps losing them.
A kendama has two parts: the ken (a wooden handle with three cups and a spike) and the tama (a 60 mm ball), tethered by about 40 cm of string. Tracking them through freestyle trick footage is a small, unusually clean instance of a hard problem — two similar-scale objects, physically coupled, moving fast enough to smear at a 60 fps shutter, repeatedly occluding each other and the player's hands.
A YOLOv8n detector fine-tuned on my own clips reaches 0.907 mAP50 on a held-out split, but on video it still returns a usable box in only about two-thirds of frames — it misses precisely the fast, blurred, small instances that matter most. Rather than answer that with more labels alone, the pipeline encodes three facts a kendama player already knows: the tama falls under gravity whenever nobody is holding it, the two objects can never be further apart than the string, and the ken is attached to a hand. Those constraints, plus a second detection pass aimed only at the gaps, lift per-frame coverage from 67% to 94%.
The recovered trajectory is then worth something on its own. It drives an automatic camera operator that crops and follows the trick — dynamic zoom, a speed-coupled widen, and a slew-rate limit so it never moves faster than a human could — and a catch detector that turns pixel paths into the numbers a player actually wants: attempts, catch rate, airtime, throw height in metres.
Everything below runs on real output from the pipeline. The demos replay an actual 13.5-second clip and its actual tracking data; the camera figure re-runs the real algorithm in your browser, so switching a mechanism off breaks it in exactly the way it breaks offline.
01 / The problemWhy an ordinary detector is not enough here.
Four properties of the footage, each with a measured consequence:
- The tama is small. At the framing players actually shoot, its median box
min-side runs 31–41 px across the three dev clips. The detector was originally trained at
imgsz 640while inference ran at 1280 — and at 640 that same ball is under 20 px, near the floor of what YOLO's finest detection head resolves. - Everything is blurred. A 60 fps phone shutter is roughly 1/120 s. A thrown tama crosses more than its own diameter in that time, so the object the detector is asked to find is a streak, not a ball — and it is exactly the fast frames that carry the trick.
- The two objects are confusable and coupled. Ken and tama are similar scale and string-tethered, so an identity-agnostic tracker flip-flops between them — the single worst failure, because a swapped label corrupts every downstream measurement.
- Scene scale changes inside a clip. The camera sits on the ground while the player lunges toward and away from it. Any gate expressed in pixels — string length, speed limits, gate radii — silently means something different at the start of a clip than at the end.
The through-line: the detector is the weak link, and it fails non-randomly. It drops out during fast motion, occlusion, and distance — which is to say, during the interesting parts. A model that is 90% accurate on a shuffled validation split can still be blind for a quarter-second at a time on video, and a quarter-second is a whole trick.
02 / The pipelineDetect, then argue with the detector.
One pass produces raw detections and track IDs. Everything after it is post-processing that knows what a kendama is:
The ordering matters. Stationary-distractor suppression runs before tracking, because a motionless object misclassified as a tama will otherwise capture a track ID and hold it. Smoothing runs last, because it should smooth the final answer rather than feed its own output back into the gates.
03 / The detectorA small model and a labeling flywheel.
YOLOv8n, two classes, fine-tuned from COCO weights at imgsz 1280. The interesting
part is not the architecture — it is where the labels come from. Both Streamlit apps double as
annotation tools, so every clip anyone processes makes the next model better:
- Seed clicks. Tap the ken and tama on a handful of frames to lock the tracker for that run. Those taps are also point labels.
- SAM2 video propagation. Roughly five taps per clip are used as prompts for SAM2's video tracker, which follows both objects through every frame in between — one clip yields hundreds of tight boxes instead of ten. Boxes pass size, aspect and per-frame velocity sanity gates, and each new tap re-anchors the tracker so drift stays bounded to one segment.
- Hard negatives. The detection-check step asks the user to confirm each find as ken, tama, or not the kendama. That last verdict is the valuable one: it is what teaches the model to stop flagging shoes, spigots and dark round rocks.
Training is wrapped in a promote-or-archive loop: train, score the new weights and the
incumbent on the same validation split, write both to a registry, and only replace
best.pt if it actually won. Nothing is ever overwritten. That registry is reproduced
verbatim in §07 — including both times the rule got it wrong, once by being
too permissive and once by being too strict.
04 / Physics priorsThree things a player knows that a detector doesn't.
Each of the following is a few dozen lines. Together they are the difference between a trajectory with holes in it and one you can measure.
4.1 Gravity, calibrated from the ball itself
During a gap, the tracker predicts forward and searches near the prediction. For the tama that
prediction should curve downward — it is in free flight whenever nobody is holding it. The
original code used a hardcoded 0.5 px/frame², which is not a physical constant at
all: gravity in pixels depends on how many pixels a metre spans and on the frame rate, and both
are knowable from the data.
A competition tama is 60 mm. The median min-side of confident tama boxes therefore gives
pixels-per-metre directly, and g = 9.81 · px_per_m / fps². On this clip that is
1.50 px/frame², three times the old constant — which had been pointing the search window
hundreds of pixels above a falling ball.
4.2 The string, softened
Ken and tama cannot be further apart than the string, so the maximum plausible separation is estimable from the data — the 95th percentile of observed ken–tama distance, plus a 30% margin. On this clip: 209 px, about 6.4 tama diameters.
The first version used that as a hard veto, and it backfired in two specific situations. If the ken track was briefly wrong, the check would reject the genuinely correct tama for being too far from a phantom. And on throws toward the camera, ken and tama sit at genuinely different depths, so their apparent separation exceeds the flat-image estimate while nothing is actually wrong. Making it a rank penalty instead of a veto — candidates past 0.8× the estimate are down-weighted, and only past 1.5× are they dropped — bought 3–4 percentage points of tama coverage on two dev clips with teleports and false positives unchanged.
4.3 Depth normalization
Every pixel gate assumes the player stays at a constant distance from the camera, which is false — the camera is on the ground and the player lunges. The detections themselves carry the fix: the tama's box min-side is its diameter, so its running size is a per-frame depth signal. Smoothed with a 2 s rolling median (long enough to reject the fast dips motion blur causes) and a 0.5 s mean, clamped to [0.5, 2.0], it rescales the string length, speed limits and gate radii every frame.
This is a mixed result and worth stating plainly. On IMG_0098 — the clip where the player genuinely changes distance — ken teleports fall from 1.60 to 0.27 per 100 pairs, an 83% cut with coverage unchanged. On the constant-depth clip the same change is neutral to slightly worse for the tama. Averaged over both clips the three arms sit within noise of each other. The mechanism is correct for the failure it targets, and it is not free elsewhere; it ships because the failure it fixes is the one that corrupts measurements.
4.4 Stationary-distractor suppression
A real kendama is essentially never still. A fixed object that gets persistently misclassified — a drain cover, a dark knot in a fence — forms a tight, motionless, long-lived cluster of detections, and that signature is separable from a genuinely held-still kendama by its per-frame median step. Detections matching it are removed before tracking, so they never get to claim a track ID.
4.5 Frame-rate awareness
Every gate in the system is either a frame count (how long to keep interpolating, how wide a gap
to bridge) or a speed in px/frame — and both change meaning with capture rate. At 120 fps a
60-frame gap is half a second rather than two, and the kendama covers a quarter of the per-frame
distance. All of them, plus the tracker's track_buffer, are rescaled from the 30 fps
they were tuned at, so slow-mo footage works without retuning. Pure distances — gate radii, string
length — are correctly left alone. This matters more than it sounds: a high frame rate caps
exposure time, which is the direct fix for the motion blur causing most missed tama detections in
the first place.
05 / The auto-cameraA crop that behaves like an operator.
Once there is a trajectory, the obvious product is a camera. The crop follows the midpoint of ken and tama, but a crop that simply centres on a fast-moving target is unwatchable — it snaps, it jitters, and it loses whichever object left frame. Four mechanisms fix that, and the figure below lets you switch each one off:
- Dynamic zoom — the crop widens to contain the ken+tama box plus an 18% cushion, so a big throw pulls the frame open instead of leaving the ball outside it.
- Speed-coupled zoom — it also widens with the kendama's speed, so that its on-screen velocity stays low enough for a slow pan to keep up. The speed signal is smoothed non-causally, which means the widen begins slightly before the peak: the camera anticipates, as an operator who knows the trick would.
- Slew-rate limits — a hard cap on how far the crop can pan or zoom per frame. Zoom-out is allowed to be four times faster than zoom-in: widen fast for safety, recover slowly for smoothness.
- Person-aware framing — MediaPipe Pose biases the crop 30% toward the player, so you get a shot of a person doing a trick rather than a disembodied ball. It also calms the camera down: with the bias off, mean pan speed rises from 262 to 331 px/s, because the crop is chasing the ball literally instead of hanging off a body that moves less. A cushion clamp sits underneath as a backstop, pulling the crop back whenever the kendama would otherwise reach the frame edge.
Instrumenting the ablations turned up something I had assumed the other way round. The crop has two reasons to widen — contain the ken+tama box and the kendama is moving fast — and on this clip the containment term never binds: even at full extension the box plus its cushion stays under the base crop height. Every widening you see is the speed term. The mechanism I would have described as the safety net is the one doing the work, and the one I would have called the main rule is dormant.
06 / Catch detectionFrom pixel paths to a number a player cares about.
This part uses no machine learning at all — it is trajectory arithmetic, and it is where the project stops being a tracking demo and becomes a training tool. Real units come from the ball: a 60 mm tama gives pixels-per-metre, and fps gives seconds.
The naive approach — call it a catch whenever ken and tama are close — fails immediately, because a tama dangling on its string already sits several diameters from the ken and drifts near it constantly. The working definition separates three states:
- Attempt — a real upward velocity burst, either object exceeding 0.9 m/s, that subsequently reaches genuine separation (2.2 diameters). Velocity alone is a blip; separation alone is just string-hang.
- Catch — separation collapsing to contact (≤1.6 diameters) and staying there for at least 0.25 s. A caught tama sits about a diameter from the ken centre; a hanging one sits several.
- Miss — relative speed dying out without contact, or a 3 s timeout.
07 / EvaluationWhat to measure when you have almost no ground truth.
Twelve versions of the tracker were originally compared by watching annotated videos. That works right up until two versions trade errors — version 10 traded false positives for recall, and eyeballing genuinely could not say whether it netted out positive. So the project got a metrics script that needs no labels at all, computing four proxies over every archived run:
- Real coverage — fraction of frames with an actual detection (not filled).
- Total coverage — fraction with any position, real or recovered.
- Teleports per 100 — consecutive real detections implying an impossible jump. This is the false-positive proxy: it is what a track hopping onto a shoe looks like in numbers.
- Jitter — residual after smoothing, the noise proxy.
Where sparse verified labels do exist — from the app's own detection-check step — the script also reports true hit rate, median centre error, and how often a track sits on a confirmed false positive.
7.1 The version-10 regression
The second-pass re-detection was a good idea implemented without discipline: it accepted any in-class detection above 0.10 confidence within 200 px of the prediction, applying none of the safeguards the main pass enforces — no string constraint, no colour consistency, no velocity check, no hand logic. Worse, it wrote its results back marked as real detections, which the gap-filler then treated as trusted anchors and built bridges to.
The cost was measurable the moment there was something to measure it with — teleports per 100 pairs went from 0.87 across v5–v9 to 8.80 — and visible in the footage: on the hard clip, a 0.36-confidence "tama" sitting on the player's back shoe, 730 px from a correctly-boxed ken, well beyond the 605 px string estimate the main pass was enforcing. Applying the same string check, the same colour distance, and a 0.20 threshold kept most of the recall the pass had bought and brought teleports back to 2.80.
The general lesson: a stage that bypasses your invariants and then labels its output as trustworthy is worse than the same stage being merely wrong. It launders low-confidence guesses into anchors that everything downstream builds on.
7.2 The promotion gate that got fooled
The retraining loop promotes a new model when it beats the incumbent on validation mAP50, both scored on the same split. In July a new model scored 0.924 against the incumbent's 0.324 and was promoted automatically. It was a clear regression.
The validation split had filled up with frames from friends' clips — a different domain, and by then the majority of the labels. The new model was genuinely better there and much worse on my own footage, where held-out tama coverage collapsed from 0.82 to 0.41. A detection-level metric on a shifted distribution said "ship it"; the actual task said the opposite.
The fix was to stop promoting on mAP alone. Model selection now has to clear a task-level gate on held-out clips — tracking coverage must not regress, verified false-positive hits must stay at zero, teleports must stay under threshold — and mAP is treated as a sanity check rather than a decision rule. The next model was promoted only after passing that gate; the one after it was rejected despite competitive mAP. The registry below is the audit trail, unedited.
7.3 Where the system actually stands
08 / LimitationsWhat this does not do.
- It is one player's domain. The training footage is one person, mostly one dark tama, ground-level selfie angle, four locations, one New England winter. As a personal tool that is a feature — overfitting to your own clips is the point. It will degrade on other players, pale tamas, and indoor gyms, and that is a data problem, not an architecture one. The friend-facing app exists to broaden it.
- The post-processing is a hand-rolled Kalman smoother. Forward prediction with gating, plus backward information, plus averaging, is structurally a constant-acceleration Kalman filter and an RTS smoother — implemented as three passes with about twenty coupled constants. It works and every line is understood, but each new failure mode is currently answered with another constant, and the constants interact. A proper filter would expose four parameters instead of twenty and give uncertainty for free; the domain rules would survive as gating on top.
- Coverage is not correctness. A recovered position is a physically plausible guess, not a measurement. The metrics separate the two deliberately, and the sparse verified labels exist precisely because the proxies can be gamed by a filter that confidently invents smooth trajectories.
- Catch detection is validated on very few events. Four attempts on one clip is a demonstration, not an evaluation. The thresholds are physically motivated but tuned by inspection, and no labelled catch dataset exists yet.
MethodsHow this page was made.
Every number, trace and box on this page is real output from the pipeline, exported from the
project's own artifacts — trajectory_IMG_0043.json for the tracking data,
eval_report/metrics.csv for the version history, models/registry.jsonl
for the training log, and trick_stats.py for the catch events. Nothing is
illustrative or hand-drawn.
The demo clip is the source video, downscaled to 540×960 and re-encoded to H.264; boxes, trails and camera crops are composited live in a canvas from the trajectory data rather than baked into the video, which is why they can be toggled. The camera figure is a faithful JavaScript port of the offline renderer's arithmetic — the same Gaussian smoothing, the same asymmetric slew limits, the same cushion clamp — and it validates itself on load against the Python result for the default settings, reporting the maximum per-frame disagreement in the figure header.
That check is worth one honest footnote. The camera centre agrees to within a rounding step of the exported data across all 808 frames, and the crop size agrees exactly on all but a single frame, where it differs by two pixels. The cause is instructive rather than a bug: the zoom passes through a slew-rate limiter, a sequential filter with a branch in it, and NumPy and JavaScript accumulate the preceding Gaussian convolution in a different order. On one frame the difference between them — around 10−13 px — falls exactly on the branch, the two sequences separate by one limiter step, and they re-converge the next frame. Any port of a branching sequential filter has this property; it is only visible here because the check is reported instead of assumed.
No framework, no build step. One HTML file, one generated data file, one script, one video.