- Side
- Build
- Type
- Open source
- My role
- Design and engineering
- When
- Feb 2026 to May 2026
- Status
- Live
Corporate Signal Translator
A gesture classifier that runs entirely in your browser. The model and the runtime that executes it come to 1.7 MB, nothing is sent to a server, and a small decision layer makes it dependable.
- classifier plus the TensorFlow.js runtime that runs it
- 1.7 MB
- Checked against the data
- server calls for inference; camera frames never leave the device
- 0
- Checked against the data
- of model weights, 17,098 parameters
- 67 KB
- Checked against the data
- frames that must agree before a gesture fires
- 8
- Checked against the data
- decision-engine tests passing
- 45
- Checked against the data
Stack
- 1.7 MBclassifier and runtime
- 0server calls
Corporate Signal Translator watches your hand through the webcam and translates what it sees into the language of meetings. An open palm says “Let’s put a pin in that for now.” A thumbs up says “I am fully aligned with this initiative.” It runs entirely in the browser, with no server and no API keys, and it can read the phrase aloud.
It is a joke with real machine learning underneath. I built it in early 2026, and it was also my BTech final-year major project at Black Diamond College of Engineering and Technology.
1.7 MB, all of it on your device
The interesting constraint is where it runs. There is no inference server. The classifier, the runtime that executes it, and any personal retraining all live in the browser tab, and the app is deployed as static files.
| Piece | Size | Where it comes from |
|---|---|---|
| TensorFlow.js runtime | 1.6 MB, 248 KB compressed | Loaded lazily, after the interface is up |
| Gesture classifier | 67 KB of weights | Served with the app |
| Hand tracking | Separate | Google’s MediaPipe Hands, from a CDN |
That design decided several things at once:
- Privacy by construction. Camera frames are processed on the device and never uploaded. There is nothing to secure on a server because there is no server.
- No running cost. Inference is free at any number of users, because each user brings their own compute.
- A tight budget. Everything has to fit inside a frame budget on the main thread, so the model is small, predictions run inside
tf.tidy()to release memory every frame, and one dummy prediction warms the model up before the camera starts. - Resilience to slow networks. The CDN scripts load with a 15-second timeout and two retries, and the heavy runtime arrives in its own chunk so the page is usable before it finishes.
The problem: a classifier that flickers
The first version used hand-written rules. The second replaced them with a neural network, which recognised more gestures but introduced a different failure: per-frame predictions jumped between classes. A fist with the thumb slightly visible was read as a thumbs up. Holding one gesture fired the same phrase again and again. The model was mostly right, and the app was still unusable.
Landmarks, not pixels
The classifier never sees an image. MediaPipe Hands finds 21 points on the hand in each frame. I subtract the wrist position from every point, so the hand’s location in the frame stops mattering, and flatten the result into 63 numbers. A small network (63 inputs, two hidden layers of 128 and 64 units with dropout, 10 outputs) maps those numbers to a gesture.
The whole model is 17,098 parameters, about 67 KB of weights. It is trained offline on 8,000 synthetic hand poses: hand-built skeletons for each gesture with random noise added. TensorFlow.js itself is the heavy part, at 1.6 MB, so it loads lazily after the interface is up.
A decision layer instead of a better model
The fix for the flicker was not more training. It was a deterministic layer between the model and the user, which decides whether a prediction is allowed to become an action.
- Landmarks21 points from MediaPipe, made relative to the wrist
- ClassifySmall TensorFlow.js network, 10 classes
- Confidence floorBelow 0.60, the frame counts as no gesture
- Intent lockIgnore all frames for 2.5 s after a gesture fires
- Geometry checksHand rules overrule the model for known confusions
- Unanimous voteThe last 8 frames must all agree
- DeduplicateFire only if it differs from the current gesture
The geometry checks are the interesting part. For a thumbs up to count, the thumb tip has to be more than 1.3 times as far from the wrist as any other fingertip; otherwise the frame is treated as a fist. Similar rules separate a “call me” sign from a fist. They are explainable, cheap and testable, which a retrained model would not have been.
When the hand leaves the frame, the whole engine resets, so you can make the same gesture again deliberately. The speech output has its own 500 ms throttle on top.
A feature I removed
Briefly, pinching your thumb and index finger toggled the voice. A stricter version required all six conditions at once, including all three remaining fingers curled, a held duration of 700 to 900 ms, and suppression while other gestures were active. Real hands still triggered it by accident, so I removed it and put the button back. The changelog records that as the lesson: a simple control was more reliable than a clever one.
Fixing one specific confusion
Expanding from five gestures to ten introduced a new problem: the “call me” sign and a closed fist were confused for each other. I fixed it from both sides. The synthetic training data gained left and right hands, palm-out and knuckle-out, and the geometry checks gained rules that can promote a fist to “call me”, or reject a fist that is not compact enough to be one.
Personal training in the browser
Anyone can retrain the model on their own hand without sending data anywhere. Training mode records about 30 samples per gesture over three seconds, requires at least ten for every gesture, trains for 50 epochs in the browser, saves the model to IndexedDB and swaps it in without a reload. On the next visit the personal model loads first, with the default as a fallback.
What I’d do differently
- Measure on real hands. The default model is trained and evaluated only on synthetic poses. There is no accuracy figure for real hands, so I don’t quote one.
- Normalise for hand size. Features are made relative to the wrist but not scaled, so the geometry thresholds assume a typical distance from the camera.
- Test the real module. The 45 decision-engine tests run against a JavaScript mirror of the logic rather than importing the TypeScript engine itself.
- Finish the tie-break. A rule for near-equal predictions is declared in the config but not implemented yet.