Madhav Sharma
Side
Build
Type
Open source
My role
Design and engineering
When
Feb 2026 to May 2026
Status
Live

Corporate Signal Translator

A gesture classifier that runs entirely in your browser. The model and the runtime that executes it come to 1.7 MB, nothing is sent to a server, and a small decision layer makes it dependable.

classifier plus the TensorFlow.js runtime that runs it
1.7 MB
Checked against the data
server calls for inference; camera frames never leave the device
0
Checked against the data
of model weights, 17,098 parameters
67 KB
Checked against the data
frames that must agree before a gesture fires
8
Checked against the data
decision-engine tests passing
45
Checked against the data

Stack

  • TypeScript
  • React
  • MediaPipe Hands
  • TensorFlow.js
  • Web Speech API
  • Vite
  • Vercel
corporate-signal-translatorruns in the browser
  1. 1.7 MBclassifier and runtime
  2. 0server calls
The decision layer, illustrated: 21 hand landmarks feed the classifier, every frame casts a vote, and a gesture only fires when the last eight agree.

Corporate Signal Translator watches your hand through the webcam and translates what it sees into the language of meetings. An open palm says “Let’s put a pin in that for now.” A thumbs up says “I am fully aligned with this initiative.” It runs entirely in the browser, with no server and no API keys, and it can read the phrase aloud.

It is a joke with real machine learning underneath. I built it in early 2026, and it was also my BTech final-year major project at Black Diamond College of Engineering and Technology.

1.7 MB, all of it on your device

The interesting constraint is where it runs. There is no inference server. The classifier, the runtime that executes it, and any personal retraining all live in the browser tab, and the app is deployed as static files.

Piece Size Where it comes from
TensorFlow.js runtime 1.6 MB, 248 KB compressed Loaded lazily, after the interface is up
Gesture classifier 67 KB of weights Served with the app
Hand tracking Separate Google’s MediaPipe Hands, from a CDN

That design decided several things at once:

  • Privacy by construction. Camera frames are processed on the device and never uploaded. There is nothing to secure on a server because there is no server.
  • No running cost. Inference is free at any number of users, because each user brings their own compute.
  • A tight budget. Everything has to fit inside a frame budget on the main thread, so the model is small, predictions run inside tf.tidy() to release memory every frame, and one dummy prediction warms the model up before the camera starts.
  • Resilience to slow networks. The CDN scripts load with a 15-second timeout and two retries, and the heavy runtime arrives in its own chunk so the page is usable before it finishes.

The problem: a classifier that flickers

The first version used hand-written rules. The second replaced them with a neural network, which recognised more gestures but introduced a different failure: per-frame predictions jumped between classes. A fist with the thumb slightly visible was read as a thumbs up. Holding one gesture fired the same phrase again and again. The model was mostly right, and the app was still unusable.

Landmarks, not pixels

The classifier never sees an image. MediaPipe Hands finds 21 points on the hand in each frame. I subtract the wrist position from every point, so the hand’s location in the frame stops mattering, and flatten the result into 63 numbers. A small network (63 inputs, two hidden layers of 128 and 64 units with dropout, 10 outputs) maps those numbers to a gesture.

The whole model is 17,098 parameters, about 67 KB of weights. It is trained offline on 8,000 synthetic hand poses: hand-built skeletons for each gesture with random noise added. TensorFlow.js itself is the heavy part, at 1.6 MB, so it loads lazily after the interface is up.

A decision layer instead of a better model

The fix for the flicker was not more training. It was a deterministic layer between the model and the user, which decides whether a prediction is allowed to become an action.

  1. Landmarks21 points from MediaPipe, made relative to the wrist
  2. ClassifySmall TensorFlow.js network, 10 classes
  3. Confidence floorBelow 0.60, the frame counts as no gesture
  4. Intent lockIgnore all frames for 2.5 s after a gesture fires
  5. Geometry checksHand rules overrule the model for known confusions
  6. Unanimous voteThe last 8 frames must all agree
  7. DeduplicateFire only if it differs from the current gesture
Every frame passes through these steps. Only a gesture that survives all of them reaches the screen or the speaker.

The geometry checks are the interesting part. For a thumbs up to count, the thumb tip has to be more than 1.3 times as far from the wrist as any other fingertip; otherwise the frame is treated as a fist. Similar rules separate a “call me” sign from a fist. They are explainable, cheap and testable, which a retrained model would not have been.

When the hand leaves the frame, the whole engine resets, so you can make the same gesture again deliberately. The speech output has its own 500 ms throttle on top.

A feature I removed

Briefly, pinching your thumb and index finger toggled the voice. A stricter version required all six conditions at once, including all three remaining fingers curled, a held duration of 700 to 900 ms, and suppression while other gestures were active. Real hands still triggered it by accident, so I removed it and put the button back. The changelog records that as the lesson: a simple control was more reliable than a clever one.

Fixing one specific confusion

Expanding from five gestures to ten introduced a new problem: the “call me” sign and a closed fist were confused for each other. I fixed it from both sides. The synthetic training data gained left and right hands, palm-out and knuckle-out, and the geometry checks gained rules that can promote a fist to “call me”, or reject a fist that is not compact enough to be one.

Personal training in the browser

Anyone can retrain the model on their own hand without sending data anywhere. Training mode records about 30 samples per gesture over three seconds, requires at least ten for every gesture, trains for 50 epochs in the browser, saves the model to IndexedDB and swaps it in without a reload. On the next visit the personal model loads first, with the default as a fallback.

What I’d do differently

  • Measure on real hands. The default model is trained and evaluated only on synthetic poses. There is no accuracy figure for real hands, so I don’t quote one.
  • Normalise for hand size. Features are made relative to the wrist but not scaled, so the geometry thresholds assume a typical distance from the camera.
  • Test the real module. The 45 decision-engine tests run against a JavaScript mirror of the logic rather than importing the TypeScript engine itself.
  • Finish the tie-break. A rule for near-equal predictions is declared in the config but not implemented yet.

More work on the same problems