Hand Controls: Learning Human Input Through 130 Lines of Computer Vision
Input devices usually disappear into the background until you try building one yourself. A mouse click feels trivial until you reduce it to distances between fingertips, noisy camera frames, and a script trying to decide whether a pinch was intentional.
This started as a learning project, not an attempt to build a full accessibility system. I wanted to understand how far basic hand tracking could go with simple logic before adding complexity. That constraint shaped the whole project.
Geometry Instead of Heavy Models
The core system is simple. Webcam frames come in through OpenCV. MediaPipe tracks hand landmarks. From there, most decisions are just geometry. Thumb and index close together triggers a left click. Index and middle close while the thumb stays extended triggers a right click. Thumb position relative to the wrist controls scrolling.
What I liked about this approach was how little machinery it needed. No gesture classifier. No training data. Just landmark relationships and threshold checks.
- Wrist position maps to cursor movement across the screen
- Fingertip distances drive click gestures
- Real time landmark overlays made debugging gesture thresholds practical
One small technical detail I found surprisingly interesting was coordinate mapping. Hand movement in camera space does not feel natural unless the image is mirrored and screen coordinates scale correctly. That mattered more than I expected.
The Decisions That Kept It Small
I kept everything in one script. About 130 lines. That was deliberate. For a project like this, abstraction too early would have hidden what was actually happening.
I also used hardcoded thresholds. Fifty pixels for a pinch. Larger spacing for separating gestures. That is not robust, but it made experimentation fast. For a first pass, fast feedback mattered more than configurability.
What Broke or Fell Short
The rough edges showed up quickly. Fixed thresholds depend on hand size and camera distance. Lighting affects detection. Rapid gestures can fire repeated clicks because there is no debounce logic. Using the wrist alone for cursor movement also introduces jitter.
If I rebuilt it, I would smooth cursor motion, add gesture cooldowns, and make thresholds adaptive instead of hardcoded. A palm centroid would probably be a better control point than the wrist.
I also learned that adding more gestures is not automatically better. Beyond a few commands, recognition conflicts start to appear. Simplicity was doing more work than I first realized.
What It Represented
This was one of those projects where the code itself was not the main point. The point was realizing interaction systems can be prototyped with surprisingly little code if you understand the primitives underneath.
It was a small project. It taught me a lot anyway.
Sometimes 130 lines is enough to change how you think about interfaces.