← All posts

Building a Voice Assistant Before I Knew What an Assistant Really Needed

2023Python · Voice AI

Friction hides in small actions. Opening the same tabs every day, searching the same things, reaching for the keyboard when voice should be enough. I built this as a first serious attempt at making software respond like a conversation, even if the conversation was primitive.

The project was small. About a hundred lines. One file. That was part of the point. I wanted to understand how speech recognition, command parsing, and text to speech fit together without hiding behind frameworks.

Wiring the Loop

Everything revolved around one loop. Listen. Parse. Execute. Speak. Repeat. The interesting part was not any single feature. It was getting the loop to feel continuous.

I used speech_recognition with Google's API for input, pyttsx3 for output, and a simple if and elif command router for behavior. Saying "play" triggered YouTube through pywhatkit. Saying "wikipedia" pulled one sentence summaries. "Open" mapped into hardcoded sites and Google Workspace shortcuts.

Looking back, one small detail I still like is the separation between listen(), speak(), and execute_command(). It made a beginner project feel modular instead of tangled.

Simple Decisions That Mattered

The command parser was basic substring matching. Not clever. But it forced a design question I still think about. When is simple enough actually enough.

The error handling in listen() taught me a lot. If recognition failed, it retried. That made the assistant feel more alive, but it also exposed a flaw. I used recursive retries. It worked, but it was the wrong long term structure.

What Broke First

Voice assistants look impressive until ambiguity shows up. Then the cracks are obvious. "Time" could match unintended phrases. Wikipedia disambiguation was clumsy. The joke command told exactly one joke forever.

I also hardcoded too much. Websites were fixed. Responses were fixed. There was no context, no memory, no wake word, no real conversation. It was closer to voice triggered utilities than an assistant.

If I rebuilt it now, I would replace substring matching with intent parsing, remove recursive retries, and add a lightweight event driven architecture instead of an infinite listening loop.

What It Actually Taught Me

I started this thinking I was learning voice interfaces. I was really learning edge cases. Software feels intelligent only until the first misunderstood input.

That changed how I think about building. Features are easy to count. Failure modes are where the engineering starts.

It was a small project. It still changed how I debug.

View all projects →View on GitHub ↗