Key takeaways
- Our narration used to cost a fraction of a cent per line. Small enough to ignore on paper, big enough to skip on a busy day.
- Anything with a price per use turns into a small decision, and a small decision made forty times a week is how a daily publishing habit quietly dies.
- It now runs on the graphics card already sitting in the desk, for nothing. The paid service is the fallback rather than the default.
- The part that made it safe to run unattended is not the voice model. It is the check that listens back.
- Every line is transcribed and compared to the script before it is used. If the machine did not say what we wrote, that take is thrown out and rolled again.
We publish video most days. Some of it is narrated by a clone of Michael's voice, built from a recording he made once for something else entirely. The clone is good. The clone was also rented, and rented turned out to matter more than good.
The problem was never the money
The per line cost was tiny. Nobody was going to notice it on a statement.
What we noticed was the hesitation. Every voiceover became a small yes or no. Is this piece worth spending on. Should we wait and do a batch. On a light day the answer is yes, obviously, go. On a heavy day the answer is later, and later usually means never.
That is the real tax on anything with a price per use inside a daily process. It is not the balance going down. It is the decision that keeps reappearing, and the way a small recurring decision eventually gets answered no by default, quietly, without anybody choosing to stop.
Michael put it plainly. He was tired of paying to generate it. So we moved it in house.
The voice was picked by ear, not by settings
The clone is built from a twenty two second reference, cut out of a raw recording he had already made months earlier. Same take, same microphone, same room as the paid clone was built from, which is why the two intercut without anybody hearing a seam. Then we made five variants at different settings and Michael listened to all of them next to the paid one.
The flattest, steadiest read won. Not the liveliest, not the one with the most energy in it, and that surprised us. His words were that it sounded most like him all the way through, which is the honest test. A take that starts strong and drifts is worse than a take that is even, because the drift is the part a listener notices.
So the settings are locked, and there is a note attached to them saying not to add energy without asking. That is worth writing down anywhere a machine imitates a person. The version that sounds most impressive in isolation is frequently not the version that sounds most like them.
The gate is the whole build
Here is the part that matters if you are automating any generative step at all.
These models are stochastic. The same sentence at the same settings comes out clean four times and stumbles on the fifth. One ordinary phrase about building a follow up flow tripped four of the five variants we tested. Not a rare edge case. A normal sentence.
If a person is listening, fine, they roll it again. If nobody is, a stumble ships, and the pipelines using this run unattended overnight.
So every chunk gets transcribed back by a speech recognition model, and the transcript is compared word for word against the script it was supposed to read. If the mismatch is above a threshold, the seed changes and it tries again, up to four times, keeping the best attempt it has seen. Cost is a few seconds per chunk and it removed the entire category of failure.
Two honest notes about that threshold, because it is easy to read it as a quality score.
It is not one. The transcriber itself mishears somewhere between five and ten percent of perfectly clean audio, so the number we accept is a floor rather than a bar. Setting it tighter would reject good takes for the listener's mistakes rather than the speaker's.
And a chunk that still fails after four tries is usually not a bad take. It is a bad sentence. Something in the phrasing is hard to say, and the fix is rewriting the line rather than rolling a fifth time.
Cheaper to run, more expensive to set up
Being straight about this matters, because the pitch for running things locally is usually only half told.
Getting it working took a day of unglamorous problems. Four separate ones, all environmental, none about the voice. The best was that the graphics toolkit could not cope with a folder path containing a space, which is most folders on a normal Windows machine, so the whole environment had to move elsewhere on the drive. Another was a bundled library file that crashed against the installed driver.
None of that is hard. All of it is time, and none of it appears in the tutorial. So the trade is real: you pay once in setup hours to stop paying forever in per use fees and, more importantly, in hesitation.
What this means for your business
Go looking for the metered steps inside your own daily work. Not the big subscriptions, the small ones. The per message fee, the per minute transcription, the per lookup enrichment. Anything cheap enough to ignore and frequent enough to skip.
Then ask two questions. Can hardware you already own do this well enough, and what is the setup cost honestly. Sometimes the answer is no and you keep paying, which is a fine answer once you have actually asked.
And if you automate anything generative, build the reject test before you build the volume. A generative step with nobody watching needs an automatic way to know it produced garbage, otherwise scale is just a faster way to publish mistakes. That principle is the same whether the machine is writing a voiceover, a first draft, or a reply to a customer, and it is a large part of what marketing automation San Jose businesses should be asking for when somebody offers to automate their content.
Want this built for you
We build CRM systems, websites and automation for small businesses, and we would rather own a step than rent it. Start at optechsol.llc.