Key takeaways
- On August 31 we found that the AI scoring step in our sales applicant voice screen had never run. The request was being rejected every time, and a fallback produced a believable number anyway.
- We had described the shape of the answer with two limits, a whole number from 0 to 25 and a list of at most three items. The service accepts a narrower set of rules than the standard it is based on, and it rejected both.
- The code took its fallback path and doubled the half of the score it could measure on its own. A candidate nobody had assessed got a believable 92 out of 100.
- The fix moved the limits into plain descriptions, clamps the answer in code, and the note the system writes now says in plain words when only half of the score ran.
- The question for any tool that scores or rates something for you: how would you know if the clever part never ran?
I build a lot of marketing automation for San Jose businesses, and the thing I trust least in any of it is a number that always shows up. A number that always shows up might be working. It might also be a fallback that never fails because it never tries.
On August 31 I caught one of mine doing exactly that.
How the voice screen scores someone
Our careers page has a second step. After the written application, the applicant opens a page where one of our AI agents asks eight questions out loud and records the answers. About six minutes, one take. Three of the eight questions do the real work: reading our actual cold call opener aloud, handling a live objection, and explaining what we do to a plumber.
The score is built in two halves on purpose. Half is measured: the speech engine reports how confident it was in each word, and the code works out words per minute and the filler rate, 50 points. The other half is judged. One call to a language model reads all eight transcripts and scores grammar and substance, another 50 points.
It never rejects anybody. It moves the applicant's card exactly one stage, writes a note with the score and the transcripts, and raises one task. I play the audio and decide about the interview myself. The score is a recommendation, and the note says so.
The request that never worked
When you ask a language model for a structured answer, you describe the shape you want back. Ours asked for two whole numbers, a short note, two short lists and a summary. To keep the numbers honest I put limits on them: a whole number from 0 to 25. To keep the lists short I capped them at three items.
Those limits are standard in the format the service uses. The service itself accepts a narrower set of rules than that standard, and it rejected both of mine. The number limit was not supported. The list cap was not supported. Every scoring request failed, from the day it shipped.
What the fallback did
When the judging call fails, the code has a fallback, and the fallback was a reasonable choice when I wrote it. If the judged half is missing, the total would cap at 50, and 50 reads as a bad candidate rather than a failed call. So the code scales the measured half up to 100 instead.
Which means a candidate whose answers nobody had assessed got a believable 92 out of 100. The note landed in the CRM looking like every other note, and nothing in it said the clever half had not run.
I only found it because I tested the exact request body against the live service on its own, outside the app, and read the error. Four hundred, limits not supported. The app had been printing that same error into a log nobody reads, then carrying on.
The fix
Three changes, all small.
The limits moved out of the schema and into the plain description of each field. "Whole number from 0 to 25." "At most 3 items." The model reads those and follows them, and the service no longer rejects the request.
The code stops trusting the answer. Whatever comes back gets clamped to the range and trimmed to three items, because an out of range language score would have quietly pushed the total past 100.
And the note the system writes says plainly when only half of the score ran. If the language review did not happen, the note says so, right under the number, so the number cannot pass as a full assessment.
One more thing changed, in how I test rather than in the code. When a step has a fallback, I no longer accept the end to end result as proof the step worked. I check for something only the real path produces, here the written review that only the model can write. And if a call that should take the better part of half a minute comes back in two seconds, I treat that as a warning, not a win.
What this means for San Jose businesses running marketing automation
Lead scoring, AI call summaries, review sentiment, "likely to close" flags. More and more of the automation sold to local businesses has a clever step in the middle, and almost all of it has a fallback for when that step fails. The fallback is the dangerous part, because it is built to look fine.
So before you ask whether a tool's scores are accurate, ask the question that comes first: how would I know if the clever part never ran?
Three things worth doing this week:
- Pick one automated score or summary you act on. Find out what the tool shows when the AI step fails. If the answer is "the same thing," you cannot tell the two apart.
- Look at the spread. If every lead, call or applicant lands in the same tight band, something in the middle may not be running.
- Find one output only the real path can produce, a written reason, a quote from the call, a named concern, and check that it is actually there before you trust the number next to it.
Marketing automation in San Jose is worth paying for when the numbers mean something. If a number cannot fail, it is not telling you anything.
Want to build this yourself?
Join the On Point Tech Academy at optechsol.llc/academy. It's free, and we build live every Tuesday and Thursday.
The On Point Tech Academy costs nothing and never asks for a card. We build live every Tuesday and Thursday, 12:30 to 1:30 PT.
Join the Academy free ← Back to all Field Notes