They already had voice notes. They still typed.
I went into the visits sure I knew the product. Therapists drown in documentation. Voice is faster than typing. Put a good model on the floor, capture the session, emit a SOAP note. That story sells.
Then I sat in the gym.
The EHR they already pay for has voice SOAP. They do not use it. Detail is wrong. Presence is wrong. You cannot run a session and narrate it for a model at the same time. The good therapists jot a few words during care and finish the note later, often at home. Their idea of a perfect world is paid documentation time on the schedule, not a smarter microphone.
The job was not "listen." It was "clean up."
What they actually wanted: the therapist types a seed. Cues, load, form, what happens next visit. The model turns that into a note that would survive an audit. Completeness against the seed. Abstain if it was not documented. Do not dump the whole chart into SOAP.
That is a different product. It is also a worse demo. Nobody claps for a cleanup pass. They clap for ambient magic.
Ambient magic is what they had already rejected.
I have shipped Voice AI that answers real phone calls, so I am not allergic to the modality. I am allergic to putting it where the user has already voted with their behavior. A feature the vendor shipped and the floor ignores is not an adoption problem. It is a job-to-be-done problem you refused to see.
Two metrics, not one
The other thing the visits taught me: accuracy is the wrong single number.
A model can be precise on the sentences it writes and still fail the note, because it omitted what the therapist actually said. Completeness against the seed is the product metric. Hallucination is the safety metric. Verbosity is already a fail. "Not documented" is a feature.
If you pick the model off a general intelligence score, you will ship a fluent note that is missing the load, the side, or the reason the next visit exists. Fluent and incomplete is how you get a chart that cannot back the billed unit.
Clinic notes are not a writing task. They exist to justify care. Start from what was said on the floor. Do not start from a blank SOAP and a microphone.
What I now refuse to build first
I still want voice in the stack, later, for the parts of the job that are actually speech: a between-visit check-in, a quick "what did you do since last time," a patient who will not type.
I will not lead with ambient capture on the gym floor. The users already told me, with a feature they have and ignore.
The lesson is older than AI. Watch what people do when the vendor feature is already there. If they type anyway, your voice demo is not insight. It is costume.