When AI Watches the Clock, Who Watches the Patient?
Kaiser nurses say AI monitoring of call times and tone of voice is changing how care gets delivered. Here is what that actually signals for any operator using AI to measure performance.
The Signal #049 — Dakota’s read on the AI news that actually matters to people running a business.
There is a version of AI adoption that looks great on a dashboard and quietly breaks the thing you were trying to protect.
That is what a group of Kaiser nurses are describing right now. And whether you run a healthcare org or not, the dynamic they are pointing to is one every operator using AI to measure people should understand.
What happened
CalMatters reported that Kaiser Permanente nurses who handle advice and triage calls are raising serious concerns about how AI is being used to monitor their work. Seven current and former nurses told the outlet that spending more than 15 minutes on a call routinely draws criticism from management or triggers a performance evaluation meeting. Call time, they said, factors into monthly performance scores.
Beyond call length, nurses described software that tries to predict on a daily basis whether they are being unproductive or failing to answer calls quickly enough. AI systems have also been used to rate their empathy and tone of voice.
One nurse, Raquel Alvarez Sanchez, described staying on a call with a suicidal patient for more than an hour while waiting for police to arrive. She knew the whole time that the call would hurt her average handle time for weeks and could lead to management questions. Another nurse, speaking anonymously, said she cut short a conversation with an elderly woman who had just received a terminal cancer diagnosis because she feared a reprimand for going off script.
“I think at some point all of the nurses have been talked to about their average handle time,” Sanchez said. “The only thing I can think of is they’re doing it for profit.”
Kaiser, for its part, says it does not use average handle time to assess performance and deploys AI with patient safety in mind. The California Nurses Association is currently bargaining a new contract for 25,000 nurses, with AI listed as a central issue. Kaiser is the largest private employer in California, serving more than 9 million people in the state.
Why it matters for operators
This story is not really about healthcare. It is about what happens when you point AI at a proxy metric instead of the actual outcome you care about.
Every operation has a version of this tension. A SaaS support team measuring ticket close rate. A law firm tracking billable hours per associate. A fulfillment center measuring picks per hour. The metric is easy to capture. The outcome it is supposed to represent, genuine quality, is harder to see.
AI makes proxy metrics faster, cheaper, and more granular than ever before. You can now monitor tone of voice, response time, and predicted productivity at a scale that was not practical two years ago. That is a real capability. The risk is that when people know they are being scored on a proxy, they optimize for the proxy. Not because they are lazy or dishonest, but because the system is telling them that is what matters.
The nurse who shortened her call with a terminal cancer patient was not being careless. She was responding rationally to the incentive in front of her. That is the part worth sitting with.
What most people get wrong
Most operators treat AI monitoring as a neutral upgrade to accountability. More data, more visibility, better management. And in some contexts, that is true.
The mistake is assuming that measuring something more precisely is the same as measuring the right thing. It is not.
AI surveillance (automated systems that track worker behavior and output in real time) is genuinely useful for catching patterns that humans miss, flagging compliance issues, or identifying where processes break down. Used well, it surfaces problems early. Used poorly, it trains your team to perform for the sensor rather than for the customer or patient.
The other mistake is deploying measurement without a clear theory of what good actually looks like. If your best employees are the ones a monitoring system would flag most often, that is a signal the system is miscalibrated, not a signal those employees are underperforming.
California lawmakers are currently considering bills that would protect doctors and nurses from retaliation for overriding automated care recommendations. That kind of policy discussion does not happen unless a meaningful number of people feel the automation is already overriding their judgment in practice.
The lesson here
AI is very good at measuring what you point it at. It is not good at knowing whether what you pointed it at is the right thing.
Before you add automated performance monitoring to any role, it is worth asking one question first. If your people optimize perfectly for this metric, does the actual outcome get better or worse? If you cannot answer that with confidence, the measurement system is not ready.
For a longer look at how to think through AI in your operations before you build it in, xovionlabs.com is a good place to start.