At Health Evolution Connect (Sept. 28), Dr. David Kirk, CMO of Regard, and Dr. Renee Allenbaugh, Regional CMO at Penn Highlands Healthcare, led a Brass Tacks session titled "Your EHR is a database, your scribe is a typist, what's the strategy?".
Key takeaways
- Clinicians can dismiss a correct AI finding just as easily as they can accept a wrong one.
- Hospitals already decide what nurses can do without a doctor's order, and AI needs the same clear limits.
- High acceptance rates and fading complaints are early signs that clinicians have stopped reviewing.
- Long, bloated notes are hard to check, and any error a clinician signs off on gets carried forward.
- Training programs have to teach residents and APPs to question recommendations, or they won’t develop the judgment needed to catch.
A correct finding the physician doubted
Dr. Kirk opened with a story from a recent go-live. A hospitalist CMIO reviewing Regard's diagnosis recommendations for a patient he'd treated for several days dismissed one, a renal mass, as an AI hallucination. Dr. Kirk's team asked the tool to explain its reasoning.
It pointed to an admission CT ordered to check for colitis or appendicitis. The summary at the top of the report ruled out both, but the body described a large renal mass that no one had addressed, and the patient was about to be discharged.
Findings like this often go unnoticed because charts have grown faster than the time clinicians have to read them. Dr. Kirk recalled an ICU shift when three patients came up from the ED at once, each with prior admissions and outside records. With the ED pressing to free up beds, he reviewed less than a third of their charts, and he said that pressure drove much of the burnout in his group.
The CMIO's reaction shows the other side of the problem. A tool can catch what clinicians miss, but only if they trust it enough to check, so oversight has to account for both.
Scope of practice, applied to software
Dr. Allenbaugh sees around 25 complex patients a day in a rural hospital, and she values a tool that reads the chart alongside her. What worries her is output that sounds right but is missing one crucial ingredient: context, saying, "The human brain kind of looks at things that sound correct, and it says, 'Oh, that must be correct.'" The risk is highest for a clinician interrupted mid-note or a learner who doesn't yet know what to look for.
Dr. Kirk pointed out that hospitals already manage this risk with people. They build in second checks for high-stakes work and let some roles act on their own for lower-risk tasks. A pharmacist double-checks his heparin math, for example, while nurses can consult wound care or speech therapy without a physician order. "Much like we decided what nurses could do versus what doctors could do," he said, health systems will decide which lower-risk tasks AI can own. Dr. Allenbaugh added that trust in any tool is unlikely to ever be absolute, so those checks have to keep running long after go-live.
Dr. Kirk also warned about anchoring. Clinicians are taught not to treat the ED's working diagnosis as final, but a model can latch onto an earlier label in the chart and then cite it as supporting evidence. He said more organizations are now validating diagnoses with methods other than large language models (LLMs)for that reason.
Oversight gets harder as accuracy improves
Dr. Kirk questioned the assumption that better accuracy means less risk, stating that "the more accurate it is, the more likely you are to hit the yes button without thinking."
These tools also record how clinicians respond to each recommendation, which gives leaders a way to spot the problem. A clinician who accepts every recommendation unchanged, every time, isn't reading them. As Dr. Kirk put it, "AI doesn't keep bad doctors from doing bad doctor things," but the data does show leaders when it's happening.
Dr. Allenbaugh listens for silence, sharing, "...when do my people stop telling me things that are wrong? That's a huge red flag for me." A drop in complaints may mean the tool improved, or that clinicians stopped checking it. For her, finding a reliable way to audit for the difference is still a work in progress. Dr. Kirk has seen how misleading silence can be. At one hospital, clinicians spent years pushing to get EEGs for ICU patients having seizures, and then the complaints stopped. When he looked into it, EEG use was the lowest it had ever been, and the team told him, "We gave up."
Every signed note needs an owner
An orthopedic surgeon in the audience said his AI-assisted discharge summaries repeat information so often that he can't audit them. Dr. Kirk's view was that "being concise and bringing the right things at the right time is as important as bringing more."
A behavioral health clinician described a sharper risk. Psychiatrists often leave details out of notes to protect a patient's privacy, and those details now appear anyway. A patient who says "I just wanna die" in jest can have the phrase quoted throughout the note, so a low-risk patient reads as high-risk to every clinician who opens the chart next.
Dr. Allenbaugh raised the liability question. Many physicians add a disclaimer to dictated notes saying the text wasn't reviewed. With AI, an error is no longer a misheard word. It can be a wrong diagnosis that spreads through the chart, and she questioned how that will hold up in malpractice cases. "We're not gonna be able to rely on our little paragraph," she said, "which I don't think works anyway."
Training for judgment
One attendee worried about the next generation of reviewers. If junior roles are automated, new clinicians lose the repetition that builds instinct. Dr. Kirk added a related problem he called cognitive spoofing, where a first-year cardiology fellow's AI-assisted note reads as well-reasoned as an attending's, making it harder for supervisors to tell what a learner knows.
Both panelists had recent cases showing why that instinct matters. Dr. Kirk's patient had severe bradycardia and a low potassium that didn't fit the rest of the labs, so he repeated it before replacing aggressively, and the repeat came back slightly high. Dr. Allenbaugh's patient had a high potassium flagged as hemolyzed in the EHR, though only the number carried forward. In her view, the value of these tools is the time they return: "The benefit of AI is giving the doctors time to smell."
Until the research catches up, each health system is setting its own standard for careful review. For Dr. Kirk, that standard starts with the signature, because signing a note means vouching that it's true, whoever or whatever drafted it. Acceptance data and feedback channels show whether clinicians are meeting that standard, and training prepares the next generation to meet it.



