EPISODE 003 · SEPTEMBER 22, 2026
What Your AI Does Not Tell You
Jivesh Sharma, M.D. · With Talia, an AI cohost · 18:40
A clinical summary can sound accurate and still leave out something important. Jivesh Sharma, M.D., and Talia examine omissions, source records, meaningful human review, and the decisions around healthcare AI.
Listen without the screen
The narration stands on its own. The visual edition adds explanatory diagrams.
Narrated using an AI-generated version of Jivesh Sharma, M.D.’s own voice, with his authorization. Talia is an AI cohost. HealthIT.com is the host’s own resource, part of the Nexgen Precision portfolio.
Reading the evidence
The patient-message scenario is fictional and contains no real patient information. This is an educational discussion of workflow design. The host’s proposals are editorial recommendations, not a validated clinical protocol or individual medical advice.
Sources checked September 22, 2026. The studies evaluate particular models, inputs and review methods; they do not certify every current product. This episode does not update the separate product-evidence snapshot.
Sources and context
- Van Veen et al. — Adapted large language models can outperform medical experts in clinical text summarization ↗
Nature Medicine, 2024
Favorable physician ratings on specified summarization tasks do not establish universal product safety or improved patient outcomes.
- A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation ↗
npj Digital Medicine, 2025
Simulated primary-care consultations; authors affiliated with Tortus AI. Error consequences were reviewer assessments, not observed patient outcomes. Omissions and hallucinations have different denominators and must not be treated as directly comparable rates.
- NIST AI 600-1 — Generative Artificial Intelligence Profile ↗
July 2024
Risk-management guidance, not clinical approval, legal certification or validation of a product.
Episode chapters
- 00:00:00 · What your AI doesn’t tell you
- 00:00:25 · HealthIT.com
- 00:00:35 · Who decides what gets attention?
- 00:00:50 · Present in the chart. Absent from the decision.
- 00:01:41 · What should a summary preserve?
- 00:01:52 · Both statements can be true.
- 00:02:51 · What does the evidence support?
- 00:02:57 · Benefits and errors need separate evaluation.
- 00:04:14 · Did the model receive the message?
- 00:04:23 · Follow information to the decision.
- 00:05:25 · What should the clinician see?
- 00:05:29 · Make the boundaries inspectable.
- 00:06:40 · Did we save work or create review work?
- 00:06:47 · Measure the whole workflow.
- 00:07:53 · Draft, suggest, or act?
- 00:08:01 · Define authority with concrete verbs.
- 00:09:07 · What does human review actually catch?
- 00:09:14 · Give the reviewer room to disagree.
- 00:10:24 · Test the intended use.
- 00:10:29 · Include difficult cases deliberately.
- 00:11:51 · One score can conceal the tradeoff.
- 00:11:57 · Measure benefit and failure separately.
- 00:13:06 · Confidence in what?
- 00:13:14 · Check the answer and the available information.
- 00:14:15 · One evaluation is not a permanent answer.
- 00:14:22 · Keep changes and failures traceable.
- 00:15:34 · Where does the patient fit?
- 00:15:40 · Leave room for the patient to correct the picture.
- 00:16:44 · Ask to see a difficult case.
- 00:16:47 · Questions for the team building the tool.
- 00:17:44 · Save attention. Preserve the ability to question.
- 00:17:51 · Give attention back to the patient.
- 00:18:30 · Thank you for listening.
Full transcript
What your AI doesn’t tell you
00:00:00 · Jivesh Sharma, M.D.
Imagine opening a patient's chart and finding a beautifully written summary. It is concise. It is accurate about the scan. It tells you what you expected to see. Then you discover a message that never made it into the summary. The patient says something has changed. The AI did not invent a fact. It left one out. Today, I want to explore what happens when software starts deciding where our attention goes, and how we can make that help us care for people.
HealthIT.com
00:00:25 · Show introduction / closing
HealthIT.com, with Doctor Jivesh Sharma. We explore the evidence behind healthcare AI and what it means for the people giving and receiving care.
Who decides what gets attention?
00:00:35 · Talia · AI cohost
I'm Talia, your AI cohost. Today's patient example is fictional. Jivesh, you recently wrote that the scarce resource in medicine may be attention. We have spent years trying to give clinicians more information. Why does that stop being enough?
Present in the chart. Absent from the decision.
00:00:50 · Jivesh Sharma, M.D.
Because information has to reach a person at a moment when that person can use it. A result can be present somewhere in the record and still be absent from the decision. Think about a physician preparing for a visit. There may be a previous note, a report from an outside facility, a medication list, and several messages. Each item has a date. Each comes from a particular source. Some of them may disagree. The task is to understand what matters now, and what still needs to be resolved. A good summary can help enormously. It can bring relevant information together and make the first few minutes of a visit more productive. I want that benefit. But shortening a record requires selection. Once we use the summary as our starting point, that selection influences the questions we ask. It may also influence the questions we never think to ask. That is the part I want us to make visible.
What should a summary preserve?
00:01:41 · Talia · AI cohost
But leaving things out is the purpose of a summary. If we keep every detail, we have recreated the original chart. How do you distinguish useful compression from an omission that should concern us?
Both statements can be true.
00:01:52 · Jivesh Sharma, M.D.
Start by asking what decision the summary is meant to support. A short scheduling handoff and a preparation note for a complex oncology visit have different jobs. The information required for one may be insufficient for the other. We should not evaluate both against an abstract idea of a good summary. For our fictional example, imagine that a scan report is reassuring. A later patient message describes a worsening symptom. Both statements can be true. The symptom might be related to the cancer, to treatment, or to something else. The summary should not settle that question for us. What matters is that the newer information could change the conversation and what the clinician decides to assess. If that message disappears, the remaining sentences can all be accurate. A reviewer who only checks whether the scan was described correctly may call the summary a success. A reviewer who asks whether it preserved the information needed for this visit may reach a different conclusion. That difference is why correctness alone is an incomplete test.
What does the evidence support?
00:02:51 · Talia · AI cohost
Have researchers measured these omissions? I would also like to hear where AI summaries have performed well.
Benefits and errors need separate evaluation.
00:02:57 · Jivesh Sharma, M.D.
They do. A twenty twenty-four study by Dave Van Veen and colleagues in Nature Medicine evaluated adapted language models on several clinical summarization tasks. Physician readers often rated the best-adapted models' summaries as equivalent to, or better than, the comparison summaries written by medical experts. The study also examined safety problems in both human and model outputs. Those results show that these models can produce useful summaries. They do not establish that every summary system improves patient outcomes. A twenty twenty-five study in Nature Partner Journals, Digital Medicine looked at notes generated from simulated primary-care consultations. The research team was affiliated with the AI documentation company Tortus. They assessed both missing information and unsupported additions, including their possible clinical impact. Both types included clinically important errors. A larger share of the hallucinations was rated major. How often an error occurs and what it could do to a patient are separate questions. These studies used particular models, inputs, and review methods. They are not a scorecard for every product available today. When we evaluate a tool, we need to examine what it gets right, what it leaves out, and how those errors could affect the work we intend it to do.
Did the model receive the message?
00:04:14 · Talia · AI cohost
Suppose our patient message is missing. It is tempting to say the model failed. But could the problem have happened before the model ever saw the chart?
Follow information to the decision.
00:04:23 · Jivesh Sharma, M.D.
Absolutely. I would trace the message through the process before assigning a cause. First establish whether it was available in the source record at the time, and whether it was included in the material retrieved for the summary. Check whether a date filter excluded it or it was attached to a different encounter. If the model never received the message, a better prompt may not fix the problem. Then look at the generation itself. Perhaps the input contained the message, but the summary left it out. Perhaps the system was asked to focus only on imaging. Perhaps the information was included in a longer draft and then lost during a second shortening step. Each possibility calls for a different change. Finally, inspect what the clinician actually sees. An important item may be present in the generated output but hidden behind a collapsed section or a confusing label. A workflow can fail even when the underlying text contains the right information. That is why I want an evaluation that follows information all the way to the decision. Testing the model in isolation cannot answer every question about the finished product.
What should the clinician see?
00:05:25 · Talia · AI cohost
What should a clinician be able to see or check when opening one of these summaries?
Make the boundaries inspectable.
00:05:29 · Jivesh Sharma, M.D.
I would want the scope stated plainly. What records were included, and up to what time? If the system did not receive outside records or patient messages, that limitation should be visible where I use the summary. A generic disclaimer elsewhere is much less useful. I would also want direct access to the sources behind consequential statements. A link to the actual report is more useful than an explanation that merely sounds reasonable. The date matters, too. A statement from six months ago should not silently appear to describe the patient's current status. Where the sources disagree, I would prefer to see the disagreement. For example, if a medication appears active in one place and discontinued in another, the system should not quietly choose a version without making that choice inspectable. Resolving the discrepancy may require a conversation with the patient or another member of the care team. None of these features can prove that nothing important was omitted. A source link helps me inspect what was included. It does not automatically reveal everything that was excluded. We still need tests that compare the summary with the original material, and a process for learning when users discover a miss.
Did we save work or create review work?
00:06:40 · Talia · AI cohost
If a physician has to reread the whole chart to check the summary, where is the time saving? We could be adding another task.
Measure the whole workflow.
00:06:47 · Jivesh Sharma, M.D.
That is a real objection. I would not accept a time-saving claim that counts only how quickly the first draft appears. We should count the review, the corrections, the extra clicks, and the work created when an error is discovered later. We should also ask who is doing that work. A task can become easier for one person because it became harder for somebody else. The answer will depend on the use. In a lower-consequence administrative task, a team may be able to define narrow boundaries and a limited review process. In a consequential clinical task, the review requirements may be much more demanding. We should decide that deliberately, with the people responsible for the work. I would start with a narrow purpose and test whether the combined human and software process is actually better than the current process. Does it preserve the information people need? Does it reduce total effort? Does it introduce a failure that is difficult to detect? If we cannot answer those questions, we have demonstrated a fast draft. We have not yet demonstrated a better workflow.
Draft, suggest, or act?
00:07:53 · Talia · AI cohost
The same summary might be used to prepare a visit, suggest an action, or trigger an action automatically. Does that change how we should evaluate it?
Define authority with concrete verbs.
00:08:01 · Jivesh Sharma, M.D.
It changes the stakes. A draft that a clinician can inspect has one role. A system that places an order, changes a queue, or closes a task has another. We should specify exactly what the system is allowed to do, and what must happen before an action takes effect. Consider our fictional message again. If the summary is incomplete, a clinician may still discover the message elsewhere. If that same summary is also used to decide that no follow-up is needed, the omission has become part of an action. The opportunity to catch it may have narrowed. My preference is to describe authority in concrete verbs. The system can collect these records. It can draft this note. It can propose this action. It cannot complete that action without the specified review. The exact boundary will vary, but it should be understandable to the people using and supervising the tool. This is also where a stop needs to be real. A flag that nobody sees is not a useful escalation. A stop that leaves an urgent item stranded is not a complete safety process. Someone needs to own the next step, including what happens when the usual reviewer is unavailable.
What does human review actually catch?
00:09:07 · Talia · AI cohost
We often hear that a human remains in the loop. What would make that meaningful to you, beyond having an approval button at the end?
Give the reviewer room to disagree.
00:09:14 · Jivesh Sharma, M.D.
I would look at what the reviewer can actually do. They need access to the original information and a clear understanding of the action they are approving. They should be able to change or reject the output without losing the work. The review also needs enough time and an understood route for raising a problem. Those are questions about the design of the work. A person's presence alone does not answer them. If the system has already filtered the information, and the reviewer sees only the filtered version, the reviewer may have little basis for spotting an omission. We should not confuse the existence of an approval step with evidence that the step catches important errors. The National Institute of Standards and Technology's generative AI risk profile discusses risks involving confabulation and human interaction with AI. I treat it as risk-management guidance, not a certification that a product is safe. For a particular clinical workflow, we still need direct evidence of how the people and the system perform together. The practical goal is a reviewer who can meaningfully disagree. That includes making a disagreement useful to the next evaluation, rather than letting it disappear after a single corrected note.
Test the intended use.
00:10:24 · Talia · AI cohost
A clinic wants to try one of these tools. How should it test the summaries before relying on them?
Include difficult cases deliberately.
00:10:29 · Jivesh Sharma, M.D.
First, define the task. Which people will use the summary, which source records will it receive, and which decision is it meant to support? Write down what it is not intended to do. That makes it possible to recognize when the use has expanded beyond the original evaluation. Then build a set of cases that reflects the intended work. Include straightforward examples and difficult ones. For early demonstrations, synthetic cases can help test a specific failure without exposing patient information. A meaningful assessment for clinical use may also require appropriately governed data that represents the actual setting. Synthetic examples alone cannot establish clinical performance. For each case, ask qualified reviewers to identify the information that matters for the defined task. Compare the generated summary with that reference and with the current workflow. Review incorrect additions, missing information, contradictions, and the possible consequences of each. If reviewers disagree, record why. A disagreement about relevance can reveal that the task itself was poorly specified. I would also test simple changes that should not alter the clinical meaning. Move the important message to a different part of the input. Include a later correction. Make one source unavailable. These are proposed stress tests, not a universal validation standard. Their purpose is to reveal where a polished demonstration may be relying on unusually favorable conditions.
One score can conceal the tradeoff.
00:11:51 · Talia · AI cohost
Would you reduce that to a single accuracy score? Or do we need a different way to describe whether the tool is helping?
Measure benefit and failure separately.
00:11:57 · Jivesh Sharma, M.D.
A single score is convenient, but it can hide the tradeoff we care about. I would want to know how often important information is omitted, how often unsupported information is added, and how the errors differ in possible consequence. I would keep those results separate from the time it takes to review and correct the output. I would also look at variation. Does the tool behave differently with a longer record, a missing document, a less common presentation, or information arriving in a different format? An average can look reassuring while a particular kind of case remains difficult. The evaluation should report the limits of what was tested, including groups or situations that were not adequately represented. Then there is the human part. Did the reviewer notice the consequential error? How long did that take? Did the summary help the person ask a better question, or did it distract them from one? Those questions are harder to measure than generation speed, but they are closer to the benefit we actually want. The results should support a decision about a defined use. They should not become a broad claim that the system is safe for anything involving a patient record.
Confidence in what?
00:13:06 · Talia · AI cohost
Many products display a confidence score or ask the model to say when it is uncertain. Would that address the missing-information problem?
Check the answer and the available information.
00:13:14 · Jivesh Sharma, M.D.
It can help only if we know what that signal means and whether it works for the intended use. A model can produce a confident answer from incomplete input. Asking it how confident it feels does not tell us whether an important document failed to arrive. I would separate uncertainty about the answer from uncertainty about the information available. One question is whether the system can interpret what it received. Another is whether it received what it needed. Our fictional patient message illustrates the second problem particularly well. For a designed workflow, some useful checks may be explicit. Did the expected source arrive? Is the information current enough for this task? Are there unresolved conflicts? A failed check can trigger review without asking the language model to diagnose its own blind spot. These checks will not catch every problem. But they make some assumptions visible and testable. I would rather see a clear statement that a source was unavailable than a smooth answer that silently depends on the missing source containing nothing important.
One evaluation is not a permanent answer.
00:14:15 · Talia · AI cohost
What happens when the vendor updates the model, or the clinic changes how it records information? When should the team test it again?
Keep changes and failures traceable.
00:14:22 · Jivesh Sharma, M.D.
Keep a record of what was evaluated. That includes the model, the prompt, the source selection, and the interface the user actually saw. A change in any of those can matter. A new model version is obvious. A change to a date filter or the location of patient messages may be less obvious, but it can affect what reaches the summary. When a consequential miss is discovered, preserve enough information to understand it under the organization's privacy and access controls. Ask where the failure occurred and whether the evaluation set would have detected it. Add an appropriate test, fix the relevant part, and check that the change has not introduced a different problem. I would also define a fallback before it is needed. If the tool becomes unavailable, or if a serious concern emerges, how does the team continue the work? Who decides whether to pause the affected use? Who checks that unresolved items have reached a person? This does not mean every small change requires the same process. It means the response should match the consequence, and the team should retain a usable way to detect, investigate, and respond to changes.
Where does the patient fit?
00:15:34 · Talia · AI cohost
We have talked about the physician, the software, and the people evaluating it. Where does the patient fit into this picture?
Leave room for the patient to correct the picture.
00:15:40 · Jivesh Sharma, M.D.
The patient is not simply a collection of inputs. The visit is also an opportunity to learn something that is not in the record, or to discover that the record gives the wrong impression. A summary should help us prepare for that conversation. It should not make us feel that the conversation has already happened. In our example, the later message matters because it describes a change. But patients can also clarify their priorities, explain why a medication list is wrong, or tell us that an earlier plan was never carried out. The system should leave room for those corrections to change the picture. I would like an AI-assisted workflow to make that easier. The clinician might arrive better prepared, with unresolved questions visible and relevant sources available. The patient might spend less time repeating information and more time explaining what matters now. Those are benefits worth pursuing. We should evaluate whether they actually occur, rather than assume that a shorter note produces them. The aim is to support the relationship and the work of care. If the summary becomes a reason to listen less carefully, we have used a potentially helpful tool in a way that defeats that purpose.
Ask to see a difficult case.
00:16:44 · Talia · AI cohost
What should a listener ask at their next product demonstration?
Questions for the team building the tool.
00:16:47 · Jivesh Sharma, M.D.
Ask the vendor or the internal team to show a case where the first summary was incomplete. Have them explain what was missing, how it was discovered, and what changed afterward. A useful answer includes the limits of the evaluation and what the team still does not know. Then ask to see the path from source information to the action the user takes. That should make clear which records are included, how the reviewer inspects the source, and what happens when information conflicts or fails to arrive. If the system can do more than draft, show the boundary between a suggestion and an action. Finally, I would ask how the team knows the combined workflow is better. Include review effort, corrections, and the experience of the people doing the work. A faster draft is valuable when it contributes to a better result. It should not be the only result we measure. Those questions do not replace a formal evaluation. They help move the discussion toward the actual work. They also make it easier to distinguish a team that understands its tool's limitations from a demonstration that has been arranged to avoid them.
Save attention. Preserve the ability to question.
00:17:44 · Talia · AI cohost
So the summary should help the clinician prepare, while keeping the original record easy to check. The patient may still have something new to say.
Give attention back to the patient.
00:17:51 · Jivesh Sharma, M.D.
Exactly. I want AI to help us find what matters and spend more of our attention on the patient. To do that well, we have to evaluate the information that disappears, as well as the sentences that appear on the screen. The example we used today was fictional, and the workflow checks are a way to frame evaluation. They are not a guarantee of safety. The studies and supporting sources are listed in the episode notes. If your team is working with AI summaries, I would be interested in one specific question. What important information is hardest to keep visible in your workflow? Please discuss the pattern without sharing patient details.
Thank you for listening.
00:18:30 · Show introduction / closing
You've been listening to HealthIT.com. This episode uses an AI-generated version of my voice. Explore the sources, and follow for our next conversation.