DistilINFO
Recent Posts
HomeProviderHealth System CIOs Tame AI Token Costs

Health System CIOs Tame AI Token Costs

Health System

Healthcare’s rush into generative AI has created a cost center many organizations only notice after the bill arrives: the token, the basic unit AI platforms use to bill for every prompt sent and every response generated, driving mounting health system AI token costs as adoption accelerates.

Why Health System AI Token Costs Climb So Quickly

As adoption scales from pilots to enterprisewide rollouts, health system leaders say spending can climb quickly, largely because the employees using these tools have no idea a meter is running. “If you don’t manage this well, you will have runaway costs, because people don’t understand the cost,” said Luis Taveras, PhD, executive vice president and chief digital and information officer at Philadelphia-based Jefferson Health.

Why Access Discipline Matters More Than Contract Price

That gap in understanding is why Dr. Taveras and CIOs at several other health systems said the biggest lever they’re pulling isn’t a cheaper contract with an AI vendor, it’s discipline over who gets access to which tool, and for what.

Jefferson Health’s Tiered Approach to Health System AI Token Costs

At Jefferson, that discipline starts with a ladder. Employees with routine questions are told to stick with Google. Those who need more are pointed toward Jefferson’s basic Microsoft Copilot license, bundled into its existing Microsoft agreement at no extra cost, and then to Copilot Plus, a paid tier billed as a fixed monthly fee. Only employees who need more firepower than that get access to Anthropic’s Claude, or other frontier models billed by the token.

A Concrete Example of the Tradeoffs Involved

Dr. Taveras pointed to a recent example: a colleague needed to analyze a large batch of files and asked for access to Cowork, Anthropic’s agentic desktop app. Jefferson granted it that morning, and by the end of the day, the employee had finished work that otherwise would have taken about two weeks, at a token cost of $75. “If I think about this, a highly compensated person, two weeks’ worth of work in four hours, that’s not a bad return,” Dr. Taveras said. “But if he does that every day, the number is going to get pretty high.”

How Jefferson Manages Usage and Model Verbosity

Jefferson has started allocating tokens by user type rather than giving everyone the same amount, tracking heavier users, such as its investment analytics team, differently from occasional users, and sending people reports on their own usage. Certain tools are unlocked only after employees go through an education process on what the tools cost.

Why Model “Verbosity” Adds to the Bill

“The other thing we have to be careful with is what I call controlling the verbosity of these models. They’re very verbose, you ask a simple question, it comes down and gives you an answer, and then it says, ‘Can I do this? Can I do that for you?'” Dr. Taveras said. “If you’re using one where the meter is running, it’s self-serving for the model to keep asking you for more, because you pay a lot more for what comes down than what you pay for what you send up.”

How UVA Health and Emory Approach Health System AI Token Costs

Charlottesville, Va.-based UVA Health is also being deliberate about which employees get access to which AI capabilities, rather than opening every tool to everyone at once. “Our focus at UVA Health is to balance innovation with stewardship,” said CIO Sonney Sapra. “We are evaluating the appropriate levels of access, matching tools to specific business and clinical needs, measuring value and outcomes, and understanding utilization patterns before expanding broadly.”

Emory’s Model-Matching and Vendor Transparency Requirements

At Atlanta-based Emory Healthcare, Chief AI Officer Nabile Safdar, MD, said his team applies a similar filter before a tool is turned loose enterprisewide. “One of the primary levers we use to manage those costs is selecting the lowest-intensity model that is appropriate for the task at hand,” he said. Emory also requires AI vendors to open up their financial operations data before signing on, so the team can better understand and forecast potential token consumption and associated costs.

Shifting From Cost Per Token to Cost Per Outcome

Eric Kirkendall, MD, chief medical information officer and CIO for academic health at Atrium Health Wake Forest Baptist, said the model-matching question has become a standard part of governance conversations, not “what’s the best model,” but what’s the least expensive one that reliably gets the job done. “Many organizations are still focused on reducing the cost of individual prompts. While that matters, we’ve found that the larger opportunity is ensuring AI is applied to the right problems, with the right model, and with clear measures of value before usage scales,” he said.

Why a High Token Bill Can Still Be Justified

“A high token bill may be entirely justified if the workflow reduces clinician burden, improves patient engagement, accelerates research or eliminates manual administrative work,” Dr. Kirkendall said. “We’re starting our journey to focus less on cost per token and more on cost per outcome delivered.”

How Stanford and Houston Methodist Evaluate AI Costs Upfront

At Stanford Health Care, Chief Information and Digital Officer Michael Pfeffer, MD, said cost enters the conversation before a project is even greenlit. “We start with a deep understanding of the problem we’re trying to solve to determine the best possible IT solution, which may or may not be AI,” he said. “If it does require AI, then we use our FURM assessment framework to understand the costs relative to the value of the solution before proceeding.”

Houston Methodist’s Governance Committee Model

At Houston Methodist, Chief Innovation Officer Roberta Schwartz, PhD, said a standing internal committee vets proposed AI use cases for return on investment before the system decides whether to build, buy or partner on a given tool, then keeps watching after it goes live. “Like any technology we implement, agentic AI is subject to ongoing governance, monitoring and optimization to ensure it continues to perform as intended and deliver value safely and effectively,” she said.

Brown University Health’s Pricing Strategy for Health System AI Token Costs

Adam Landman, MD, chief digital information officer at Providence, R.I.-based Brown University Health, takes a different tack on exposure: pricing structure itself. Most of the AI capabilities Brown has deployed at scale run on flat per-user pricing rather than token-based consumption, which shields the system from utilization surprises. That won’t hold for every tool, particularly the most powerful task-executing agents, where the underlying compute is genuinely expensive.

Why FinOps Visibility Matters

“For consumption-priced capabilities, we start with a small pilot group and usage limits rather than broad access. That caps our exposure while we learn, and it generates real utilization data so we can forecast annual cost with confidence before scaling,” Dr. Landman said. “Underneath all of it, the enabling capability is FinOps. Real-time visibility into consumption and spend is essential, if costs move unexpectedly, we want to know within hours to days, so we can act while the number is still small.”

Why Some Health Systems Are Eyeing AI Sovereignty

Dr. Landman also pointed to a longer-term shift some organizations are beginning to weigh: moving away from per-token vendor pricing altogether. “Token-based pricing is challenging for any organization because cost scales with adoption, and adoption is the goal,” he said. “That’s one reason many organizations are exploring ‘AI sovereignty,’ running open-weight models on dedicated infrastructure, where cost is driven primarily by the hardware investment rather than each individual transaction.”

Why Governance, Not Price, Will Determine the Winners

For now, most of the leaders interviewed agreed that the winners of this phase of AI adoption won’t be the systems that land the cheapest per-token rate. “The organizations that will manage AI costs most effectively won’t necessarily be those with the cheapest tokens,” Dr. Kirkendall said. “They’ll be the ones with the strongest governance, the best reuse of established capabilities and the clearest understanding of where AI creates measurable value.”

What This Health System AI Token Costs Debate Means Going Forward

With seven health systems all converging on the same core lesson, that disciplined access management matters more than vendor pricing alone, this collective experience offers a practical governance template for other health systems navigating the same rapid scaling from pilots to enterprisewide AI deployment. Given the range of strategies represented, from Jefferson’s tiered access ladder to Brown’s flat-rate pricing preference to Stanford’s FURM-based upfront evaluation, health systems building their own cost-management frameworks may find value in combining elements from multiple approaches rather than adopting a single model wholesale.

What to Watch Going Forward

As more health systems move from pilot programs into full enterprise AI deployment, industry observers will likely watch whether the shift from “cost per token” to “cost per outcome” that Dr. Kirkendall described becomes a more widely adopted evaluation standard industrywide. Given Dr. Landman’s observations about “AI sovereignty” and open-weight models running on dedicated infrastructure, health system AI token costs management may increasingly extend beyond access governance into fundamental questions about infrastructure ownership as adoption continues to scale across the industry.

For more healthcare industry updates, insights and news, visit DistilINFOClick here to subscribe to stay informed.

No comments

Sorry, the comment form is closed at this time.