Stop Asking What AI Can Do. Ask What It Should Never Touch, and What Doesn't Even Need It.

The real skill isn't finding more tasks for AI. It's triaging correctly across three very different buckets.

Most AI roadmaps chase capability creep: whatever the technology can technically do next. That instinct is the wrong one. The harder, more useful question is where an organization’s scarce human judgment needs protecting, where AI earns its keep, and where a health plan already owns a cheaper, more reliable answer. Getting “top of license” right means deliberately sorting every use case into one of three buckets before a single line of code gets written, not maximizing how much of the operation AI touches. Skip that step, and the same mistake follows: AI budget gets wasted, and trust gets eroded, because every use case gets treated the same way.

Bucket One: Where the Decisions Stay Human

Bucket one holds the decisions where the human has to make the call, not just review one: clinical necessity determinations, sensitive member communications, network contracting terms, and appeals that turn on ambiguous plan language. None of these are AI-appropriate, and the reason isn’t caution for its own sake. The judgment is complex, sensitive, and locked to context that shifts case by case. That means weighing a member’s full circumstances and reasoning through ambiguity the way an experienced clinician or negotiator would, then being accountable for a call that has no single right answer. A model doesn’t do any of that. A medical director signing off on a denial isn’t providing a signature. They’re applying clinical judgment nothing else in the organization can substitute for. Get it wrong, and the harm lands on the member whose denial letter mishandled their actual circumstances, not on a scorecard somewhere. Regulation and reputational risk are why plans keep these decisions human, but the deeper truth is that the reasoning and the accountability behind it can’t genuinely be replicated by AI. Automate one of these calls, and the organization risks more than an error: it erases the reason the role, and often the license behind it, existed in the first place.

Beyond internal judgment, regulation and accreditation standards back this up too. State insurance regulations and CMS rules for Medicare Advantage commonly require that adverse clinical determinations be made or reviewed by a licensed physician or appropriate clinical peer, and NCQA and URAC accreditation standards hold plans to similar requirements for utilization management decisions. The regulator or the accreditor has already drawn the line in these cases, and a plan’s job is to build the AI system to respect it, not relitigate.

Consider an AI system that auto-created a provider record after an NPI match. The matching itself was pattern-based and high-volume, exactly the kind of work AI should be doing. What crossed into bucket one was different: the model decided, unsupervised, that a new provider location was real and under contract. That was a credentialing judgment call, and only a person could own it. The test for this first bucket has nothing to do with whether AI is capable of the task. It's whether the organization is willing to let go of the accountability a person currently holds.

Bucket Two: Where AI Earns Its Keep, and Which AI Earns It

Bucket two is different: high-volume, pattern-based work with enough variability that fixed rules break down. Extracting structured data from an unstructured clinical note, summarizing a record for a CM nurse, prioritizing an appeals queue by likely urgency, drafting first-pass correspondence for someone to review, flagging an anomalous claims pattern for a fraud, waste, and abuse analyst to chase down: all of it belongs here. AI is worth the investment in this kind of work because it handles nuance at scale that a static rule can’t, provided a person still owns the decision the AI is informing. Most healthcare leaders who’ve deployed this kind of AI report the investment paying off, even though the industry hasn’t settled on one consistent way to measure how much.

Landing in bucket two only answers half the question. Most health plans make a second, quieter mistake right after they clear the first: they treat AI as one thing, when it's actually a spectrum. A traditional predictive model, a propensity score or a classifier trained on structured claims fields, is explainable, cheap to run, and easy to monitor. That's the right amount of AI for a well-defined, structured-data task like flagging readmission risk. A generative model or an LLM is built for unstructured text instead: summarizing a clinical note, drafting correspondence, pulling meaning out of a free-text appeal. It's more flexible and handles nuance a predictive model can't, but it carries more variance and needs stronger review and tighter guardrails around it. An agentic system sits above both. It takes multi-step action across systems on its own, the way a provider-record pend bot that auto-resolves NPI mismatches does, which makes it the most capable tier and the highest-risk one. It belongs only on tasks where the downstream consequences of an autonomous action are fully mapped out, deliberately bound in advance, and checked by a person at points that carry real judgement. That’s exactly where the credentialing example earlier went wrong.

Matching the right tier to a task is its own discipline, separate from deciding the task belongs in bucket two at all. Reach for an agentic, autonomous build when a simple predictive model would do, and the plan has over-engineered: more cost, more complexity, and unsupervised-action risk bolted onto a task that never needed independent judgment in the first place. Reach for a rigid, narrow model on a task with real linguistic nuance, and the plan under-delivers instead, quietly pushing the error-correction work back onto staff. Getting this right takes the same payer-operational fluency bucket one demands. It’s not enough to know a task is AI-appropriate. A plan has to know how much autonomy the task can tolerate, what type of model fits it, and what hedges its downstream consequences.

Bucket Three: Where You Already Have the Answer

Some tasks are deterministic, rules-based work, the kind a business rules engine already handles, or a module sitting inside the plan’s core admin or claims platform already does: simple eligibility routing, a benefit accumulator update, standard EOB generation, a straightforward pend code. No AI required, and no AI wanted. A striking number of AI initiatives are solving a problem a configuration change would have solved for free. That wastes budget, and it adds model risk (hallucination, drift, the overhead of monitoring a system) to a decision that never needed any of it.

This bucket goes beyond the simple, deterministic tasks. Provider data management, utilization management queuing, claims editing: plenty of complex workflow problems like these are already solved, by mature platforms vendors have spent years purpose-building for health plan operations. AI’s job here is extending what those platforms already do, and only for the piece of the problem that sits beyond their built-in capability.

In practice, plans make their most expensive mistake here, and it shows up in two forms at once. The first: buying or building an AI point solution for something a business rules engine, or a module already sitting in the plan’s core admin or claims platform, already does. The second: reaching for the most sophisticated AI tier available for a task a simpler bucket-two model would have handled just as well, at lower cost and lower risk. Avoiding both takes payer-operations expertise paired with deep vendor and technical expertise, the kind that knows what capability already lives in a client’s platform and what tier of AI a given task calls for. Get both right, and no use case gets more technology than it needs.

Where This Leaves You

Top of license means protecting human judgment where judgment is the point, deploying the right tier of AI where real variability demands it, and refusing to spend AI money, or take on AI risk, on a problem the plan’s existing systems already solve, not pushing AI into every workflow it’s capable of touching.

i2 Health’s use case evaluation starts with this triage. The technology recommendation comes second, not first. Running that triage well depends on more than good instincts. It depends on whether your organization actually has the data quality, governance, and change-management muscle to sort these calls correctly at scale, not just in theory. i2 Health's AI Readiness Assessment measures exactly that: five questions, about two minutes, no email required. Get your score →


Frequently Asked Questions

Q: We’re already being pitched agentic AI for tasks a simpler model, or a rules engine, could handle. How do we tell the difference?

Ask what tier of AI the task needs, not what tier the vendor sells. A structured task like flagging readmission risk needs a predictive model: explainable, cheap, easy to monitor. Unstructured text work, like summarizing a note or drafting correspondence, needs a generative model with review built around it. An agentic system belongs only where the consequences of an autonomous action are fully mapped and bounded, and most of what gets sold as agentic doesn't clear that bar. If a pitch skips to the most capable tier without naming what happens when it acts on bad information, ask before the contract, not after.

Q: What does "human in the loop" actually mean for health plan AI, and when is it not enough?

It's not enough when the human is there to catch mistakes after the system acts. In the NPI matching example, the data matching was a legitimate automation task, but deciding that a new provider location was real and contracted is a credentialing judgment call. That stayed a person's job until the system quietly took it over, and no one noticed because nothing forced them to. A person in the loop means the judgment call requires a human action to complete. A person available to review afterward is a safety net, and safety nets depend on someone noticing in time.