Approval that never quite finishes

The National Commission into the Regulation of AI in Healthcare published 44 recommendations on 10 September, drawing on more than 12,000 contributors, according to the MHRA’s announcement. The proposal that made the headlines is “L-plates” for new AI models, and our news report covers what the Commission asked for. The idea underneath it matters more. Read in full, the 119-page report turns authorisation from a verdict a product receives once into a set of conditions it has to keep satisfying for as long as it is in use. That is a regulatory design for software that keeps changing after it ships, and every UK regulator now dealing with AI has exactly that problem.

Strategic Insight: The L-plate is the memorable part, but the lasting change is the trade behind it. The report proposes that regulators accept thinner evidence before launch in return for binding monitoring commitments afterwards. Once that trade exists in one sector, it becomes the obvious answer to put to regulators in every other.

Why does one-off approval fail for AI?

Traditional product approval assumes the thing being approved stays the same. A device is tested, cleared, and then sold in the form that was assessed. The regulator’s main tools after that point are incident reports and recalls, which only activate once something has already gone wrong.

AI systems break that assumption in three ways. They get retrained and updated. They behave differently depending on where they are deployed and on whom. And their accuracy can drift as the data they meet in practice moves away from the data they were validated on. Neil Lawrence, the Cambridge professor who led the Commission’s technology working group, put the consequence plainly: “we can’t rely on a single point of approval and assume the job is done.”

The real story is the trade, not the learner driver

Our news coverage set out the four headline asks: staged authorisation, lifelong monitoring, a searchable public safety record, and stronger enforcement. What the press release does not convey is how tightly the full report links them. Staged authorisation only works if monitoring is strong enough to catch problems during the supervised phase. Monitoring only works if the regulator can act on what it finds. Enforcement only lands if the public and clinicians can see the signals too.

The report is explicit that pre-market testing is becoming a poor predictor of how AI devices perform in real clinical settings, and it asks the MHRA to shift the weight of evidence towards the post-market phase. It goes further, suggesting that stronger ongoing oversight could reduce what manufacturers must prove before launch. That is the trade: faster entry, paid for with obligations that never expire.

MetricValueStrategic implication
People who fed into the Commission’s workMore than 12,000A mandate broad enough that government will find it hard to shelve the direction, even if details change
Recommendations in the full report44The L-plate is one of dozens; change control, foundation model disclosure and deployer readiness carry equal weight
Call-for-evidence respondents who said current post-market surveillance is not sufficient65%The existing safety net is widely seen as inadequate even by those operating within it
Share who said the current framework restricts innovationTwo-thirds of industry, half of healthcare providersBoth sides of the market want change, which makes reform politically easier than usual

Strategic Reality: The same shift is under way in UK financial services, where the FCA’s AI Live Testing programme helps firms that are ready to use AI in live markets work through evaluation, live monitoring and governance. Two regulators independently moving towards evidence gathered after deployment is a pattern, not a coincidence.

How the Commission’s model actually works

Three mechanisms in the report fit together, and they are worth separating because each one could travel to other sectors on its own.

Staged authorisation. A device aimed at a new clinical use would launch within a narrow, closely supervised scope, on the strength of the MHRA’s review of early evidence and an agreed set of risk controls and reporting duties. Its authorisation widens once it passes evidence thresholds fixed in advance. The report insists these stages should be “temporary only”, with a clear route to full authorisation.

Monitoring plans judged as part of approval. Rather than treating post-market surveillance as something that happens later, the report suggests the MHRA should evaluate a manufacturer’s monitoring plan as part of the application itself. A credible plan becomes part of the case for approval. It also floats routine performance reporting after deployment, and defined thresholds that would trigger escalation when performance deteriorates.

Boundary-based change control. This is the least discussed proposal and the most significant. Recommendation 6 asks for change management that allows “clear boundaries and guardrails to define the allowable scope of change rather than requiring specific, planned changes”.

Critical Context: The established model for approving AI updates in advance requires the manufacturer to spell out the changes it plans to make. The FDA’s guidance on predetermined change control plans for AI-enabled devices asks for a description of planned modifications, and Article 43(4) of the EU AI Act spares high-risk systems that keep learning from reassessment only where the changes were pre-determined at the initial conformity assessment. The Commission’s report says this approach struggles with adaptive technology and proposes approving an operating envelope instead of a list of edits.

That distinction is why the model is portable. Pre-specified change lists suit a device that is retrained on a schedule. They do not suit a generative AI system whose behaviour shifts with prompts, context and the underlying model it is built on. An envelope approach, where the regulator approves the limits and the monitoring that proves the system stays within them, could work for anything from a diagnostic tool to a credit decisioning model.

Success factors that are easy to miss

  • Thresholds fixed before deployment: The report asks for evidence thresholds to be set before a device’s scope widens, and for performance thresholds that would signal meaningful deterioration. Without numbers agreed in advance, “monitoring” becomes a record of what happened rather than a trigger for action. Our analysis of policing’s AI evidence pipeline shows what deployment looks like when nobody has set that number.
  • Conditions on the deployer, not only the product: The report expects staged authorisation to be supported through arrangements like partnerships with NHS providers able to show “sufficient AI readiness”. The regulated unit becomes the product in a specific setting, which is a significant change from approving products in the abstract.
  • Visibility of upstream dependencies: The report proposes an opt-in master file so developers of foundation models can share technical information confidentially with the regulator, plus an expectation that manufacturers disclose when their product depends on a general-purpose model. A monitoring regime is blind if it cannot see that the model underneath has changed.
  • Public signal, not just regulator signal: The proposed searchable safety record and earlier sharing of safety signals with providers and clinicians mean monitoring data would not stay inside the regulator.

The implementation reality

The model depends on infrastructure that mostly does not exist yet. Somebody has to collect performance data from live clinical systems, compare it against thresholds, and route warnings to the manufacturer, the deploying trust and the MHRA. The report acknowledges that the regulator may face a much larger volume of post-market data and suggests it could use AI tools of its own to process it.

Some of the capability exists commercially. Haris Shuaib, chief executive of Newton’s Tree, which tracks how AI performs in live NHS settings, said their experience is that “performance can change once systems are in routine use”. But turning that into a national regime is a different scale of task. Jennifer Dixon of the Health Foundation, the Commission’s research partner, said “the real test will be whether the NHS has the capacity, skills and systems in place to implement and monitor AI applications safely and effectively at the scale now needed.”

Hidden Cost: Lifecycle regulation turns a one-off compliance project into a permanent operating cost. Manufacturers will need monitoring pipelines, reporting staff and escalation processes for the life of every product, and deploying organisations will need people who can read the results. Budgets built around a launch milestone will not cover it.

Who carries the weight when approval never finishes?

A one-off approval concentrates the regulatory burden on the manufacturer before launch. A lifecycle model spreads it across everyone who touches the system afterwards, and that is where the organisational friction will come from.

The report is careful on this point. It argues that responsibility for a risk should sit with whoever is in the best position to manage it, and that clinicians and other users “should not be asked to carry accountability for risks they cannot plausibly mitigate”. Recommendation 24 asks DHSC and healthcare organisations across the UK to make the division of responsibility clear at every stage of an AI product’s life. Our earlier analysis of the Commission’s call for evidence found that liability and safety monitoring after approval were already the two gaps surfacing repeatedly in submissions. The final report treats both as design problems for the whole system.

Stakeholder groupPrimary impactsSupport needsSuccess metrics
AI suppliersEarlier market entry under a restricted scope, in exchange for permanent monitoring and reporting dutiesClear evidence thresholds, workable guidance on defining change boundariesTime from staged to full authorisation; no enforcement actions
Deploying organisations (NHS trusts, and their equivalents in other sectors)Become a formal condition of authorisation; must show they are ready to deploy and overseeGovernance toolkits, monitoring capacity, trained staffPassing readiness assessments; incidents caught internally before external escalation
Frontline professionalsUse systems that are formally still being proven; need to know when to question outputsTransparent labelling of staged tools, training, clarity on accountabilityConfidence surveys; appropriate override rates
Regulators outside healthcareFace pressure to explain why their approach differs from a published, evidence-based templateShared methods for thresholds and monitoring across regulatorsConsistency of post-deployment expectations across sectors

What actually signals success

For a manufacturer, success under this model is not clearing the bar. It is moving from a staged to a full authorisation on schedule, with a monitoring record that shows the product stayed inside its agreed boundaries. That record becomes a commercial asset, because it is the evidence that lets the next deployment start from a wider scope.

For a deploying organisation, success is demonstrating readiness before the regulator or supplier asks. Organisations that can already show governance, monitoring and escalation for their existing AI tools will be the partners suppliers want for staged launches, which means earlier access to new tools.

🎯 Success Factor: The organisations that benefit first will be those that already treat AI monitoring as an operational function with named owners and agreed thresholds. Staged authorisation rewards deployers that can prove they would notice a problem, and it passes over those that cannot.

How should organisations prepare, inside healthcare and out?

Nothing in the report binds anyone yet. The Commission is a non-statutory advisory body, and a cross-government response is still to come. But the direction is consistent with what the MHRA’s earlier call-for-evidence findings already showed, and with the live testing sandbox the MHRA opened with a Manchester NHS trust on the same day. Preparing now costs little and positions organisations well, whichever details survive.

💡 Implementation Framework: Building for lifecycle oversight

Phase 1: Inventory and baseline (next three months)

  • List every AI system in use, including those embedded in wider software
  • Record the underlying model and supplier for each, and how you would learn it had changed
  • Capture current performance on the measures that matter in your setting

Phase 2: Thresholds and ownership (three to six months)

  • Agree, in writing, the performance level at which each system would be paused or escalated
  • Assign a named owner for monitoring each system, separate from the team that deployed it
  • Map who is accountable at each stage, from supplier to frontline user

Phase 3: Evidence and reporting (six to twelve months)

  • Produce regular performance reports, even where no regulator yet requires them
  • Test the escalation route with a simulated degradation
  • Require suppliers to disclose model dependencies and change boundaries in contracts

Priority actions for different starting points

For organisations just starting with AI

  1. Buy with monitoring in mind: Ask suppliers how performance is tracked after go-live and what data you will receive. A supplier with no answer is pricing a one-off approval world.
  2. Set thresholds before the pilot: Decide what error rate or performance drop would stop the tool, before results arrive and create pressure to accept them.
  3. Keep the scope narrow: Mirror the staged model internally. Start with a limited use case and widen it only once agreed evidence is in.

For organisations with AI already deployed

  1. Audit for silent change: Check whether any deployed system has been updated or had its underlying model swapped without a fresh assessment on your side.
  2. Close the reporting gap: Move from incident-only reporting to periodic performance reporting, which is the direction the report recommends for AI devices.
  3. Clarify accountability: Document who answers for an AI-influenced decision that goes wrong, and make sure that person has the authority to act on warnings.

For advanced implementations

  1. Define your own envelopes: For adaptive or fine-tuned systems, specify the operating boundaries and the tests that prove the system stays within them. This is the work the report expects regulators to ask for.
  2. Offer yourself as a staged-launch partner: In healthcare, readiness is likely to become a condition of early access. Elsewhere, the same capability strengthens the case with any regulator running live testing.
  3. Engage with the policy response: The cross-government response will settle the details. Organisations with real monitoring data will have more influence over the thresholds than those with opinions.

Four problems the blueprint does not solve

Challenge 1: The enforcement side depends on government acting

Staged entry is the part industry will welcome. The part that makes it safe, including enhanced enforcement mechanisms that the report says could extend to financial penalties, depends on government accepting the recommendations and giving the MHRA the means to use them. If the response adopts the faster entry routes but delays the enforcement tools, the balance the Commission designed tips towards risk.

Mitigation Strategy: Treat the recommendations as a package when responding to consultation or planning investment. Deploying organisations should build their own escalation and pause rights into supplier contracts now, rather than waiting for statutory powers that may arrive later or in a weaker form.

Challenge 2: Temporary stages have a habit of becoming permanent

The report insists staged pathways should be temporary. In practice, a restricted authorisation that generates revenue gives a supplier little urgency to complete the evidence for full approval, and a regulator with limited capacity may not push. Tools could sit in a supervised state for years, with patients and clinicians unclear what that status means.

Mitigation Strategy: Deployers should ask for the specific evidence thresholds and target dates attached to any staged tool, and set internal review points against them. Where a tool stays in a staged state beyond its expected window, that should trigger a procurement review.

Challenge 3: Readiness will not be evenly spread

If early access is tied to demonstrated AI readiness, it will favour organisations with established digital teams and governance. The report itself acknowledges that readiness varies considerably between healthcare providers. The likely result is that well-resourced organisations receive new tools first, and the gap between the best-equipped and the rest widens.

Mitigation Strategy: Organisations with limited capacity should pool monitoring and governance through regional or sector partnerships rather than attempt it alone. Policymakers adopting the model in any sector should pair readiness conditions with support for those that fall short, or accept that access will be unequal.

Challenge 4: Boundaries are hard to define for systems nobody fully controls

An operating envelope sounds tidy for a diagnostic model with a defined task. It is much harder for products built on general-purpose models owned by a handful of large developers, often outside the UK. The report flags this dependency as a resilience and sovereignty risk. A boundary is only as good as the ability to detect when the model underneath has shifted, and downstream suppliers often have limited visibility of that.

Mitigation Strategy: Require model dependency disclosure and change notification in contracts today, independent of any regulatory outcome. For high-stakes uses, favour suppliers that can show how they would detect an upstream model change, and test that claim.

Reality Check: A formal response to the report is still pending, so none of these mechanisms applies yet, and some will need new guidance or powers before they can operate. Organisations should expect years, not months, before lifecycle authorisation is routine in healthcare, and longer before other regulators adopt it formally. The preparation, however, is useful immediately.

What this means beyond the NHS

The Commission has produced a detailed, public design for regulating AI across its working life. Healthcare is where it starts because the stakes are visible and the MHRA already regulates devices. The problem it answers is not specific to healthcare. Financial services, legal services, policing and public sector procurement all face AI systems that change after sign-off, and the ICO’s work on a statutory data protection sandbox shows other regulators are already rethinking how AI can be tested under real conditions.

Three factors will decide whether organisations benefit or are caught out:

  1. Monitoring as a standing function: Organisations that already measure AI performance in live use, against agreed thresholds, will meet lifecycle expectations with modest adjustment. Those relying on launch-time testing will be rebuilding under deadline.
  2. Contracts that reach upstream: Change boundaries and model dependencies need to be written into supplier agreements, because a regulator will ask the deployer questions only the supplier can answer.
  3. Clear accountability: A lifecycle model spreads responsibility across suppliers, deployers and users. Organisations that decide who owns which risk before an incident will handle one far better than those that decide afterwards.

Rethinking what “approved” means

Until now, the question buyers asked of an AI product was whether it had been approved. Under a lifecycle model, the better question is whether it is still performing within the limits it was approved for, and who would know if it stopped. That changes what boards should ask for: not a certificate on file, but a current performance report and a named person responsible for acting on it.

It also changes how progress should be measured. The number of AI tools deployed says little. The share of deployed tools with agreed thresholds, live monitoring and a tested escalation route says far more about whether an organisation can use AI safely at scale, and whether a regulator would trust it with a tool still wearing L-plates.

Strategic Insight: The Commission’s report is framed as healthcare policy, but it reads as a working specification for regulating any AI system that keeps changing. Organisations in other sectors that build to it now are likely to find their own regulators asking for something very similar within a few years.

Your next steps

Immediate actions (this week):

  • Read Recommendations 5, 6, 14 and 17 of the report alongside our news summary
  • Identify which of your AI systems could change without you being told
  • Name one person responsible for post-deployment AI performance

Strategic priorities (this quarter):

  • Set written performance thresholds for each deployed AI system
  • Add model dependency disclosure and change notification to new supplier contracts
  • Map accountability for AI-influenced decisions from supplier to frontline user

Long-term considerations (this year):

  • Produce periodic performance reports for AI systems before any regulator requires them
  • Track the cross-government response and any equivalent moves by your own sector regulator
  • Decide whether to position your organisation as a partner for staged or live-testing launches

Source: Independent Commission led by NHS doctors sets out blueprint to accelerate safe AI adoption in healthcare (Medicines and Healthcare products Regulatory Agency, GOV.UK, 10 September 2026). Recommendation wording, survey figures and detail on staged authorisation, change control and post-market surveillance verified against the Commission’s full report.

This strategic analysis was written by Resultsense, a UK-focused AI news and analysis publication. We will be watching the cross-government response to see whether the enforcement powers arrive alongside the faster entry routes, and whether other UK regulators start borrowing the model. Read more analysis at Insights, or get in touch.