Beyond the Cloud: Indirect Control as the New Paradigm for Medical Device Software

Bernhard Kappe
Bernhard Kappe
TIR115 Bernhard Thought Leadership Image Banner

The Problem the Cloud Made Visible

For decades, medical device quality rested on a quiet assumption: the manufacturer controls the device. It controls the hardware, the software, and every change to either. Validation, change control and submissions were all built on that assumption.

Public cloud computing broke it. A device function running in the cloud depends on infrastructure the manufacturer does not operate, and the provider changes that infrastructure continuously, sometimes without notice.

AAMI TIR115:2026, Guidance for the appropriate use of public cloud computing to enable medical device functions, is the industry’s answer for the cloud. But in writing it, our working group found something larger. The real subject wasn’t the cloud. It was indirect control: what happens when a device depends on computing someone else runs and changes. And that pattern is now everywhere.

The Indirect Control Paradigm

Four ideas run through TIR115. Stated generally, they apply to any dependency a manufacturer doesn’t fully control.

  1. Acknowledge indirect control. Name the things your device depends on but doesn’t control. The benefits are real, and so are the risks. Pretending you have direct control is the most common failure.
  2. Take a risk-based approach. Analyze what can go wrong when the dependency changes or fails, mitigate what you can, and be willing to conclude that a function shouldn’t depend on it at all.
  3. Draw the device boundary. Separate the Medical Device System (MDS), which is legally part of the device, from the Medical Device Digital Environment (MDDE), which the device relies on but which is not part of it. Generic services usually sit in the MDDE; clinically specific components sit in the MDS.
  4. Manage the dependency as a supplier. TIR115 treats cloud infrastructure as a procured service managed under ISO 13485 supplier controls, not as software of unknown provenance (SOUP). The general lesson: a dependency you can’t control still needs a relationship you can manage.

Applied to the cloud, these ideas are now guidance. Applied elsewhere, they are a way of thinking that teams can use today, before any technology-specific guidance exists. TIR115 itself points the way: its list of out-of-scope topics (Annex E) names applying its principles to smartphones, clinical-system integration and AI/ML as possible future work.

At a Glance

Indirect Control Bernhard Article

TIR115 addresses the lower right: generic infrastructure under indirect control. Third-party LLMs doing clinical work sit in the upper right, the quadrant existing frameworks cover least.

Smartphones: The Device in the Patient’s Pocket

When a medical device app runs on a patient’s own phone, the manufacturer controls the app and very little else. The operating system updates on the phone owner’s schedule. New hardware ships every year. Permissions, background execution rules, Bluetooth behavior and app store policies all change without the manufacturer’s input. TIR115’s SpiroBreath case study (Annex D) recalls an earlier era, when a manufacturer could tell a clinic by agreement not to install OS updates until it had tested them. Patients’ own phones offer no such lever.

The paradigm maps cleanly:

  • Boundary. TIR115’s connected insulin pump example (Annex C) draws the line this way: the patient’s phone hardware and OS are MDDE; the manufacturer’s app is MDS.
  • Risk. An OS update that throttles background activity can silently delay an alert. That hazard belongs in the risk file with a mitigation, such as detecting the condition and warning the user.
  • Control you can exercise. A supported-platform policy, monitoring of OS beta releases, and compatibility testing before each major release replace the control you don’t have.

Many teams already do some of this informally. The paradigm makes it deliberate and documented, which is what reviewers and auditors look for.

Third-Party LLMs: The Hardest Case

For the cloud, the paradigm has a comfortable answer. Delegate generic infrastructure to the provider and keep the clinical function under your own control. As Randy Horton has put it, you can hand the cloud provider low-level functions, but you shouldn’t let it decide when it’s a heart attack. TIR115 makes the same point with a wearable defibrillator (section 5.1): real-time detection can’t depend on a cloud resource, while retrospective analysis of the same data can.

A third-party large language model breaks that comfort. When a device uses a model it calls through an API, the component under indirect control may be doing the clinical work itself: summarizing a record, drafting a recommendation, interpreting a patient’s words. That puts something inside the MDS that the manufacturer still can’t fully control. The cloud case mostly avoids that combination: TIR115 expects supplier-controlled cloud resources to sit in the MDDE, and treats one inside the MDS as possible but less likely (Annex D, FAQ 1).

Three properties make it harder still:

  • Behavior changes without code changes. A model update can shift outputs while your software stays identical.
  • Outputs aren’t fully deterministic. The same input may not produce the same answer, so verification means testing distributions of behavior, not single results.
  • Versions are retired. Providers deprecate models on their own schedules, forcing change on yours.

The paradigm still gives teams a place to start:

  • Acknowledge that the model is a dependency under indirect control, even when it sits inside the device.
  • Treat the provider as a supplier, with terms covering version pinning, notice of changes and deprecation timelines where the provider allows it.
  • Build evaluation suites that run against every model version before it reaches patients, and monitor behavior in use.
  • Decide deliberately which functions a third-party model should perform, and which it shouldn’t.

This thinking complements, rather than replaces, FDA’s frameworks for AI-enabled devices, such as predetermined change control plans. Those frameworks address changes the manufacturer plans. Indirect control addresses changes someone else makes.

Risk Management Patterns for Indirect Control

The four ideas say what to think about. Teams also need repeatable ways to analyze and monitor risk when a dependency can fail or change underneath them. For the cloud, those patterns are well established. For smartphones, they are taking shape. For agentic AI and LLMs, the industry, including us, is still working them out.

Cloud: Bottom-Up DFMEAs with Transparent Failure Modes

Top-down hazard analysis starts from what can harm a patient. On its own, it struggles with a system built from dozens of services, each of which can be slow, unavailable, wrong or changed by its provider. What works for cloud systems is to pair it with a bottom-up design FMEA (DFMEA) for each module and service.

  • One DFMEA per module or service. For each, list how it can fail: unavailable, slow, returning stale or wrong data, completing only partway, or behaving differently after a provider change.
  • Effects stated as residual risk. Record each failure mode’s effect after its mitigations, such as retries, fallbacks, redundancy and alerts. What flows up to the system-level analysis is the risk that remains, not the raw failure.
  • Transparent failure modes. Design each service so its failures are detected and reported, never silent. A service that fails visibly can be contained by the services around it; one that fails quietly passes bad data downstream.

This follows the same logic as FDA’s guidance on interoperable medical devices and on multiple function device products: analyze each interface or function on its own, then assess how its failure affects the functions around it. It is also how cloud-native design already works. Health checks, timeouts, circuit breakers, bulkheads and observability are mitigations that make failure modes visible and keep them contained. The DFMEA turns those engineering patterns into risk management evidence, and because each service has its own analysis, the risk file follows the architecture. When a provider changes a service, the team updates that service’s DFMEA and checks whether the residual effect it passes up has changed.

Smartphones: Field-Level Self-Validation

On a patient’s phone, the manufacturer can’t test every combination of hardware, OS version and settings before release. The pattern that generalizes is field-level self-validation: the app checks, while running, that the conditions it depends on still hold, such as the OS version, permissions, background execution, Bluetooth connection and notification delivery, and tells the user or care team when they don’t. Monitoring those checks across the installed base shows when a new OS release starts causing trouble, often before support calls do. The same DFMEA approach applies, with the phone and its OS treated as a dependency whose failure modes the app must detect.

Agentic AI and LLMs: Patterns Still Forming

For agentic AI and LLMs, there are not yet generalizable patterns as mature as the cloud DFMEA or smartphone self-validation. Several are emerging from our own work with multiagent systems and from conversations across the industry:

  • Give agents design controls as context. In our work with multiagent systems, providing design controls context (approved user needs, requirements, risk analysis, architecture and test scenarios) substantially reduces risk. Left to themselves, agents often go beyond what was asked or behave in unexpected ways; given approved design controls, they stay within them. The more structured those controls are, the better this works: well-formed user stories, Gherkin scenarios, and structured language for non-functional requirements and risk, all kept as code under configuration management.
  • Surround the model with deterministic checks. Linters, trace checks and automated tests can verify what an agent produced more reliably than another model’s judgment.
  • Separate duties between agents. The agent that writes a requirement should not be the one that implements or verifies it.
  • Define an operating envelope and test against it statistically. Because outputs vary, verification means characterizing behavior across many runs, with evaluation suites rerun whenever the model changes.
  • Monitor in use. Run a new agentic function alongside the existing process first, and measure how often it goes wrong at volume before relying on it.

These patterns apply most directly to agentic AI used to build device software. How far they carry over to an LLM performing a clinical function inside the device is still open. [An industry working group on risk management for agentic AI is forming to take up both questions.]

Four Questions to Ask Now

Whatever your device depends on, the same four questions apply:

  1. What does our device depend on that we don’t control?
  2. What happens to patients if that dependency changes or fails?
  3. Where is the line between our device and its environment, and is it documented the same way everywhere?
  4. How do we find out when the dependency changes, and who decides what to do about it?

Teams that can answer these clearly for the cloud are well placed to answer them for phones and AI models next.

How Orthogonal Can Help

Orthogonal co-chaired the AAMI working group that wrote TIR115. We help MedTech teams apply its principles wherever indirect control shows up, starting with a cloud readiness assessment and extending to smartphone apps and third-party AI models.

TIR115 addresses public cloud computing and lists these extensions as out of scope (Annex E). Applying its principles to smartphones and AI models reflects the authors’ views, not AAMI guidance.


AAMI TIR115:2026: Guidance for the Appropriate Use of Public Cloud Computing to Enable Medical Device Functions

Published by the Association for the Advancement of Medical Instrumentation (AAMI), TIR115 provides guidance for managing public cloud computing used to support regulated medical device functions.

Related Posts

Article

AAMI Publishes TIR115: New Guidance for Using Public Cloud Computing in Medical Devices

Article

Your Next AI Algorithm May Not Be the Hard Part

Article

How MedTech Teams Must Evolve for the Agentic AI Era

Article

Why Ecosystem Design Controls Are Really About Moving Faster With the Right Rigor