Skip to content
Cloud Security DeskSearch
Menu

Research noteAI systems

Threat-model the system around the model

The model endpoint is one component. The consequential paths often run through retrieval stores, orchestration identities, evaluation data, and operator tools.

Demonstration publication. The scenario and all numerical data are illustrative, not observed research findings.

By
Umair Akbar and Ahmed Elshekh
Published
Reading time
6 minutes
Coverage
AWS · Azure · Google Cloud

A concise research note for extending cloud threat models around production AI workloads.

At a glance

Key findings

  • The orchestration identity often spans more systems than the model endpoint.
  • Retrieval and evaluation data need separate provenance boundaries.
  • Operator tools can turn model output into privileged action.

Follow the action path

Start with what the application can cause, not only what the model can produce. Map each tool call, queue, function, data store, and human approval that turns output into an effect.

Separate data roles

Training, retrieval, evaluation, conversation, and operational telemetry serve different purposes. Give each a named owner, provenance expectation, retention decision, and permitted set of consumers.

Questions for the review

Ask questions that connect model behavior to cloud control evidence.

  • Which identity performs retrieval?
  • Where are tool arguments logged?
  • Can an operator replay a sensitive request?
  • What stops output from invoking an unintended action?

References

  1. NIST AI Risk Management Framework
  2. OWASP Top 10 for LLM Applications

From the desk

About the authors

This demonstration publication is attributed to Umair Akbar and Ahmed Elshekh, the publication’s owners and chief editors.

Owner & Chief Editor

Umair Akbar

Owner & Chief Editor

Ahmed Elshekh

Questions answered

  1. What does “Threat-model the system around the model” investigate?

    The model endpoint is one component. The consequential paths often run through retrieval stores, orchestration identities, evaluation data, and operator tools.

    Supporting context

    A concise research note for extending cloud threat models around production AI workloads.

  2. What is the publication’s central conclusion?

    The orchestration identity often spans more systems than the model endpoint.

    Supporting context

    Retrieval and evaluation data need separate provenance boundaries. Operator tools can turn model output into privileged action.

  3. Who should use this analysis, and for what decision?

    The research note is most useful to practitioners evaluating AI systems across AWS, Azure, and Google Cloud. It is designed to support a concrete review or operational decision, not to replace environment-specific testing.

  4. What mechanism or pattern does the analysis explain?

    Ask questions that connect model behavior to cloud control evidence.

    Supporting context
  5. What evidence supports the analysis?

    The publication explains its evidence or method in “Separate data roles” and cites 2 numbered references that readers can inspect.

    Supporting context

    Training, retrieval, evaluation, conversation, and operational telemetry serve different purposes. Give each a named owner, provenance expectation, retention decision, and permitted set of consumers.

  6. What are the scope boundaries or limitations?

    The scenario and numerical values are illustrative, not observed provider benchmarks or measured customer findings. The publication demonstrates a review method and must not be treated as a prevalence estimate.

  7. What should cloud-security practitioners do next?

    Start with what the application can cause, not only what the model can produce. Map each tool call, queue, function, data store, and human approval that turns output into an effect.

    Supporting context
  8. Which cloud systems and security topics are in scope?

    The publication covers AI systems with explicit scope across AWS, Azure, and Google Cloud.

  9. Who wrote the publication, and when was it updated?

    Umair Akbar and Ahmed Elshekh wrote the research note, published on July 14, 2026. The estimated reading time is 6 minutes.