
An architecture analysis for platform and DevSecOps engineers who run claude-code-action, run-gemini-cli or Copilot on public issues and pull requests, built from the 2025 and 2026 disclosures and GitHub's own documentation. It maps which triggers admit untrusted text, what the job can reach, which channels carry data out, and the read-only agent job plus scoped write job that contains the risk.
At a glance
Key findings
- An agent that reads an issue title, a pull request body or a comment is reading text an outside contributor wrote, so the design question is not how to phrase the prompt but what a followed instruction can reach: the job's GITHUB_TOKEN scopes, its model API key, any id-token write permission, and the Actions cache. [2][3][4]
- Scoping one job is not the boundary. The Clinejection chain Adnan Khan disclosed starts in a triage job with a restricted token and reaches release credentials by flooding the Actions cache until GitHub evicts the legitimate entries, then poisoning the keys a nightly publish workflow restores. [1][12]
- A public summary or log is an output channel. The Flatt chain read a low-privilege token from the publicly visible run summary, used it to edit a trusted user's issue, and reached a tag-mode workflow that held id-token write and the OIDC request credentials. [3][19]
- The structural fix is job separation: a read-only agent job that holds no publish secrets, writes buffered as an artifact and applied by a separate job with one narrow write scope, and no cache writes from low-trust triggers. The buffer is only as strong as the validator that reads it. [8][10][11]
- GitHub changed the platform defaults in 2026: read-only cache tokens for untrusted triggers, cache-mode went generally available on September 10, 2026, and workflow execution protections disable pull_request_target by default on public repositories without an applicable event policy from November 2, 2026. [13][14][15]
The job an outside contributor can steer
An AI coding agent wired into GitHub Actions is a program that reads text an outside contributor wrote and then calls tools with your repository's credentials. The useful way to hold it is as an untrusted-input processor: assume the issue title, the pull request body and the comment can carry instructions the model will follow, and build the workflow so that a followed instruction reaches nothing worth stealing. The protection is the shape of the job around the model, not a cleverer system prompt. Microsoft's analysis of Claude Code's action makes the same point, advising that the system prompt be treated as a defense-in-depth layer whose job is to reduce noise, make the agent more predictable and block simple exploits, and Anthropic's own security documentation says its environment scrubbing reduces but does not eliminate prompt injection risk. [4][5]
Three questions set the blast radius. Which triggers let outside text reach the model: issues, comments, pull request titles and bodies, pull_request_target, workflow_run, edited events and bot actors. What the job can reach once the model acts: the GITHUB_TOKEN scopes, a request for an OIDC token through id-token write, stored API keys, and write access to the Actions cache. And which channels can carry data back out: issue and pull request comments, commits, the run summary, the logs, a fetch to a pre-approved domain, and the cache itself. Answer the three for a given workflow and the structural fix follows: a read-only agent job with no publish secrets, writes buffered and applied by a separate narrowly scoped job, and no cache writes from low-trust events. [2][4][8]
The reason a short answer like "the triage job has no secrets" fails is that the job is not the boundary. Two disclosed chains make the point. In the Cline chain, a triage job with a restricted token could poison a cache that a later, privileged nightly workflow restored and trusted. [1] In the Flatt Security research against Claude Code's action, a low-privilege token leaked through the publicly visible run summary was used to edit a trusted user's issue and pivot into a workflow that held id-token write. [3] This guide follows the untrusted text from the trigger that admits it, through what the job holds, to the channels that carry it out, then builds the split that contains it and the platform defaults that back the split up.
Untrusted issue text crossing into a job that holds tokens
The failure mode is one job that both reads outside text and holds the credentials; the fix is a partition the writes pass through, not the reads. [2][3]

Source. Conceptual illustration based on the Comment and Control and Flatt Security disclosures. [2][3]
Method. Conceptual hand-authored illustration. Shapes and counts are illustrative and do not represent a specific workflow.
Accessible table and figure data
| Element | What it represents |
|---|---|
| Issue and comment cards | Issue title, body and comments written by an outside contributor |
| Red arrows | Untrusted text crossing into the job that runs the model |
| One CI job | A single job that both reads the text and holds credentials |
| Credential chips | Model key, GITHUB_TOKEN and OIDC request token in that one job |
| Solid partition and cabinet | Where the boundary belongs, with the publish secrets behind it |
| Element | What it represents |
|---|---|
| Issue and comment cards | Issue title, body and comments written by an outside contributor |
| Red arrows | Untrusted text crossing into the job that runs the model |
| One CI job | A single job that both reads the text and holds credentials |
| Credential chips | Model key, GITHUB_TOKEN and OIDC request token in that one job |
| Solid partition and cabinet | Where the boundary belongs, with the publish secrets behind it |
How untrusted text reaches the model
Start by naming every trigger that can place outside text in the model's context, because the triggers differ sharply in what they hand the job. A pull request from a fork is the gentlest: with the exception of the GITHUB_TOKEN, secrets are not passed to a workflow triggered from a fork, and that token is read-only on fork pull requests. [18] The dangerous triggers are the ones that run with the base repository's privileges while still carrying attacker-authored content. A pull_request_target or a workflow_run executes with the base repository's secrets and a writable token, which is the whole reason the pwn request pattern exists. [16][17] GitHub's documentation is blunt that a workflow started by workflow_run can read secrets and use a write token even when the run that triggered it could not. [18]
Issues and comments are the triggers people underestimate. A workflow on the issues or issue_comment event runs in the context of the default branch with the repository's own token at whatever scopes the workflow declared, and anyone with a GitHub account can fire it by opening an issue or leaving a comment. The issue_comment event also fires for comments on pull requests, so a comment trigger reaches both. [18] Two refinements widen the surface. An attacker who gains an issues write token can edit a trusted user's issue after a privileged workflow is set to read it, turning trusted content into attacker content; Anthropic later mitigated this by ignoring issues and comments edited after the trigger. [3] And payloads hidden in HTML comments are invisible in rendered markdown but still reach a model that reads the raw text, which is how one disclosed Copilot chain ran without the victim seeing the instruction. [2][4]
Bot and app actors deserve their own line. A GitHub App has implicit read access to public repositories and can open issues and pull requests on a repository where it is not installed, so an actor check that trusts any account ending in [bot] can be bypassed by anyone who registers an app. That was the exact bypass RyotaK reported against claude-code-action's agent mode, fixed by adding a human-actor check. [3] The lesson that holds across every row is that write access to the repository is not the same as authorship of the content: an approval gate answers who may start the agent, never who wrote the text the agent is about to read.
The table below is a conceptual synthesis of the documented behavior, not a measurement. Read it as the first inventory step for a specific repository: for each trigger your workflows use, write down who can fire it and what the resulting job can reach.
Which triggers admit outside text and what each job reaches
The triggers that run with base secrets and a write token are the ones anyone can fire, so inventory them first. [16][18]

Source. Conceptual synthesis of GitHub documentation on events, secure use, dependency caching and execution protections, with the bot-actor bypass from Flatt Security. [3][12][15][16][18]
Method. Conceptual. Each row states documented behavior as of October 2026; it is a synthesis, not a measurement.
Accessible table and figure data
| Trigger | Who can fire it | What the job can reach |
|---|---|---|
| issues (opened, edited) | Anyone with an account | Repo token at declared scopes, default-branch cache read-only |
| issue_comment | Anyone with an account | Same as issues; also fires on pull request comments |
| pull_request from a fork | Any forker | Read-only token, no secrets, cache scoped to the merge ref |
| pull_request_target | Any forker | Base secrets and a write token; blocked by default on public repos from Nov 2, 2026 |
| workflow_run after a fork PR | Any forker upstream | Secrets and a write token even if the upstream run had none |
| Bot and app actors | Anyone who registers an app | Implicit read on public repos; can open issues and pull requests |
| Trigger | Who can fire it | What the job can reach |
|---|---|---|
| issues (opened, edited) | Anyone with an account | Repo token at declared scopes, default-branch cache read-only |
| issue_comment | Anyone with an account | Same as issues; also fires on pull request comments |
| pull_request from a fork | Any forker | Read-only token, no secrets, cache scoped to the merge ref |
| pull_request_target | Any forker | Base secrets and a write token; blocked by default on public repos from Nov 2, 2026 |
| workflow_run after a fork PR | Any forker upstream | Secrets and a write token even if the upstream run had none |
| Bot and app actors | Anyone who registers an app | Implicit read on public repos; can open issues and pull requests |
Case records from 2025 and 2026
Four disclosures published between February and June 2026, from reports dating back to October 2025, map cleanly onto four failure modes, and each is worth reading as a design lesson rather than as news. Check identifiers against the primary record before repeating them. The cross-vendor Comment and Control research lists bug bounty report numbers and no CVE. The Gemini CLI trust-model advisory, GHSA-wpqr-6v78-jr5g, was published on April 24, 2026 with a CVSS 3.1 score of 10.0 and no CVE; CVE-2026-12537, which Google's CVE numbering authority published on June 24, 2026 at CVSS version 4.0 base 10.0, references that advisory, while the repository advisory page itself still shows no known CVE. [2][6][7]
Comment and Control, published April 15, 2026, is the clearest statement of the base failure mode: an agent that reads a pull request title, issue body or comment and runs tools in the same runtime that holds the secrets. Aonan Guan and colleagues demonstrated the same pattern against three agents, exfiltrating the host repository's own secrets back through GitHub comments, commits and logs rather than any outside server. The Copilot variant bypassed environment filtering, secret scanning and a network firewall at once by reading the environments of the parent and MCP server processes, which the filter on the bash subprocess did not cover, base64-encoding the tokens past the scanner, and pushing them as a commit to the allowlisted github.com. Anthropic rated its finding None after initially accepting a critical severity, and blocked the ps command; the researcher noted that reading /proc/*/environ achieves the same result, which is why a blocklist is not a boundary. [2]
Clinejection, disclosed February 9, 2026, is the lateral-movement lesson. Cline's issue triage workflow ran the Claude action with allowed_non_write_users set to * and a broad tool list, so any account could trigger it, and its only direct secret was the model key. [1] The pivot did not need a secret in that job at all. Because any workflow could then read and write the Actions cache, and an issues-triggered run shares the default branch cache scope with the nightly publisher, the attacker could flood the cache past its 10 GB limit to force eviction of the legitimate entries, then claim the vacated keys with poisoned node_modules. In the chain Khan described, the nightly publish would then restore the poisoned cache and expose VSCE_PAT, OVSX_PAT and NPM_RELEASE_TOKEN, which carry production rights because the two extension marketplaces tie tokens to the publisher and the npm token covered the one package used for both nightly and production releases. [1][12] Khan tested only on a mirror and reports that a different actor used his proof of concept against Cline and obtained the publication credentials. After the public fix, a credential rotation that deleted the wrong npm token left the exposed one live, and an unauthorized cline@2.3.0 with an added postinstall script reached npm eight days later. [1][22]
The Flatt Security research, published June 1, 2026, is the chaining lesson, and Microsoft's June 5, 2026 write-up is the sandbox-gap lesson. RyotaK showed that a low-privilege issues: write token recovered from the publicly visible run summary could edit a trusted user's issue, and that a tag-mode workflow holding id-token write exposed ACTIONS_ID_TOKEN_REQUEST_TOKEN and ACTIONS_ID_TOKEN_REQUEST_URL, which together let an attacker replay the OIDC exchange for a privileged installation token. [3] Microsoft Threat Intelligence found that the action's Read tool ran in-process and could read /proc/self/environ outside the Bubblewrap sandbox that scrubbed the Bash subprocess, recovering the unscrubbed model key; a prompt that cut the first seven characters of the value slipped it past both the model's refusal and GitHub's secret scanner. [4] The chart records how long each fix took where both the report date and an end date are published.
Days from report to fix for the disclosed agent CI issues
Reported fixes landed within days to about six weeks, but two of these intervals end at a bug-bounty resolved state rather than a confirmed code fix. [1][2][3][4]

Source. Calculated from report and fix or resolution dates published by the researchers, reviewed October 9, 2026. [1][2][3][4]
Method. Whole-day differences between the published report date and the published end date. The Cline (report 2026-01-01 to fix 2026-02-09), claude-code-action bypass (2026-01-12 to 2026-01-16) and Read tool (2026-04-29 to 2026-05-05) intervals end at a code fix. The Copilot (2026-02-08 to 2026-03-09) and Claude Code Security Review (2025-10-17 to 2025-11-25) intervals end at a bug-bounty resolved state, which is not proof of a fix.
Accessible table and figure data
| Disclosure | Days to fix or resolve |
|---|---|
| Cline triage workflow (fix) | 39 |
| claude-code-action bot bypass (fix) | 4 |
| Claude Code Read tool (fix) | 6 |
| Copilot agent (bounty resolved) | 29 |
| Claude Code Security Review (bounty resolved) | 39 |
| Disclosure | Days to fix or resolve |
|---|---|
| Cline triage workflow (fix) | 39 |
| claude-code-action bot bypass (fix) | 4 |
| Claude Code Read tool (fix) | 6 |
| Copilot agent (bounty resolved) | 29 |
| Claude Code Security Review (bounty resolved) | 39 |
Tokens, secrets and caches inside the job
Once untrusted text can steer the model, the exposure is whatever the job holds, whether or not a secret is passed as an input. The GITHUB_TOKEN carries only the scopes the workflow declares, and once any permission is specified every unlisted permission is set to none, so a top-level permissions: {} and a per-job contents: read is a real reduction. [27] It is also short lived: the token expires when the job finishes, with a maximum of six hours on GitHub-hosted runners, so what matters in an incident is what it did during the run, not whether it is still valid afterward. [20] A model API key is a different matter: it is a long-lived secret that stays valid after the run ends. Anthropic's guidance makes the same distinction for GitHub credentials, warning that a static personal access token does not rotate between runs and could be recovered over time through prompt injection, and recommending the auto-generated secrets.GITHUB_TOKEN for workflows that admit users without write access. [5]
The sharpest escalation is id-token write. It does not grant access to anything by itself, but it lets the job request an OIDC token and exposes the two request variables an attacker needs to replay the exchange against whatever external trust the job can reach. [19] Treat id-token write as a credential in its own right and keep it out of any job that reads outside text. The same logic applies to a pre-approved fetch domain: Claude Code auto-approved any path under huggingface.co for its WebFetch tool, outside any --allowedTools restriction, which turned an allowlisted domain into an out-of-band exfiltration channel until version 2.1.163 fixed it. An allowlisted destination is still a way out. [26]
The Actions cache is the channel most workflows forget, and it works in two directions. Anyone who can open a pull request can read the base branch caches, so the cache is an exfiltration surface; and because cache contents are not signed or verified, a poisoned entry becomes code execution in any later run that restores it. [12] The eviction mechanics are what make poisoning practical: the default limit is 10 GB per repository, entries unused for over seven days are removed, and GitHub evicts least-recently-used entries once the limit is passed, so an attacker who can write junk can force out a legitimate key and claim it. [12]
The output side deserves the same inventory as the input side, because a secret the model reads is only a problem if it can leave. Enumerate every channel the job can write to: issue and pull request comments, commits and branches, the run summary, the step logs, a fetch to any pre-approved domain, and the cache. Each appears in a disclosed chain or advisory. The Comment and Control variants returned credentials through comments, a committed file and the Actions log. [2] The Flatt chain read a token straight from the publicly visible run summary. [3] A pre-approved fetch domain was reported as an out-of-band channel. [26] An attacker does not need a connection to an outside server when the result can be written back where the repository already lets it be read. The fragment below keeps the agent job read-only and turns cache access off for it; the next section moves the publish secrets out entirely.
name: issue-triage
on:
issues:
types: [opened]
# Nothing is granted that is not re-granted per job.
permissions: {}
jobs:
triage:
runs-on: ubuntu-latest
permissions:
contents: read
issues: read
# No id-token, no write scopes, no cache for a job that reads outside text.
cache-mode: none
steps:
- uses: actions/checkout@1111111111111111111111111111111111111111 # v7.0.1
with:
persist-credentials: false
- uses: anthropics/claude-code-action@2222222222222222222222222222222222222222 # v1.0.248
with:
# The model key is itself a secret; keep publish secrets out of this job.
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
github_token: ${{ secrets.GITHUB_TOKEN }}
# A prompt selects agent mode. Actors without write access are skipped
# unless allowed_non_write_users is set; set it only on a job scoped like this.
prompt: "Read issue #${{ github.event.issue.number }} with mcp__github__get_issue and propose labels."
claude_args: '--allowedTools "mcp__github__get_issue"'
Split read, reason and write into separate jobs
The containment that holds under injection is to separate the job that reads from the job that writes. GitHub's own Agentic Workflows project is built around this split and is the clearest documented reference for it: the agent runs with read-only permissions and cannot write to external state, its proposed actions are buffered as an artifact, a threat-detection job with no write permission inspects the artifact and emits a pass or fail verdict, and only then do separate safe-output jobs with narrowly scoped write tokens apply the changes. [8][9] Its documentation, checked on October 10, 2026, gives no overall release stage and labels several individual features experimental or preview, so treat it as a reference design to study rather than a supported product. [8] A fully compromised agent in this shape cannot create a pull request or post a comment directly, because it never holds the token to do so. The agent job can still read repository state through a read-only server; the write lives on the far side of a validation step.
The split is a real boundary, but two gh-aw advisories show exactly where it leaks, and both lessons transfer to any hand-built version. In the first, the safe-output validator spread the raw agent item and checked only the schema-declared fields, so any undeclared field the agent added reached the applier verbatim, which let a steered agent overwrite another user's comment by id or retarget a pull request's base branch. The advisory lists v0.86.1 as patched and proposes building the item from the declared schema fields only. [10] The lesson is that the write job must validate against a closed schema and read nothing the schema did not declare. In the second, the artifact that the safe-output jobs saved captured their raw process output, including a credential-bearing git configuration command, so a CI trigger token that never passed through the agent was recoverable by anyone who could download the artifact. [11] Every artifact is data that leaves the run, so treat each one as a surface that can carry secrets.
A workable split for a hand-built workflow puts the agent and the applier in different jobs connected by an artifact, as in the fragment below. The agent job holds no write scopes and uploads a single file of proposed actions. The apply job holds exactly one write scope, downloads the file, and runs an applier that validates a closed schema before it touches the API. Keep the artifact free of anything secret, since it is downloadable by anyone with read access to the run.
jobs:
agent:
runs-on: ubuntu-latest
permissions:
contents: read # no write token in the job that reads outside text
cache-mode: none
steps:
- uses: actions/checkout@1111111111111111111111111111111111111111 # v7.0.1
with:
persist-credentials: false
- run: ./run-agent.sh > proposed-actions.json
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
- uses: actions/upload-artifact@3333333333333333333333333333333333333333 # v7.0.2
with:
name: proposed-actions
path: proposed-actions.json
apply:
needs: agent
runs-on: ubuntu-latest
permissions:
contents: read # check out the applier from the default branch
issues: write # the one write scope this job holds
cache-mode: none
steps:
- uses: actions/checkout@1111111111111111111111111111111111111111 # v7.0.1
with:
persist-credentials: false
- uses: actions/download-artifact@4444444444444444444444444444444444444444 # v8.0.2
with:
name: proposed-actions
- run: ./apply-actions.sh proposed-actions.json # reject any field the schema did not declare
env:
GH_TOKEN: ${{ github.token }}A read-only agent job with writes applied by a separate job
The agent job holds the model key and read scopes only; the write token lives in a job that runs after a validation step. [8][9]

Source. Conceptual architecture based on the GitHub Agentic Workflows security model, with the validator and artifact failure modes from its advisories. [8][9][10][11]
Method. Conceptual. Edges are authored to follow the documented job separation; the design generalizes to a hand-built split.
Accessible table and figure data
| Job or object | Holds | Role |
|---|---|---|
| Issue or PR text | Nothing | Untrusted input from an outside author |
| Agent job | Model key, read scopes | Reads the text, proposes actions, no write token |
| Buffered artifact | Proposed actions as data | Carries the plan between jobs; keep it free of secrets |
| Threat-detection job | No write permission | Inspects the artifact and gates the writes |
| Safe-output job | One scoped write token | Applies the approved actions after validating a schema |
| Job or object | Holds | Role |
|---|---|---|
| Issue or PR text | Nothing | Untrusted input from an outside author |
| Agent job | Model key, read scopes | Reads the text, proposes actions, no write token |
| Buffered artifact | Proposed actions as data | Carries the plan between jobs; keep it free of secrets |
| Threat-detection job | No write permission | Inspects the artifact and gates the writes |
| Safe-output job | One scoped write token | Applies the approved actions after validating a schema |
Permissions and approvals that hold under injection
With the split in place, tighten what each job holds and how a write gets approved. Set permissions: {} at the top of the workflow and add scopes only in the jobs that need them, so the read job cannot be widened by a default. [17] Replace the static model key where you can: Anthropic's action supports workload identity federation, which exchanges the workflow's OIDC token for a short-lived Anthropic token so there is no long-lived ANTHROPIC_API_KEY to store or leak, at the cost of granting id-token write to that job. [24] That cost collides with the rule above: the request credentials in that job can mint an OIDC token for any audience, which is the replay the Flatt chain used. [3][19] A reasonable reading is that federation fits jobs that read only trusted input, while a job that reads public text is better served by a scoped, separately monitored API key. Put publish secrets behind a GitHub environment with required reviewers so a job cannot read them until a person approves the run, and keep the agent job out of that environment entirely.
Scope the apply job as narrowly as the task allows. A labeling workflow needs issues: write and nothing else; a comment workflow needs the same; a workflow that opens a pull request needs contents: write and pull-requests: write and should get no issue access. The point of the narrow scope is that if the applier is tricked into acting on a field it should have rejected, the damage is bounded by the one scope it holds rather than by the full set the agent might otherwise have carried. Keep each write job to a single purpose so its token stays small.
- Keep
allowed_non_write_usersand a wildcardallowed_botsoff unless the job has almost no permissions; Anthropic flags both as significant risks and notes that an allowed bot is not checked for repository access, so on a public repository any app can match. [5] - Leave
show_full_outputanddisplay_reportat their disabled defaults. Full output prints tool results into Actions logs that are public on public repositories, and the step summary, whichdisplay_reportcontrols, carried the token in the Flatt chain before Anthropic turned the summary off by default. [3][5] - Run low-trust agents with a strict tool allowlist. Gemini CLI, which the run-gemini-cli action runs, evaluates tool allowlisting even under its yolo mode from version 0.39.1, and its trust guidance asks for a minimal set such as read and search tools when the input is untrusted. [6][23]
- Require human review before a bot's changes run workflows. Copilot's cloud agent only lets write-access users trigger it, never shows comments from users without write access to the model, and by default holds its pull request workflows until a person with write access clicks Approve and run workflows. [21]
Platform defaults that changed in 2026
GitHub spent 2026 moving several of these controls from advice into platform defaults, and the dates matter because some are enforced automatically. On June 26, 2026 the Actions service began issuing read-only cache tokens on the default branch for events that someone without write access can trigger, such as pull_request_target, issue_comment and fork-driven workflow_run, which closes the Clinejection-style poisoning path for workflows that do not opt out. [13] On September 10, 2026 the cache-mode key reached general availability, letting a workflow or job declare read, write, write-only or none; the low-trust default is read, and an explicit write-capable mode on an untrusted trigger brings the poisoning risk back and raises a warning annotation. [12][14]
The checkout action changed too. On June 18, 2026 actions/checkout v7 went generally available and refuses the common pwn request patterns by default, failing when a pull_request_target or pull-request-driven workflow_run tries to check out a fork's head or merge ref. The enforcement was backported to supported major versions on July 20, 2026, but a workflow pinned to a specific SHA, as supply-chain guidance recommends, does not receive the backport and must be upgraded deliberately. The change also does not cover a run step that fetches untrusted code with git or the gh CLI, nor the issue_comment trigger. [16]
The largest shift is workflow execution protections, generally available on September 17, 2026, which let an administrator define actor and event allowlists evaluated before a run starts. For public repositories with no applicable event policy, GitHub introduced a default rule that disables pull_request_target, running first in evaluate mode and enforced automatically on November 2, 2026. A repository that still needs the trigger has to allow it explicitly in an event policy before that date. The mechanics of the trigger and this default change are covered separately; the point for agent workflows is that any agent wired to pull_request_target on a public repository will stop running unless its owner acts. [15]
Test the workflow with hostile inputs
Most of these defects only appear when the workflow meets the input it was built to resist, so cause each case on purpose in a disposable repository that holds only throwaway credentials, and record the date and the action, agent and runner versions you tested. The aim is not to prove the model refuses, which it may do inconsistently, but to prove that a successful injection reaches nothing that matters.
- Fire each trigger your workflows use with a payload in the field it reads: an issue title, an issue body, a pull request title, a comment, and the same instruction hidden inside an HTML comment so you confirm the sanitizer strips it.
- Edit an issue created by a trusted account after the workflow is set to read it, and confirm the agent processes the original text, not the edit.
- Open an issue or pull request from a bot or app actor and from an account with no write access, and confirm the agent does not run where it should not.
- From inside the agent job, attempt to read a publish secret, request an OIDC token, write a default branch cache entry, and reach a domain outside the allowlist. Each should fail.
- Feed the apply job an artifact carrying an undeclared field and a forged size, and confirm it rejects both rather than passing them to the API. [10]
- Read the run summary and the logs after each test and confirm no secret value appears, including a value with its first characters removed. [4]
When not to run an agent on public input
Some workflows should not run an agent on untrusted input at all, and the test is whether the content itself decides a consequential action. When the task requires the model to choose which tool to run based on text an attacker can write, no job split contains it, because the decision lives inside the untrusted content; the realistic options are to limit the agent to drafting output for a person to approve, or not to run it on that trigger. Microsoft frames the same limit as a rule that an AI workflow should never hold untrusted input, access to sensitive systems and the ability to change state or communicate externally at the same time. [4] The published patterns for containing injection make the same point in the general case, and this is their concrete application to continuous integration.
For the workflows that do belong on public input, the order of operations is fixed. Inventory the triggers and write down who can fire each and what its job reaches. Strip the read job to the model key and read-only scopes, with no id-token write and no cache write. Buffer every write as an artifact and apply it from a separate job that holds one narrow scope and validates a closed schema. Put publish secrets behind an environment with required reviewers. Turn on the 2026 platform defaults rather than relying on them silently: explicit cache-mode on low-trust jobs, checkout v7 or the backport, and an event policy for pull_request_target before November 2, 2026. Then run the hostile-input tests and record what passed.
If one rule has to survive a review, use this one: the job that reads text from outside the repository must never be the job that holds a credential worth stealing. Everything above is a way to keep those two things in different jobs.
Method and provenance
Source-led analysis of original security research (Khan, Guan, Flatt Security, Microsoft Threat Intelligence), GitHub and vendor documentation, GitHub security advisories read through the GitHub and CVE Services APIs, and GitHub changelogs. Sources were reviewed on October 9, 2026 and re-checked in an independent fact-check on October 10, 2026.
No GitHub repository, workflow or cloud account was configured or run for this article. Trigger behavior, cache limits and platform dates are bounded to the cited documentation and advisories as of the review date, and fast-moving action and agent defaults can change after publication. Two charted intervals end at a bug-bounty resolved state rather than a confirmed code fix, as noted on the figure.
AI assistance. AI assisted the research synthesis, drafting, diagram planning and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.
Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.
References
- Clinejection: compromising Cline's production releases by prompting an issue triager Adnan Khan. Published . Accessed .
- Comment and Control: prompt injection to credential theft in Claude Code, Gemini CLI and GitHub Copilot Aonan Guan. Published . Accessed .
- Poisoning Claude Code: one GitHub issue to break the supply chain GMO Flatt Security. Published . Accessed .
- Securing CI/CD in an agentic world: Claude Code GitHub action case Microsoft Threat Intelligence. Published . Accessed .
- claude-code-action security documentation Anthropic. Accessed .
- Gemini CLI and run-gemini-cli trust model (GHSA-wpqr-6v78-jr5g) Google. Published . Accessed .
- CVE-2026-12537: unauthenticated remote code execution in Gemini CLI CI/CD workflows CVE Program. Published . Accessed .
- GitHub Agentic Workflows security architecture GitHub. Accessed .
- GitHub Agentic Workflows safe outputs reference GitHub. Accessed .
- gh-aw safe-output validator forwards undeclared fields (GHSA-jxrq-hq57-gwwm) GitHub. Published . Accessed .
- gh-aw safe-output artifacts may expose CI trigger tokens (GHSA-8h78-hpm7-29gg) GitHub. Published . Accessed .
- Dependency caching reference GitHub. Accessed .
- Read-only Actions cache for untrusted triggers GitHub. Published . Accessed .
- Control GitHub Actions cache access with cache-mode GitHub. Published . Accessed .
- Workflow execution protections in GitHub Actions generally available GitHub. Published . Accessed .
- Safer pull_request_target defaults for GitHub Actions checkout GitHub. Published . Accessed .
- Secure use reference for GitHub Actions GitHub. Accessed .
- Events that trigger workflows GitHub. Accessed .
- OpenID Connect reference for GitHub Actions GitHub. Accessed .
- About the GITHUB_TOKEN GitHub. Accessed .
- Risks and mitigations for GitHub Copilot cloud agent GitHub. Accessed .
- Unauthorized npm publish of cline@2.3.0 (GHSA-9ppg-jx86-fqw7) Cline. Published . Accessed .
- run-gemini-cli trust guidance for GitHub Actions Google. Accessed .
- claude-code-action setup and workload identity federation Anthropic. Accessed .
- gh-aw confused-deputy check omits pull_request_target (GHSA-r8gh-v7wv-8g7h) GitHub. Published . Accessed .
- Out-of-band data exfiltration via pre-approved HuggingFace domain in WebFetch (GHSA-fg94-h982-f3mm, CVE-2026-54316) Anthropic. Published . Accessed .
- Workflow syntax for GitHub Actions GitHub. Accessed .