☰
jena.run
/
♡

My AWS keys weren't compromised, but an "AssumeRole" event was logged.

A GitHub Action That Triggered a Security Incident Response in the Early Morning

My AWS keys weren't compromised, but I got a "AssumeRole" error


It’s not really unusual for a security alert to go off in the early morning.

Most of the time, it’s something like this.

FailedLogin
PortScan
WAFBlock
ImpossibleTravel
KnownScanner

I check it, and if it’s a false positive, I shut it down.

But this alert was a little strange.

AWS detected an STS call for a production role, even though it wasn’t a typical deployment time.

AssumeRoleWithWebIdentity
Role: deploy-prod
Source: GitHub Actions OIDC
Environment: production

My first thought was simple:

"Did someone deploy something in the middle of the night?"

I checked the Slack deployment channel.

There was nothing there.

No GitHub deployment either.

I asked the developer on call.

“Did you happen to run a prod deployment?”

The answer came right away.

“No.”

From that moment on, it was no longer just a system failure response—it became an incident response.


One strange thing

At first, I suspected a credential leak.

This is a common scenario in AWS incidents.

Access Key Leak
      |
      v
Credential Abuse
      |
      v
AWS API Calls

However, in this case, it wasn’t a long-term Access Key starting with “AKIA....”

CloudTrail showed “AssumeRoleWithWebIdentity.”

In other words, GitHub Actions used an OIDC token to obtain temporary credentials from AWS STS.

GitHub Actions’ OIDC is actually a pretty good architecture designed to avoid storing long-term AWS credentials in GitHub Secrets. GitHub also recommends that when using OIDC, you configure the cloud provider to verify trust conditions, such as specific repositories or branches.

So at first, I found it even more puzzling.

The AWS key hadn’t been leaked, so how did it gain access to a production role?


I started by tracing back through CloudTrail

First, I checked which APIs were called during that session.

To simplify, the example looked something like this.

02:13:41 AssumeRoleWithWebIdentity
02:13:43 GetCallerIdentity
02:13:46 DescribeClusters
02:13:49 ListBuckets
02:13:51 GetAuthorizationToken
02:14:02 DescribeServices

At this point, I became even more certain that this wasn’t a standard automated deployment.

In a normal deployment, the call patterns are fairly predictable.

GetCallerIdentity
GetAuthorizationToken
PushImage
UpdateService
DescribeServices

However,

ListBuckets
DescribeClusters

there were some outliers mixed in.

It was closer to an enumeration than a deployment.

The attacker didn’t immediately start breaking things after gaining access to AWS.

First,

“How far can I go?”

.


Next up was GitHub.

When it comes to OIDC sessions, GitHub Actions is the starting point.

I dug through the workflow for the problematic repository.

After just a few minutes, I spotted a strange YAML file.

name: PR Validation

on:
  pull_request_target:
    types: [opened, synchronize, reopened]

permissions:
  contents: read
  id-token: write

jobs:
  validate:
    runs-on: [self-hosted, linux]

    steps:
      - uses: actions/checkout@v4
        with:
          ref: ${{ github.event.pull_request.head.sha }}

      - run: npm ci

      - run: npm test

At first glance, it doesn’t look that strange.

When a PR is submitted, it installs dependencies and runs tests.

That’s standard procedure for CI.

The problem was that three of them were chained together.

pull_request_target
        +
Untrusted PR Checkout
        +
id-token: write

Even the GitHub documentation warns that if a privileged workflow—such as `pull_request_target`—is run after checking out the PR head (e.g., for build or test), an attacker could execute the code included in the PR within a privileged context. This is a pattern commonly referred to as a “pwn request.”


A single PR was essentially a license to execute code.

What the attacker had to do wasn’t complicated.

Fork the repository.

Create a pull request.

Then, they embed code that runs during the build process within a change that appears to be a legitimate code modification.

For example, the attack surface doesn’t necessarily have to be the application source code.

package.json
Makefile
setup.py
build.gradle
scripts/
test config
custom build hook

It just needs to be executed by CI.

The process ultimately looks like this.

Attacker Fork
      |
      v
Pull Request
      |
      v
Privileged Workflow
      |
      v
Checkout Untrusted Code
      |
      v
Code Execution on Runner

id-token: writewas also included here.

In GitHub Actions, a job with this permission can request an OIDC JWT.

In other words, there was no need to steal the AWS access key.

CI itself was in a position to obtain legitimate temporary AWS credentials on behalf of the attacker.


But shouldn’t AWS have prevented this?

That’s right.

So I checked the AWS IAM trust policy.

And that’s when the second problem arose.

To simplify the example, the structure was like this.

{
  "Effect": "Allow",
  "Principal": {
    "Federated": "arn:aws:iam::123456789012:oidc-provider/token.actions.githubusercontent.com"
  },
  "Action": "sts:AssumeRoleWithWebIdentity",
  "Condition": {
    "StringEquals": {
      "token.actions.githubusercontent.com:aud": "sts.amazonaws.com"
    },
    "StringLike": {
      "token.actions.githubusercontent.com:sub": "repo:example-org/*:*"
    }
  }
}

It was too broad.

I can understand the intent.

“Let’s allow access to the deployment role for our organization’s GitHub repositories.”

It’s convenient for operations.

There’s no need to modify the IAM policy every time a new repository is added.

However, from a security perspective, it’s a completely different story.

Trusted Organization
        !=
Trusted Execution Context

GitHub also explicitly states that when configuring OIDC, cloud providers should not accept tokens unconditionally but should restrict access based on predictable conditions such as repositories, branches, and environments. It also recommends applying protection rules when using a production environment.

In our environment, GitHub’s privileged workflows and AWS’s broad trust policies aligned perfectly.

Looking at each setting individually, they made sense from an operational standpoint.

But when combined, they created a vulnerability.


But that wasn’t the end of it.

It’s easy to assume that once you’ve identified this much, you’ve found the root cause.

Root Cause Found
Incident Closed

But this is where things actually start to get tedious for senior-level engineers.

It’s not just the fact that an attacker gained access to AWS—what matters is how far they were able to penetrate.

deploy-prod I checked what permissions were attached to the role.

ECR.

ECS.

Some S3.

Some read permissions for Secrets Manager.

There were also a few Kubernetes-related permissions.

In other words, the question had changed.

“Did you join AWS?”

But rather

“What did you view and what could you change using these credentials?”

That’s the question.


Cut them off immediately

We can’t proceed by first completing a thorough forensic investigation and then blocking access.

We need to reduce the blast radius first.

The first step was to temporarily block the GitHub OIDC trust for the production role.

Disable Trust
     |
     v
Block New Sessions

Next, we disabled the problematic workflow.

We also isolated the self-hosted runner group.

We also suspended all deployments during that timeframe.

The key point is that we can’t just rotate the credentials and call it a day.

In the case of an OIDC-based attack, the long-term AWS key may no longer exist.

If the root cause remains unresolved, the attacker can simply obtain a new token.


And there was a self-hosted runner.

This is where the situation escalated even further.

The workflow was not hosted on a GitHub-hosted runner

self-hosted

.

With a GitHub-hosted runner, you can expect a structure that uses a new VM for each job, but with a self-hosted runner, the same machine may be reused depending on the operational setup.

GitHub also explicitly states that special caution is required with self-hosted runners, as running untrusted code on them can lead to a persistent compromise of the environment.

In other words, we could not assume that the attacker had simply obtained the AWS credentials during the first CI job and stopped there.

For example, let’s imagine the attacker left something like this on the runner.

/tmp/
Runner workspace
Shell profile
Background process
Docker socket
Build cache
Git hooks
Credential helper

If the next normal deployment runs on the same runner, the attacker could wait for much stronger credentials.

The structure is as follows.

Malicious PR
     |
     v
Compromise Runner
     |
     v
Wait
     |
     v
Trusted Production Job
     |
     v
Steal New Credentials

Therefore, we decided to scrap the runner and create a new one, rather than “cleaning it up and reusing it.”


Now we have to comb through the entire organization.

It’s pointless to fix just one instance after finding it.

It’s highly likely that similar workflows have been copied to other repositories.

In fact, in large companies or organizations with many repositories, it’s common for CI workflows to be copied and pasted, causing the same security issue to spread to dozens of places.

So, we searched the entire organization for the following patterns.

pull_request_target

id-token: write

runs-on: self-hosted

pull_request.head.sha

repository: ${{ github.event.pull_request.head.repo.full_name }}

gh pr checkout

git fetch

We gave high priority to these combinations in particular.

Untrusted Input
      +
Code Execution
      +
Privileged Token

The focus is not on individual keywords, but on identifying where trust boundaries intersect.


We also took another look at third-party actions.

When investigating incidents, it’s easy to get carried away.

“Since I’m already looking at all GitHub Actions anyway, I might as well check this out too.”

As I dug through the workflows, I found many instances like this.

- uses: vendor/example-action@v3

Only the version tag was pinned.

It’s convenient, but tags aren’t immutable.

In fact, during the 2025 “tj-actions/changed-files” supply chain attack, the attacker replaced multiple version tags with malicious commits, exposing CI/CD secrets in the workflow logs; GitHub Advisory documented this as an incident that could have affected approximately 23,000 repositories.

GitHub recommends full-length commit SHA pinning as the safest method when using third-party Actions. Currently, the method for pinning Actions to immutable releases is also full SHA pinning.

Therefore,

- uses: vendor/example-action@v3

instead

- uses: vendor/example-action@8f4b7c2e6a9d...

the transition to this format was also included.


The cause of the incident was not just a single line of YAML

There is one conclusion that incident reviews are most wary of:

“The developer wrote the workflow incorrectly.”

If we reach this conclusion, the issue almost inevitably resurfaces.

The actual cause was a combination of several valid decisions.

Convenient PR Workflow
        |
        v
pull_request_target
        |
        +
Fast CI Infrastructure
        |
        v
Self-hosted Runner
        |
        +
Passwordless Cloud Auth
        |
        v
GitHub OIDC
        |
        +
Easy Repository Management
        |
        v
Broad IAM Trust Policy

Each had its own justification.

The problem was that all these convenience features were placed within the same trust boundary.

Security incidents rarely result from a single critical configuration error.

They occur when five instances of “This much should be okay” pile up.


So we changed the structure itself.

First, we completely separated the PR workflow from the deployment workflow.

Untrusted PR
      |
      v
Isolated CI
      |
      v
No Cloud Credentials

Furthermore, production deployments are executed only within a separate trusted context.

Protected Branch
      |
      v
Approved Environment
      |
      v
OIDC
      |
      v
Production Role

In GitHub’s privileged workflow, we do not check out and execute untrusted PR code. GitHub also recommends avoiding the execution of untrusted code during events—such as `pull_request_target`—where secrets or privileged tokens can be used.


OIDC isn’t “safe just because it doesn’t use keys”

When migrating to OIDC, people often say things like this:

“It’s safer now that the AWS key isn’t on GitHub anymore.”

That’s only half true.

The risk associated with static credentials is certainly reduced.

But this raises a new question:

“Who is authorized to receive a token?”

In OIDC, the trust policy becomes more of a secret than the credentials themselves.

A good architecture looks like this.

Repository
   +
Branch
   +
Environment
   +
Audience
   +
Approval
   |
   v
Short-lived AWS Session

A poor architecture looks like this.

Anything from GitHub Org
           |
           v
      Production AWS

Just because it’s passwordless doesn’t mean it’s trustless.


permissions:You shouldn’t just slap it on top of the workflow haphazardly.

We also eliminated the pattern of granting “id-token: write” to the entire workflow after an incident.

Keep the default as small as possible.

permissions:
  contents: read

And only add it to deployment jobs that truly require AWS authentication.

jobs:
  deploy:
    permissions:
      contents: read
      id-token: write

GitHub also recommends specifying “GITHUB_TOKEN” permissions as the minimum necessary and expanding them only for the jobs that require them.

From a security engineer’s perspective, even though a single line like this may seem insignificant, it determines the blast radius.


The hardest part of incident response isn’t determining “whether we’ve been breached or not.”

The development team, of course, asks this first.

“So, were we actually compromised?”

From the security team’s perspective, this is the hardest question to answer.

The absence of evidence is not the same as there being no compromise.

So the answer is usually not binary.

Confirmed
Likely
Possible
Unlikely
Ruled Out

We assess it based on the scope.

In a case like this,

AWS role assumption is Confirmed.

Resource enumeration is confirmed.

Verify access to Secrets Manager via CloudTrail.

Whether the secret was actually used requires a separate investigation.

Self-hosted runner persistence is Possible.

Source code modifications are verified via GitHub audit logs and commit history.

Modifications to production workloads are verified using CloudTrail, EKS audit logs, and other sources.

Incident response ultimately involves deciding how conservatively to respond given this imperfect evidence.


As you become a senior engineer, you spend more time making decisions than actually hacking.

Early in a security engineer’s career, it’s easy to think that the technical aspects are the hardest.

Exploit
Reverse Engineering
Cloud Security
Malware Analysis
Detection Engineering

Of course, they are difficult.

However, at the senior level, the real challenge lies in questions like these:

Is the situation serious enough to halt all production deployments right now?

Should we scrap all 30 runners?

Should we rotate the credentials for 200 repositories?

Do we need to disclose this to our customers?

To what extent should we block the development team’s releases?

Should we gather more evidence, or take down the infrastructure right now?

Even when there is a technically correct answer, there is often no operationally correct answer.

The security team’s role is not simply

Find Vulnerability

but

Understand Risk
      |
      v
Limit Blast Radius
      |
      v
Preserve Evidence
      |
      v
Restore Trust
      |
      v
Prevent Recurrence

.


And what remained in the end was simpler than expected

The initial entry point for this incident wasn’t even a major 0-day vulnerability.

AWS itself wasn’t compromised either.

Nor was GitHub OIDC vulnerable.

Nor was the “self-hosted runner” technology itself vulnerable.

For the most part, each system functioned as intended.

GitHub issued a token because the workflow requested it.

AWS issued a role because the trust policy allowed it.

The runner executed the code because the workflow instructed it to do so.

The problem lay somewhere in between.

GitHub trusted the workflow.

AWS trusted GitHub.

The runner trusted the PR.

The workflow trusted the runner.

Nobody questioned the entire chain.

In security, the risk often doesn’t come from “untrusted systems.”

It’s systems that trust each other too much.

And perhaps the job of a senior security engineer ultimately boils down to scrutinizing those trust relationships one by one.

While looking at a single line in CloudTrail at 2 a.m.

“But why does this one have this permission?”

끝까지 읽었네… 메밀 승인 🐾
+