Skip to content
KeyDrift
Scan for free
All posts

Build logs and CI artifacts: the leak surface nobody rotates

Echoed env vars, defeated masking, cached working directories, and 90-day retention: how CI infrastructure outlives the secret you rotated.

KeyDrift13 min read

Build logs and CI artifacts: the leak surface nobody rotates

Deleting the commit does not empty the pipeline. This maps how secrets persist in CI/CD logs, artifacts, and caches long after the offending commit is gone — with concrete hygiene settings for the major systems.

Somewhere behind most projects there is a log line that prints an environment variable, an artifact that contains a .env file, or a cache holding a working directory from a year ago. None of it is in version control, so none of it is caught by commit-time scanning. And unlike a leaked repository file, nobody rotates these objects, because nobody thinks of them as holding credentials at all.

That is the defining property of this surface: persistence without ownership. A key pasted into source has an obvious remediation path. A key echoed into a build log on a Tuesday in March has none — it sits in someone else's infrastructure, indexed and searchable, for as long as the platform's retention policy allows. This guide walks the mechanisms: what logs record, where masking fails, what artifacts and caches carry, and the settings that bound the exposure. Where a leaked key is found, the response is rotation, per the fix guides; this article is about shrinking the window before and after.

What build logs actually record

Logs accumulate secrets through several mundane mechanisms, none of which look like a mistake while they are happening.

Shell tracing. set -x (or Set-PSDebug -Trace in PowerShell) echoes every executed line, including variable expansions. A deploy script traced for debugging prints the exact commands that authenticate — with the credential expanded inline. The tracing flag is added during one broken build and rarely removed deliberately.

Debug output. Scripts print environments "just to see what's available": printenv or env | sort, dumped config objects, HTTP clients logging request headers. Any of these captures Authorization: Bearer … values verbatim.

Verbose tooling. Package managers, IaC tools, and deployment CLIs all have debug modes that log more than intended. A Terraform run with verbose logging can emit provider blocks including sensitive arguments. Test frameworks print fixtures. Database migration tools log connection strings — which for most managed databases are credentials, username and password included.

Failure dumps. When a step crashes, some runners print the full environment of the failed process in the error report. The crash you debugged in April may have published every variable the job held.

The common thread: the secret never touched git. It moved from the platform's secret store into process memory into a log line, entirely within the CI system.

Masking works — until you transform the value

Every major CI has a masking story, and every masking story has the same structural limit: masking matches the literal registered value. Transformations produce strings the matcher has never seen.

GitHub Actions documents this directly: values from the secrets context and values registered via ::add-mask:: are replaced in logs, but if a secret is used to produce another value, that derived value is not masked (Using secrets in GitHub Actions). The canonical failure is encoding:

- name: Debug deploy (do not do this)
  env:
    DEPLOY_TOKEN: ${{ secrets.DEPLOY_TOKEN }}
  run: |
    echo "token is: $DEPLOY_TOKEN"
    # ^ masked: prints "***"

    echo "$DEPLOY_TOKEN" | base64
    # ^ NOT masked: the base64 encoding is a different string.
    # Anyone with the log and one decode command has the token.

    jq -n --arg t "$DEPLOY_TOKEN" '{token: $t}'
    # ^ NOT masked again: JSON-escaping changes the bytes.

The same gap exists wherever masking is literal-matching. The practical rules that follow from it:

  1. Never encode, encrypt, wrap, or template a secret in a step whose logs you have not reviewed.
  2. If a pipeline must derive values from secrets, register the derived form explicitly — GitHub exposes ::add-mask::VALUE for exactly this (Workflow commands).
  3. Treat set -x as incompatible with steps that touch secrets. Trace locally against dummies, not in CI against real ones.

Masking format constraints matter too. GitLab CI, for instance, only masks variables meeting specific requirements — minimum length, single-line, restricted character sets — so short or structured values silently go unmasked (GitLab variable docs). A variable you believe is masked because the box is checked may not be.

Log and variable handling across major CI systems

The systems differ less in capability than in defaults. The table summarizes documented behavior worth verifying against your own setup:

SystemMasking behaviorLog visibility controlsRetention controls
GitHub ActionsLiteral matching of secrets context and add-mask registrations; derived values unprotected (docs)Repo-scoped; forks and PR runs see redacted secrets by defaultLogs and artifacts follow the retention setting — default 90 days, configurable per workflow/repo/org (docs)
GitLab CIMasks CI/CD variables that meet format constraints; others pass through (docs)Job logs visible per project permissions; protected environments gate variable exposureInstance/admin-configured log expiration
CircleCIContext and project variables handled per current platform docs; verify masking for your planLog access tied to org/project roles; contexts restrict which jobs see which valuesArtifacts expire on a configurable schedule (30-day default class) (docs)
JenkinsCredentials Binding plugin masks bound credentials in console output (docs)Console access per auth matrix; builds often world-readable inside teams by defaultUnlimited by default — logs persist until workspace/disk cleanup

Two rows deserve emphasis. Jenkins's unlimited default means decade-old consoles outlive every rotation you have ever performed. And GitHub's 90-day default cuts the other direction — it bounds exposure but also means a leak discovered late is still inside its retention window, fully searchable.

Artifacts, caches, and working directories

Logs are half the surface. The other half is stored files.

Artifacts snapshot working directories. Upload steps glob more than authors expect: upload: . captures dotfiles, including .env files that exist only on the runner because the platform injected them. A test-results artifact that sweeps the workspace can carry the very credentials the tests used. Artifact contents are downloadable by anyone with read access to the repo's actions — a much larger audience than the people who approved the secret.

Caches outlive branches. Dependency and build caches keyed loosely (or restored with fallback keys) can propagate a polluted working directory across jobs and months. Cache poisoning aside, the benign case is bad enough: a cached directory containing a local-only config file resurfaces in a different pipeline context long after everyone forgot the file existed.

Retention is the clock on all of it. Defaults differ by platform and setting tier, and the maximum is usually longer than teams assume. The operational takeaway is to treat retention as a security parameter, not a storage cost knob: shorter retention for artifacts produced by steps that handle credentials, and explicit excludes for anything resembling environment files.

A minimal hardening block for GitHub Actions illustrates the knobs:

jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      # Scope the upload; never glob the whole workspace.
      - name: Upload test results
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: test-results
          path: |
            junit.xml
          retention-days: 7        # default would be 90
          if-no-files-found: ignore

      - name: Deploy
        env:
          DEPLOY_TOKEN: ${{ secrets.DEPLOY_TOKEN }}
        run: ./scripts/deploy.sh   # script must not set -x or echo env

And the audit equivalent for existing history: enumerate recent workflow runs, download their artifacts, and grep them the same way you would grep a checkout — the formats are identical, because the files inside artifacts are ordinary project files.

# After downloading artifacts for review:
grep -rEo 'sk_live_[A-Za-z0-9]{10,}|AKIA[0-9A-Z]{16}|sb_secret_[A-Za-z0-9_]{10,}' .

(All patterns match synthetic examples fine; they are format detectors, not validators.)

Why this surface never gets rotated

Ask what happens after a repo leak and you get a checklist: revoke, rotate, purge history. Ask what happens after a log leak and the honest answer is usually nothing — not because teams don't care, but because the object holding the secret is invisible to every process built around credentials.

Three structural reasons:

No inventory connects logs to keys. Rotation procedures start from "where is this key used?" — code references, platform config. A log line is outside every such map. Nobody searching for usages of STRIPE_SECRET_KEY in a monorepo will find March's traced deploy script.

Discovery lags exposure by the length of retention. The leak is noticed when someone happens to read an old log — which, given retention windows measured in months, can be long after the key has been changed for unrelated reasons, or worse, exactly as long-lived as the key itself. AWS-style access keys that go unrotated for years pair perfectly with a CI appliance whose logs go unread for years.

Ownership is split. The pipeline belongs to whoever wrote the workflow; the logs belong to the platform; the key belongs to whichever team owns the provider account. A leak in the seams has no obvious first responder. The pragmatic fix is procedural: treat any confirmed log-or-artifact disclosure exactly like a source disclosure — rotate immediately via the provider-specific procedures, then apply hygiene so the class of leak stops repeating.

A hygiene baseline

Condensing the above into the settings worth applying this week:

  • Remove set -x and environment dumps from any job with real credentials; provide a sanitized --dry-run mode for debugging instead.
  • Never encode or template secrets in logged steps; register derived forms with your CI's masking mechanism if derivation is unavoidable.
  • Scope artifact uploads to explicit paths; exclude dotfiles and environment files; cap retention-days aggressively for credential-adjacent jobs.
  • Review cache keys for cross-branch fallbacks; exclude local-only config from cached directories.
  • Set the shortest workable retention for logs on repos where pipelines hold elevated credentials; verify your platform's actual default rather than assuming.
  • Audit historical artifacts and logs once, then automate the ongoing check — the same artifact-level scanning logic that watches your deploy output extends naturally to CI output, and continuous monitoring exists because one audit ages out at the next deploy. For the sibling surfaces: what ships in bundles and why history deletions never delete keys complete the post-git map.

Structuring scripts so logs stay safe

Redaction failures trace to scripts that print reflexively. Three patterns harden them at the source:

1. Quiet defaults with explicit verbosity. Invert the usual arrangement — tools log one-line progress by default, and full detail requires an explicit flag plus a conscious decision:

./scripts/deploy.sh            # one line: "deployed rev a1b2c3 → staging"
./scripts/deploy.sh --verbose  # detailed output; refuses to run when
                               # DEPLOY_TOKEN is set (guard below)

2. The verbose-guard. Make debug modes incompatible with real credentials mechanically:

if [[ "$VERBOSE" == "1" && -n "$DEPLOY_TOKEN" ]]; then
  echo "refusing: --verbose with live credentials" >&2
  exit 2
fi

3. Mask-before-transform helpers. When derivation is unavoidable, register derived forms immediately and locally (the masking gap):

derived=$(printf '%s' "$DEPLOY_TOKEN" | base64)
echo "::add-mask::$derived"    # GitHub Actions; equivalent per platform

Code review can verify all three patterns in seconds, which matters more than their sophistication: logging hygiene that requires reading diffs carefully never survives review pressure. Teams adopting these report a secondary benefit worth naming — scripts become testable, since quiet-mode output is stable enough to assert against in CI, closing the loop between pipeline reliability and pipeline secrecy.

The vendor half of the surface

Self-hosted assumptions break quietly on hosted CI, where logs live in someone else's systems under policies you influence but don't control. The standing questions for every platform in your pipeline:

  • Who can read historical logs? Org roles define the audience; contractor and offboarding hygiene determine how that audience changes over time. Support staff access during incidents is real and worth knowing about rather than discovering.
  • What retention applies by default and by plan? Defaults differ (the 90-day Actions default versus CircleCI's artifact schedule versus Jenkins's unlimited local consoles); plan tiers sometimes change them. Verify against current documentation per platform (table above).
  • Where do logs get forwarded? Log aggregation, observability platforms, and chat integrations copy pipeline output into systems with their own retention and access models — a leak surface one hop removed from the CI system everyone audits.
  • What do vendors do with failure dumps? Some platforms capture environment context in crash diagnostics; review provider privacy documentation during procurement, not after incidents.

None of this argues against hosted CI — it argues for treating log residency as part of the threat model. The same credential leaked to self-hosted files versus a multi-tenant aggregator has different audiences, different persistence, and different cleanup options; pipelines should know which world they operate in before deciding what's safe to echo.

Forwarded logs extend the retention question to every aggregation layer downstream of CI: observability platforms, log warehouses, chat integrations posting build summaries. Each destination carries its own retention defaults, access models, and export paths — and each becomes searchable storage for whatever the pipeline echoed. The audit question generalizes cleanly: for every place build output lands, who can search it, and for how long? Answers belong in the same inventory as the CI platforms themselves.

Log hygiene composes with the rest of this series rather than standing alone: the same substitution leaks that put keys into bundles put them into verbose builds (the founding mechanism), and log discoveries rotate through the same provider procedures as any other finding. Teams running continuous monitoring on deployed artifacts typically extend the habit to scheduled artifact-log reviews within a quarter — the retrieval muscle generalizes.

FAQ

If I rotate a key, do old logs become harmless? Mostly yes — the logged value no longer authenticates anything. Two caveats: connection strings and session tokens age differently from API keys, and the pattern that caused the leak keeps producing fresh leaks. Rotation closes yesterday's incident; hygiene closes tomorrow's.

Our CI is self-hosted and logs never leave our network. Does that reduce the risk? It shrinks the audience from "anyone with repo read access" to "anyone with runner access" — a narrower set that still spans everyone who can read runner output, plus every future compromise of the runner host. Internal exposure is exposure on a delay.

Are PRs from forks a special risk? Yes, in a precise way: fork pull requests run with restricted secret access by design on hosted platforms, so the fork itself cannot exfiltrate secrets. The risk is misconfiguration — a workflow that grants secrets to fork PRs to make something work converts every drive-by contribution into a potential collection run.

What should I do if I confirm a credential in a stored log right now? Rotate the credential first — treat it as disclosed to everyone with access to that system. Then delete or shorten-retention the affected logs and artifacts where the platform allows, and record the finding in your incident notes so patterns repeat visibly.


Pipeline logs and artifacts fail silently: no scanner watches them by default, and their exposure window is measured by retention policy, not by your attention. KeyDrift's free scan — no signup — audits what your deployed build actually serves, which is the other half of the same rule: audit artifacts, not intentions.

KeyDrift reads the JavaScript your app actually serves and finds the Supabase, Stripe, OpenAI and AWS keys that should never have left your server. Run a free scan.

KeyDrift is a Veristria product. More about KeyDrift.