Where detection belongs in the pipeline
Pre-commit, pre-push, CI, pre-deploy, post-deploy: what each gate actually sees, what it costs in friction, and the layered answer that survives real teams.
Where detection belongs in the pipeline
For teams designing leak detection into their delivery process: the five candidate gates compared by coverage and friction, working configurations for each, and why any single gate — even the right one — fails alone.
"Add secret scanning to the pipeline" is advice shaped like a decision but sized like a slogan. Pipelines aren't one place; they're a sequence of gates, each observing different objects at different moments under different incentives. A scanner at the wrong gate isn't useless — it's blind to specific leak classes by construction. A scanner at every gate with equal strictness drowns developers in friction until they route around it.
This article maps the five gates honestly: what each can see given its position, what it structurally cannot, the configuration that makes it work, and the failure mode that eventually visits unlayered setups. The through-line: coverage is a property of layers, not gates — each gate exists because another one gets bypassed, misses a class, or fires too late. Companion reading: why point-in-time checks expire covers the when argument; this covers where.
The five gates at a glance
| Gate | Observes | Catches best | Friction | Structural blind spots |
|---|---|---|---|---|
| Pre-commit (local hook) | Staged files, pre-history | Pasted literals before they exist anywhere shared | Seconds per commit | Bypassed with --no-verify; machine-local |
| Pre-push / server-side push protection | Pushed refs, at the forge | Everything local tools missed, enforced centrally | One blocked push when triggered | Only git content; only partnered patterns reliably |
| CI (build job) | Full checkout + history + built output | Team-wide backstop; build-output substitutions | Minutes per build | Runs after the leak already entered history |
| Pre-deploy (CD gate) | Final artifacts | Substitution leaks materialized at build (mechanics) | Deploy-blocking on findings | Nothing about already-live surfaces |
| Post-deploy + scheduled | Live URLs, payloads, chunks | Drift between deploys; regressions; non-deploy changes | None (async alerting) | Detection, not prevention |
Read the last column as a unit: no row covers the others' blind spots. That's the entire argument for layering, made tabular.
Gate one: pre-commit — the cheapest no
The local hook scans staged files in milliseconds, converting would-be commits into editor sessions. It's the only gate positioned to cost the author nothing but attention:
# .husky/pre-commit
#!/usr/bin/env sh
. "$(dirname -- "$0")/_/husky.sh"
# Scans ONLY staged changes; redaction keeps keys out of hook output logs.
gitleaks protect --staged --redact -v
Setup takes minutes (npm i -D husky && npx husky init, drop the script above), and the ruleset travels with the repository so every contributor inherits it. Two design constraints keep it durable:
- Budget: seconds. Scan staged diffs, not whole trees or history. Hooks that take longer than a coffee sip get disabled culturally if not literally.
- Scope: additions only. The commit moment can't remediate history; demanding more ensures the hook argues with developers instead of informing them.
The honest limitation: client-side enforcement is advisory by nature. git commit --no-verify skips everything, machines lack the hook entirely until tooling installs it, and contributors on unfamiliar setups generate their first commits ungated. Pre-commit reduces accident volume; it enforces nothing.
Gate two: pre-push and server-side protection
Moving enforcement server-side converts advisory into mandatory. Forge-level push protection evaluates pushed content against pattern databases and rejects offending pushes outright — GitHub's implementation blocks partnered-format credentials at push time across eligible repositories (push protection docs), with coverage and gaps mapped honestly in the dedicated analysis.
Server-side placement buys three properties no client hook can:
- Non-bypassability — there is no
--no-verifyflag for someone else's server. - Universality — every contributor, every machine, including the ones where hooks never installed.
- Pattern currency — the forge maintains detector updates continuously; locally maintained rule files rot.
Its blind spot mirrors its strength: jurisdiction ends at git content on that host. Push protection knows nothing of values that reached built artifacts without transiting a commit — which is precisely the dominant modern leak path, and precisely why the next two gates exist.
Gate three: CI — the team-wide backstop
The build job sees everything at once: complete checkout, optionally full history, and — critically — the compiled output that no earlier gate observes. As the team's invariant enforcer, it belongs in every repository regardless of local-tool adoption:
# .github/workflows/secret-scan.yml
name: secret-scan
on: [pull_request, push]
jobs:
leaks:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with: { fetch-depth: 0 } # full history for the audit scan
- name: Repository scan (tree + history)
uses: gitleaks/gitleaks-action@v2
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
# The step most pipelines skip — and the one that finds substitution leaks:
- name: Build
run: npm ci && npm run build
- name: Scan built output
run: npx --yes gitleaks detect --no-git --source ./dist --redact
That final step deserves its comment. Framework conventions inline environment values during builds (prefix mechanics), meaning a repository can pass every source-oriented check while dist/ carries a live credential. The build directory is the first object that resembles what users receive; scanning it moves detection one gate closer to reality.
CI-gate caveats worth engineering around: results arriving here mean the leak already sits in history (response shifts toward rotation plus rewrite per the git-history guide); and CI's own secrets handling needs the masking discipline covered in the CI-artifacts guide, lest the scanner's home become a leak surface.
Gate four: pre-deploy — the final artifact check
Between "build succeeded" and "traffic routes," a deploy gate examines exactly what will serve: bundles, chunks, payload scripts, source-map decisions. This is detection's last preventive position — findings here block publication rather than describe exposure.
Mechanically it reuses gate-three tooling against the release artifact, but placement changes semantics: a finding now means this release carries the credential, rollback is cheap, and the fix deploys on normal cadence. Teams wiring this as a required CD check report the practical benefit differently than CI placement: because it inspects the assembled artifact across all environments' inputs, it catches environment-specific accidents — staging credentials promoted inside production builds — that source scans cannot see even in principle.
For deployed-surface checking beyond the build directory, KeyDrift's CLI makes the same assertion scriptable directly against rendered output:
keydrift https://staging.example.com --fail-on critical
# exit 0 clean · exit 1 findings · exit 2 could-not-scan (per /docs/scanner)
Placed as the deployment's final step, exit codes gate promotion without custom plumbing.
Gate five: post-deploy and scheduled — monitoring the live surface
The pipeline ends; drift doesn't. After promotion, several forces mutate served output without any new build: cache invalidation layers shifting chunk availability, runtime-payload changes, preview deployments appearing with inherited variables, hotfixes bypassing ceremony, and the slow regeneration of old mistakes by AI-assisted iteration (the migration loop). Post-deploy gates observe once; the live surface keeps changing afterward.
This gate's shape differs from all others: asynchronous, continuous, diff-based. KeyDrift's monitoring model fits here — re-checks riding deploys plus scheduled sweeps, alerting on transitions (new findings, regressions) rather than re-reporting known state (scanning versus monitoring). Its outputs feed response rather than blocking: findings trigger the rotation playbook and path repair, with severity routed by the same tiers scanners use (rules reference).
Two design notes keep the gate trustworthy. Regression severity: a finding that returns after remediation signals the causal loop survived — the prompt that reintroduced it, the convention that taught it — and deserves escalation beyond its format's base tier. And signal discipline at scale: post-deploy channels observing continuously must apply the same filtering rigor as every other detector, or volume trains operators toward muting precisely when incidents need readers (the false-positive economics).
Why any single gate fails
Each gate's bypass story, tabulated — the case for layering in one view:
| If you run only… | The failure arrives via… |
|---|---|
| Pre-commit hooks | --no-verify, unhooked machines, non-git surfaces |
| Push protection | Unpartnered formats; values that never transit commits (bundles, logs) |
| CI source scans | Build-time substitution into artifacts; post-merge timing |
| Pre-deploy scans | Changes after promotion; hotfix paths; preview environments |
| Post-deploy monitoring | Prevention gap — detection arrives after exposure begins |
The layered answer composes them by role: prevention at pre-commit and push (cheap rejection), verification at CI and pre-deploy (artifact ground truth), assurance post-deploy (continuous observation of what prevention and verification can't fully cover). Each layer presumes the others will eventually leak something past them; none is trusted absolutely; together they reduce both leak probability and detection latency to levels no single gate reaches.
Budget-wise, the composition is cheaper than it sounds: hooks are minutes, CI steps are minutes, the CD assertion is one command, and post-deploy monitoring starts with a free first scan. The expensive configuration — the one teams actually end up paying for — is discovering which uncovered gate the incident used.
An adoption sequence for a typical team
Layering sounds expensive until sequenced properly; each gate lands on its own timetable, owned by whoever already holds that stage:
| Week | Gate added | Owner | Success looks like |
|---|---|---|---|
| 1 | CI repository scan (tree + history) | Whoever owns CI | First findings triaged; baseline established |
| 1–2 | Pre-commit hooks (opt-in rollout) | Each developer | Hook installed by default via bootstrap script |
| 2 | Push protection enabled at forge | Org admin | Blocked-push telemetry visible |
| 3 | Built-output scan appended to CI build job | CI owner | Substitution leaks become catchable |
| 4 | Pre-deploy assertion (CD final step) | Deploy owner | Exit-code gate blocking criticals |
| 4+ | Post-deploy monitoring + scheduled sweeps | Security/on-call rotation | Alert routing with SLAs active |
The ordering follows dependency, not importance: CI runs first because everything else keys off patterns it establishes; hooks follow so individual machines converge on team defaults; artifact stages come after source stages stabilize because they inherit severity tiers and allowlists (rules design). A team of one compresses this to two sittings — CI plus monitoring cover most solo risk surface — while larger teams find the table maps onto existing ownership without new hires.
Each row's "success" column doubles as its acceptance test: gates that never produced a finding, blocked a push, or fed an alert need investigation rather than congratulation — either coverage is misconfigured or signal quality needs the tuning pass. Gates earn trust through demonstrated catches and through quiet correctness between them; both belong in the rollout record.
Budget honestly at each step: weeks one through three cost hours total using free tooling; week four's monitoring starts free (first scan) before any paid watching begins. The expensive failure mode isn't any line item — it's stopping after week two feeling covered while artifact stages, where modern leaks actually materialize, remain dark.
Metrics for the pipeline itself. Gates need their own instrumentation, or they decay invisibly. Five numbers, each owned by the gate's owner from the adoption table:
| Metric | Gate | Healthy signal |
|---|---|---|
| Hook disable/bypass rate | Pre-commit | Near zero; spikes signal friction to fix, not enforcement to tighten |
| Blocked pushes per month | Push protection | Nonzero and reviewed — every block is a near-miss with a story |
| CI findings trend | CI scan | Declining after onboarding; sudden rises correlate with changes worth examining |
| Deploy blocks / overrides | Pre-deploy | Blocks rare and justified; overrides require named approval |
| Detection latency | Post-deploy monitoring | Time-from-introduction-to-alert trending down; regressions alerted within one cycle |
Review them monthly in existing engineering meetings rather than security rituals — five minutes of dashboard reading. The metrics also arbitrate tuning honestly: a gate producing zero findings and zero blocks for a quarter is either covering a genuinely clean surface (verify once) or misconfigured (fix immediately), and only the numbers distinguish those cases.
One meta-metric closes the loop: mean time from finding to rotation across all gates. Every layer above exists to shrink that number; when it grows, look first at alert quality (the fatigue economics) before blaming responders.
Handling findings at each gate
Detection without disposition planning just relocates chaos to a new stage. Each gate's findings have a natural owner and response:
| Gate | Finding means | First responder | Response |
|---|---|---|---|
| Pre-commit | You just almost committed a key | The author | Replace with env reference or placeholder; nothing shipped |
| Push protection | A push was blocked org-wide | Author + tooling owner | Fix locally; treat as near-miss telemetry |
| CI scan | Credential entered history | Author + reviewer | Rotate if valid; rewrite per history guide |
| Built-output / pre-deploy | Substitution materialized a secret | Service owner | Block deploy; rotate; fix consumption path |
| Post-deploy monitor | Live surface serves (or served) a credential | On-call / security rotation | Incident runbook; regression gets escalated severity |
The table's quiet insight: response urgency increases toward the right — earlier gates catch intent, later gates catch disclosure. Teams that route every finding to the same queue erase that gradient, treating blocked pushes like live exposures and burning the attention that real incidents need.
FAQ
Won't all these gates create constant friction? Friction scales with findings, not with gates: a clean pipeline pays seconds for hooks and minutes for CI steps that mostly confirm cleanliness. Poorly tuned detectors — flagging every public key (the tier problem) — generate friction at any placement; fix signal quality rather than deleting layers.
We have pre-commit hooks org-mandated. Is that enough? It's the weakest single-layer posture available, precisely because it's the most bypassable gate. Add server-side push protection next (free at the forge for supported patterns); treat everything beyond as closing specific blind spots the table names.
Should the CD gate block deploys or just warn? Block on critical-severity findings (live secret-tier formats), warn otherwise. Blocking everything trains teams to override the gate; warning everything trains them to ignore it. Severity routing is what makes enforcement sustainable.
Where do AI-generated contributions fit? Everywhere above — agents produce commits like anyone else, minus the familiarity that motivates bypasses. The gates that matter most for agent-era codebases are the non-advisory ones: push protection and artifact-level checks, since neither depends on agent cooperation or local setup.
Layer your pipeline, then watch the surface it ships to: run KeyDrift's free scan — no signup — and put the post-deploy layer on today.