DevSecOps Operations
SLA-as-Code: Automating Security Deadlines in CI/CD
Move beyond code blocking in 2026. Learn how SLA-as-Code and InstaSLA automate security remediation deadlines across GitHub repos to fix existing debt
By InstaSLA Superadmin · Published · 13 min read

SLA-as-Code: Automating Security Deadlines in CI/CD
Policy-as-code has become the standard way for platform teams to express security and compliance rules. The rules live in Git, get reviewed like any other change, are tested, and are evaluated automatically whenever something is provisioned, configured or deployed. That model is excellent at answering yes/no questions: is this S3 bucket encrypted, does this image come from an approved registry?
It is much weaker at the question every security team eventually runs into: what should happen when the answer is "no" and the business can't stop shipping?
The numbers suggest most organizations are losing this race. Verizon's 2026 DBIR found that exploitation of vulnerabilities is now the leading initial access vector, at 31% of breaches, overtaking credential abuse at 13%. The same report shows the median time to fully remediate a vulnerability from CISA's Known Exploited Vulnerabilities (KEV) catalog rising from 32 to 43 days, and only 26% of KEV entries were fully remediated in 2025, down from 38% the year before.
Blocking every build is not the answer, and neither is a warning nobody reads. The missing piece is a clock. This article covers how to turn static policy into automated, time-bound remediation deadlines in your CI/CD pipeline. We call the approach SLA-as-Code.
Where Policy-as-Code Succeeds, and Where It Stops
Policy-as-code means writing security, compliance and operational rules as version-controlled, testable code that a policy engine evaluates automatically. The best-known engine is Open Policy Agent (OPA) and its Rego language. OPA graduated in the CNCF in February 2021, and it remains a CNCF graduated project even after Apple hired several of its core maintainers from Styra in 2025. According to the project's FAQ, governance and licensing were unchanged by that move.
Typical policies are simple to state: every storage bucket must be encrypted, every container image must come from an approved registry. Violations are flagged or blocked automatically, which gives you consistent enforcement across environments.
Mature teams also know not to switch on blocking on day one. Gatekeeper, the OPA-based Kubernetes admission controller, has three enforcement actions: deny (the default) rejects the request, warn admits it but returns a warning to the client, and dryrun records violations without any admission feedback. Starting in dryrun or warn and moving to deny once the noise is understood is the recommended rollout path.
That works well for a new policy on a fresh cluster. It breaks down for existing debt:
- If a scanner finds a medium-severity vulnerability in a legacy microservice, blocking the pipeline immediately can cause real business disruption.
- If the pipeline only warns, developers under feature pressure will ignore it.
warnhas no expiry date. Nothing in the policy says when a warning must become a failure, so warnings pile up into security debt.
The gap is not detection or enforcement. It is time.
Deadlines Are Getting Shorter and More Contextual
Regulators and standards bodies have converged on the same idea: a remediation deadline should be explicit, and increasingly it should depend on risk context rather than a severity label alone.
| Framework | What it sets | Who it applies to |
|---|---|---|
| CISA BOD 26-04 | Remediation in 3, 14 or 60 days, or deferral to the next system upgrade, depending on four risk variables | Federal civilian agencies; FedRAMP providers are aligned to it |
| PCI DSS 6.3.3 | Critical and high-risk patches installed within one month of release; other patches on a timeline you define | Entities handling payment card data |
| EU Cyber Resilience Act, Article 14 | Early warning within 24 hours, notification within 72 hours, final report within 14 days of a fix, for actively exploited vulnerabilities | Manufacturers of products with digital elements sold in the EU |
| SOC 2 | No prescribed timeframe; you define your SLAs and must show you meet them | Service organizations audited against the Trust Services Criteria |
CISA BOD 26-04: from flat deadlines to risk tiers
The most significant recent change is CISA's Binding Operational Directive 26-04, "Prioritizing Security Updates Based on Risk", issued on June 10, 2026. It supersedes and revokes both BOD 19-02 and BOD 22-01. Under BOD 22-01, every KEV entry got the same clock: 14 days for CVEs assigned after 2021, six months for older ones.
The new directive evaluates each asset-vulnerability pair on four variables: public exposure, KEV status, exploit automation and technical impact. Those inputs place the pair in one of five tiers. Depending on the combination, the deadline is 3, 14 or 60 calendar days, or the fix can wait for the next major upgrade. The three-day tier also requires forensic triage to check whether the asset was already compromised.
Two details are worth noting. First, the design is meant to concentrate effort: in CISA's pilot at one agency, roughly 60% of vulnerability instances qualified for deferral and about 1% needed action within three days, according to an IDC analyst quoted by FedTech. Second, the rollout is phased: agency processes were due by August 7, 2026, and the full remediation timelines take effect on December 7, 2026. FedRAMP has tied its new vulnerability detection and reporting rules to the same December 7 date.
You are probably not a federal agency, but the lesson transfers. A deadline is a function of exposure, exploitation evidence and impact. If your SLA policy can only express "high = 14 days", it cannot express what the most current federal guidance now does.
The EU Cyber Resilience Act: a regulatory clock that is already running
Since September 11, 2026, manufacturers must report actively exploited vulnerabilities and severe incidents: an early warning within 24 hours of becoming aware, a full notification within 72 hours, and a final report no later than 14 days after a corrective measure is available. Reports go through ENISA's Single Reporting Platform, which became operational on the same date. The remaining CRA requirements apply from December 11, 2027.
Compliance guidance stresses that the reporting duty covers products already on the EU market, not only new ones. This is a different clock from your internal remediation SLA, but the two connect: if you cannot quickly see which of your products contain a known-exploited component, and who owns it, you cannot meet a 24-hour window. Confirm with counsel how the rule applies to your products.
What "SLA-as-Code" Means
SLA-as-Code treats a remediation deadline as a versioned, machine-checkable part of your delivery system rather than a paragraph in a policy PDF. A finding is not just open or closed. It moves through a lifecycle, and the pipeline behaves differently at each stage:
| State | Trigger | Pipeline behavior |
|---|---|---|
| Detected | An alert opens | The clock starts and an owner is assigned |
| Within window | Age is under the SLA | Deploy proceeds |
| At risk | Two days or less remain | Deploy proceeds with a warning; escalation notice sent |
| Breached | Age exceeds the SLA and a fix exists | Deploy is blocked |
| Resolved or accepted | Fixed, or covered by a documented exception with an expiry date | Deploy proceeds |
The key move is that a warning is no longer permanent. It is a countdown that ends in a hard block unless someone acts. That turns a prioritization argument into a date everyone can see.
Traditional ticketing struggles here. Opening a Jira ticket for every violation and tracking due dates by hand does not scale, and stale tickets are exactly how compliance breaches happen. The deadline needs to attach automatically to the finding and to a named owner.
Architecture: Three Layers
A workable design separates three responsibilities:
- Detection. Your scanners and GitHub's own alerting produce findings (Dependabot alerts, code scanning, and so on).
- Ownership and deadlines. Something assigns each finding an owner, applies the SLA window, sends reminders and keeps the history. This is where a tool like InstaSLA fits. Its documented feature set covers syncing GitHub security alerts through a GitHub App, assigning them to a person, team or repository owner, severity-based deadlines with visible breach risk, fix campaigns that group duplicate advisories across repositories, email escalation policies, and exports of alert state, owner, due date, remediation history and accepted risk for audit purposes.
- Enforcement. The pipeline consults the SLA state and decides whether a deploy may proceed. Enforcement is not part of InstaSLA's documented feature set, so in the example below it is a small piece of policy-as-code that you own, using OPA and GitHub Actions.
Keeping enforcement separate has a practical benefit: the gate stays simple, auditable and portable, and the deadline data can be as rich as your ownership tooling allows.
Building It
1. Codify the SLA matrix
Start with a matrix you can defend to an auditor. InstaSLA's own quick start suggests critical in 3 days, high in 7, medium in 30 and low in 90 as a starting point to adjust for customer commitments and internal policy. Those numbers sit comfortably inside PCI DSS's one-month rule for critical and high patches.
Two refinements are worth making early:
- Exploitation evidence should override severity. A medium-severity flaw on the KEV list, on an internet-facing service, deserves a shorter clock than a critical one buried in an internal tool. BOD 26-04 formalizes exactly this.
- Store the matrix in Git. Changes to deadlines then go through review like any other change, which is also useful audit evidence.
2. Apply risk context from repository metadata
Not every repository carries the same risk. A payment gateway needs tighter windows than an internal admin tool used by five people. GitHub's custom properties let your organization attach metadata such as handles_pii or internet_facing to repositories, and the REST API exposes them so a pipeline can read them at run time.
3. Assign owners before deadlines start
A deadline without an owner is just noise. Map alerts to accountable teams by repository, path or package before turning on escalation, so the person who receives the reminder can actually ship the fix. InstaSLA's quick start describes repository, path, team and manual owner mappings for this purpose.
4. Write the policy
The policy below reads open Dependabot alerts, applies the matrix, halves every window for repositories flagged as handling personal data, and produces warn and deny messages. It only counts alerts that already have a patched version available, because the clock is about shipping an available fix. That mirrors PCI DSS, where the one-month countdown starts when a patch is released, not when the vulnerability is discovered.
package sla
import rego.v1
# Baseline remediation windows, in days. Adjust to your own policy.
base_windows := {"critical": 3, "high": 7, "medium": 30, "low": 90}
# Repos that handle PII get every window halved (rounded down).
windows := {s: floor(d / 2) | some s, d in base_windows} if input.repo.handles_pii
windows := base_windows if not input.repo.handles_pii
day_ns := 86400 * 1000000000
age_days(alert) := (time.parse_rfc3339_ns(input.now) - time.parse_rfc3339_ns(alert.created_at)) / day_ns
# Only alerts that have a fix available can breach: the clock is about
# shipping a fix, not about vulnerabilities nobody can patch yet.
open_alerts contains alert if {
some alert in input.alerts
alert.state == "open"
alert.patched
}
# Expired SLA: the pipeline turns the warning into a hard block.
deny contains msg if {
some alert in open_alerts
age_days(alert) > windows[alert.severity]
msg := sprintf(
"SLA breached: %s alert #%d (%s) is %.1f days old, window is %d days",
[alert.severity, alert.number, alert.package, age_days(alert), windows[alert.severity]],
)
}
# Approaching expiry: surface a warning while there is still time to act.
warn contains msg if {
some alert in open_alerts
remaining := windows[alert.severity] - age_days(alert)
remaining >= 0
remaining <= 2
msg := sprintf(
"SLA at risk: %s alert #%d (%s) expires in %.1f days",
[alert.severity, alert.number, alert.package, remaining],
)
}The current time is passed in as input.now rather than read inside the policy. That keeps the policy deterministic, so you can unit-test it with opa test against fixed dates.
With sample alerts, a 15-day-old high-severity alert against a 7-day window produces:
SLA breached: high alert #12 (lodash) is 15.2 days old, window is 7 daysA critical alert one day from expiry produces a warning instead of a failure:
SLA at risk: critical alert #13 (axios) expires in 1.0 days5. Wire the gate into the deploy workflow
The gate is a job that fetches open alerts through the GitHub REST API, evaluates the policy and fails if anything is in the deny set. The deploy job depends on it.
name: deploy
on:
push:
branches: [main]
permissions:
contents: read
jobs:
sla-gate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@<full-commit-sha> # pin to a full SHA
- uses: open-policy-agent/setup-opa@<full-commit-sha> # pin to a full SHA
- name: Build policy input from open Dependabot alerts
env:
# Token with "Dependabot alerts: read" (GitHub App or fine-grained PAT)
GH_TOKEN: ${{ secrets.SLA_GATE_TOKEN }}
run: |
# Risk profile comes from a repository custom property
PII=$(gh api "repos/${GITHUB_REPOSITORY}/properties/values" \
--jq '[.[] | select(.property_name=="handles_pii") | .value] | first // "false"')
gh api --paginate \
"repos/${GITHUB_REPOSITORY}/dependabot/alerts?state=open&per_page=100" \
| jq -s --arg now "$(date -u +%Y-%m-%dT%H:%M:%SZ)" --argjson pii "$PII" '
{now: $now,
repo: {handles_pii: $pii},
alerts: [ .[][] | {number, state, created_at,
severity: .security_advisory.severity,
package: .dependency.package.name,
patched: (.security_vulnerability.first_patched_version != null)} ]}' > input.json
- name: Evaluate SLA policy
run: |
opa eval -d policy/sla.rego -i input.json 'data.sla' --format json > result.json
jq -r '.result[0].expressions[0].value.warn[]? | "::warning::" + .' result.json
jq -r '.result[0].expressions[0].value.deny[]? | "::error::" + .' result.json
jq -e '(.result[0].expressions[0].value.deny // []) | length == 0' result.json > /dev/null
deploy:
needs: sla-gate
runs-on: ubuntu-latest
steps:
- run: echo "Deploying..."A few implementation notes:
- The Dependabot alerts endpoint needs a token with read access to Dependabot alerts, so use a GitHub App installation token or a fine-grained personal access token stored as a secret.
- Because of that, put the gate on
pushor deploy workflows rather than fork pull requests. Workflows triggered by Dependabot onpull_requestrun with a read-only token and cannot read user-defined secrets. - Replace the
<full-commit-sha>placeholders with real commit hashes. The reason is covered below.
6. Escalate before the deadline, not after
A gate that only fires at breach time feels like an ambush. Send a due-soon digest to owners and an overdue notice to the owner and their lead, and check delivery logs so you know reminders arrived. InstaSLA's quick start describes exactly this due-soon and overdue escalation pattern with delivery logs.
For backlog burn-down, GitHub's own security campaigns, generally available since April 2025, let security teams group alerts and set a fix timeframe, with Copilot Autofix suggesting fixes for supported code scanning alerts. Campaigns coordinate a burst of remediation work; SLAs make sure the backlog does not reform.
7. Capture the evidence
Auditors do not just want to see that vulnerabilities were detected. They sample records to check that the SLA you claim matches what actually happened. Keep, per finding, the owner, the due date, the completion date, and any accepted-risk record. If your ownership tooling exports these rows by date range, severity and SLA state, most of your evidence collection becomes a report rather than a scramble.
Rolling It Out Without a Revolt
Borrow the Gatekeeper playbook of moving from observation to enforcement:
- Observe first. Run the gate in report-only mode (for example, with
continue-on-error: trueon the evaluation step) and share the results. This is the equivalent ofdryrun. - Set a starting line for legacy debt. Give existing findings a documented burn-down schedule instead of an instant breach, so the first day of enforcement is not a wall of red.
- Enforce the top severities first. Turn on blocking for critical and high, then extend.
- Make exceptions expire. Risk acceptances should carry an owner, a reason and an end date. An exception with no expiry is just a warning by another name.
Pitfalls to Design For
- Decide when the clock starts. From alert creation, advisory publication or patch release? Pick one, document it and be consistent. PCI DSS uses patch release.
- Don't block on unfixable findings. If no patched version exists, the right response is mitigation or isolation, not a broken pipeline. BOD 26-04 itself defines remediation broadly: any action that eliminates the vulnerability, including isolation, mitigation or removal.
- Choose fail-open or fail-closed on purpose. If the alerts API is unreachable, should the deploy proceed? There is no universal answer, but there must be a written one.
- Check contract wording on business days. If customer commitments specify business days, your calendar logic must match.
- Protect the gate itself. This one deserves its own section.
Your Security Tooling Is Part of the Attack Surface
The March 2026 Trivy incident is a useful reminder that the scanner in your pipeline runs with your pipeline's secrets. According to the official advisory, on March 19, 2026 an attacker using compromised credentials published a malicious Trivy release, force-pushed 76 of 77 version tags in aquasecurity/trivy-action to credential-stealing code, and replaced all seven tags in aquasecurity/setup-trivy. The exposure window for trivy-action was roughly 12 hours.
Analyses of the incident point to an earlier compromise in late February, where a pull_request_target misconfiguration led to a stolen token, and credential rotation afterward was incomplete. The follow-on attack then reused what survived. Tags can be moved after the fact, so any workflow that referenced a tag automatically ran whatever code the tag pointed to.
The defense is unglamorous: pin actions to full commit SHAs, because tags can be force-pushed and commit hashes cannot. Since August 2025 GitHub has let administrators enforce SHA pinning through the allowed-actions policy and block specific actions or versions outright. Apply the same rule to the SLA gate above. A gate that enforces your security deadlines is only trustworthy if its own dependencies are pinned and reviewed.
One more observation from that incident: it was a malicious-code compromise of trusted tooling, not a conventional vulnerability in your own code, and the exposure window was measured in hours. A remediation SLA measured in days works on a different timescale. SLA-as-Code manages the lifecycle of known findings; it complements supply chain hygiene rather than replacing it.
Conclusion
Policy-as-code answers "is this allowed?" SLA-as-Code answers "how long can this exception live?" Combining them closes the gap between hard blocks that stop delivery and warnings that never get read.
The building blocks are all available today:
- Rego or a similar engine to express the rule and test it.
- A GitHub-native ownership and deadline layer to assign owners, send reminders and produce evidence.
- A small, pinned, reviewed gate in the deploy workflow that turns an expired deadline into a failed build.
The direction of travel in regulation is the same. Deadlines are getting shorter for the vulnerabilities that matter most, longer for the ones that don't, and more dependent on exposure and exploitation evidence. Teams that encode their deadlines in code can adapt by editing a file and opening a pull request.
References
- Verizon DBIR 2026 coverage: TechRepublic, Help Net Security
- CISA BOD 26-04: directive, implementation guidance, FedTech Magazine, Action1, Qualys
- FedRAMP: Public Notice 0014
- EU Cyber Resilience Act: European Commission reporting page, Element
- PCI DSS 6.3.3: TrustedSec
- SOC 2 remediation SLAs: Pixee
- Open Policy Agent: CNCF graduation announcement, OPA project FAQ on the Apple move
- Gatekeeper enforcement actions: OneUptime, Google Cloud Policy Controller docs
- Trivy supply chain incident: GitHub advisory GHSA-69fq-xp46-6x23, ARMO, Sysdig
- GitHub: SHA pinning enforcement, security campaigns GA, Dependabot alerts REST API, repository custom properties API
- InstaSLA: features, quick startCISA BOD 26-04: How to Meet the New 3-Day Remediation SLA in GitHub