DevSecOps Operations

SLA-as-Code: Automating Security Deadlines in CI/CD

Move beyond code blocking in 2026. Learn how SLA-as-Code and InstaSLA automate security remediation deadlines across GitHub repos to fix existing debt

By InstaSLA Superadmin · Published · 13 min read

automated vulnerability remediationCI/CD security automationPolicy-as-Code DevSecOpsSLA-as-Codeautomate security compliance GitHubcontinuous compliance enforcementPolicy as Code remediationautomated security SLAsOPA Rego compliancesecurity policy enforcementautomated remediation deadlinesmanaging security technical debtshift-left security governanceopen policy agent SLAsDevSecOps compliance automationautomated fix deadlinesGitHub security automationcontinuous compliance pipelinesautomated vulnerability lifecycledeveloper security workflows
SLA as Code Automating Security Deadlines in CICD

SLA-as-Code: Automating Security Deadlines in CI/CD

Policy-as-code has become the standard way for platform teams to express security and compliance rules. The rules live in Git, get reviewed like any other change, are tested, and are evaluated automatically whenever something is provisioned, configured or deployed. That model is excellent at answering yes/no questions: is this S3 bucket encrypted, does this image come from an approved registry?

It is much weaker at the question every security team eventually runs into: what should happen when the answer is "no" and the business can't stop shipping?

The numbers suggest most organizations are losing this race. Verizon's 2026 DBIR found that exploitation of vulnerabilities is now the leading initial access vector, at 31% of breaches, overtaking credential abuse at 13%. The same report shows the median time to fully remediate a vulnerability from CISA's Known Exploited Vulnerabilities (KEV) catalog rising from 32 to 43 days, and only 26% of KEV entries were fully remediated in 2025, down from 38% the year before.

Blocking every build is not the answer, and neither is a warning nobody reads. The missing piece is a clock. This article covers how to turn static policy into automated, time-bound remediation deadlines in your CI/CD pipeline. We call the approach SLA-as-Code.

Where Policy-as-Code Succeeds, and Where It Stops

Policy-as-code means writing security, compliance and operational rules as version-controlled, testable code that a policy engine evaluates automatically. The best-known engine is Open Policy Agent (OPA) and its Rego language. OPA graduated in the CNCF in February 2021, and it remains a CNCF graduated project even after Apple hired several of its core maintainers from Styra in 2025. According to the project's FAQ, governance and licensing were unchanged by that move.

Typical policies are simple to state: every storage bucket must be encrypted, every container image must come from an approved registry. Violations are flagged or blocked automatically, which gives you consistent enforcement across environments.

Mature teams also know not to switch on blocking on day one. Gatekeeper, the OPA-based Kubernetes admission controller, has three enforcement actions: deny (the default) rejects the request, warn admits it but returns a warning to the client, and dryrun records violations without any admission feedback. Starting in dryrun or warn and moving to deny once the noise is understood is the recommended rollout path.

That works well for a new policy on a fresh cluster. It breaks down for existing debt:

  • If a scanner finds a medium-severity vulnerability in a legacy microservice, blocking the pipeline immediately can cause real business disruption.
  • If the pipeline only warns, developers under feature pressure will ignore it.
  • warn has no expiry date. Nothing in the policy says when a warning must become a failure, so warnings pile up into security debt.

The gap is not detection or enforcement. It is time.

Deadlines Are Getting Shorter and More Contextual

Regulators and standards bodies have converged on the same idea: a remediation deadline should be explicit, and increasingly it should depend on risk context rather than a severity label alone.

FrameworkWhat it setsWho it applies to
CISA BOD 26-04Remediation in 3, 14 or 60 days, or deferral to the next system upgrade, depending on four risk variablesFederal civilian agencies; FedRAMP providers are aligned to it
PCI DSS 6.3.3Critical and high-risk patches installed within one month of release; other patches on a timeline you defineEntities handling payment card data
EU Cyber Resilience Act, Article 14Early warning within 24 hours, notification within 72 hours, final report within 14 days of a fix, for actively exploited vulnerabilitiesManufacturers of products with digital elements sold in the EU
SOC 2No prescribed timeframe; you define your SLAs and must show you meet themService organizations audited against the Trust Services Criteria

CISA BOD 26-04: from flat deadlines to risk tiers

The most significant recent change is CISA's Binding Operational Directive 26-04, "Prioritizing Security Updates Based on Risk", issued on June 10, 2026. It supersedes and revokes both BOD 19-02 and BOD 22-01. Under BOD 22-01, every KEV entry got the same clock: 14 days for CVEs assigned after 2021, six months for older ones.

The new directive evaluates each asset-vulnerability pair on four variables: public exposure, KEV status, exploit automation and technical impact. Those inputs place the pair in one of five tiers. Depending on the combination, the deadline is 3, 14 or 60 calendar days, or the fix can wait for the next major upgrade. The three-day tier also requires forensic triage to check whether the asset was already compromised.

Two details are worth noting. First, the design is meant to concentrate effort: in CISA's pilot at one agency, roughly 60% of vulnerability instances qualified for deferral and about 1% needed action within three days, according to an IDC analyst quoted by FedTech. Second, the rollout is phased: agency processes were due by August 7, 2026, and the full remediation timelines take effect on December 7, 2026. FedRAMP has tied its new vulnerability detection and reporting rules to the same December 7 date.

You are probably not a federal agency, but the lesson transfers. A deadline is a function of exposure, exploitation evidence and impact. If your SLA policy can only express "high = 14 days", it cannot express what the most current federal guidance now does.

The EU Cyber Resilience Act: a regulatory clock that is already running

Since September 11, 2026, manufacturers must report actively exploited vulnerabilities and severe incidents: an early warning within 24 hours of becoming aware, a full notification within 72 hours, and a final report no later than 14 days after a corrective measure is available. Reports go through ENISA's Single Reporting Platform, which became operational on the same date. The remaining CRA requirements apply from December 11, 2027.

Compliance guidance stresses that the reporting duty covers products already on the EU market, not only new ones. This is a different clock from your internal remediation SLA, but the two connect: if you cannot quickly see which of your products contain a known-exploited component, and who owns it, you cannot meet a 24-hour window. Confirm with counsel how the rule applies to your products.

What "SLA-as-Code" Means

SLA-as-Code treats a remediation deadline as a versioned, machine-checkable part of your delivery system rather than a paragraph in a policy PDF. A finding is not just open or closed. It moves through a lifecycle, and the pipeline behaves differently at each stage:

StateTriggerPipeline behavior
DetectedAn alert opensThe clock starts and an owner is assigned
Within windowAge is under the SLADeploy proceeds
At riskTwo days or less remainDeploy proceeds with a warning; escalation notice sent
BreachedAge exceeds the SLA and a fix existsDeploy is blocked
Resolved or acceptedFixed, or covered by a documented exception with an expiry dateDeploy proceeds

The key move is that a warning is no longer permanent. It is a countdown that ends in a hard block unless someone acts. That turns a prioritization argument into a date everyone can see.

Traditional ticketing struggles here. Opening a Jira ticket for every violation and tracking due dates by hand does not scale, and stale tickets are exactly how compliance breaches happen. The deadline needs to attach automatically to the finding and to a named owner.

Architecture: Three Layers

A workable design separates three responsibilities:

  1. Detection. Your scanners and GitHub's own alerting produce findings (Dependabot alerts, code scanning, and so on).
  2. Ownership and deadlines. Something assigns each finding an owner, applies the SLA window, sends reminders and keeps the history. This is where a tool like InstaSLA fits. Its documented feature set covers syncing GitHub security alerts through a GitHub App, assigning them to a person, team or repository owner, severity-based deadlines with visible breach risk, fix campaigns that group duplicate advisories across repositories, email escalation policies, and exports of alert state, owner, due date, remediation history and accepted risk for audit purposes.
  3. Enforcement. The pipeline consults the SLA state and decides whether a deploy may proceed. Enforcement is not part of InstaSLA's documented feature set, so in the example below it is a small piece of policy-as-code that you own, using OPA and GitHub Actions.

Keeping enforcement separate has a practical benefit: the gate stays simple, auditable and portable, and the deadline data can be as rich as your ownership tooling allows.

Building It

1. Codify the SLA matrix

Start with a matrix you can defend to an auditor. InstaSLA's own quick start suggests critical in 3 days, high in 7, medium in 30 and low in 90 as a starting point to adjust for customer commitments and internal policy. Those numbers sit comfortably inside PCI DSS's one-month rule for critical and high patches.

Two refinements are worth making early:

  • Exploitation evidence should override severity. A medium-severity flaw on the KEV list, on an internet-facing service, deserves a shorter clock than a critical one buried in an internal tool. BOD 26-04 formalizes exactly this.
  • Store the matrix in Git. Changes to deadlines then go through review like any other change, which is also useful audit evidence.

2. Apply risk context from repository metadata

Not every repository carries the same risk. A payment gateway needs tighter windows than an internal admin tool used by five people. GitHub's custom properties let your organization attach metadata such as handles_pii or internet_facing to repositories, and the REST API exposes them so a pipeline can read them at run time.

3. Assign owners before deadlines start

A deadline without an owner is just noise. Map alerts to accountable teams by repository, path or package before turning on escalation, so the person who receives the reminder can actually ship the fix. InstaSLA's quick start describes repository, path, team and manual owner mappings for this purpose.

4. Write the policy

The policy below reads open Dependabot alerts, applies the matrix, halves every window for repositories flagged as handling personal data, and produces warn and deny messages. It only counts alerts that already have a patched version available, because the clock is about shipping an available fix. That mirrors PCI DSS, where the one-month countdown starts when a patch is released, not when the vulnerability is discovered.

package sla

import rego.v1

# Baseline remediation windows, in days. Adjust to your own policy.
base_windows := {"critical": 3, "high": 7, "medium": 30, "low": 90}

# Repos that handle PII get every window halved (rounded down).
windows := {s: floor(d / 2) | some s, d in base_windows} if input.repo.handles_pii

windows := base_windows if not input.repo.handles_pii

day_ns := 86400 * 1000000000

age_days(alert) := (time.parse_rfc3339_ns(input.now) - time.parse_rfc3339_ns(alert.created_at)) / day_ns

# Only alerts that have a fix available can breach: the clock is about
# shipping a fix, not about vulnerabilities nobody can patch yet.
open_alerts contains alert if {
  some alert in input.alerts
  alert.state == "open"
  alert.patched
}

# Expired SLA: the pipeline turns the warning into a hard block.
deny contains msg if {
  some alert in open_alerts
  age_days(alert) > windows[alert.severity]
  msg := sprintf(
    "SLA breached: %s alert #%d (%s) is %.1f days old, window is %d days",
    [alert.severity, alert.number, alert.package, age_days(alert), windows[alert.severity]],
  )
}

# Approaching expiry: surface a warning while there is still time to act.
warn contains msg if {
  some alert in open_alerts
  remaining := windows[alert.severity] - age_days(alert)
  remaining >= 0
  remaining <= 2
  msg := sprintf(
    "SLA at risk: %s alert #%d (%s) expires in %.1f days",
    [alert.severity, alert.number, alert.package, remaining],
  )
}

The current time is passed in as input.now rather than read inside the policy. That keeps the policy deterministic, so you can unit-test it with opa test against fixed dates.

With sample alerts, a 15-day-old high-severity alert against a 7-day window produces:

SLA breached: high alert #12 (lodash) is 15.2 days old, window is 7 days

A critical alert one day from expiry produces a warning instead of a failure:

SLA at risk: critical alert #13 (axios) expires in 1.0 days

5. Wire the gate into the deploy workflow

The gate is a job that fetches open alerts through the GitHub REST API, evaluates the policy and fails if anything is in the deny set. The deploy job depends on it.

name: deploy
on:
  push:
    branches: [main]

permissions:
  contents: read

jobs:
  sla-gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@<full-commit-sha> # pin to a full SHA
      - uses: open-policy-agent/setup-opa@<full-commit-sha> # pin to a full SHA

      - name: Build policy input from open Dependabot alerts
        env:
          # Token with "Dependabot alerts: read" (GitHub App or fine-grained PAT)
          GH_TOKEN: ${{ secrets.SLA_GATE_TOKEN }}
        run: |
          # Risk profile comes from a repository custom property
          PII=$(gh api "repos/${GITHUB_REPOSITORY}/properties/values" \
            --jq '[.[] | select(.property_name=="handles_pii") | .value] | first // "false"')

          gh api --paginate \
            "repos/${GITHUB_REPOSITORY}/dependabot/alerts?state=open&per_page=100" \
          | jq -s --arg now "$(date -u +%Y-%m-%dT%H:%M:%SZ)" --argjson pii "$PII" '
              {now: $now,
               repo: {handles_pii: $pii},
               alerts: [ .[][] | {number, state, created_at,
                                  severity: .security_advisory.severity,
                                  package: .dependency.package.name,
                                  patched: (.security_vulnerability.first_patched_version != null)} ]}' > input.json

      - name: Evaluate SLA policy
        run: |
          opa eval -d policy/sla.rego -i input.json 'data.sla' --format json > result.json
          jq -r '.result[0].expressions[0].value.warn[]? | "::warning::" + .' result.json
          jq -r '.result[0].expressions[0].value.deny[]? | "::error::" + .' result.json
          jq -e '(.result[0].expressions[0].value.deny // []) | length == 0' result.json > /dev/null

  deploy:
    needs: sla-gate
    runs-on: ubuntu-latest
    steps:
      - run: echo "Deploying..."

A few implementation notes:

6. Escalate before the deadline, not after

A gate that only fires at breach time feels like an ambush. Send a due-soon digest to owners and an overdue notice to the owner and their lead, and check delivery logs so you know reminders arrived. InstaSLA's quick start describes exactly this due-soon and overdue escalation pattern with delivery logs.

For backlog burn-down, GitHub's own security campaigns, generally available since April 2025, let security teams group alerts and set a fix timeframe, with Copilot Autofix suggesting fixes for supported code scanning alerts. Campaigns coordinate a burst of remediation work; SLAs make sure the backlog does not reform.

7. Capture the evidence

Auditors do not just want to see that vulnerabilities were detected. They sample records to check that the SLA you claim matches what actually happened. Keep, per finding, the owner, the due date, the completion date, and any accepted-risk record. If your ownership tooling exports these rows by date range, severity and SLA state, most of your evidence collection becomes a report rather than a scramble.

Rolling It Out Without a Revolt

Borrow the Gatekeeper playbook of moving from observation to enforcement:

  1. Observe first. Run the gate in report-only mode (for example, with continue-on-error: true on the evaluation step) and share the results. This is the equivalent of dryrun.
  2. Set a starting line for legacy debt. Give existing findings a documented burn-down schedule instead of an instant breach, so the first day of enforcement is not a wall of red.
  3. Enforce the top severities first. Turn on blocking for critical and high, then extend.
  4. Make exceptions expire. Risk acceptances should carry an owner, a reason and an end date. An exception with no expiry is just a warning by another name.

Pitfalls to Design For

  • Decide when the clock starts. From alert creation, advisory publication or patch release? Pick one, document it and be consistent. PCI DSS uses patch release.
  • Don't block on unfixable findings. If no patched version exists, the right response is mitigation or isolation, not a broken pipeline. BOD 26-04 itself defines remediation broadly: any action that eliminates the vulnerability, including isolation, mitigation or removal.
  • Choose fail-open or fail-closed on purpose. If the alerts API is unreachable, should the deploy proceed? There is no universal answer, but there must be a written one.
  • Check contract wording on business days. If customer commitments specify business days, your calendar logic must match.
  • Protect the gate itself. This one deserves its own section.

Your Security Tooling Is Part of the Attack Surface

The March 2026 Trivy incident is a useful reminder that the scanner in your pipeline runs with your pipeline's secrets. According to the official advisory, on March 19, 2026 an attacker using compromised credentials published a malicious Trivy release, force-pushed 76 of 77 version tags in aquasecurity/trivy-action to credential-stealing code, and replaced all seven tags in aquasecurity/setup-trivy. The exposure window for trivy-action was roughly 12 hours.

Analyses of the incident point to an earlier compromise in late February, where a pull_request_target misconfiguration led to a stolen token, and credential rotation afterward was incomplete. The follow-on attack then reused what survived. Tags can be moved after the fact, so any workflow that referenced a tag automatically ran whatever code the tag pointed to.

The defense is unglamorous: pin actions to full commit SHAs, because tags can be force-pushed and commit hashes cannot. Since August 2025 GitHub has let administrators enforce SHA pinning through the allowed-actions policy and block specific actions or versions outright. Apply the same rule to the SLA gate above. A gate that enforces your security deadlines is only trustworthy if its own dependencies are pinned and reviewed.

One more observation from that incident: it was a malicious-code compromise of trusted tooling, not a conventional vulnerability in your own code, and the exposure window was measured in hours. A remediation SLA measured in days works on a different timescale. SLA-as-Code manages the lifecycle of known findings; it complements supply chain hygiene rather than replacing it.

Conclusion

Policy-as-code answers "is this allowed?" SLA-as-Code answers "how long can this exception live?" Combining them closes the gap between hard blocks that stop delivery and warnings that never get read.

The building blocks are all available today:

  • Rego or a similar engine to express the rule and test it.
  • A GitHub-native ownership and deadline layer to assign owners, send reminders and produce evidence.
  • A small, pinned, reviewed gate in the deploy workflow that turns an expired deadline into a failed build.

The direction of travel in regulation is the same. Deadlines are getting shorter for the vulnerabilities that matter most, longer for the ones that don't, and more dependent on exposure and exploitation evidence. Teams that encode their deadlines in code can adapt by editing a file and opening a pull request.

References

Related articles