DevSecOps Operations

SLA-as-Code: Turning Policy-as-Code Failures into Enforceable Security Deadlines

Control the surge of AI-generated IaC misconfigurations in Terraform and K8s. Use InstaSLA to enforce strict pre-deployment SLAs for platform teams

By InstaSLA Superadmin · Published · 11 min read

platform engineering securityIaC security SLAsecure Terraform GitHubAI generated cloud misconfigurationsDevSecOps cloud securityInfrastructure as Code governanceplatform team SLA routingpre-deployment IaC securityKubernetes manifest misconfigurationsautomated IaC remediationCSPM misconfiguration managementAI coding assistant cloud riskTerraform security guardrailsShift Left IaC remediationcloud infrastructure SLAautomated Terraform fixesInfrastructure as Code complianceLLM generated Terraform bugsIaC shift-left enforcementblocking cloud misconfigurations
xcZxoy9E

SLA-as-Code: Turning Policy-as-Code Failures into Enforceable Security Deadlines

Policy-as-Code answers one question very well: is this change allowed right now? It is much less helpful with the question security and engineering teams argue about every week: this finding is real, but the release matters too, so by when must it be fixed, who owns it, and what happens if the date passes?

That gap is getting more expensive. Verizon's 2026 Data Breach Investigations Report found exploited vulnerabilities to be the most common initial access vector for the first time, accounting for 31% of breaches, up from 20% the year before. The same dataset shows the median time to fully remediate a CISA Known Exploited Vulnerability (KEV) rising from 32 to 43 days, and only 26% of KEV vulnerabilities fully remediated, down from 38%. Meanwhile, Mandiant's M-Trends 2026 estimates the mean time to exploit at roughly minus seven days, meaning exploitation often begins before a patch exists. That is an average pulled down by zero-days, but the direction is unmistakable: open-ended warnings do not keep up.

This article describes an approach we'll call SLA-as-Code: treating the remediation deadline as versioned, reviewable data that a policy engine evaluates in your pipeline, instead of a paragraph in a PDF. (The term is our shorthand for the practice, not an established industry standard.)

Where Policy-as-Code runs out of road

Policy as code means writing security, compliance and operational rules as version-controlled, testable code that is evaluated automatically when infrastructure is provisioned, configurations change or software is deployed. A rule like "every S3 bucket must be encrypted" or "images must come from approved registries" stops being a wiki page a developer has to interpret and becomes a check that either passes or fails.

The best-known general-purpose engine for this is Open Policy Agent (OPA), where policies are written in a language called Rego. A few facts worth knowing if you are betting a pipeline on it in 2026:

  • OPA is a graduated project of the Cloud Native Computing Foundation (CNCF), a status it has held since 2021.
  • OPA 1.0, released in December 2024, made the newer Rego syntax the default, so if and contains are now mandatory in rule definitions. Older tutorials that use the pre-1.0 style will not run unchanged.
  • In August 2025, OPA's core maintainers moved from Styra to Apple. The project stays under CNCF governance and the maintainer list is unchanged apart from employer, and several Styra tools are being brought into the community organization. It is still worth watching release cadence and maintainer diversity, as with any project that depends heavily on one sponsor.

It also helps to be precise about what OPA is. It does not scan your code. It takes structured input (say, the JSON output of a dependency scanner plus some repository metadata) and a policy, and returns a decision. Detection comes from tools like Dependabot, code scanning or Trivy; OPA decides what to do about the results.

Teams adopting this usually start in warning mode, surfacing violations without blocking, to build trust and tune rules before moving important ones to enforcement. That works for new code. The friction appears in the middle ground:

  • A medium-severity vulnerability in a legacy microservice appears. Blocking every pipeline immediately is a business disruption nobody signed up for.
  • If the pipeline only warns, the warning competes with feature deadlines, and it loses.

Without a date attached, warnings do not get resolved; they accumulate into security debt. What is missing is not another rule but a time dimension: a policy that says "this is allowed to exist until Tuesday the 14th."

Deadlines are becoming part of compliance itself

You are not the only one moving in this direction. Regulators and frameworks are increasingly specifying when, and increasingly making the answer depend on context rather than a single severity score.

SourceWhat it says about timing
CISA BOD 26-04 (issued June 10, 2026)Replaces the uniform KEV deadlines of BOD 22-01 (14 days for post-2021 CVEs) with a risk-tiered model built on four variables: public exposure, KEV status, exploit automation and technical impact. The tiers run from three days plus forensic triage for the highest-risk combination, through 14 and 60 days, to deferral until a future upgrade. Agency processes were due by August 7, 2026, and the full timelines apply from December 7, 2026. It binds federal civilian agencies directly; others can adopt it voluntarily.
FedRAMP VDR and VER rulesFedRAMP's Notice 0014 makes the new Vulnerability Detection and Response and Vulnerability Evaluation and Reporting rules mandatory by December 7, 2026, with a grace period to March 7, 2027. Timeframes depend on potential agency impact, internet reachability and likely exploitability, not just severity.
EU Cyber Resilience ActSince September 11, 2026, manufacturers of products with digital elements sold in the EU must file an early warning within 24 hours of becoming aware of an actively exploited vulnerability, a full notification within 72 hours, and a final report within 14 days of a corrective measure. The main obligations follow on December 11, 2027.
PCI DSS 4.0.x, Requirement 6.3.3Sets a one-month window from release for critical security patches. Published summaries disagree on whether v4.0.1 also covers high-severity patches (one says it was narrowed to critical, another still lists critical or high), so confirm the wording with your assessor.
NIST SSDF (SP 800-218)Its Respond to Vulnerabilities group covers assessing, prioritizing and remediating vulnerabilities. Draft Rev. 1 (SSDF 1.2) was published in December 2025.

SOC 2 and ISO 27001 are different: they do not prescribe a number of days. Auditors check whether you met the timelines in your own policy. Which means having a written SLA is only half the job; you also need to be able to prove you kept it.

One more data point on why context matters: in CISA's own pilot at one agency, roughly 60% of vulnerability instances qualified for deferral while only about 1% required action within three days, as reported by FedTech. A flat 30-day SLA over-serves the 60% and under-serves the 1%.

What SLA-as-Code means in practice

The idea is to give each concern in the lifecycle a clear owner, rather than asking one tool to do everything:

JobTypical tooling
Detect vulnerabilitiesDependabot, code scanning, Trivy, Snyk, and similar scanners
Decide whether a finding is within its allowed windowOPA/Rego (or any policy engine) reading scanner output plus repository context
Own, track and escalate each finding until it is fixedA GitHub-native SLA layer such as InstaSLA
Enforce at the point of changeGitHub Actions, required status checks and rulesets

On the third row, precision matters. InstaSLA installs as a GitHub App, syncs security alert metadata, lets you assign each alert to a person, team or repository owner, applies severity-based deadlines with breach status, groups duplicate advisories across repositories into fix campaigns, sends due-soon and overdue escalations, and exports remediation evidence including accepted risks. It is an ownership, deadline and evidence layer. It does not replace your scanners, and its public documentation does not claim that it blocks builds or deployments. The gate lives in your pipeline, which is exactly where a policy engine fits.

A blueprint

1. Write the SLA matrix down as data

Start with a small matrix and put it in Git, where it gets pull request review like any other policy. InstaSLA's quick-start suggests these starting values, to be adjusted for customer commitments and internal policy:

SeverityRemediation window
Critical3 days
High7 days
Medium30 days
Low90 days

Treat these as a baseline, not a standard. Your contracts, your regulators and your risk appetite should move them.

2. Let context move the clock

Not every repository carries the same risk. A public payment gateway deserves a tighter window than an internal tool used by five people. BOD 26-04 makes the same point at national scale: exposure and known exploitation, not the CVSS number alone, decide urgency.

In GitHub you can carry that context on the repository itself using custom properties, which are readable through the REST API. Define, for example, an internet_facing true/false property and a data_class single-select property at the organization level, then let the policy read them. You can layer in KEV membership as well, since CISA publishes the catalog in machine-readable form.

3. Enforce it in GitHub Actions

Below is a minimal policy. It halves the window for internet-facing or PII-handling repositories, warns when 80% of the window has elapsed, and fails when a deadline has passed. It uses current Rego syntax and was checked with OPA v1.20.2 against sample alert data; treat it as a starting point and test it against your own.

package sla

# Days allowed per severity. Starting points only: tune to customer contracts and regulators.
base_days := {
	"critical": 3,
	"high": 7,
	"medium": 30,
	"low": 90,
}

# Multiply every window by this factor. Internet-facing or PII-handling repos get half the time.
default factor := 1

factor := 0.5 if input.repo.internet_facing == true

factor := 0.5 if input.repo.data_class == "pii"

# An unrecognised severity is treated as medium rather than silently exempted.
allowed_days(alert) := ceil(object.get(base_days, alert.security_advisory.severity, 30) * factor)

age_days(alert) := floor((time.now_ns() - time.parse_rfc3339_ns(alert.created_at)) / 86400000000000)

open_alerts contains alert if {
	some alert in input.alerts
	alert.state == "open"
}

# Past the deadline: the pipeline fails on any entry here.
breached contains msg if {
	some alert in open_alerts
	age_days(alert) > allowed_days(alert)
	msg := sprintf(
		"SLA breached: alert #%d (%s, %s) is %d days old; the window is %d days",
		[alert.number, alert.security_advisory.severity, alert.security_advisory.ghsa_id, age_days(alert), allowed_days(alert)],
	)
}

# 80% of the window used: the pipeline warns but passes.
due_soon contains msg if {
	some alert in open_alerts
	age_days(alert) <= allowed_days(alert)
	age_days(alert) >= 0.8 * allowed_days(alert)
	msg := sprintf(
		"Due soon: alert #%d (%s, %s) is %d of %d days old",
		[alert.number, alert.security_advisory.severity, alert.security_advisory.ghsa_id, age_days(alert), allowed_days(alert)],
	)
}

And a workflow that feeds it the repository's open Dependabot alerts:

name: sla-gate
on:
  pull_request:
  push:
    branches: [main]
  schedule:
    - cron: "0 6 * * *"   # keeps the clock honest even when nobody pushes

permissions:
  contents: read

jobs:
  sla-gate:
    runs-on: ubuntu-latest
    steps:
      # Pin third-party actions to full commit SHAs in production.
      - uses: actions/checkout@v4
      - uses: open-policy-agent/setup-opa@v2
        with:
          version: latest
      - name: Collect open alerts and repository context
        shell: bash          # explicit bash enables pipefail, so a failed API call fails the job
        env:
          GH_TOKEN: ${{ secrets.SLA_GATE_TOKEN }}
          REPO: ${{ github.repository }}
        run: |
          gh api --paginate "repos/$REPO/dependabot/alerts?state=open&per_page=100" | jq -s 'add // []' > alerts.json
          gh api "repos/$REPO/properties/values" > props.json
          jq -n --slurpfile a alerts.json --slurpfile p props.json '
            { alerts: $a[0],
              repo: (($p[0] | map({(.property_name): .value}) | add // {})
                     | .internet_facing = (.internet_facing == "true")) }' > input.json
      - name: Evaluate SLA policy
        shell: bash
        run: |
          opa eval -d policy/ -i input.json -f json 'data.sla' > result.json
          jq -r '.result[0].expressions[0].value.due_soon[]? | "::warning::" + .' result.json
          jq -r '.result[0].expressions[0].value.breached[]? | "::error::" + .' result.json
          jq -e '.result[0].expressions[0].value.breached | length == 0' result.json > /dev/null

Two details trip people up. First, the built-in GITHUB_TOKEN has not been able to list Dependabot alerts, a long-standing limitation, so check the current docs and expect SLA_GATE_TOKEN here to be a GitHub App installation token or a fine-grained token with the "Dependabot alerts" read permission, as the REST docs specify. Second, make your deploy job depend on this one (needs: sla-gate) or mark it as a required status check, otherwise it is only advice.

This gate deliberately governs the existing backlog through the clock. For vulnerabilities a pull request newly introduces, pair it with a diff-based check such as GitHub's dependency review action, which can fail a PR outright at a severity you choose.

4. Route to owners and escalate before the breach

A deadline nobody owns is just a number. Ownership should be resolved before escalation starts: by repository, by path or package where a monorepo makes the repository owner the wrong answer, or manually as a fallback. InstaSLA supports repository, path, team and manual owner mappings and email escalation policies, for example a due-soon digest to owners and an overdue notice to the owner and their lead, with delivery logs so you can check that a reminder was actually sent.

Duplicates are the other tax on remediation. One vulnerable package can raise the same advisory in dozens of repositories, and grouping them into a campaign turns dozens of tickets into one piece of coordinated work. GitHub also has its own security campaigns with due dates and Copilot Autofix, which cover code scanning and secret scanning alerts; they are a good complement for those alert types.

5. Keep the evidence

Every step above produces a record: what was found, who owned it, what the deadline was, when it closed, and whether a risk was formally accepted. Export that history for auditors and customer security reviews. It is the kind of record SOC 2 and ISO 27001 audits ask for, and it supports the SSDF's Respond to Vulnerabilities practices for assessing, prioritizing and remediating findings.

Design decisions that decide whether this works

Let the fix ship. If an expired SLA blocks every deployment, it also blocks the deployment that resolves the breach. Build an explicit path for remediation changes, such as a protected security-fix label or a rule that permits pull requests touching only dependency manifests and lockfiles. This is a design choice, not a standard, but without it the gate can deadlock a team during an incident.

Decide when the clock starts. GitHub's alert creation time is convenient, but a fix may not exist yet. PCI DSS counts from patch release. Choose one definition per severity, write it down, and consider a separate mitigation deadline for vulnerabilities with no available fix.

Give exceptions an expiry. Risk acceptance is legitimate, but an exception with no end date is just a permanent ignore. Each one needs an owner, a reason and an expiry after which the finding returns to the queue.

Treat the SLA as a ceiling, not a target. One analysis of BOD 26-04 argues that its three-day window should be read as a maximum, not an operational goal for the highest-risk cases. If a vulnerability is KEV-listed and internet-facing, mitigation, such as taking the interface off the public internet, should not wait for a deadline.

Fail closed on missing data. The sample policy treats an unknown severity as medium, and the workflow fails if the API call fails. A gate that silently passes when its data source is unavailable is worse than no gate, because it produces false confidence.

Keep the two definitions aligned. If the pipeline gate and your SLA tracker each hold a copy of the matrix, they will eventually drift. Review both in the same change, and on a schedule.

Where to start

  1. Agree on a four-row severity matrix and the clock-start rule, and put it in Git.
  2. Turn on ownership and deadline tracking for your alerts so every open finding has an owner and a due date.
  3. Add the gate in warning-only mode for a few weeks, watch how many findings are already past due, then flip it to enforcing.

Conclusion

Policy-as-Code was built to decide whether something is allowed now. Security debt needs a second answer: until when. Regulators are converging on the same idea, moving from single flat deadlines toward time limits that depend on exposure, exploitation and impact, and the exploitation data says the old comfort of 30-day windows is fading.

SLA-as-Code is the engineering response: deadlines as versioned data, a policy engine that evaluates them where code ships, and an ownership and evidence layer that keeps every finding attached to a person and a date. It will not remove the debt. It makes the debt visible, time-bound and provable, which is what turns a warning into a commitment.From SECURITY.md to Merge Gates: A 2026 Guide to Enforcing GitHub Vulnerability SLAs

Related articles

Managing Remediation SLAs for AI-Generated IaC | InstaSLA | InstaSLA