To build a vulnerability management program, define its scope and owners, get credentialed coverage of every asset class, score findings by exploitability, exposure and business impact, fix them against agreed SLAs, verify each fix, and report a small set of metrics that show risk going down. Vulnerability management is the continuous cycle of finding, prioritizing, fixing and verifying weaknesses in software, configurations and identities. A program turns that cycle into an operating model. It has a charter, a RACI, a scoring rule, an exception process and dashboards that both engineering and leadership trust. It covers servers, endpoints, cloud workloads, containers, applications and identities, not just Patch Tuesday.
Key takeaways
- ✓A vulnerability management program is an operating model, not a scanner, and it needs scope, owners, SLAs, an exception process and a verification loop.
- ✓Asset inventory and ownership come before scanning, because a finding with no owner never gets fixed.
- ✓Prioritization should combine exploit intelligence such as CISA KEV and EPSS with internet exposure, runtime reachability, asset criticality and compensating controls, rather than relying on CVSS alone.
- ✓The metrics that matter most are exposure reduction and SLA attainment by team, not the raw count of open findings.
- ✓Every closed finding needs rescan or runtime evidence, and every exception needs an approver and an expiry date.
What a vulnerability management program is, and what it needs first
A vulnerability management program is the documented, repeatable way an organization discovers, prioritizes, remediates and verifies security weaknesses. It only works once assets, owners and scanner access are in place. NIST frames vulnerability management as a capability to manage risk created by defects present in software on the network. A program adds the people, policy and measurement around that capability.
| A vulnerability management program is | It is not |
|---|---|
| A continuous cycle with defined cadences and SLAs | A quarterly scan followed by a PDF report |
| Risk-based, with priority driven by exploitability, exposure and business impact | Sorting by CVSS and working top-down |
| Owned jointly: security sets policy and asset owners fix | The security team patching everything itself |
| Measured by exposure reduction and verified fixes | Measured by the number of findings closed |
| Inclusive of cloud, containers, apps, SaaS and identities | Limited to Windows and Linux servers |
Prerequisites checklist
- ✓Asset inventory: one source of truth for servers, endpoints, network devices, cloud accounts, clusters, images, repos and SaaS tenants. Reconcile it at least weekly against cloud provider APIs and DHCP or EDR data.
- ✓Ownership model: every asset carries an owner tag (a team, not a person) and an escalation contact. Unowned assets go to a default queue that the VM lead reviews weekly.
- ✓System classification: every asset has a criticality tier and a data sensitivity label, using the schema below.
- ✓Exposure visibility: a map of what is internet-facing, what sits behind VPN or partner links, and how network segments connect.
- ✓Scanner credentials: service accounts or agents for authenticated scanning, stored in a vault such as HashiCorp Vault or AWS Secrets Manager and rotated on a schedule.
- ✓CMDB quality: owner and tier fields are populated for at least Tier 0 and Tier 1 assets before tickets route automatically.
- ✓Software composition: SBOMs in CycloneDX or SPDX format for in-house applications and images, so third-party library flaws map to real services.
Asset criticality schema
| Tier | Definition | Examples | Tag |
|---|---|---|---|
| Tier 0 | Compromise gives control of the estate | Active Directory, Entra ID, IAM admin roles, CI/CD runners, secrets stores | criticality=t0 |
| Tier 1 | Revenue or regulated data | Payments API, customer database, production Kubernetes clusters | criticality=t1 |
| Tier 2 | Internal business systems | HR apps, internal wikis, staging environments | criticality=t2 |
| Tier 3 | Low impact, easily rebuilt | Dev sandboxes, lab hosts, kiosks | criticality=t3 |
Coverage beyond servers
Each asset class across the hybrid estate needs its own detection method:
| Asset class | Detection method |
|---|---|
| Endpoints and servers | Agent-based or credentialed scans for OS and application CVEs |
| Network devices | Firmware versions via authenticated SNMPv3 or SSH checks |
| Cloud workloads and services | CSPM checks for misconfigurations such as public storage buckets and open security groups |
| Containers and Kubernetes | Image scanning in the registry and at runtime, plus cluster configuration checks |
| Web apps and APIs | SAST, DAST and software composition analysis for vulnerable libraries |
| Identities | Unused admin roles, wildcard IAM policies, stale service account keys and weak AD configurations |
| Third-party and SaaS | Vendor advisories and SaaS tenant configuration reviews |
How to build a vulnerability management program step by step
You build the program in six steps: charter, coverage, intake, prioritization, remediation and validation, each with a defined cadence, output and owner.
- Write the charter and define scope. Draft a two-page charter covering objectives, in-scope asset classes, SLAs by priority, roles and the exception process. Build it from the asset inventory, regulatory requirements (PCI DSS, SOC 2, ISO 27001) and risk appetite. The CISO and the heads of infrastructure and engineering sign it, and you review it annually.
- Get discovery and scan coverage. Deploy agents, credentialed scanners, cloud API connectors and image scanning, and track coverage by asset type rather than as one number. Scan cloud and containers continuously, servers at least weekly, and internet-facing assets externally every day. The output is a findings feed tied to asset IDs.
- Set up intake and triage. Normalize findings from every source, deduplicate on asset ID plus CVE, enrich them with owner and tier from the CMDB and with threat intel (CISA KEV, FIRST EPSS, NVD), and drop findings on decommissioned assets. Automate this daily, with a 30-minute human triage of new P1 candidates.
- Prioritize with a scoring model. Score every finding with the model below.
- Remediate against SLAs. Auto-create tickets in the owning team’s queue with the priority, SLA due date and fix guidance. Check fixes against the change calendar and review SLAs weekly.
- Validate, close and report. Close a ticket only after a rescan with the same authenticated method, or a runtime check, confirms the fix. Trigger the rescan automatically within 24 hours of a fix being marked done, and attach evidence: fixed package version, image digest, scan ID and ticket link.
| Step | Success criterion | Roles |
|---|---|---|
| Charter | Every in-scope asset class has a named owning team | CISO (accountable), VM lead (author), infrastructure and engineering leads (consulted) |
| Coverage | 100% of Tier 0 and Tier 1 assets have authenticated, agent or sensor coverage | VM team, IT ops (agent rollout), cloud platform team (connectors) |
| Intake | Fewer than 5% of findings lack an owner after enrichment | VM analyst (triage), CMDB owner (data fixes) |
| Remediation | 90%+ of P1 and P2 findings closed within SLA | Asset owners (accountable), IT ops and platform engineers (responsible), VM team (tracking) |
| Validation | Reopen rate under 5% | VM team (validation), owners (evidence) |
Unauthenticated scans only see what a port banner reveals, so they miss most local package and library flaws. If you are still choosing scanners, compare cloud, container, and network scanners before you commit to one model.
A sample 100-point scoring model
Risk-based vulnerability management replaces severity-only sorting with a score that reflects how likely and how damaging exploitation would be:
| Factor | Points | Scoring rule |
|---|---|---|
| Exploit intelligence | 0–30 | Listed in CISA KEV = 30; EPSS ≥ 0.1 or weaponized exploit = 20; EPSS 0.01–0.1 = 10; none = 0 |
| Internet exposure | 0–20 | Directly internet-facing = 20; reachable via VPN or partner link = 10; internal isolated = 0 |
| Runtime reachability | 0–15 | Vulnerable package loaded and executing = 15; unknown = 10; installed but never loaded = 3 |
| Asset criticality | 0–20 | Tier 0 = 20; Tier 1 = 15; Tier 2 = 8; Tier 3 = 3 |
| Privilege and blast radius | 0–15 | Runs as root or holds admin or wildcard IAM = 15; access to sensitive data = 10; limited = 3 |
| Compensating controls | −15 to 0 | Tested WAF rule or virtual patch = −10; verified segmentation = −5 |
Bands: P1 = 80+, P2 = 60–79, P3 = 40–59, P4 = below 40. Use the CVSS base score only as a tiebreaker inside a band and as a floor: a CVSS 9.0+ finding on a Tier 0 asset never drops below P2.
SLAs and remediation options
| Priority | Typical profile | SLA | Action |
|---|---|---|---|
| P1 | KEV-listed, internet-facing, Tier 0–1 | 72 hours to 7 days | Emergency change or same-day mitigation; VM lead tracks daily |
| P2 | High exploit probability, internal or Tier 2 | 30 days | Next patch or release cycle |
| P3 | Moderate risk, limited exposure | 90 days | Standard maintenance window |
| P4 | Not loaded, isolated or Tier 3 | 180 days or next rebuild | Quarterly backlog review |
The asset owner owns every band. Remediation options, in order of preference:
- Patch or upgrade the package or OS.
- Rebuild the container image from a patched base and redeploy, rather than patching running containers.
- Fix the Terraform or CloudFormation source via pull request so drift does not return.
- Remove the unused package, service or IAM permission.
- Apply a compensating control (WAF rule, segmentation, feature flag) while a fix is scheduled.
- Decommission the asset.
Pre-remediation checklist: confirm the owner, test the fix in staging, write a rollback plan, approve the change ticket and agree on the validation method before anyone touches production.
Worked example: one CVE from detection to closure
In this illustrative case, a Java payments service runs on Amazon EKS behind an internet-facing Application Load Balancer. Its image contains log4j-core 2.14.1, which is vulnerable to CVE-2021-44228 (Log4Shell) and listed in CISA KEV.
- Detection: the registry scan flags the image, and a runtime check confirms the library is loaded.
- Scoring: KEV 30 + internet-facing 20 + loaded 15 + Tier 1 15 + service account with read access to a customer S3 bucket 10 = 90, so the finding is P1 with a 72-hour SLA.
- Routing: the CMDB tag owner=payments-team sends a Jira ticket with the due date to that team.
- Mitigation: within hours, a WAF rule blocks
${jndi:patterns as a temporary control. - Remediation: on day two, the team bumps to log4j-core 2.17.1, rebuilds the image and redeploys through the pipeline.
- Exception: a second copy sits in a legacy batch job on an internal EC2 host tied to a vendor release three weeks out. The team removes JndiLookup.class from the jar, and the CISO approves a 21-day exception.
- Validation: a rescan of the new image digest shows 2.17.1, and the runtime check shows the old library no longer loads, so the ticket closes with evidence.
- Reporting: P1 MTTR is recorded at two days, and the exception appears on the executive view with its expiry date.
Governance, ownership, and exception handling
Governance defines who sets policy, who fixes, who can accept risk and how missed SLAs escalate, so findings do not stall between teams.
| Activity | Security / VM team | Infrastructure & IT ops | Cloud / platform | App owners | CISO & leadership |
|---|---|---|---|---|---|
| Policy, SLAs, scoring model | R | C | C | C | A |
| Scan and sensor coverage | A | R | R | I | I |
| Triage and prioritization | R/A | C | C | C | I |
| Remediation | C | R | R | A | I |
| Exceptions and risk acceptance | R | C | C | R (request) | A |
| Executive reporting | R | I | I | I | A |
Escalation path for missed SLAs
- On the breach date, the ticket is flagged and the owner and their manager are notified.
- Seven days past due, the engineering director is notified.
- Thirty days past due, the CISO reviews it: the team fixes, files an exception or the change is forced.
- Every month, a risk committee reviews breaches by team and business unit.
Exception and risk acceptance rules
Every exception record must contain these fields:
- ✓Finding ID, affected asset and current priority score
- ✓Business justification for not fixing now
- ✓Compensating controls, with evidence that they were tested
- ✓Residual risk score after controls
- ✓Requester, asset owner and approver
- ✓Expiry date and review cadence
| Priority | Approver | Maximum duration |
|---|---|---|
| P1 | CISO plus business owner | 30 days |
| P2 | Security director | 90 days |
| P3 and P4 | VM lead | 180 days |
Review active exceptions monthly, and configure the ticketing system to reopen the finding automatically when an exception expires.
Metrics and dashboards that show the program is working
The right metrics show whether exposure is shrinking and fixes are sticking, broken down by asset type and team rather than reported as a single finding count. The targets below are starting points to tune to your environment.
| Metric | Formula | Starting target |
|---|---|---|
| Credentialed coverage by asset type | Assets with a successful authenticated, agent or sensor scan in the last 7 days ÷ in-scope assets of that type | ≥95% servers, ≥90% endpoints |
| Critical assets without credentialed coverage | Tier 0–1 assets lacking authenticated coverage ÷ all Tier 0–1 assets | 0% |
| MTTR by priority | Sum of (verified close date − detection date) ÷ closed findings, per band | P1 ≤ 7 days, P2 ≤ 30 days |
| SLA attainment by team | Findings closed within SLA ÷ findings due, per team | ≥90% for P1, P2 |
| Aging past SLA | Open findings past due, bucketed 1–30, 31–60, 61–90 and 90+ days | 90+ bucket trending to zero |
| Time to detect vs. time to remediate | Disclosure or deploy to first detection; detection to verified fix | Detection within 24–72 hours on internet-facing assets |
| Reopen rate | Findings reopened within 90 days ÷ findings closed | <5% |
| Remediation success rate | Fixes verified on first rescan ÷ fixes attempted | ≥95% |
| Exception volume and expiry | Active exceptions by priority; % expired without review | 0% expired unreviewed |
| Exposure reduction | Sum of priority scores for open P1, P2 findings on Tier 0–1 assets, month over month | Down every quarter |
Dashboards by audience
| View | Cadence | What it shows |
|---|---|---|
| VM and SOC team | Daily | New P1 and P2 findings from the last 24 hours, new KEV entries matched to your assets, failed credentialed scans and coverage gaps by asset type, reopened findings and validation failures |
| IT operations and engineering | Weekly | Open findings per team sorted by SLA due date, top fixes by risk removed (for example, one base image upgrade that closes 300 findings across 40 services), pending exception requests and upcoming expiries |
| Executives | Monthly or quarterly | Exposure reduction trend on Tier 0–1 assets, SLA attainment by business unit, active exceptions with residual risk and expiry dates, coverage percentage and current maturity stage |
Tools and integration architecture
A working program connects discovery, enrichment, scoring, ticketing, remediation and validation tools into one loop keyed on a shared asset ID. Before shortlisting products, map which of these jobs your current vulnerability management tools already cover.
| Category | Job in the program | Examples |
|---|---|---|
| Asset inventory / CMDB | Source of truth for owner and tier | ServiceNow CMDB, AWS Config, cloud provider asset APIs |
| Host and network scanners | OS, application and firmware CVEs | Credentialed network scanners, endpoint agents |
| Cloud security (CSPM / CNAPP) | Misconfigurations, workload CVEs, identity risk | Amazon Inspector, AWS Security Hub, Microsoft Defender for Cloud, Google Security Command Center, CNAPP platforms |
| Container and IaC scanning | Image CVEs, Terraform and CloudFormation misconfigurations | Registry scanning, CI/CD pipeline checks |
| Application security | Code flaws and vulnerable libraries | SAST, DAST, SCA with CycloneDX or SPDX SBOMs |
| Threat intelligence | Exploit likelihood signals | CISA KEV, FIRST EPSS, NVD |
| Patch and configuration management | Deploy fixes at scale | Microsoft Intune, Configuration Manager, Ansible, AWS Systems Manager Patch Manager |
| Ticketing and paging | Route work and track SLAs | Jira, ServiceNow, PagerDuty |
| EDR and cloud detection | Spot exploitation attempts against open findings | EDR agents, cloud detection and response |
Integration flow
- Scanners, agents, sensors and cloud connectors send findings to a central platform.
- The platform deduplicates on asset ID plus CVE and enriches findings with CMDB owner and tier.
- KEV, EPSS, exposure and runtime reachability data feed the scoring model.
- Tickets open automatically in Jira or ServiceNow with priority and SLA date, and P1s also page through PagerDuty.
- Owners remediate through patch tooling, image rebuilds or IaC pull requests.
- A completed ticket triggers a rescan or runtime check, which closes or reopens the finding.
- Open and closed data flow into the dashboards.
Where do cloud findings break the flow? Short-lived containers and autoscaled instances often vanish before a weekly scan runs, and a single vulnerable base image fans out into hundreds of duplicate findings. Group findings by image digest and fix the source image, not each running copy.
How Upwind fits into a vulnerability management program
Upwind adds runtime evidence to the prioritization and validation steps. Its eBPF sensors show which vulnerable packages are loaded, reachable and internet-exposed, so the runtime reachability factor in your scoring model comes from observation rather than guesswork. Upwind holds a 4.8/5 rating from 88 reviews on Gartner Peer Insights as of October 2026, where one reviewer wrote: “Within minutes of connecting Upwind, we were able to understand our most critical vulnerabilities.” Teams running mostly on AWS Fargate or other node-less serverless platforms should note that the sensor needs node-level access and does not run there.
- ✓Prioritizes CVEs by reachability, loaded packages, internet exposure and exploitability
- ✓Covers EKS, GKE, AKS and OKE, plus CSPM, CIEM, API and data security in one platform
- ✓Scans IaC, plugs into CI/CD and provides SBOM visibility
- ✓Traces runtime findings back to the line of code or configuration that introduced them
- ✓Routes findings to Jira, ServiceNow, PagerDuty and AWS Security Hub
Rolling out the program: 30/60/90 days, maturity stages, and mistakes to avoid
Most teams can stand up a working program in 90 days by fixing inventory and ownership first, then adding risk scoring, then automating routing and validation.
- Days 1–30: sign the charter, reconcile the asset inventory, tag owners and tiers on Tier 0–1 assets, deploy credentialed scanning and cloud connectors, and publish SLAs.
- Days 31–60: turn on deduplication and enrichment, adopt the scoring model with KEV and EPSS feeds, start auto-ticketing for P1 and P2, and launch the exception process.
- Days 61–90: add runtime reachability for cloud workloads, automate rescans on ticket close, ship the three dashboards and hold the first monthly risk review.
| Stage | Traits | Common blockers | Next step |
|---|---|---|---|
| Reactive | Ad hoc scans, CVSS-only sorting, spreadsheets | No inventory, no owners | Build the inventory and assign owners |
| Repeatable | Scheduled credentialed scans, CVSS-based SLAs, tickets | Noisy backlog, manual routing | Add KEV, EPSS and asset criticality |
| Risk-based | Scoring model, governed exceptions, team-level SLAs | Missing context for cloud and containers | Add runtime reachability and auto-routing |
| Automated | Auto-ticketing, pipeline gates, automatic validation, exposure metrics | Trust in automation, change control | Auto-remediate low-risk classes such as base image updates |
Common mistakes
- ✗Scanning before ownership is assigned, which fills a queue nobody works.
- ✗Relying on CVSS alone, which treats an unloaded library on a lab host like an exploited flaw on a payments API.
- ✗Reporting finding counts instead of exposure reduction, which rewards closing easy tickets.
- ✗Setting SLAs with no escalation path, so breaches carry no consequence.
- ✗Closing tickets without a rescan, which inflates MTTR results and hides reopened issues.
- ✗Granting exceptions with no expiry date, which turns them into permanent risk.
Start with the charter, inventory and owners, add a transparent scoring model, and measure verified fixes rather than activity. A vulnerability management program built this way earns trust from engineering and leadership, because it fixes the risks that matter first and proves it with evidence.
FAQ
What is a vulnerability management program?
A vulnerability management program is a documented, repeatable operating model for discovering, prioritizing, remediating, and verifying security weaknesses across servers, endpoints, cloud workloads, containers, applications, and identities. It includes scope, owners, SLAs, a scoring model, an exception process, and dashboards.
Why do asset inventory and ownership come before scanning?
The program only works once assets, owners, and scanner access are in place. A finding with no owner rarely gets fixed, so every asset should have an owning team, a criticality tier, and enough CMDB quality for tickets to route automatically.
How should vulnerabilities be prioritized instead of using CVSS alone?
The article recommends a risk-based scoring model that combines exploit intelligence such as CISA KEV and FIRST EPSS with internet exposure, runtime reachability, asset criticality, privilege or blast radius, and compensating controls. CVSS should be used only as a tiebreaker within a priority band and as a floor for the highest-risk assets.
Which vulnerability management metrics matter most?
The most useful metrics show whether exposure is going down and whether fixes are completed on time. The article highlights credentialed coverage by asset type, critical assets without coverage, MTTR by priority, SLA attainment by team, aging past SLA, reopen rate, remediation success rate, exception expiry, and exposure reduction on Tier 0 and Tier 1 assets.
How do you verify that a vulnerability is actually fixed?
A finding should close only after a rescan using the same authenticated method, or a runtime check, confirms the fix. The article recommends triggering the rescan automatically within 24 hours of a ticket being marked done and attaching evidence such as the fixed package version, image digest, scan ID, and ticket link.
