20 cloud security best practices for 2026

20 cloud security best practices for 2026

Santerra Holler October 09, 2026

20 cloud security best practices for 2026

The 20 cloud security best practices for 2026 fall into five areas. Teams should enforce least privilege for human and machine identities, classify and encrypt data, keep backups immutable, harden the software supply chain and Kubernetes, secure AI workloads, detect threats at runtime, and normalize controls across AWS, Azure and GCP. All of it should be measured with clear KPIs.

Traditional security programs were not built for the cloud. Infrastructure appears from a Terraform plan in minutes, machine identities outnumber employees, and AI workloads add models, vector databases and GPU clusters to the attack surface. Perimeter firewalls, quarterly scans and annual access reviews assume static data centers. They miss the ephemeral containers, short-lived credentials and API-driven changes where most common cloud security risks now begin.

This guide, updated in October 2026, turns those realities into cloud security best practices you can act on. Each practice explains what to do, how to implement it and what can go wrong. A phased roadmap at the end shows where to start.

Key takeaways

  • ✓Identity, both human and non-human, is the primary control plane in cloud environments, so least privilege, federation and just-in-time access come first.
  • ✓Vulnerability and misconfiguration findings should be prioritized by runtime evidence, such as whether a package is loaded and whether the asset is reachable, rather than by severity score alone.
  • ✓Immutable, cross-account backups with tested restore objectives are the most reliable defense against ransomware and destructive cloud attacks.
  • ✓AI workloads, vector databases and agents need the same inventory, least-privilege and runtime monitoring controls as any other production workload.
  • ✓A cloud security program improves only when it tracks KPIs such as mean time to remediate, MFA coverage and backup restore success.

Three shifts shape cloud security this year:

  • Identity-led attacks: attackers log in with stolen keys and tokens far more often than they exploit firewalls.
  • Alert overload: configuration scanners produce thousands of findings, and teams need runtime context to know which ones are exploitable.
  • AI expansion: models, agents and GPU clusters introduce new data flows and new privileged identities.

The table below summarizes all 20 practices by the risk each one addresses, its control type (preventive, detective or responsive), the effort involved and the key tools or processes.

# Best practice Risk addressed Control type Difficulty Key tools / processes
1 Least privilege with CIEM Excessive permissions, privilege escalation Preventive Medium CIEM, IAM Access Analyzer, SCPs
2 Govern non-human identities Machine credential sprawl Preventive High Workload identity federation, IMDSv2
3 Federation, phishing-resistant MFA, JIT access Account takeover, standing admin rights Preventive Medium SSO, FIDO2, PIM
4 Zero trust segmentation Lateral movement, public exposure Preventive Medium Private endpoints, NetworkPolicies
5 Discover and classify data Shadow data, residency violations Detective Medium DSPM, region guardrails
6 Encryption and key management Data exposure, weak cryptography Preventive Medium KMS, HSM, BYOK/HYOK
7 Secrets management and rotation Hardcoded credentials Preventive Low Secrets vaults, secret scanning
8 Immutable, isolated backups Ransomware, destructive attacks Responsive Medium Object Lock, Vault Lock
9 IaC scanning and policy as code Misconfigurations at deploy time Preventive Low OPA, Kyverno, drift detection
10 Software supply chain controls Malicious packages, tampered builds Preventive High SBOM, cosign, SLSA provenance
11 Runtime-prioritized vulnerability management CVE backlog, exploitable flaws Preventive / Detective Medium Reachability analysis, CISA KEV
12 Container and Kubernetes hardening Container escape, cluster takeover Preventive Medium Pod Security Standards, admission control
13 AI workload security Prompt injection, model and data leakage Preventive / Detective High AI-BOM, AI-SPM, runtime monitoring
14 Centralized logging Blind spots, lost evidence Detective Low CloudTrail, Cloud Audit Logs, OCSF
15 Runtime detection and detection engineering Active intrusions, alert fatigue Detective High eBPF sensors, MITRE ATT&CK
16 Incident response playbooks and tabletops Slow containment Responsive Medium SOAR playbooks, exercises
17 Multi-cloud normalization Inconsistent controls across providers Preventive High Landing zones, one control catalog
18 Compliance control mapping Audit failures, regulatory fines Detective Medium Continuous compliance evidence
19 Shared responsibility and security culture Ownership gaps, human error Preventive Low Security champions, paved roads
20 KPI-driven program management Unmeasured risk Detective Low Metrics dashboards, SLAs

Identity, access and network best practices

Stolen or over-permissioned credentials give attackers a direct path through cloud accounts, so these practices limit what each human and machine identity can reach.

1. Enforce least privilege with entitlement management

Cloud permissions pile up. Developers request broad roles to unblock work, and nobody removes them afterward. Cloud Infrastructure Entitlement Management (CIEM) compares the permissions an identity was granted with the permissions it actually uses.

  • ✓Right-size roles from usage data: AWS IAM Access Analyzer and last-accessed data, GCP IAM Recommender, and Entra ID access reviews.
  • ✓Remove wildcard actions ("Action": "*") and wildcard resources from production roles.
  • ✓Flag toxic combinations. For example, iam:PassRole plus ec2:RunInstances lets a user launch an instance with a more privileged role attached.
  • ✓Set ceilings with AWS SCPs and permission boundaries, Azure Policy deny assignments and GCP organization policies.

For example, a deployment role with AdministratorAccess that used only 14 actions over 90 days can be scoped to those 14 actions. The most common mistake is treating cleanup as a one-off project. Permissions drift back within months unless reviews recur.

2. Govern non-human and workload identities

Non-human identities include service accounts, CI/CD tokens, Lambda execution roles, Kubernetes service accounts and AI agents. They rarely have owners and often hold long-lived keys.

  • ✓Replace static access keys with federation: EKS Pod Identity or IRSA, Azure Workload Identity, GKE Workload Identity, and GitHub Actions OIDC to cloud roles.
  • ✓Tag every machine identity with an owner. Disable identities that have no owner or have been unused for 90 days.
  • ✓Enforce IMDSv2 on EC2 to block SSRF-based credential theft.

In a typical breach, an attacker finds an SSRF flaw in a web app, queries the instance metadata service, retrieves the role’s temporary credentials and lists every S3 bucket the over-permissioned role can read. Practices 1 and 2 together break that chain.

3. Use federation, phishing-resistant MFA and just-in-time access

Human access should flow through one identity provider, such as Entra ID or Okta, over SAML or OIDC. Avoid local IAM users for people.

  • ✓Require FIDO2 security keys or passkeys for every administrator, and retire SMS codes.
  • ✓Grant just-in-time (JIT) elevation through tools such as Entra Privileged Identity Management or AWS IAM Identity Center, time-boxed to 1–4 hours with approval.
  • ✓Apply just-enough-administration (JEA) so elevated sessions cover only the task at hand.
  • ✓Lock root and global admin accounts behind hardware keys, and alert on every break-glass login.

4. Apply zero trust network segmentation

Zero trust means no request is trusted because of where it comes from. In the cloud, that translates into default-deny networking:

  • ✓Never open SSH (22) or RDP (3389) to 0.0.0.0/0. Use AWS Systems Manager Session Manager, Azure Bastion or Google IAP instead.
  • ✓Reach managed services through AWS PrivateLink, Azure Private Link or Private Service Connect.
  • ✓Apply default-deny Kubernetes NetworkPolicies per namespace, and use mTLS through a service mesh.
  • ✓Filter egress traffic to block data exfiltration and command-and-control callbacks.

Data protection and resilience best practices

These practices make sure sensitive data is known, encrypted under keys you control, free of embedded secrets and recoverable after an attack.

5. Discover and classify sensitive data continuously

You cannot protect data you have not found. Data Security Posture Management (DSPM) discovers data stores, classifies PII, PHI and cardholder data, and maps which identities and workloads can reach them. Pay special attention to shadow copies such as snapshots, dev buckets and analytics exports.

For residency, pin regulated data to approved regions. Use the aws:RequestedRegion condition in SCPs, Azure Policy allowed locations and the GCP resource location constraint. For example, a production database snapshot copied into a dev account and shared publicly exposes customer records even though production itself was locked down.

6. Encrypt with disciplined key management

Default encryption at rest is table stakes. What matters is who controls the keys.

  • ✓Use customer-managed keys in AWS KMS, Azure Key Vault or Cloud KMS for regulated data, with automatic annual rotation.
  • ✓Separate key administrators from key users in key policies.
  • ✓Choose your control level. BYOK imports your key material into the provider’s KMS. HYOK keeps keys in your own HSM through services such as AWS KMS External Key Store or Google Cloud External Key Manager, which gives more control but adds a dependency on your availability.
  • ✓Tokenize card numbers to shrink PCI DSS scope, and enforce TLS 1.2 as a minimum, with TLS 1.3 preferred.
  • ✓Inventory cryptographic usage now, track post-quantum standards such as ML-KEM (FIPS 203), and design for crypto agility.

7. Eliminate hardcoded secrets and rotate automatically

Store credentials in AWS Secrets Manager, Azure Key Vault, Google Secret Manager or HashiCorp Vault. Never put them in code, container images or unencrypted Terraform state. Run secret scanning in pre-commit hooks and CI. Rotate database credentials on a schedule, for example every 30 days, or use dynamic secrets that expire per session.

In a typical incident, a developer pushes an access key to a public repository, automated scrapers harvest it, and the attacker launches GPU instances for cryptomining until the bill reveals the leak.

8. Build immutable, isolated backups for ransomware recovery

Ransomware operators who gain admin access delete snapshots before they encrypt anything. Backups must survive a compromised account.

  • ✓Follow the 3-2-1-1 rule: three copies, on two media types, with one off-site and one immutable.
  • ✓Enable S3 Object Lock in compliance mode, AWS Backup Vault Lock or Azure immutable vaults.
  • ✓Copy backups into a separate backup account with its own credentials, and deny delete actions from production.
  • ✓Set restore objectives per tier. For example, tier-1 systems might target an RPO of 15 minutes and an RTO of 4 hours.
  • ✓Test restores at least quarterly and record the success rate.

Workload, supply chain and AI security best practices

These practices secure what you build and run, including infrastructure code, build pipelines, containers, Kubernetes clusters and the models that serve production traffic.

9. Shift left with IaC scanning and policy as code

Scan Terraform, CloudFormation, Bicep and Helm charts in every pull request. Encode rules as policy as code with OPA/Rego and Conftest in CI, and enforce them in clusters with Gatekeeper or Kyverno. Fail builds on critical issues such as public buckets or unencrypted volumes.

Detect drift when someone changes resources in the console, and fix the code rather than the live resource. Otherwise the next deployment reintroduces the flaw.

10. Secure the software supply chain

Attackers increasingly target build systems and dependencies instead of production systems.

  • ✓Generate an SBOM (software bill of materials) in SPDX or CycloneDX format for every build.
  • ✓Sign artifacts with Sigstore cosign, and verify signatures at admission before pods run.
  • ✓Produce SLSA provenance attestations that record what was built, where and from which commit.
  • ✓Pin dependencies by version and hash, and pin container images by digest rather than :latest.
  • ✓Run builds on isolated, ephemeral runners with short-lived OIDC credentials.
  • ✓Proxy packages through a private registry, and block typosquatted or dependency-confusion packages with scoped namespaces.

11. Prioritize vulnerabilities by runtime reachability

Severity scores alone create backlogs no team can clear. Rank CVEs by real-world context:

  • ✓Is the vulnerable package actually loaded at runtime?
  • ✓Is the workload reachable from the internet after network controls are applied?
  • ✓Is the CVE listed in CISA’s Known Exploited Vulnerabilities catalog, or does it have a high EPSS score?
  • ✓Does the workload hold a privileged identity or access to sensitive data?

Set patch SLAs by context. For example, fix critical, reachable and exposed issues within 7 days, high-severity issues within 30 days and the rest within 90 days.

Rebuild and redeploy container images instead of patching running containers. Patched-in-place containers drift from their image and lose the fix on the next restart.

12. Harden containers and Kubernetes

Kubernetes adds a control plane, identities and misconfigurations of its own. Use these hardening steps as a baseline:

  • ✓Use minimal or distroless base images, run as non-root and mount read-only root filesystems.
  • ✓Enforce the “restricted” Pod Security Standard and block privileged pods and hostPath mounts.
  • ✓Limit cluster-admin RBAC bindings, keep the API server endpoint private and encrypt secrets in etcd.
  • ✓Benchmark clusters against the CIS Kubernetes Benchmark and enable audit logging.
  • ✓Monitor runtime behavior with lightweight eBPF sensors that observe system calls and network flows from the kernel.

13. Secure AI workloads, models and vector databases

AI services such as Amazon Bedrock, Azure OpenAI and Vertex AI, along with self-hosted models on GPU nodes, bring new assets and new data paths. Secure them in five areas:

  • ✓Inventory: maintain an AI bill of materials (AI-BOM) covering models, endpoints, notebooks, agents and vector databases.
  • ✓Vector databases: require authentication, avoid public endpoints and filter retrieval results by the source document’s permissions.
  • ✓Prompts and outputs: log them with PII redaction, and treat model output as untrusted input to defend against prompt injection.
  • ✓Agents: give agent tool permissions least privilege, and require human approval for destructive actions.
  • ✓Models and GPUs: verify model weights by hash, prefer safetensors over pickle formats, isolate GPU node pools and watch for cryptomining.

The OWASP Top 10 for LLM Applications and MITRE ATLAS provide threat models for this work.

Detection and response best practices

These practices assume prevention will eventually fail, so they make sure you see attacks early and contain them fast.

14. Centralize logging across every cloud

Collect these sources at a minimum:

  • ✓An organization-wide AWS CloudTrail trail, including data events for sensitive buckets.
  • ✓Azure Activity Logs plus Entra ID sign-in and audit logs.
  • ✓GCP Cloud Audit Logs, including Data Access logs, routed through an organization-level sink.
  • ✓VPC flow logs, DNS logs and Kubernetes audit logs.

Write everything to a dedicated log archive account with Object Lock. Normalize events to a common schema such as OCSF so a single query works across providers. As an example retention policy, keep 90 days searchable and one year archived.

15. Deploy runtime detection and engineer detections deliberately

Posture tools show what could happen. Runtime detection shows what is happening. Build detections for high-signal cloud use cases and map each one to the MITRE ATT&CK Cloud matrix:

  • ✓An unexpected shell or reverse shell inside a container (Execution).
  • ✓A cryptominer process, or connections to mining pools (Impact).
  • ✓Credentials taken from instance metadata and then used from an external IP (T1552.005).
  • ✓Disabling CloudTrail or other cloud logs (T1562.008).
  • ✓A new access key or admin role created after an anomalous login (T1098).

Manage detections as code in Git, with tests. Run weekly hypothesis-driven hunts in cloud logs, for example looking for unusual AssumeRole chains or spikes in GetSecretValue. Reduce false positives by adding context rather than disabling rules. An alert on an internet-facing workload with access to sensitive data deserves a page. The same behavior in an isolated test namespace may not.

16. Maintain incident response playbooks and test them

Cloud incidents move faster than ticket queues, so playbooks must be written before they are needed. Each one should include:

  1. Triggers, severity levels and a named owner for each role.
  2. Containment steps per resource. Revoke sessions with a deny policy that uses aws:TokenIssueTime, apply a quarantine security group, cordon the Kubernetes node and rotate exposed keys.
  3. Evidence preservation: take EBS snapshots and capture memory before you terminate anything.
  4. Communication and regulatory clocks, such as GDPR’s 72-hour breach notification and DORA incident reporting.
  5. Recovery steps and a blameless post-incident review.

Run tabletop exercises quarterly and a full technical simulation at least once a year. Update the playbooks after every real incident.

Governance, multi-cloud and measurement best practices

These practices tie the technical controls together, so they apply consistently across providers, satisfy auditors and improve over time.

17. Normalize security across AWS, Azure and GCP

Most enterprises run more than one cloud, and each provider names and enforces controls differently. Build one consistent cloud security architecture on four pillars:

  • ✓Policy normalization: define one control catalog, such as “no public storage,” and implement it as SCPs and AWS Config rules, Azure Policy and GCP organization policies.
  • ✓Identity federation: connect one identity provider to all three clouds, with consistent role names.
  • ✓Centralized logging: send everything to one SIEM or security data lake (practice 14).
  • ✓Landing zones: use AWS Control Tower, Azure landing zones and Google Cloud foundation blueprints, with mandatory owner, environment and data-classification tags.

18. Map controls to compliance frameworks

Map each technical control to the frameworks you report against, then collect evidence continuously instead of before audits. The table below shows the main mappings, and CIS Benchmarks for each provider supply the configuration-level checks.

Practice cluster NIST CSF 2.0 ISO 27001:2022 SOC 2 PCI DSS 4.0 HIPAA / GDPR / DORA
Identity and access (1–4) PR.AA A.5.15–5.18, A.8.2, A.8.5 CC6.1, CC6.3 Req. 7, 8 164.312(a), (d) / Art. 32 / Art. 9
Data and encryption (5–7) PR.DS A.8.10–8.12, A.8.24 CC6.1, CC6.7 Req. 3, 4 164.312(e) / Art. 32, 44–49 / Art. 9
Backup and resilience (8) PR.DS, RC.RP A.8.13, A.5.30 A1.2, A1.3 Req. 12.10 164.308(a)(7) / Art. 32 / Art. 11–12
Workloads and supply chain (9–13) ID.RA, PR.PS A.8.8, A.8.25–8.28 CC7.1, CC8.1 Req. 6, 11 164.308(a)(1) / Art. 25 / Art. 28
Detection and response (14–16) DE.CM, RS.MA A.8.15–8.16, A.5.24–5.26 CC7.2, CC7.4 Req. 10, 12.10 164.312(b) / Art. 33 / Art. 17–19

19. Clarify shared responsibility and build security culture

Responsibility shifts with the service model:

  • IaaS (EC2, Azure VMs): you own the OS, patching, network rules, identities and data.
  • PaaS (RDS, App Service): the provider runs the OS and runtime, and you own configuration, IAM and data.
  • SaaS (Microsoft 365): you still own identities, sharing settings and data.

Put a security champion in each engineering team. Publish paved-road Terraform modules with secure defaults, and train on cloud-specific scenarios rather than generic phishing alone. Review overlap between the types of cloud security tools you run, since tool sprawl splits ownership and context.

20. Run the program on measurable KPIs

Track a small set of metrics monthly and report trends to leadership. The targets below are examples to adapt to your risk appetite.

KPI Example target Practices measured
Mean time to remediate critical, internet-exposed misconfigurations Under 72 hours 4, 9, 17
Admin identities with phishing-resistant MFA 100% 3
Privileged identity review frequency Monthly for admin roles, quarterly for all others 1, 2
Long-lived machine access keys Trending to zero 2, 7
Patch SLA compliance by severity 95% or more within SLA 11, 12
Backup restore test success rate 100%, tested quarterly 8
Mean time to detect and contain runtime incidents Detect in minutes, contain in under 1 hour 15, 16

How Upwind supports runtime-first cloud security

Upwind grounds posture, identity and vulnerability findings in runtime evidence from lightweight eBPF sensors, combined with agentless discovery for broad cloud inventory. That context shows which risks are real, which supports the prioritization in practices 1, 11 and 15. Users rate Upwind 4.8/5 from 88 reviews on Gartner Peer Insights as of October 2026, with reviewers highlighting runtime visibility and reduced alert fatigue. A few reviewers say Upwind’s GCP support could be improved.

  • ✓Vulnerability prioritization by reachability, loaded packages, internet exposure and exploitability.
  • ✓CIEM that flags over-permissioned roles, unused permissions and toxic combinations, with least-privilege policy recommendations.
  • ✓Cloud detection and response that correlates kernel-level telemetry with identity and configuration context into Threat Stories with timelines and root cause.
  • ✓AI-SPM, AI-DR and AI-BOM for AI workloads, models and agents.
  • ✓IaC scanning, CI/CD integration, SBOM visibility and tracing runtime findings back to the line of code.

A phased roadmap for implementing the 20 practices

Implement the 20 cloud security best practices in order of risk reduction per unit of effort. Start with identity and visibility, then move to advanced detection and AI controls.

  1. Foundational (first 90 days): federate human access and enforce phishing-resistant MFA (3). Centralize logging (14). Close public exposure (4). Move secrets into vaults (7). Make backups immutable and isolated (8). Start tracking the KPIs in practice 20.
  2. Intermediate (3–9 months): right-size permissions with CIEM (1). Govern machine identities (2). Classify data and enforce residency (5). Adopt customer-managed keys (6). Gate IaC in CI (9). Prioritize CVEs by runtime context (11). Harden Kubernetes (12). Write and test playbooks (16). Map controls to frameworks (18).
  3. Advanced (9–18 months): enforce supply chain signing and provenance (10). Secure AI workloads (13). Build a detection engineering program mapped to MITRE ATT&CK (15). Normalize controls across clouds (17).

FAQ

Why do traditional security programs fall short in cloud environments?

Traditional programs assume static infrastructure, annual access reviews and perimeter defenses. In the cloud, infrastructure is created through APIs in minutes, credentials are short-lived, containers are ephemeral and machine identities often outnumber employees, so attacks commonly begin in places older security models miss.

What is the top cloud security priority in 2026?

Identity is the primary control plane in cloud environments, so the first priority is enforcing least privilege for both human and non-human identities, combined with federation, phishing-resistant MFA and just-in-time access.

What does runtime-prioritized vulnerability management mean?

It means ranking vulnerabilities by real-world context instead of severity alone, such as whether the vulnerable package is loaded at runtime, whether the workload is reachable from the internet, whether the CVE is actively exploited and whether the asset has privileged access or sensitive data.

Why are immutable, isolated backups so important for cloud security?

Attackers with admin access often delete snapshots before encrypting systems. Immutable, cross-account backups help ensure recovery still works after a compromised production account, especially when restore objectives are defined and restores are tested quarterly.

How should teams secure AI workloads in the cloud?

Teams should inventory models, agents, notebooks and vector databases with an AI-BOM, enforce least privilege for AI identities and tools, avoid public vector database endpoints, log prompts and outputs with PII redaction, verify model weights by hash and monitor GPU workloads at runtime.

Contents
Add the Upwind RSS Feed to Slack
Connect the Upwind RSS Feed to your Slack.
Follow the how-to here.
Threat RSS
Add the Upwind RSS Feed to Slack
Connect the Upwind RSS Feed to your Slack.
Follow the how-to here.
Main RSS