Cloud runtime security explained: detecting threats while workloads run

Cloud runtime security explained: detecting threats while workloads run

Santerra Holler October 07, 2026

Cloud runtime security explained: detecting threats while workloads run

Cloud runtime security is the practice of monitoring and protecting cloud workloads (containers, Kubernetes pods, virtual machines and serverless functions) and the cloud control plane while they run. It detects malicious behavior as it happens, instead of only scanning code, images or configurations before deployment. Sensors and log connectors collect live telemetry, such as system calls, process launches, network connections, Kubernetes audit events and cloud API activity. They compare that activity against rules and learned baselines, then trigger alerts or containment when something deviates. Pre-deployment checks cannot see an attacker who abuses a workload or identity that already passed them. Runtime detection is the layer that shows security teams what is actually happening in production.

Key takeaways

  • ✓Cloud runtime security detects threats in running workloads and cloud accounts by watching live behavior, such as process execution, network flows and API calls, rather than static configurations.
  • ✓Effective runtime detection correlates three layers: process activity on hosts and containers, Kubernetes control plane events, and cloud service and IAM activity.
  • ✓eBPF sensors give kernel-level visibility into system calls and network traffic with lower overhead than traditional kernel modules, but they need node access that platforms like AWS Fargate do not provide.
  • ✓Runtime programs fail most often because of over-alerting, so teams should start in detection-only mode, baseline normal behavior and tune suppressions before they automate response.
  • ✓Mean time to detect, mean time to respond, alert fidelity and sensor coverage by workload type are the core metrics for judging a runtime security program.

What is cloud runtime security?

Cloud runtime security is the set of controls that observe, detect and respond to threats in cloud environments during execution, after code has been built and deployed. The CNCF Cloud Native Security Whitepaper divides the cloud native lifecycle into four phases: develop, distribute, deploy and runtime. Runtime is the only phase where you see real attacker behavior rather than potential weaknesses.

The term covers more than workload protection. Workload runtime security watches what happens inside a host or container. Cloud runtime security extends that view to the orchestration layer and the cloud provider’s APIs, because an attacker who steals an access key never needs to touch a container at all.

Detection layer Telemetry source Example detections
Process / workload System calls, process trees, file access and network sockets (via eBPF or kernel modules) Shell spawned by a web server, binary written to /tmp and executed, outbound connection to a mining pool
Kubernetes Kubernetes audit logs, API server events, admission decisions kubectl exec into a production pod, new ClusterRoleBinding to cluster-admin, privileged pod created in kube-system
Cloud service / API AWS CloudTrail, Azure Activity Log, Google Cloud Audit Logs, VPC Flow Logs Instance role credentials used from an external IP, CloudTrail logging disabled, mass S3 GetObject calls

Each layer catches attacks the others miss. A reverse shell is invisible in CloudTrail, and a stolen IAM key used from a laptop is invisible to a container sensor. When you correlate the three layers, isolated alerts link up into one sequence of attacker actions.

Why does runtime security matter in cloud-native environments?

Cloud-native workloads change faster than static scans can track, and many exploitable conditions only exist while software is running. A Kubernetes cluster may replace thousands of pods a day, autoscaling groups create and destroy VMs by the hour, and serverless functions live for seconds. A weekly scan describes an environment that no longer exists.

Static checks also cannot answer the questions that decide real risk:

  • Is the vulnerable package actually loaded into memory, or just present on disk?
  • Is the workload reachable from the internet after security groups, load balancers and network policies apply?
  • Is this IAM role actually used, and by which process?
  • Is someone exploiting a zero-day that no scanner has a signature for yet?

How runtime security compares with other approaches

Approach When it runs What it sees Strengths Limitations
Static analysis (SAST, image and IaC scanning) Build and CI/CD Source code, dependencies, Dockerfiles, Terraform Cheap to fix issues early; blocks known CVEs before deployment No execution context; high volume of unreachable findings
Dynamic analysis (DAST) Test or staging Application responses to crafted requests Finds exploitable web flaws such as injection Limited coverage; does not see host, container or cloud activity
Posture management (CSPM, KSPM) Periodic or event-driven snapshots Cloud and cluster configuration Finds public buckets, open security groups, excessive permissions Point-in-time; shows possible exposure but not live activity
Endpoint detection and response (EDR) Continuously on VMs and servers Host processes and files Mature host detection on long-lived machines Built for laptops and servers; usually lacks pod, namespace, deployment, service account and cloud identity context
Runtime security Continuously in production Processes, syscalls, network flows, audit and API logs Detects active attacks and zero-days; confirms which risks are real Requires sensors or log pipelines, tuning and response ownership

What threats does cloud runtime security detect?

Cloud runtime security detects attacks that only reveal themselves through behavior, such as code execution, privilege escalation, persistence, credential abuse, lateral movement and data exfiltration in running environments. The main categories are:

  • Remote code execution: exploitation of a running application, such as a deserialization flaw, that leads to unexpected process launches.
  • Container escape: attempts to break isolation through privileged containers, mounted Docker sockets, hostPath volumes or kernel exploits.
  • Cryptojacking: unauthorized mining processes that consume CPU and connect to mining pools.
  • Credential theft and misuse: reading service account tokens, querying the instance metadata service (IMDS) or harvesting secrets from environment variables.
  • Lateral movement: pivoting between pods, nodes and cloud accounts with stolen credentials.
  • Defense evasion: deleting logs, disabling CloudTrail or running fileless payloads from memory.
  • Data exfiltration: unusual volumes of object storage reads or outbound transfers to unknown destinations.

The four scenarios below follow one example attack from initial access to cloud account abuse.

Scenario: reverse shell in a web container

An attacker exploits a vulnerable Java library in an internet-facing pod. The JVM spawns /bin/sh, which opens a TCP connection to an external IP on port 4444 and redirects stdin and stdout to the socket. A process-level rule fires on “shell spawned by java in a container” plus “outbound connection from a process with a socket bound to its standard input.”

Scenario: cryptojacking after initial access

The attacker downloads an XMRig binary with curl into /tmp, marks it executable and runs it. Detection triggers on write-then-execute from a writable path, a process name or hash that matches known miners, and DNS lookups for mining pool domains. In this example, node CPU jumps from a typical 30% to near 100%.

Scenario: container breakout attempt

A pod running with privileged: true lets the attacker call mount on the host’s root filesystem or use nsenter to enter the host’s PID namespace. Syscall-level detection catches setns, unshare and suspicious mounts from inside a container, while a Kubernetes-layer rule flags the privileged pod spec itself.

Scenario: stolen service account token and IAM credentials

The attacker reads /var/run/secrets/kubernetes.io/serviceaccount/token and queries the IMDS endpoint at 169.254.169.254 for the node’s IAM role credentials. Minutes later, Kubernetes audit logs show that token listing secrets across namespaces. CloudTrail shows the role’s temporary keys calling sts:GetCallerIdentity and s3:ListBuckets from an IP outside your VPC. Only correlation across all three layers ties these events back to the original container compromise.

How cloud runtime security detects threats while workloads run

The detection pipeline has six stages. It collects telemetry, adds context, checks events against rules and baselines, and sends high-confidence findings to response actions:

  1. Collect: sensors capture kernel events from nodes, while connectors ingest Kubernetes audit logs and cloud provider logs such as CloudTrail.
  2. Normalize: events are converted into a common schema with timestamps, process IDs, container IDs and source identities.
  3. Enrich: each container ID is mapped to its pod, deployment, namespace, node, cloud account and IAM role, then joined with vulnerability, exposure and data-sensitivity context.
  4. Detect: a rule engine and anomaly models evaluate the enriched stream in near real time.
  5. Correlate: related detections are grouped into one incident with a timeline.
  6. Respond: alerts go to the SOC, a SIEM or a SOAR playbook, and approved containment actions run automatically or after analyst approval.

Kernel telemetry and eBPF

eBPF (extended Berkeley Packet Filter) lets small, verified programs run inside the Linux kernel without loading a kernel module. Security sensors attach eBPF programs to tracepoints such as sys_enter_execve and sys_enter_connect, to kprobes on kernel functions, and to LSM hooks, then push events to user space through ring buffers. The kernel verifier rejects unsafe programs, so eBPF is more stable than legacy kernel modules. On kernels with BTF (BPF Type Format) support, CO-RE (“compile once, run everywhere”) programs work across kernel versions without per-kernel compilation.

Rules, baselines and enforcement

Rules encode known-bad patterns as conditions on event fields. In Falco’s syntax, a typical condition reads: spawned_process and container and proc.name in (bash, sh) and proc.pname in (java, node, python). Behavioral baselines catch novel attacks that have no rule yet. They record what each workload normally does over a learning window, for example 7 to 14 days, and then flag new binaries, new outbound destinations or new API calls.

Enforcement blocks activity instead of only reporting it. Seccomp profiles block disallowed syscalls, AppArmor and SELinux restrict file and capability access, and Cilium Tetragon can send SIGKILL to a process from inside the kernel the moment a policy matches.

Response actions

  • ✓Kill the offending process or restart the pod.
  • ✓Isolate the pod with a deny-all Kubernetes NetworkPolicy while preserving it for forensics.
  • ✓Cordon and drain the node, or quarantine it in a dedicated security group.
  • ✓Revoke active IAM sessions and rotate leaked keys or service account tokens.
  • ✓Snapshot disks and capture memory before the workload is terminated.

Deployment tradeoffs: agent-based, agentless and sidecars

Agentless scanning reads cloud APIs and disk snapshots. It deploys in minutes and covers every account, but it is point-in-time and cannot see processes or network connections as they happen. Agent-based sensors, usually deployed as a Kubernetes DaemonSet (one pod per node), give continuous visibility but need privileged node access and a supported kernel. Many teams use agentless scanning for broad inventory and posture, and sensors for live detection on the workloads that matter most.

AWS Fargate and GKE Autopilot restrict privileged containers and host access, so node-level eBPF sensors cannot run there. For those workloads, the options are sidecar-based agents (which add CPU and memory to every pod), language-level instrumentation, or cloud logs. Measure sensor overhead in staging under production-like load and set explicit resource requests and limits on the DaemonSet.

Editor’s tip: Inventory kernel versions across node pools before rollout. Older or custom kernels without BTF are the most common reason eBPF sensors fail to start, and they tend to run the legacy workloads you most need to watch.

Cloud runtime security tools and where they fit in the stack

Cloud runtime security tools include open-source detection engines, kernel enforcement mechanisms and commercial platforms that combine runtime sensors with posture, identity and vulnerability data. Common open-source building blocks include:

  • Falco: CNCF project for syscall-based detection of suspicious container and host behavior, using a rules language.
  • Cilium Tetragon: eBPF-based process and network monitoring with in-kernel enforcement.
  • SELinux, AppArmor and seccomp: mandatory access controls and system call filtering on the host.
  • Open Policy Agent (OPA) Gatekeeper and Kyverno: Kubernetes admission control that blocks risky pod specs before they start.
  • Cilium, Calico, Istio and Linkerd: network segmentation, mTLS and traffic visibility between services.
  • Prometheus, Grafana and Jaeger: metrics, visualization and tracing that help explain anomalies.

Runtime detection is one layer in a wider stack. A cloud SIEM stores and correlates logs across sources, while SOAR automates response playbooks.

Category Role Relationship to runtime security
Image scanning and secrets detection Find CVEs and hardcoded keys before deployment Runtime confirms which CVEs are loaded and reachable
Admission control Block non-compliant pods at deploy time Prevents risky configurations that runtime would otherwise have to watch
CSPM / CIEM Find misconfigurations and excessive permissions Runtime shows which permissions are actually used and abused
CWPP Protect hosts, containers and serverless workloads Provides the workload layer of runtime security
SIEM / SOAR Central log analysis and automated response Consumes runtime detections and triggers containment playbooks
CNAPP Unified platform for posture, workload and identity security Combines runtime evidence with posture data to prioritize risk

The Cloud Security Alliance describes how a CNAPP helps organizations evaluate cyber risk across multiple cloud technologies and providers. With one data model across layers, you can correlate events without stitching together exports from five tools.

How Upwind approaches cloud runtime security

Upwind builds its CNAPP on cloud runtime security. It pairs agentless discovery through read-only cloud scanners and APIs with lightweight eBPF sensors on VMs, containers and Kubernetes clusters, including EKS, GKE, AKS and OKE. The sensors observe system calls, network flows, API activity and in-memory execution at the kernel level. Upwind correlates that telemetry with IAM actions and cloud configuration into Threat Stories that include a timeline, root-cause analysis and response recommendations. Upwind holds a 4.8/5 rating from 88 reviews on Gartner Peer Insights as of October 2026. Reviewers repeatedly cite runtime visibility and detection accuracy, and a few note that its GCP support could be improved.

  • ✓CVE prioritization based on whether a package is loaded, reachable, internet-exposed and exploitable
  • ✓Containment through Microsoft Sentinel playbooks: container isolation, process termination and node quarantine
  • ✓Routing to ServiceNow, Jira, PagerDuty and AWS Security Hub
  • ✓Agentic Pack AI agents that investigate threats, validate exposure and generate fixes

Implementing cloud runtime security: rollout, tuning, metrics and best practices

Successful runtime programs roll out in phases and tune heavily before they automate response. They measure fidelity and coverage, not alert volume. These five phases double as a maturity model:

  1. Foundational logging: enable an organization-wide CloudTrail trail (or Azure Activity Log and Google Cloud Audit Logs), Kubernetes audit logging and VPC Flow Logs, with retention that meets your investigation and compliance needs.
  2. Visibility: deploy sensors to production clusters and VMs, and build an inventory of what runs, what talks to what and which identities are used.
  3. Detection-only mode: run rules and baselines without blocking, for example for two to four weeks, and measure noise per rule.
  4. Response automation: automate low-risk actions first, such as network isolation of a non-critical pod, then expand to process kills and credential revocation with approval gates.
  5. Continuous tuning: review rule performance on a fixed cadence, retire rules that never fire true positives, and add rules from incidents and threat intelligence.

Managing false positives

  • ✓Customize rules per environment: a shell in a dev namespace is routine; a shell in a payments namespace is not.
  • ✓Scope suppressions narrowly: suppress by image digest, namespace and process together, never by rule alone, and give each suppression an owner and an expiry date.
  • ✓Map severity to context: raise severity when the workload is internet-exposed, holds sensitive data or runs with a privileged IAM role.
  • ✓Re-learn baselines after major releases: otherwise new legitimate behavior floods the queue.

Metrics that show whether it works

Metric How to calculate it
Mean time to detect (MTTD) Average time from first malicious event to alert
Mean time to respond (MTTR) Average time from alert to containment
Alert fidelity True positive alerts ÷ total alerts, tracked per rule
Coverage by workload type Percentage of nodes, clusters, VMs and serverless functions with runtime telemetry
Enforcement coverage Percentage of production workloads under blocking policies such as seccomp or in-kernel kill rules

Compliance and evidence collection

Runtime telemetry is also your forensic record. Containers disappear on restart, so capture process trees, command lines, network connections, file changes and the identity behind each action before termination. Snapshot volumes and store evidence in a separate, write-once account to preserve chain of custody. File integrity monitoring from runtime sensors supports change-detection requirements such as PCI DSS requirement 11.5, and audit logs let you reconstruct an incident minute by minute for auditors and regulators.

Common implementation mistakes

  • ✗Turning on every default rule at once and burying the SOC in alerts.
  • ✗Leaving detections without a clear owner across security, platform and application teams.
  • ✗Writing policies so broad they suppress real attacks, or so strict they break deployments.
  • ✗Ignoring ephemeral workloads such as CronJobs, CI runners and serverless functions.
  • ✗Analyzing workload, Kubernetes and cloud logs in separate tools, which hides multi-stage attacks.

Best practices in priority order

  1. Monitor internet-facing workloads and those with privileged IAM roles or sensitive data first.
  2. Prioritize high-signal detections: shells in production containers, IMDS credential access, privileged pod creation, new cluster-admin bindings and cloud credentials used outside expected networks.
  3. Reduce the attack surface with Kubernetes security best practices such as least-privilege RBAC and Pod Security Standards.
  4. Harden images with Docker container security best practices, such as minimal base images, non-root users and read-only root filesystems. Hardened images make runtime anomalies easier to spot.
  5. Correlate process, Kubernetes and cloud signals in one incident view, then automate containment for the scenarios you have tested.

Static scanning and posture checks tell you what could go wrong. Cloud runtime security tells you what is going wrong right now, in which workload and under which identity. Start with logging and visibility, tune in detection-only mode, measure fidelity and coverage, and expand enforcement once your rules have proven accurate.

FAQ

What is cloud runtime security?

Cloud runtime security is the practice of monitoring and protecting cloud workloads and the cloud control plane while they run. It uses live telemetry such as system calls, process launches, network connections, Kubernetes audit events and cloud API activity to detect malicious behavior in production.

Why isn’t static scanning enough for cloud security?

Static scanning and posture checks show potential weaknesses before deployment, but they cannot see active attacks, stolen identities or malicious behavior that happens after a workload is running. Runtime security adds execution context and shows what is actually happening in production.

What threats can cloud runtime security detect?

It detects behavior-based threats such as remote code execution, container escape attempts, cryptojacking, credential theft, lateral movement, defense evasion and data exfiltration. Examples include a shell spawned by a web server, execution of a binary from /tmp, or cloud credentials used from an external IP.

How does cloud runtime security detect threats while workloads run?

It collects telemetry from workload sensors and cloud log connectors, normalizes and enriches the data with Kubernetes and IAM context, evaluates it against rules and behavioral baselines, then correlates related signals into incidents. High-confidence findings can then trigger alerts or containment actions.

What are the limits of eBPF-based runtime security sensors?

eBPF sensors provide kernel-level visibility with lower overhead than traditional kernel modules, but they require node access and a supported kernel. Platforms such as AWS Fargate and GKE Autopilot restrict privileged host access, so teams may need to rely on sidecars, language-level instrumentation or cloud logs instead.

Contents
Add the Upwind RSS Feed to Slack
Connect the Upwind RSS Feed to your Slack.
Follow the how-to here.
Threat RSS
Add the Upwind RSS Feed to Slack
Connect the Upwind RSS Feed to your Slack.
Follow the how-to here.
Main RSS