kubernetes

Kubernetes Security Best Practices for Production

By Shubhankar Tripathi • • 5 min read

Kubernetes Security Best Practices for Production

bitcodematrix.com | Kubernetes Series

1. Introduction

Kubernetes is secure only to the degree that it is configured to be. Out of the box it favours flexibility: Pods can talk to every other Pod, containers often run as root, Secrets are only base64-encoded, and a mis-scoped ServiceAccount can quietly become a path to cluster-wide access. Production security is therefore a collection of deliberate choices made at every layer, not a single feature that you switch on.

This article pulls together the controls covered individually in earlier articles of this series into one practical, prioritised view. It explains the defence-in-depth model, walks through each layer, gives working examples, and ends with a production readiness checklist you can adapt to your own environment.

After reading it you should be able to:

  • Explain why Kubernetes security must be layered and name the layers.

  • Harden the API server, authentication, authorization, audit logging and etcd.

  • Lock down Pods with Pod Security Standards, securityContext, ServiceAccount settings and NetworkPolicy.

  • Reduce supply-chain risk with image scanning, signing and admission policies.

  • Apply a repeatable checklist and troubleshoot common security-related failures.

2. The Defence-in-Depth Model

No single control is perfect, so security is built in layers. If an attacker gets past one layer, the next one limits the damage. A common way to organise Kubernetes security is the "4C" idea (Cloud, Cluster, Container, Code), which this article expands into five layers for clarity.

Five nested security layers from cloud to application code

Figure 1: The five layers of Kubernetes security

An important rule follows from the figure: a weakness in an outer layer undermines everything inside it. A perfectly hardened Pod does not help if the API server is reachable from the internet with weak credentials, and a locked-down cluster does not help if the image contains a known remote-code-execution flaw.

3. Threat Model: What Are You Defending Against?

Controls make more sense when tied to concrete threats. The table below lists common production threats and the primary defences used against them.

Threat

Typical cause

Primary defences

Compromised container

Vulnerable application or image

Non-root, read-only filesystem, dropped capabilities, seccomp, NetworkPolicy

Privilege escalation to cluster

Over-permissive RBAC or ServiceAccount token

Least-privilege RBAC, disabled token auto-mount, audit logs

Container escape to node

Privileged Pods, hostPath, host namespaces

Pod Security Standards (Restricted), sandboxed runtimes, hardened nodes

Lateral movement

Flat network, any Pod reaches any Pod

Default-deny NetworkPolicy, namespace isolation

Secret exposure

Plain Secrets in Git, broad read access

Encryption at rest, external secret stores, tight RBAC

Malicious or tampered image

Untrusted registry, no verification

Scanning, signing, digest pinning, admission policy

Resource abuse and DoS

No limits, noisy neighbours, cryptomining

ResourceQuotas, LimitRanges, runtime detection

Unpatched components

Old Kubernetes or node OS versions

Regular upgrades and patching process


4. Layer 1: Cloud, Network and Infrastructure

Everything in Kubernetes runs on infrastructure you must also protect. Whether you use a managed service or run your own clusters, apply these basics first.

  • Keep the API server private where possible. Use a private endpoint or restrict public access to known IP ranges. Never expose it openly to the internet.

  • Isolate nodes. Place nodes in private subnets, restrict inbound traffic with security groups or firewalls, and allow only the ports that are required.

  • Apply least privilege to cloud IAM. Node roles should not carry broad permissions. Use workload identity features (for example IAM roles for ServiceAccounts or equivalent) instead of long-lived cloud keys inside Pods.

  • Protect the metadata service. Block or restrict Pod access to the cloud instance metadata endpoint, because it can expose node credentials.

  • Encrypt storage and backups. Enable encryption for disks, snapshots and etcd backups.

  • Separate environments. Use separate clusters or at least separate accounts or projects for production and non-production.

5. Layer 2: Securing the Control Plane

The control plane decides what runs in the cluster, so compromising it means compromising everything. See the architecture article in this series for how the components fit together.

5.1 Authentication

Every request to the API server must be authenticated. Prefer short-lived identities issued by an external identity provider using OpenID Connect (OIDC) over static tokens or shared client certificates. Disable anonymous access unless you have a specific, understood need for it, and avoid long-lived credentials for people. The separate article on authentication versus authorization covers the mechanics in detail.

5.2 Authorization with RBAC

RBAC should follow least privilege. Grant access to named resources and verbs rather than wildcards, prefer namespaced Roles over ClusterRoles, and review bindings regularly.

  • Avoid binding the cluster-admin ClusterRole to users or groups for daily work.

  • Avoid wildcards such as resources: [""] or verbs: [""].

  • Treat permissions to create Pods, read Secrets, bind roles, and impersonate as highly sensitive, because each can be used to escalate privileges.

  • Audit regularly using kubectl auth can-i and tools that list effective permissions.

# What can this ServiceAccount do?

kubectl auth can-i --list \

  --as=system:serviceaccount:shop:orders-app -n shop


5.3 Audit Logging

Audit logs record who did what and when against the API server. Without them, incident investigation is guesswork. Enable an audit policy that captures metadata for most requests and request or response bodies only for sensitive resources, then ship the logs to a central, tamper-resistant system. On managed services, enable the provider's control plane logging option.

apiVersion: audit.k8s.io/v1

kind: Policy

rules:

  - level: None

    users: ["system:kube-proxy"]

    verbs: ["watch"]

  - level: Metadata

    resources:

      - group: ""

        resources: ["secrets", "configmaps"]

  - level: RequestResponse

    resources:

      - group: "rbac.authorization.k8s.io"

  - level: Metadata


Note that Secrets are logged at the Metadata level above so that secret values never appear in the logs.

5.4 etcd and Encryption at Rest

etcd stores all cluster state, including Secrets. Protect it by allowing only the API server to reach it, requiring mutual TLS, encrypting backups, and enabling encryption at rest for Secrets through an EncryptionConfiguration. Where available, prefer a KMS provider so the encryption keys are held outside the cluster.

apiVersion: apiserver.config.k8s.io/v1

kind: EncryptionConfiguration

resources:

  - resources: ["secrets"]

    providers:

      - aescbc:

          keys:

            - name: key1

              secret: <base64-encoded-32-byte-key>

      - identity: {}


The example uses a local key for illustration only. In production, a KMS-based provider is generally preferred, and the key must never be committed to version control. On managed Kubernetes services, check how the provider handles encryption at rest and whether you can supply your own key.

5.5 API Server Hardening

  • Keep RBAC and the Node authorizer enabled in the authorization mode list.

  • Keep the recommended admission plugins enabled and add validating policies (see section 8).

  • Use TLS everywhere, including between the API server, etcd and kubelets.

  • Restrict network access to the control plane and the kubelet API.

  • Compare your configuration against the CIS Kubernetes Benchmark and run an automated scanner such as kube-bench regularly.

6. Layer 3: Securing Nodes and the Runtime

Nodes run your containers, and a container escape lands the attacker on a node. Hardening nodes reduces how far an escape can go.

  • Use a minimal, container-optimised node OS with few packages, automatic security updates and a read-only root filesystem where supported.

  • Do not allow SSH by default. Use short-lived, audited access when it is required.

  • Harden the kubelet. Disable anonymous authentication, use Webhook authorization, and do not expose the read-only port.

  • Use a supported container runtime such as containerd or CRI-O, kept up to date.

  • Consider sandboxed runtimes (for example gVisor or Kata Containers, via RuntimeClass) for untrusted or multi-tenant workloads.

  • Enable seccomp, AppArmor or SELinux profiles so containers can use only the system calls and resources they need.

  • Replace nodes rather than patching them in place where possible, so nodes are consistent and short-lived.

# Example kubelet configuration fragment

apiVersion: kubelet.config.k8s.io/v1beta1

kind: KubeletConfiguration

authentication:

  anonymous:

    enabled: false

  webhook:

    enabled: true

authorization:

  mode: Webhook

readOnlyPort: 0

protectKernelDefaults: true


7. Layer 4: Securing Workloads

7.1 Pod Security Standards

Pod Security Standards define three profiles: Privileged, Baseline and Restricted. The built-in Pod Security Admission controller enforces them per namespace through labels. For most application namespaces in production, target the Restricted profile. The dedicated article in this series explains each profile.

apiVersion: v1

kind: Namespace

metadata:

  name: shop

  labels:

    pod-security.kubernetes.io/enforce: restricted

    pod-security.kubernetes.io/audit: restricted

    pod-security.kubernetes.io/warn: restricted


If you are introducing this to an existing cluster, start with the warn and audit modes to find violations before you switch on enforce.

7.2 securityContext

The securityContext is where you tell the runtime how tightly to confine a container. A solid production baseline looks like this:

apiVersion: v1

kind: Pod

metadata:

  name: secure-app

  namespace: shop

spec:

  automountServiceAccountToken: false

  securityContext:

    runAsNonRoot: true

    runAsUser: 10001

    seccompProfile:

      type: RuntimeDefault

  containers:

    - name: app

      image: registry.example.com/shop/app@sha256:<digest>

      securityContext:

        allowPrivilegeEscalation: false

        readOnlyRootFilesystem: true

        capabilities:

          drop: ["ALL"]

      resources:

        requests:

          cpu: 100m

          memory: 128Mi

        limits:

          memory: 256Mi

      volumeMounts:

        - name: tmp

          mountPath: /tmp

  volumes:

    - name: tmp

      emptyDir: {}


Key points about this example: the container runs as a non-root user, cannot gain extra privileges, has a read-only root filesystem with a writable /tmp provided by an emptyDir volume, drops all Linux capabilities, uses the runtime default seccomp profile, and is pinned to an image digest.

7.3 ServiceAccounts

  • Create a dedicated ServiceAccount per application instead of using the default one.

  • Set automountServiceAccountToken to false for Pods that never call the Kubernetes API.

  • Bind only the specific permissions the application needs.

  • Use short-lived, audience-bound projected tokens, which are the default in current Kubernetes versions.

7.4 Secrets Handling

A Kubernetes Secret is base64-encoded, not encrypted, by default. Treat access to Secrets as access to credentials.

  • Enable encryption at rest (section 5.4) and restrict get, list and watch on Secrets through RBAC.

  • Prefer mounting Secrets as files over environment variables, since environment variables are easily leaked through logs, crash dumps and child processes.

  • Never commit plain Secret manifests to Git. Use sealed secrets, SOPS, or an external secrets manager synchronised into the cluster.

  • Rotate credentials regularly and after any suspected exposure.

7.5 NetworkPolicy

By default every Pod can reach every other Pod. NetworkPolicy lets you change that, but only if your CNI plugin enforces it. Start every namespace with a default-deny policy and then allow only the traffic that is required.

apiVersion: networking.k8s.io/v1

kind: NetworkPolicy

metadata:

  name: default-deny-all

  namespace: shop

spec:

  podSelector: {}

  policyTypes:

    - Ingress

    - Egress

---

apiVersion: networking.k8s.io/v1

kind: NetworkPolicy

metadata:

  name: allow-frontend-to-api

  namespace: shop

spec:

  podSelector:

    matchLabels:

      app: api

  policyTypes:

    - Ingress

  ingress:

    - from:

        - podSelector:

            matchLabels:

              app: frontend

      ports:

        - protocol: TCP

          port: 8080


Remember that a default-deny egress policy also blocks DNS. Add an explicit rule that allows egress to the cluster DNS service on port 53 (UDP and TCP), otherwise name resolution will fail. For the networking details, see the CNI and DNS troubleshooting articles in this series.

7.6 Resource Governance

Resource limits are also a security control, because they contain the impact of a runaway or malicious container. Use ResourceQuotas to cap what a namespace can consume and LimitRanges to set sensible defaults, as described in the earlier article on quotas and limit ranges.

8. Admission Control and Policy as Code

Admission controllers are the enforcement point that checks every create and update request. Pod Security Admission covers a fixed set of rules, but most organisations also need custom ones: only approved registries, mandatory labels, no latest tags, required resource limits and so on. There are two mainstream ways to do this:

  • ValidatingAdmissionPolicy, a built-in, CEL-based mechanism that needs no extra webhook server. It is generally available in current Kubernetes releases.

  • Policy engines such as Kyverno or OPA Gatekeeper, which add richer features like mutation, generation and reporting.

apiVersion: admissionregistration.k8s.io/v1

kind: ValidatingAdmissionPolicy

metadata:

  name: require-approved-registry

spec:

  failurePolicy: Fail

  matchConstraints:

    resourceRules:

      - apiGroups: [""]

        apiVersions: ["v1"]

        operations: ["CREATE", "UPDATE"]

        resources: ["pods"]

  validations:

    - expression: "object.spec.containers.all(c, c.image.startsWith('registry.example.com/'))"

      message: "Images must come from registry.example.com"

---

apiVersion: admissionregistration.k8s.io/v1

kind: ValidatingAdmissionPolicyBinding

metadata:

  name: require-approved-registry-binding

spec:

  policyName: require-approved-registry

  validationActions: ["Deny"]

  matchResources:

    namespaceSelector:

      matchLabels:

        env: production


Roll new policies out in Audit or Warn mode first, review what would have been blocked, and only then switch to Deny.

9. Layer 5: Image and Supply-Chain Security

Most real-world incidents start with something you deployed: a vulnerable library, a poisoned base image, or a compromised build pipeline. Securing the supply chain reduces that risk.

  • Use minimal base images (distroless or slim variants) to reduce the attack surface and the number of vulnerabilities.

  • Scan images in CI and continuously in the registry, since new vulnerabilities are discovered after an image ships. Trivy and Grype are common open-source scanners.

  • Pin by digest rather than by mutable tags, so what you tested is exactly what runs.

  • Sign images and verify signatures at admission time, for example with Sigstore Cosign together with a policy engine.

  • Generate and store SBOMs (Software Bills of Materials) so you can quickly find affected workloads when a new vulnerability is announced.

  • Use a private registry with access control, and restrict which registries the cluster may pull from.

  • Secure the CI/CD pipeline itself: least-privilege build credentials, protected branches, and no long-lived deploy keys.

10. Monitoring, Detection and Response

Prevention will eventually fail, so you also need to see what is happening and react.

Capability

What it gives you

Audit logs

Who changed what in the API, forwarded to a central log system

Runtime detection

Alerts on suspicious behaviour in containers, such as unexpected shells or file access; Falco is a widely used open-source option

Metrics and alerting

Detection of unusual resource spikes, crash loops and failed authentication

Vulnerability reporting

Continuous scanning results mapped to running workloads

Compliance scanning

Regular CIS benchmark checks, for example with kube-bench


Prepare an incident response plan before you need it. At a minimum, know how to isolate a compromised Pod (for example by applying a restrictive NetworkPolicy label), how to preserve evidence before deleting it, how to revoke or rotate credentials, and who is responsible for each step.

11. Upgrades and Patching

Kubernetes releases a new minor version several times a year and supports only a limited number of recent versions with security patches. Running an unsupported version means known vulnerabilities stay unfixed. Plan to upgrade regularly, test in a non-production cluster first, upgrade the control plane before the nodes, and keep node images, the container runtime and cluster add-ons (CNI, CSI, ingress controllers) patched as well. Check the official Kubernetes release page for the versions that are currently supported.

12. Hands-On Walkthrough: Hardening a Namespace

This walkthrough applies several controls to a new namespace called shop. Run it on a test cluster whose CNI supports NetworkPolicy.

  1. Create the namespace with Pod Security labels enforcing the Restricted profile (see section 7.1).

  2. Apply a default-deny NetworkPolicy for ingress and egress, then add an allow rule for DNS and for the traffic your application needs.

  3. Create a dedicated ServiceAccount and a minimal Role and RoleBinding.

  4. Apply a ResourceQuota and LimitRange to the namespace.

  5. Deploy the hardened Pod from section 7.2.

  6. Verify that an insecure Pod is rejected.

kubectl create namespace shop

kubectl label namespace shop \

  pod-security.kubernetes.io/enforce=restricted

 

# Try to run a privileged Pod. It should be rejected.

kubectl run bad --image=nginx -n shop \

  --overrides='{"spec":{"containers":[{"name":"bad","image":"nginx","securityContext":{"privileged":true}}]}}'

 

# Check what the application ServiceAccount may do

kubectl auth can-i get secrets \

  --as=system:serviceaccount:shop:orders-app -n shop


Expected result: the privileged Pod is rejected with a message that it violates the Restricted policy, and the can-i check returns "no" for anything you did not explicitly grant. The exact wording of the messages can differ between Kubernetes versions.

13. Troubleshooting Security Controls

Symptom

Likely cause

What to check or do

Pod rejected: violates PodSecurity

Pod does not meet the namespace profile

Read the message; set runAsNonRoot, drop capabilities, add seccompProfile

Pod CrashLoopBackOff after hardening

Read-only filesystem or non-root user cannot write

Mount an emptyDir for writable paths; fix file ownership with fsGroup

Forbidden error from the API

Missing RBAC permission

kubectl auth can-i; check Role and RoleBinding namespace and subject

DNS fails after default-deny

Egress policy blocks DNS

Allow egress to kube-dns on port 53 UDP and TCP

NetworkPolicy has no effect

CNI does not enforce policies

Confirm your CNI supports NetworkPolicy

Admission policy blocks a deployment

Policy match or expression too strict

Read the denial message; test in Audit mode; check namespace selectors

Image pull or verification fails

Unapproved registry or signature missing

Check the registry allow-list, imagePullSecrets and signature policy


14. Production Security Checklist

Use this checklist as a starting point and adapt it to your own risk profile and compliance requirements.

Area

Checks

Infrastructure

Private or restricted API endpoint; private node subnets; metadata service restricted; encrypted disks and backups

Identity and access

OIDC or equivalent for users; no shared admin credentials; least-privilege RBAC; regular access reviews

Control plane

Audit logging on and shipped off-cluster; etcd protected; Secrets encrypted at rest; CIS benchmark reviewed

Nodes

Minimal hardened OS; kubelet hardened; seccomp enabled; automated patching or node replacement

Workloads

Restricted Pod Security profile; non-root; read-only filesystem; dropped capabilities; dedicated ServiceAccounts

Network

Default-deny NetworkPolicy per namespace; explicit allow rules; DNS egress allowed

Supply chain

Scanned, signed, digest-pinned images; approved registries only; SBOMs stored

Governance

ResourceQuotas and LimitRanges; admission policies enforced; policies tested in Audit mode first

Operations

Runtime detection; alerting; tested incident response; regular upgrades on supported versions


15. Best Practices Summary

  • Layer your defences and assume that any single control can fail.

  • Apply least privilege everywhere: RBAC, ServiceAccounts, cloud IAM and container capabilities.

  • Make secure the default: enforce Pod Security Standards and admission policies so insecure configurations cannot be deployed.

  • Start with default-deny networking and open only what is needed.

  • Treat images as untrusted until scanned, signed and verified.

  • Log, monitor and rehearse your response to incidents.

  • Automate checks in CI and with benchmark scanners so security does not rely on memory.

  • Keep Kubernetes, nodes and add-ons patched and on supported versions.

  • Introduce new controls gradually using warn and audit modes before enforcing.

16. Conclusion

Securing Kubernetes in production is not one task but a habit of closing gaps layer by layer: a protected control plane, hardened nodes, confined workloads, a restricted network, trustworthy images and good visibility. The controls described here are individually simple, and the earlier articles in this series (RBAC, ServiceAccounts, Secrets, SecurityContext, Pod Security Standards, Admission Controllers, ResourceQuotas, and the networking articles) explain each in more depth. Start with the highest-impact items, which are least-privilege access, Pod Security enforcement, default-deny networking and image verification, then work through the checklist until it becomes part of how every cluster is built and operated.