Kubernetes Security Best Practices for Production
Kubernetes Security Best Practices for Production
bitcodematrix.com | Kubernetes Series
1. Introduction
Kubernetes is secure only to the degree that it is configured to be. Out of the box it favours flexibility: Pods can talk to every other Pod, containers often run as root, Secrets are only base64-encoded, and a mis-scoped ServiceAccount can quietly become a path to cluster-wide access. Production security is therefore a collection of deliberate choices made at every layer, not a single feature that you switch on.
This article pulls together the controls covered individually in earlier articles of this series into one practical, prioritised view. It explains the defence-in-depth model, walks through each layer, gives working examples, and ends with a production readiness checklist you can adapt to your own environment.
After reading it you should be able to:
Explain why Kubernetes security must be layered and name the layers.
Harden the API server, authentication, authorization, audit logging and etcd.
Lock down Pods with Pod Security Standards, securityContext, ServiceAccount settings and NetworkPolicy.
Reduce supply-chain risk with image scanning, signing and admission policies.
Apply a repeatable checklist and troubleshoot common security-related failures.
2. The Defence-in-Depth Model
No single control is perfect, so security is built in layers. If an attacker gets past one layer, the next one limits the damage. A common way to organise Kubernetes security is the "4C" idea (Cloud, Cluster, Container, Code), which this article expands into five layers for clarity.
Figure 1: The five layers of Kubernetes security
An important rule follows from the figure: a weakness in an outer layer undermines everything inside it. A perfectly hardened Pod does not help if the API server is reachable from the internet with weak credentials, and a locked-down cluster does not help if the image contains a known remote-code-execution flaw.
3. Threat Model: What Are You Defending Against?
Controls make more sense when tied to concrete threats. The table below lists common production threats and the primary defences used against them.
4. Layer 1: Cloud, Network and Infrastructure
Everything in Kubernetes runs on infrastructure you must also protect. Whether you use a managed service or run your own clusters, apply these basics first.
Keep the API server private where possible. Use a private endpoint or restrict public access to known IP ranges. Never expose it openly to the internet.
Isolate nodes. Place nodes in private subnets, restrict inbound traffic with security groups or firewalls, and allow only the ports that are required.
Apply least privilege to cloud IAM. Node roles should not carry broad permissions. Use workload identity features (for example IAM roles for ServiceAccounts or equivalent) instead of long-lived cloud keys inside Pods.
Protect the metadata service. Block or restrict Pod access to the cloud instance metadata endpoint, because it can expose node credentials.
Encrypt storage and backups. Enable encryption for disks, snapshots and etcd backups.
Separate environments. Use separate clusters or at least separate accounts or projects for production and non-production.
5. Layer 2: Securing the Control Plane
The control plane decides what runs in the cluster, so compromising it means compromising everything. See the architecture article in this series for how the components fit together.
5.1 Authentication
Every request to the API server must be authenticated. Prefer short-lived identities issued by an external identity provider using OpenID Connect (OIDC) over static tokens or shared client certificates. Disable anonymous access unless you have a specific, understood need for it, and avoid long-lived credentials for people. The separate article on authentication versus authorization covers the mechanics in detail.
5.2 Authorization with RBAC
RBAC should follow least privilege. Grant access to named resources and verbs rather than wildcards, prefer namespaced Roles over ClusterRoles, and review bindings regularly.
Avoid binding the cluster-admin ClusterRole to users or groups for daily work.
Avoid wildcards such as resources: [""] or verbs: [""].
Treat permissions to create Pods, read Secrets, bind roles, and impersonate as highly sensitive, because each can be used to escalate privileges.
Audit regularly using kubectl auth can-i and tools that list effective permissions.
# What can this ServiceAccount do?
kubectl auth can-i --list \
--as=system:serviceaccount:shop:orders-app -n shop
5.3 Audit Logging
Audit logs record who did what and when against the API server. Without them, incident investigation is guesswork. Enable an audit policy that captures metadata for most requests and request or response bodies only for sensitive resources, then ship the logs to a central, tamper-resistant system. On managed services, enable the provider's control plane logging option.
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
- level: None
users: ["system:kube-proxy"]
verbs: ["watch"]
- level: Metadata
resources:
- group: ""
resources: ["secrets", "configmaps"]
- level: RequestResponse
resources:
- group: "rbac.authorization.k8s.io"
- level: Metadata
Note that Secrets are logged at the Metadata level above so that secret values never appear in the logs.
5.4 etcd and Encryption at Rest
etcd stores all cluster state, including Secrets. Protect it by allowing only the API server to reach it, requiring mutual TLS, encrypting backups, and enabling encryption at rest for Secrets through an EncryptionConfiguration. Where available, prefer a KMS provider so the encryption keys are held outside the cluster.
apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
- resources: ["secrets"]
providers:
- aescbc:
keys:
- name: key1
secret: <base64-encoded-32-byte-key>
- identity: {}
The example uses a local key for illustration only. In production, a KMS-based provider is generally preferred, and the key must never be committed to version control. On managed Kubernetes services, check how the provider handles encryption at rest and whether you can supply your own key.
5.5 API Server Hardening
Keep RBAC and the Node authorizer enabled in the authorization mode list.
Keep the recommended admission plugins enabled and add validating policies (see section 8).
Use TLS everywhere, including between the API server, etcd and kubelets.
Restrict network access to the control plane and the kubelet API.
Compare your configuration against the CIS Kubernetes Benchmark and run an automated scanner such as kube-bench regularly.
6. Layer 3: Securing Nodes and the Runtime
Nodes run your containers, and a container escape lands the attacker on a node. Hardening nodes reduces how far an escape can go.
Use a minimal, container-optimised node OS with few packages, automatic security updates and a read-only root filesystem where supported.
Do not allow SSH by default. Use short-lived, audited access when it is required.
Harden the kubelet. Disable anonymous authentication, use Webhook authorization, and do not expose the read-only port.
Use a supported container runtime such as containerd or CRI-O, kept up to date.
Consider sandboxed runtimes (for example gVisor or Kata Containers, via RuntimeClass) for untrusted or multi-tenant workloads.
Enable seccomp, AppArmor or SELinux profiles so containers can use only the system calls and resources they need.
Replace nodes rather than patching them in place where possible, so nodes are consistent and short-lived.
# Example kubelet configuration fragment
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
authentication:
anonymous:
enabled: false
webhook:
enabled: true
authorization:
mode: Webhook
readOnlyPort: 0
protectKernelDefaults: true
7. Layer 4: Securing Workloads
7.1 Pod Security Standards
Pod Security Standards define three profiles: Privileged, Baseline and Restricted. The built-in Pod Security Admission controller enforces them per namespace through labels. For most application namespaces in production, target the Restricted profile. The dedicated article in this series explains each profile.
apiVersion: v1
kind: Namespace
metadata:
name: shop
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/audit: restricted
pod-security.kubernetes.io/warn: restricted
If you are introducing this to an existing cluster, start with the warn and audit modes to find violations before you switch on enforce.
7.2 securityContext
The securityContext is where you tell the runtime how tightly to confine a container. A solid production baseline looks like this:
apiVersion: v1
kind: Pod
metadata:
name: secure-app
namespace: shop
spec:
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
runAsUser: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: app
image: registry.example.com/shop/app@sha256:<digest>
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
memory: 256Mi
volumeMounts:
- name: tmp
mountPath: /tmp
volumes:
- name: tmp
emptyDir: {}
Key points about this example: the container runs as a non-root user, cannot gain extra privileges, has a read-only root filesystem with a writable /tmp provided by an emptyDir volume, drops all Linux capabilities, uses the runtime default seccomp profile, and is pinned to an image digest.
7.3 ServiceAccounts
Create a dedicated ServiceAccount per application instead of using the default one.
Set automountServiceAccountToken to false for Pods that never call the Kubernetes API.
Bind only the specific permissions the application needs.
Use short-lived, audience-bound projected tokens, which are the default in current Kubernetes versions.
7.4 Secrets Handling
A Kubernetes Secret is base64-encoded, not encrypted, by default. Treat access to Secrets as access to credentials.
Enable encryption at rest (section 5.4) and restrict get, list and watch on Secrets through RBAC.
Prefer mounting Secrets as files over environment variables, since environment variables are easily leaked through logs, crash dumps and child processes.
Never commit plain Secret manifests to Git. Use sealed secrets, SOPS, or an external secrets manager synchronised into the cluster.
Rotate credentials regularly and after any suspected exposure.
7.5 NetworkPolicy
By default every Pod can reach every other Pod. NetworkPolicy lets you change that, but only if your CNI plugin enforces it. Start every namespace with a default-deny policy and then allow only the traffic that is required.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: shop
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-frontend-to-api
namespace: shop
spec:
podSelector:
matchLabels:
app: api
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: frontend
ports:
- protocol: TCP
port: 8080
Remember that a default-deny egress policy also blocks DNS. Add an explicit rule that allows egress to the cluster DNS service on port 53 (UDP and TCP), otherwise name resolution will fail. For the networking details, see the CNI and DNS troubleshooting articles in this series.
7.6 Resource Governance
Resource limits are also a security control, because they contain the impact of a runaway or malicious container. Use ResourceQuotas to cap what a namespace can consume and LimitRanges to set sensible defaults, as described in the earlier article on quotas and limit ranges.
8. Admission Control and Policy as Code
Admission controllers are the enforcement point that checks every create and update request. Pod Security Admission covers a fixed set of rules, but most organisations also need custom ones: only approved registries, mandatory labels, no latest tags, required resource limits and so on. There are two mainstream ways to do this:
ValidatingAdmissionPolicy, a built-in, CEL-based mechanism that needs no extra webhook server. It is generally available in current Kubernetes releases.
Policy engines such as Kyverno or OPA Gatekeeper, which add richer features like mutation, generation and reporting.
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: require-approved-registry
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["pods"]
validations:
- expression: "object.spec.containers.all(c, c.image.startsWith('registry.example.com/'))"
message: "Images must come from registry.example.com"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: require-approved-registry-binding
spec:
policyName: require-approved-registry
validationActions: ["Deny"]
matchResources:
namespaceSelector:
matchLabels:
env: production
Roll new policies out in Audit or Warn mode first, review what would have been blocked, and only then switch to Deny.
9. Layer 5: Image and Supply-Chain Security
Most real-world incidents start with something you deployed: a vulnerable library, a poisoned base image, or a compromised build pipeline. Securing the supply chain reduces that risk.
Use minimal base images (distroless or slim variants) to reduce the attack surface and the number of vulnerabilities.
Scan images in CI and continuously in the registry, since new vulnerabilities are discovered after an image ships. Trivy and Grype are common open-source scanners.
Pin by digest rather than by mutable tags, so what you tested is exactly what runs.
Sign images and verify signatures at admission time, for example with Sigstore Cosign together with a policy engine.
Generate and store SBOMs (Software Bills of Materials) so you can quickly find affected workloads when a new vulnerability is announced.
Use a private registry with access control, and restrict which registries the cluster may pull from.
Secure the CI/CD pipeline itself: least-privilege build credentials, protected branches, and no long-lived deploy keys.
10. Monitoring, Detection and Response
Prevention will eventually fail, so you also need to see what is happening and react.
Prepare an incident response plan before you need it. At a minimum, know how to isolate a compromised Pod (for example by applying a restrictive NetworkPolicy label), how to preserve evidence before deleting it, how to revoke or rotate credentials, and who is responsible for each step.
11. Upgrades and Patching
Kubernetes releases a new minor version several times a year and supports only a limited number of recent versions with security patches. Running an unsupported version means known vulnerabilities stay unfixed. Plan to upgrade regularly, test in a non-production cluster first, upgrade the control plane before the nodes, and keep node images, the container runtime and cluster add-ons (CNI, CSI, ingress controllers) patched as well. Check the official Kubernetes release page for the versions that are currently supported.
12. Hands-On Walkthrough: Hardening a Namespace
This walkthrough applies several controls to a new namespace called shop. Run it on a test cluster whose CNI supports NetworkPolicy.
Create the namespace with Pod Security labels enforcing the Restricted profile (see section 7.1).
Apply a default-deny NetworkPolicy for ingress and egress, then add an allow rule for DNS and for the traffic your application needs.
Create a dedicated ServiceAccount and a minimal Role and RoleBinding.
Apply a ResourceQuota and LimitRange to the namespace.
Deploy the hardened Pod from section 7.2.
Verify that an insecure Pod is rejected.
kubectl create namespace shop
kubectl label namespace shop \
pod-security.kubernetes.io/enforce=restricted
# Try to run a privileged Pod. It should be rejected.
kubectl run bad --image=nginx -n shop \
--overrides='{"spec":{"containers":[{"name":"bad","image":"nginx","securityContext":{"privileged":true}}]}}'
# Check what the application ServiceAccount may do
kubectl auth can-i get secrets \
--as=system:serviceaccount:shop:orders-app -n shop
Expected result: the privileged Pod is rejected with a message that it violates the Restricted policy, and the can-i check returns "no" for anything you did not explicitly grant. The exact wording of the messages can differ between Kubernetes versions.
13. Troubleshooting Security Controls
14. Production Security Checklist
Use this checklist as a starting point and adapt it to your own risk profile and compliance requirements.
15. Best Practices Summary
Layer your defences and assume that any single control can fail.
Apply least privilege everywhere: RBAC, ServiceAccounts, cloud IAM and container capabilities.
Make secure the default: enforce Pod Security Standards and admission policies so insecure configurations cannot be deployed.
Start with default-deny networking and open only what is needed.
Treat images as untrusted until scanned, signed and verified.
Log, monitor and rehearse your response to incidents.
Automate checks in CI and with benchmark scanners so security does not rely on memory.
Keep Kubernetes, nodes and add-ons patched and on supported versions.
Introduce new controls gradually using warn and audit modes before enforcing.
16. Conclusion
Securing Kubernetes in production is not one task but a habit of closing gaps layer by layer: a protected control plane, hardened nodes, confined workloads, a restricted network, trustworthy images and good visibility. The controls described here are individually simple, and the earlier articles in this series (RBAC, ServiceAccounts, Secrets, SecurityContext, Pod Security Standards, Admission Controllers, ResourceQuotas, and the networking articles) explain each in more depth. Start with the highest-impact items, which are least-privilege access, Pod Security enforcement, default-deny networking and image verification, then work through the checklist until it becomes part of how every cluster is built and operated.