Kubernetes Operators: The Hidden Risk of Excessive RBAC
Palo Alto Unit 42 finds that overprivileged Kubernetes operators can turn trusted automation into high-impact non-human identities.
Kubernetes operators are intended to reduce operational work. They watch custom resources, compare the declared state of an application or service with the state running in a cluster, and make changes through the Kubernetes API. That automation is useful, but it also creates a security dependency that is easy to overlook: the permissions assigned to the operator’s service account.
In research published by Palo Alto Unit 42, the organization argues that excessive operator permissions have become a significant source of risk in Kubernetes environments. Its analysis focuses on non-human identities, the permissions granted through role-based access control (RBAC), weaknesses in operator distribution channels, and the additional consequences of combining broad access with large language model (LLM)-enabled automation.
Unit 42 also released OperTraitor, an open-source, LLM-powered analysis engine that compares an operator’s documented purpose with the permissions present in its manifests. The goal is not to treat every operator as malicious, but to help defenders identify cases where an automated component can do substantially more than its stated function requires.
Why an operator compromise can spread
An operator is built around two Kubernetes mechanisms: a custom resource definition (CRD) and a controller. The CRD extends the Kubernetes API with an object representing a particular workload, policy or service. The controller runs continuously, observes that object and reconciles the cluster toward the desired state.
To perform that work, the controller uses a Kubernetes service account. That account is bound to one or more Roles or ClusterRoles, which determine what the controller can read or change. A namespace-scoped Role limits access to a particular namespace, while a ClusterRole can provide permissions across the cluster. ClusterRoleBindings can make those permissions available to the operator’s service account at cluster scope.
Unit 42’s central finding is that the consequences of an operator compromise are defined by this permission boundary. The initial compromise could result from a vulnerable dependency, a compromised container image or a hijacked node, according to the research. Whatever the entry point, an operator with broad RBAC can turn a localized foothold into access to unrelated workloads and administrative resources.
This is an identity problem as much as a software problem. The operator may be a trusted component and may not resemble an interactive user, but its service account can still read sensitive objects, alter cluster configuration or influence other workloads. A service account with more access than its function requires becomes a high-value target and a potential path for lateral movement within the cluster.
OperTraitor’s approach to operator risk
OperTraitor collects raw YAML manifests for operators installed locally or listed in OperatorHub. It extracts their RBAC settings and sends that information through an LLM configured for threat analysis. The engine then compares the effective permissions with the operator’s documented functionality and produces a normalized risk score from 1 to 10 based on the gap between required and granted privileges.
The score is an assessment aid, not proof that an operator has been compromised or that a particular permission is immediately exploitable. Security teams still need to validate the operator’s architecture, deployment model and operational requirements. Nevertheless, the approach offers a practical way to prioritize review, especially in environments containing many third-party controllers.
Unit 42 reported that slightly more than 5% of the operators it examined requested excessive privileges, including permissions that could provide an implicit route to cluster administrator-level access. The organization also said that many operator owners did not respond to responsible disclosure attempts, which in some cases may indicate that a component is no longer actively maintained. The research does not identify all of those operators in the article, so the broader finding should be treated as a warning to audit individual deployments rather than as a list of confirmed vulnerable products.
The registry and maintenance problem
The research highlights a second risk: the difference between what is easy to deploy and what is current or appropriately maintained. Unit 42 found older operator versions in OperatorHub that remained available through the Operator Lifecycle Manager (OLM), even when vendors had published newer versions through Helm charts, GitHub repositories or ArtifactHub.
That creates a practical trap for administrators. A component may be available through a familiar, integrated catalog while its more secure or better-maintained release is distributed elsewhere. The presence of an operator in a default registry should therefore not be treated as evidence that the component is current, actively maintained or minimally privileged.
For platform teams and managed service providers, this issue is particularly important because catalog-based deployment can make operator selection appear routine. A deployment workflow that does not check release age, vendor documentation, maintenance status and RBAC scope can normalize the use of stale or overprivileged software.
Case study: Prometurbo and cluster-wide secret access
One of Unit 42’s case studies involved IBM’s Prometurbo operator. The initial scan identified wildcard use in the OperatorHub version, which the researchers described as severely outdated. Unit 42 then examined a newer version available through IBM’s GitHub repository.
According to the research, that version bound the operator’s service account to a ClusterRole containing permission to get, list and watch Secret resources across the cluster. Unit 42 assessed that this was excessive for an operator that did not function as a centralized secrets manager. In a compromise scenario, such access could expose credentials, service account tokens, API keys and certificates stored in namespaces unrelated to the operator.
IBM worked with Unit 42 after disclosure and issued a subsequent release that scoped the permissions more narrowly, according to the report. IBM also published a security bulletin and assigned CVE-2026-6389, with a reported CVSS score of 8.8. Organizations using the affected operator should consult IBM’s security guidance and verify that their deployed version contains the relevant correction rather than relying on the catalog entry alone.
Case study: Datadog and the usability trade-off
Unit 42 also identified broad permissions in the Datadog operator, including cluster-wide access to Secrets and actions involving ClusterRoles and ClusterRoleBindings. The research does not present this as an undisclosed vulnerability equivalent to the Prometurbo case. Instead, it records Datadog’s explanation for the design.
Datadog told the researchers that the names of required Secrets are based on user-defined values and cannot be predicted before deployment. The vendor considered broader permissions necessary for its architecture and deployment experience. Datadog responded by documenting the RBAC settings and the mitigations it had applied, allowing customers to make an informed risk-acceptance decision.
This example illustrates why least privilege is sometimes difficult to implement, but it does not remove the need for scrutiny. A permission can be operationally justified and still create substantial impact if the associated identity is compromised. The appropriate response may be architectural separation, tighter namespace boundaries, additional monitoring or a deliberate decision to accept the residual risk.
Why agentic operators raise the stakes
Unit 42’s assessment is that the move toward agentic operators could amplify the consequences of existing RBAC weaknesses. The research describes two broad patterns. In one, an operator uses an LLM to enhance remediation or analysis. In the other, an operator acts as a bridge between the cluster and an external agent, or manages AI-agent lifecycles within Kubernetes.
If those components inherit broad permissions, an automated system may be able to read data or modify resources outside the function administrators intended. The research specifically warns that an LLM-enhanced operator with broad RBAC could access sensitive information across unintended namespaces, while an operator serving as an external-agent bridge could give that agent extensive control over cluster resources.
These are forward-looking risk assessments, not reports of a specific AI-driven compromise. Unit 42’s defensive conclusion is broader: whether an operator uses deterministic code or an LLM, the service account must be treated as a security boundary. AI does not change the need for least privilege; it can make mistakes or unexpected actions more consequential when permissions are excessive.
What organizations should do now
- Inventory non-human identities. Identify every operator, controller, service account, Role, ClusterRole and binding in each cluster. Include components installed by platform teams, application teams and MSPs.
- Review permissions against function. Compare each operator’s actual YAML manifests with vendor documentation. Pay particular attention to wildcard permissions, cluster-wide Secrets access and write actions on RBAC resources.
- Prefer namespace scope. Deploy operators within the namespaces they manage whenever the architecture permits it. Use cluster-scoped permissions only when they are demonstrably required and document the business reason.
- Validate the supply source. Check release age, maintenance status and vendor guidance before installing an operator from OLM or OperatorHub. Compare catalog versions with maintained Helm charts, ArtifactHub entries and official repositories.
- Use auditing tools. Unit 42’s OperTraitor can help prioritize reviews by comparing documented behavior with granted privileges. Its output should support, not replace, manual validation and change control.
- Monitor service-account behavior. Enable Kubernetes Audit Logs where appropriate. Establish a baseline for each operator and investigate unexpected requests, such as attempts to read Secrets in unrelated namespaces or API-server activity from an unusual source.
- Set boundaries for AI-enabled components. Restrict network access from LLM-enhanced operators and agent frameworks, prevent unnecessary connections to public or internal endpoints, and limit the data and permissions made available to the underlying model.
- Reassess after upgrades. Operator updates can change RBAC requirements. Recheck manifests after every version change and confirm that a remediation remains present in the running cluster.
Defenders should distinguish between confirmed configuration findings and assumptions about exploitation. The research identifies excessive permissions and explains how those permissions could increase impact after a compromise; it does not report that every highlighted operator was exploited. Similarly, the agentic-AI discussion describes an emerging risk model rather than a confirmed incident.
Conclusion
Kubernetes operators are valuable automation components, but their service accounts are privileged identities that deserve the same rigor applied to human administrators. Palo Alto Unit 42’s research shows how broad RBAC, stale catalog versions and uncertain maintenance can combine to expand the blast radius of a compromise.
The immediate priority is straightforward: inventory operators, verify their provenance, reduce permissions to the smallest workable scope and monitor what those identities actually do. That foundation is also essential for the next generation of AI-assisted cluster automation. Autonomous operations can reduce toil, but only when the identities behind them are constrained, observable and regularly reassessed.