
During 2026, two security advisories were published for etcd after we identified authorization bypasses affecting confidentiality and integrity. Both findings came from the same research workflow: graphify built a deterministic representation of the codebase, an LLM explored it and generated hypotheses, and those hypotheses were validated against live instances running the exact code under investigation.
The workflow combined graphify, LLM-assisted static analysis, environment provisioning and dynamic validation. The LLM made it possible to explore a codebase of etcd's size much more broadly than would have been practical through manual review alone. Previous experience with etcd helped us interpret the candidates, understand the surrounding security model and reject findings that looked plausible in isolation but did not hold up in context.
That balance became one of the most useful parts of the process. Manual analysis alone would have made this level of exploration expensive, while automated analysis without enough context would have produced far more false positives.
Summary of findings
CVE-2026-59818: CRL enforcement bypass on the gRPC listener
etcd allows operators to expose its HTTP and gRPC client endpoints through separate listeners using --listen-client-http-urls. Under this configuration, the Certificate Revocation List supplied through --client-crl-file was enforced on the HTTP listener while the gRPC listener continued accepting revoked certificates.
A client presenting the same revoked certificate could therefore be rejected by the HTTP endpoint and still authenticate successfully over gRPC.
Certificate revocation is commonly used after an identity has been compromised or should no longer be trusted. In this case, the control appeared to work correctly on one client listener while remaining ineffective on another carrying gRPC traffic.
The initial lead came from comparing the listener construction paths and observing that TLS verification settings were not propagated consistently. Dynamic testing then confirmed that the difference had a real security impact.
CVE-2026-73499: Watch API authorization bypass via open-ended range requests
The second issue involved the Watch API and etcd's RBAC enforcement. A user granted READ permission on a single exact key could open a watch using clientv3.WithFromKey(), creating an open-ended range from that key to the end of the keyspace.
Under the affected versions, events for unrelated keys inside that broader range could be delivered to the user even though the RBAC grant covered only the original key.
This candidate also emerged from comparing related execution paths. Range/Get and DeleteRange performed the expected RBAC range validation, while the Watch path behaved differently with open-ended ranges.
The issue was relevant only when authentication and RBAC were enabled. On an unauthenticated cluster, broad reads are already expected. In Kubernetes environments, however, etcd may contain Secrets, ServiceAccount tokens and other API objects, so an unexpectedly expanded READ permission can expose data outside the intended authorization scope.
Although the two vulnerabilities affected different parts of etcd, both shared the same structural pattern: related execution paths were expected to enforce similar security controls, yet one of them behaved differently.
Why we chose etcd
We chose etcd because we already had experience with the project and were familiar with its architecture, deployment model and security configuration.
That context was useful when the LLM started returning suspicious differences between execution paths. Some were genuine inconsistencies, while others were explained by the semantics of the operation, enforcement performed elsewhere or assumptions that did not match the actual security model.
The LLM was especially effective at increasing coverage. Once the code graph reduced the search space, it could compare implementations, trace call relationships and identify missing or inconsistent validation across many parts of the codebase. Reaching the same breadth through manual review would have required considerably more time.
The trade-off was volume. Many candidates looked reasonable at first sight and some came with convincing root-cause explanations, yet a significant number were discarded after deeper inspection or dynamic testing. Familiarity with etcd helped us decide which hypotheses deserved further work and which ones were artifacts of incomplete context.
Pipeline architecture
The workflow had five stages. Each one reduced the uncertainty passed to the next.

Stage 1: Building a deterministic code graph
Before starting the security analysis, we indexed the repository with graphify. The resulting graph captured functions, types, callers, callees and execution relationships for a specific commit.
This solved two practical problems.
First, it reduced context usage. Instead of feeding large parts of the repository into the model, the LLM could retrieve only the functions involved in a specific execution path. For a project the size of etcd, this made a substantial difference.
Second, it gave us deterministic information about code relationships. Function calls, reachability and caller sets came from the indexed source and could be verified independently. The LLM still had to interpret their security significance, but the underlying structure did not depend on model inference.
The graph also helped us enumerate attack surface more systematically. The LLM could identify handlers reachable from network-facing entry points, trace their interaction with persistent state and inspect where authentication, authorization and validation functions appeared across related paths.
Stage 2: Static analysis and candidate generation
We then used an LLM to analyse the relevant portions of the source and generate security hypotheses.
At this stage, a suspicious difference was enough to create a candidate. The goal was breadth: identify areas worth testing and defer stronger conclusions until later.
Traditional static analysis tools can also provide useful input at this stage. Results from tools such as Semgrep, CodeQL or SonarQube can be supplied to the LLM as additional signals from which to generate hypotheses. Their output does not need to be treated as a finding, or even assumed to be security relevant. An alert can instead become a starting point for tracing an execution path, comparing implementations or examining whether a broader security property is enforced consistently.
These tools are only one source of hypotheses. The LLM should also explore the indexed code independently and generate its own candidates. Otherwise, the search would remain bounded by the rules, queries and assumptions already encoded in the static analysis tools. Combining both approaches provides additional signals without restricting the investigation to vulnerability patterns an existing scanner already knows how to describe.
One pattern proved particularly productive: comparing sibling execution paths that appeared to share a security model and looking for differences in enforcement.
For the CRL issue, the LLM compared how TLS verification settings were propagated when HTTP and gRPC client endpoints were split across separate listeners.
For the Watch issue, the LLM compared operations accepting key ranges and traced how they reached the RBAC range-checking logic.
In both cases, static analysis identified a concrete discrepancy. It did not establish exploitability on its own. A missing check might be compensated by middleware, an interceptor, a caller or another part of the request lifecycle.
The output of this stage was therefore a queue of testable hypotheses.
Stage 3: Dynamic validation
Candidates that survived the initial review were reproduced against live etcd instances built from the same revision used to generate the graph.
For the CRL issue, the AI agent created a small PKI, issued and revoked a client certificate, generated the corresponding CRL and started etcd with separate HTTP and gRPC listeners. The same certificate was then tested against both endpoints.
The HTTP listener rejected it, while the gRPC listener accepted it. With the certificate, trust material, CRL and server remaining constant, the difference between listeners confirmed the security impact.
For the Watch issue, the AI agent enabled authentication, created an unprivileged user and granted it READ access to one exact key. The user opened a watch with WithFromKey(), while a privileged session modified unrelated keys later in the keyspace.
Those events appeared in the unprivileged client's stream, showing that the effective watch range exceeded the configured authorization grant.
Dynamic testing also eliminated many candidates. Some behaved exactly as predicted but stayed within the intended security model, while others disappeared once the full configuration and request flow were reproduced.
Stage 4: Confirmation and rejection
After reproduction, candidates passed through a stricter confirmation gate.

The behaviour had to reproduce from a clean environment, more than once, with the relevant configuration recorded. We also checked that the observed result crossed a real security boundary under realistic conditions.
This distinction mattered in cases such as the Watch issue, where broad reads only represented an authorization bypass when RBAC was enabled and the user had been granted narrower access.
The LLM also compared candidates against existing advisories and traced the runtime behaviour back to the source code.
This last step was important because the first explanation generated by the LLM was not always the correct one. In some cases, the behaviour reproduced consistently while deeper tracing showed that another function or layer better explained the result.
We therefore considered a candidate mature only when the runtime evidence and source-level explanation converged.
Persisting negative results
Rejected candidates were stored together with the reason for rejection, and the corresponding research path was marked as explored.
This prevented later iterations from rediscovering the same suspicious code after context resets. It also created a useful record of which areas had already been investigated and why specific hypotheses had failed.
Over time, these negative results became useful for understanding the model's most common failure modes.
Where false positives appeared
The most frequent false positives came from intentional behaviour that initially looked like inconsistent enforcement. Related operations often have different semantics, and similar code paths do not always imply identical security requirements.
Compensating controls were another common case. A handler could appear to lack a check that existed elsewhere, while broader inspection showed that enforcement occurred in an interceptor, wrapper or earlier stage of the request.
We also found environment-related artifacts, where the configuration used during testing changed the assumptions behind the candidate.
The most interesting cases were those where the unexpected behaviour reproduced but the initial root-cause explanation was wrong. These were harder to identify because the evidence looked strong at first: the commands worked, the result was consistent and the model produced a plausible explanation.
This is where context mattered most. Previous knowledge of etcd made it easier to notice when an explanation did not fit the surrounding architecture and to keep tracing until the source and runtime behaviour matched.
The LLM still played a central role. Manually exploring every promising authentication, TLS and authorization path across the project would have been far more expensive. The useful pattern was broad automated exploration followed by stricter contextual and dynamic validation.
Stage 5: Report generation
Only candidates that survived confirmation produced disclosure artifacts.
At that point the LLM assembled the technical description, reproduction scripts, captured evidence and code references supporting the root cause.
Keeping report generation at the end also helped avoid a common problem with LLM-assisted research: a weak hypothesis can be turned into a convincing report long before it has been properly validated.
By delaying that step, the final report reflected conclusions that had already been reproduced and understood.
What we learned
The two etcd findings showed that deterministic program analysis and LLM-assisted reasoning work well together when their roles are clearly separated.
The graph provided a precise representation of execution relationships and reduced the amount of source that had to be passed to the model. The LLM made it practical to compare many paths and identify structural inconsistencies across a large codebase. Dynamic testing then determined which of those inconsistencies had real security consequences.
Structural asymmetry was particularly productive. Both findings emerged from related paths behaving differently around a security control, and that pattern could be searched systematically across listeners, RPC methods and operations sharing an authorization model.
Persisting negative results also proved important because it prevented the research loop from repeatedly exploring hypotheses that had already been disproved.
The overall process depended on both scale and context. The LLM allowed us to inspect much more of etcd than would have been realistic through manual review alone, while our knowledge of the project helped reduce a large candidate set to the small number of behaviours that actually represented vulnerabilities.
Conclusion
This research showed that LLM-assisted vulnerability research can significantly reduce the cost of exploring large codebases when the model operates over a structured representation of the software and its hypotheses are validated against real systems.
graphify gave us deterministic execution relationships, the LLM expanded the amount of code we could investigate, and dynamic testing connected promising hypotheses to observable behaviour in etcd.
Previous experience with the project helped us interpret that output and reject candidates that were technically plausible but incorrect in context. That was important because the same pipeline that surfaced the two vulnerabilities also generated many hypotheses that did not survive validation.
The resulting workflow was simple in principle: explore broadly, validate aggressively, and keep only the findings whose source-level explanation and runtime behaviour agree.