The service mesh adoption decision record: why the proxy deployment model you chose determines your mTLS enforcement gap and your lateral movement detection failure mode

The proxy deployment model — PERMISSIVE peer authentication mode that is never migrated to STRICT because legacy services remain unmeshed through multiple planning cycles, a global retry policy configured from a getting-started guide without per-service annotations for non-idempotent or resource-exhaustion-bounded endpoints, and sidecar proxy coverage that correctly handles HTTP and gRPC service-to-service traffic while database connections and message queue consumers run outside the proxy's traffic interception and outside the mesh's telemetry — are service mesh adoption decisions that are almost never made explicitly. They emerge from an Istio installation that sets PERMISSIVE mode for gradual migration and logs the STRICT migration as a backlog item that stays on the backlog for the life of the cluster; a traffic policy copied from the official getting-started guide that applies globally because per-service customization was not part of the initial setup scope; and a mesh deployment roadmap that defines the initial scope as service-to-service HTTP observability, where the non-HTTP traffic that represents the system's data persistence layer was outside scope at deployment and never revisited. Three failure patterns: the developer tools company that deployed Istio for zero-trust networking but left PERMISSIVE mode enabled across all namespaces for 14 months, allowing a compromised CI agent to send plaintext HTTP requests to internal Kubernetes services that the PeerAuthentication policy was explicitly configured to accept; the B2B SaaS company that configured a global Envoy retry policy without resource-exhaustion awareness, turning a billing service's database connection pool exhaustion into an 800-request-per-second retry amplification cycle that restarted the billing service 11 times in 47 minutes before the retry policy could be manually disabled; and the data infrastructure SaaS that adopted Linkerd for golden-signal SLO observability but whose Kafka consumers and direct PostgreSQL connections were outside the sidecar proxy's visibility, making a 4.5-hour silent data corruption event — 47,000 records dropped by a failing event consumer — invisible to the service mesh telemetry while the SLO dashboard reported 99.94% success rate.

A 37-person developer tools company built a platform for CI/CD pipeline observability — real-time build metrics, test flakiness tracking, and deployment frequency analytics for engineering teams. Their backend ran as 14 microservices on Kubernetes, and their engineers had been discussing zero-trust networking for two quarters before they adopted Istio. The primary motivation was mutual TLS between services: their security engineer had identified that pod-to-pod traffic was unencrypted and that any compromised pod in the cluster could send HTTP requests to any other pod without authentication. The Istio installation followed the getting-started guide: they installed the control plane, configured automatic sidecar injection for their two primary namespaces, and left the PeerAuthentication policy at its default — PERMISSIVE mode, which instructed the Envoy sidecar proxies to accept both mTLS-encrypted connections from within the mesh and plaintext connections from any source that lacked a sidecar.

The PERMISSIVE mode was documented in their internal runbook as a temporary state: "PeerAuthentication is currently PERMISSIVE to allow gradual migration. After all legacy services have sidecars injected, migrate to STRICT mode." The legacy services in question were a PHP API gateway — a monolith that predated the microservice split — and a Python batch job runner that processed data exports for customers. The PHP gateway was excluded from sidecar injection because the team's version had a thread-model incompatible with the sidecar's connection interception, and the batch runner's maintainer had raised concerns about proxy overhead on long-running tasks. The STRICT mode migration was tracked in a JIRA ticket that was moved from sprint to sprint for 14 months without resolution. The Istio dashboard showed 94% of service-to-service connections as mTLS — which appeared healthy — but the 6% plaintext traffic was not a rounding error; it was the unmeshed services and any external source that had network access to production service addresses.

In month fourteen, a dependency in their CI pipeline was compromised. The company used a third-party GitHub Action for build artifact caching; the action's upstream repository was taken over through a credential-stuffing attack on the maintainer account, and a new patch version was published that exfiltrated environment variables and opened a reverse shell from the CI runner. Their CI runners ran on self-hosted Kubernetes nodes in the same cluster and VPC as their production workloads — a cost optimization that had never been assessed for its security implications. The compromised runner pod ran in the ci namespace, which did not have sidecar injection enabled.

The attacker used the compromised runner to enumerate internal services. Kubernetes ClusterIP services in the production namespace were reachable from the CI namespace: the team had relied on the Istio mTLS enforcement to serve as the zero-trust boundary and had not configured NetworkPolicies to restrict cross-namespace traffic. The attacker sent plaintext HTTP requests from the CI runner pod to the production services. Each service had an Envoy sidecar configured in PERMISSIVE mode — the sidecar accepted the plaintext connection from the CI pod, which lacked a sidecar and therefore lacked a SPIFFE service identity certificate, without complaint. The PERMISSIVE mode policy was working exactly as configured: it accepted plaintext connections from sources outside the mesh. The attacker reached the internal admin API — which had no authentication because it was designated "internal-only" and its designers had expected the mTLS enforcement to serve as the authorization gate — and exfiltrated 3,200 customer records including API keys, build configuration environment variables, and repository access tokens that customers had stored as CI secrets within the platform.

The post-incident analysis produced a single clarifying sentence: PERMISSIVE mode is not zero-trust networking. PERMISSIVE mode is zero-trust networking for connections from within the mesh combined with no-trust networking for all other connections. A mesh that is never fully migrated to STRICT mode provides the operational overhead of mTLS — sidecar resource consumption, certificate rotation, proxy configuration complexity, and troubleshooting latency — without the security guarantee that motivated the adoption. The migration from PERMISSIVE to STRICT that was completed 72 hours after containment took 2 hours; the obstacle that had deferred it for 14 months — the PHP gateway and the batch runner — was resolved using a namespace-scoped PeerAuthentication policy in STRICT mode for the two production namespaces, leaving the PHP gateway and batch runner's namespace in PERMISSIVE. This configuration option had been available in Istio since version 1.5 and was not known to the team that had set the initial cluster-wide PERMISSIVE policy. The two namespaces that could be fully meshed received STRICT enforcement; the legacy namespace remained in PERMISSIVE with a tracking item for migration that now had a deadline and a named owner. The combination of STRICT enforcement for the production namespaces and a NetworkPolicy that restricted which namespaces could initiate connections to those namespaces closed the attack path that the compromised CI agent had used.

A 43-person B2B SaaS company built an API usage analytics platform — request volume tracking, quota enforcement, and cost attribution for API-first companies. Their services were organized into three tiers: an ingestion tier that received events from customers' API gateways, a processing tier that aggregated and stored events, and a query tier that served dashboards and export APIs. They adopted Istio for reliability: their primary motivation was traffic management — circuit breakers, retries, and timeouts — to prevent cascading failures when a dependency degraded. After the initial Istio installation, the platform engineer applied a global retry policy via a VirtualService that matched all services in the production namespace. The policy was copied from an Istio blog post demonstrating retry configuration for a stateless read endpoint:

retries:
  attempts: 3
  perTryTimeout: 5s
  retryOn: "5xx,gateway-error,reset,connect-failure,retriable-4xx"

The blog post's context was a read endpoint for a product catalog service — idempotent, stateless, safe to retry on any 5xx. The team's context included billing service endpoints that were neither idempotent nor stateless: charge event recording (which wrote a row to PostgreSQL on every request), quota decrement (which decremented a counter in a shared table), and subscription status updates (which updated a row). These endpoints returned HTTP 503 when the billing service's database connection pool was exhausted — not because of a transient upstream fault, but because the service itself had no capacity to handle requests while the pool was at zero. The retryOn: "5xx" condition applied to these 503 responses exactly as it applied to any other 5xx: retry 3 times with a 5-second per-try timeout.

In month seven, a deployment of the billing service introduced a PostgreSQL connection that was opened in the request handler but never closed — a connection leak. The billing service's pool was sized at 100 connections. At 200 requests per second to the billing tier, the pool was fully exhausted 30 seconds after deployment. Every subsequent request returned HTTP 503 within 100 milliseconds, which was the pool acquisition timeout configured in the PostgreSQL client. Eight upstream services called the billing tier: the four ingestion services, the aggregation service, the quota enforcement service, the export service, and the admin service. All eight were subject to the global retry policy.

Each 503 received by any of these callers triggered 3 retries with a 5-second per-try timeout. For a 503 returned in under 100 milliseconds, each original request generated 4 total attempts before Envoy surfaced the error to the upstream. At 200 incoming requests per second to the billing tier, the effective requests-per-second reaching the billing service's connection pool was 200 × 4 = 800. The billing service's pool remained at 0 available connections. Each of the 800 attempts per second waited 100 milliseconds for a connection, received a pool acquisition timeout, and returned 503 — triggering another retry from Envoy. The 800-request-per-second load against the already-exhausted pool consumed the billing service's CPU on connection acquisition retry logic and memory on Envoy's retry bookkeeping. After 4 minutes, the billing service's memory usage exceeded its Kubernetes resource limit of 512 MiB and the pod was OOMKilled.

Kubernetes restarted the billing pod automatically. The new pod opened a fresh connection pool with 100 available connections. The queued and in-flight retry requests from 8 upstream services — Envoy's 15-second retry windows (3 retries × 5 seconds each) were still running for requests that had started before the pod restart — immediately consumed all 100 fresh connections. The pool exhausted again within 45 seconds. The billing service OOMKilled a second time 4 minutes after the restart. The cycle repeated 11 times over 47 minutes. Containment required a platform engineer to manually remove the VirtualService retry policy, which caused Envoy to surface the 503 directly to callers without retrying — reducing effective load from 800 requests per second to 200. The billing service still could not serve 200 requests per second with an empty pool, but the reduced load allowed the pod to remain alive long enough for the team to deploy the connection leak fix. The post-incident change replaced the global policy with per-service VirtualServices: non-idempotent billing endpoints received attempts: 0 (no retry); idempotent read endpoints received retryOn: "connect-failure,reset" only (transient connectivity failures); and a DestinationRule outlierDetection circuit breaker was configured for all services to open after 5 consecutive 503 responses, pausing retries to an ejected host for 30 seconds. Connect this failure to the timeout and deadline propagation decision record: the retry policy and the timeout policy share the same structural gap — both are configured globally from a getting-started guide without per-service customization, and both amplify failures when the per-attempt timeout does not account for the resource exhaustion dynamic at the destination; a per-try timeout of 5 seconds against a pool acquisition timeout of 100 milliseconds means Envoy has 49 additional seconds of retry capacity for each original request that it can apply against an already-failed service, producing an amplification cascade that the timeout was designed to prevent but instead enables.

A 54-person B2B SaaS company built a data infrastructure platform — schema management, pipeline orchestration, and data quality monitoring for engineering teams. Their backend ran as 22 services across three Kubernetes clusters. Their event-driven architecture used Apache Kafka for async event delivery between the pipeline orchestration tier and downstream consumers: the schema registry, the data quality monitoring service, and the export service each ran a dedicated Kafka consumer process that read events from topic partitions and wrote derived records to PostgreSQL. They adopted Linkerd after three incidents in six months where service degradation had been detected by customer reports rather than by internal monitoring. Linkerd's automatic golden-signal metrics — success rate, request latency, and request throughput for every HTTP and gRPC service-to-service connection — were the solution: the SRE team built SLO dashboards using Linkerd's Prometheus metrics, defined error budget burn-rate alerts, and presented the new observability posture at the next quarterly engineering review.

What the dashboards showed was correct and useful: the HTTP success rate for every meshed service was visible, the latency percentiles were tracked, and the burn-rate alerts fired correctly when HTTP error rates exceeded the SLO threshold. What the dashboards did not show was the data that moved through Kafka and PostgreSQL. Linkerd's sidecar proxy intercepts TCP connections via iptables and identifies the application protocol by reading the initial bytes of the connection. For HTTP and gRPC connections, the proxy decodes the protocol and records per-request golden signals. For connections using PostgreSQL's wire protocol or Kafka's binary protocol, the proxy identifies that the initial bytes do not match HTTP or gRPC framing, treats the connection as opaque TCP, and forwards bytes without decoding them. The Linkerd telemetry for opaque TCP connections reports bytes transferred per connection — not per-operation success rates, not error counts, not latency distributions. A Kafka consumer that reads messages and writes to PostgreSQL generates two opaque TCP connections per processing loop; whether the reads and writes succeed or fail is invisible to the Linkerd success rate metric.

In month four, a schema change was deployed to the PostgreSQL database used by the event consumer service. A column storing serialized event metadata was migrated from TEXT to JSONB. The migration converted all existing rows successfully. The event consumer service — a Go binary that consumed from Kafka and wrote to PostgreSQL using a struct with a string field for the metadata column — was not updated to produce JSONB-compatible output. Its INSERT statements passed a string value for a JSONB column; PostgreSQL rejected each INSERT with a type mismatch error. The consumer's error handling on INSERT failures incremented the Kafka consumer offset (marking the message as processed), logged a warning-level message to the container's stdout, and continued to the next message. The consumer did not update a health endpoint, did not increment a Prometheus counter for INSERT errors, and did not return a non-zero exit code. From Kubernetes' perspective, the pod was healthy — its liveness probe, an HTTP GET to /healthz, returned 200 throughout. From Linkerd's perspective, the consumer was healthy — there were no HTTP connections from or to the consumer pod for the proxy to observe.

The event consumer processed 47,000 Kafka messages over 4 hours and 17 minutes, incrementing the offset for each, inserting none, logging a warning for each, and passing every liveness probe check. The SLO dashboard showed 99.94% HTTP success rate across all services for the entire period. The failure was detected by a customer who noticed their data quality reports had stopped updating and submitted a support ticket. The support engineer identified the 4-hour gap in the report data within 20 minutes of escalation by searching the event consumer's container logs — the INSERT failure warnings were present at high frequency and immediately identifiable. The recovery required: a rollback of the schema migration, a restart of the event consumer with the reverted schema, a Kafka consumer group offset reset to re-process the 47,000 dropped messages, and a 3-day backfill of the downstream quality metrics reports, because each quality metric was a rolling aggregation over the event stream — reprocessing the events required recomputing the aggregations for all downstream reports that had ingested the gap period. Connect this failure to the observability strategy decision record: the Linkerd deployment improved observability for the HTTP traffic layer without creating observability for the data persistence layer — the layer where the actual system state was materialized; an SLO defined only in terms of HTTP success rate is not a measure of whether the system is correctly processing data; it is a measure of whether the HTTP endpoints are returning 2xx responses; a system that returns 200 OK while discarding every write it claims to have accepted satisfies the HTTP-success-rate SLO perfectly while failing its actual purpose entirely; the observability gap is not a Linkerd limitation but a strategy gap: the team's observability investment was in mesh telemetry for HTTP, and the non-HTTP data flow that constituted the core of the system's value — events consumed, records written, quality metrics computed — had no instrumentation at any layer.

Structural properties set by the service mesh adoption decision

Three structural properties are determined when a team decides — or fails to explicitly decide — how to deploy a service mesh and what its enforcement and coverage scope should be: what the mTLS enforcement mode determines about the lateral movement surface when a pod inside the cluster is compromised, what the retry policy model determines about the resource exhaustion amplification surface when a dependency fails under load, and what the mesh coverage boundary determines about the non-HTTP observability blind spot when data flows through protocols the proxy cannot decode. None of these properties are labeled as decisions in the conversations that produce them. The enforcement mode emerges from an installation that sets PERMISSIVE as a migration aid and never completes the migration. The retry policy emerges from a configuration paste from a getting-started guide whose context was a different endpoint class than the one it was applied to. The coverage boundary emerges from a deployment roadmap that scopes the mesh to HTTP traffic and never revisits whether the non-HTTP traffic carries load that is relevant to the team's SLOs.

Property 1: The mTLS enforcement mode and the lateral movement surface. The lateral movement surface in a Kubernetes cluster is the set of service-to-service communication paths that an attacker who has compromised one pod can follow to reach other services without authentication. In a cluster without a service mesh, this surface is the entire cluster network: any pod can send TCP connections to any other pod's ClusterIP address, and the receiving pod's application is responsible for its own authentication. A service mesh with STRICT mTLS enforcement closes this surface for all connections within the mesh's scope: every connection must carry a valid SPIFFE X.509 certificate issued by the mesh's certificate authority, and the certificate is only issued to sidecars injected into mesh-managed namespaces; a pod without a sidecar cannot produce a valid certificate and cannot initiate a connection to a STRICT-mode service. PERMISSIVE mode changes this property fundamentally: a pod without a sidecar can still connect to any PERMISSIVE-mode service because the sidecar accepts plaintext connections from sources that lack a certificate. The lateral movement surface under PERMISSIVE mode is not reduced relative to a cluster without a mesh; it is unchanged. The operational cost of running the mesh — sidecar resource consumption, certificate rotation overhead, Istiod control plane load, proxy configuration complexity — is fully incurred under PERMISSIVE mode, while the security benefit that justified the adoption is not delivered. The gap between the operational cost and the security benefit is invisible in standard mesh dashboards: the percentage of mTLS connections will appear high (reflecting connections between meshed services) while the plaintext path remains fully open for connections from outside the mesh. Per-namespace STRICT enforcement for namespaces that are fully meshed is the correct migration model: it delivers the zero-trust boundary for the namespaces where it can be enforced while allowing unmeshed namespaces to continue using PERMISSIVE mode during their migration. Connect this property to the incident response playbook decision record: the mTLS enforcement mode is an incident response capability as well as a prevention control — a STRICT-mode mesh that rejects unauthenticated connections logs every rejection with the source address, destination service, and reason for rejection in the Envoy access log; these rejection events are the lateral movement detection signal that would trigger a tier-1 security incident alert; in PERMISSIVE mode, plaintext connections from an attacker are accepted without logging a rejection, and the lateral movement is invisible until the application-level effect (unauthorized data access, anomalous request patterns) is detected by other means.

Property 2: The mesh retry policy model and the resource exhaustion amplification surface. The resource exhaustion amplification surface is the set of service endpoints where a mesh-level retry policy converts a failure caused by the service's own resource exhaustion into a self-reinforcing load amplification cycle. The amplification is structural: when a service endpoint returns 503 because it has no capacity to serve requests (connection pool exhausted, thread pool saturated, memory limit approaching), retrying the request applies additional load to the resource that is already exhausted. Each retry is an additional connection acquisition attempt, an additional thread reservation, an additional memory allocation. The amplification factor is (1 + numRetries) per original request per upstream caller: with 3 retries and 8 callers, a service receiving 200 requests per second under exhaustion is effectively receiving 200 × 4 × 8 / (callers sharing the incoming rate) = 800 requests per second of amplified load, assuming uniform distribution across callers. The amplification surface is maximally dangerous for services with shared-resource bottlenecks — database connection pools, external API rate limits, semaphore-bounded concurrency — because the shared resource that caused the first exhaustion failure is the resource that every retry attempt competes for. The correct model separates the retry policy decision by failure type: transient connectivity failures (connection reset, connection refused, brief upstream restart) benefit from retries because the failure is not caused by the destination's resource state and the resource is likely available on retry; resource exhaustion failures (503 from pool exhaustion, 429 rate limit, 508 loop detected) do not benefit from retries and are worsened by them. The mechanism for separating these cases is the circuit breaker: a circuit breaker that opens after a threshold of consecutive 503 responses from a specific host pauses the retry pressure on that host, giving the resource time to recover rather than amplifying the failure. The circuit breaker and the retry policy must be configured together — a retry policy without a circuit breaker applies amplification load to a failing service until the caller's overall timeout expires; a circuit breaker without appropriate retry conditions for transient failures adds latency without reducing error rate for failures that would benefit from a retry. Connect this property to the timeout and deadline propagation decision record: the retry policy's perTryTimeout and the overall request timeout interact with the resource exhaustion amplification in a specific way — a long perTryTimeout combined with a short pool acquisition timeout means Envoy retries immediately after each pool-acquisition-timeout error rather than waiting the full perTryTimeout; the effective retry cadence is determined by the destination's error response latency (100ms in the pool exhaustion case), not by the perTryTimeout; this makes the amplification rate much higher than a naive perTryTimeout analysis would suggest, and is the reason that a 5-second perTryTimeout with a 100ms pool acquisition timeout produces an 800-request-per-second amplification rather than a 200-request-per-second amplification with 5-second pauses between retry attempts.

Property 3: The mesh coverage boundary and the non-HTTP observability blind spot. The non-HTTP observability blind spot is the set of data flows in the system that are not observed by the service mesh's telemetry because they use protocols the proxy cannot decode. The blind spot is not a gap in the mesh deployment — the proxy may correctly intercept and forward every connection in the system — but a gap between what the mesh can observe and what the system actually does. For systems where the critical data flows are HTTP and gRPC service-to-service calls, the mesh telemetry is an accurate representation of system health. For systems where the critical data flows include database writes, message queue consumption, background job execution, or any other non-HTTP operation, the mesh telemetry is accurate for the HTTP layer and silent for everything else. The blind spot is proportional to the fraction of the system's value that is delivered through non-HTTP data flows: a system whose primary function is to read events from Kafka and write derived records to PostgreSQL may have HTTP endpoints that serve 100% successfully while the core function — reading and writing — is failing at 100%. The SLO dashboard built on mesh telemetry will show 100% success rate while the system is completely failing its intended purpose. The correct model combines mesh telemetry for the HTTP layer with application-level instrumentation for the non-HTTP layer: per-operation error counters in the application code for every database write, message queue read, and external API call; consumer group lag metrics from the message broker for every consumer process; per-query error rates from the database for every service's connection pool. The non-HTTP instrumentation must be configured explicitly in the application code or via a database proxy or message queue exporter — it is not automatic and does not come for free with the mesh deployment. The coverage boundary decision — which data flows are covered by the mesh telemetry and which require additional instrumentation — must be made explicitly and validated against the list of SLOs, because the SLO that a team believes the mesh telemetry is measuring may not be the SLO that is actually being measured. Connect this property to the observability strategy decision record: the mesh telemetry coverage boundary should be documented in the observability strategy as a first-class input to the SLO definition — SLOs should only be defined against signals that have complete coverage for the operation being measured; defining an SLO for "data pipeline processing success rate" using HTTP request success rate from Linkerd is measuring a proxy for the actual operation, not the operation itself; the SLO is satisfied whenever the HTTP layer is functioning correctly, which is a necessary but not sufficient condition for the data pipeline to be functioning correctly; the observability strategy should specify explicitly which operations are measured by mesh telemetry and which require dedicated application-level instrumentation, and SLOs should be defined against the deepest available signal for each operation.

The service mesh adoption decision ADR: five sections

Section 1: Control plane selection and data plane injection model. Begin the service mesh adoption decision record by specifying the control plane (Istio, Linkerd, Cilium Service Mesh, Consul Connect) and the rationale for the selection — not the features considered, but the capability gap the mesh is being adopted to fill. The primary capability driver determines the selection: if the driver is mTLS and service identity enforcement (zero-trust networking), Istio and Linkerd both provide this with different operational complexity profiles; Istio provides a richer traffic management feature set (fine-grained routing, VirtualServices, DestinationRules) at the cost of higher control plane resource consumption and configuration complexity; Linkerd provides automatic mTLS and golden-signal telemetry with a smaller control plane footprint and simpler configuration at the cost of fewer traffic management primitives. If the driver is eBPF-based network visibility with reduced per-pod overhead, Cilium Service Mesh provides service mesh capabilities without per-pod sidecars by implementing the data plane in the kernel via eBPF. Specify the injection model: automatic sidecar injection (a MutatingAdmissionWebhook that injects a sidecar container into every pod created in labeled namespaces) or manual injection (a kubectl patch or istioctl kube-inject step in the deployment pipeline for each workload). Automatic injection is the correct default for new deployments; manual injection is appropriate for workloads with strict resource constraints or for workloads that cannot tolerate the iptables rules that sidecar injection adds. Specify the namespace scope for injection: which namespaces will have automatic injection enabled at launch, which namespaces will be added to the injection scope in subsequent migration phases, and which namespaces are permanently excluded from injection (the mesh control plane's own namespace, the kube-system namespace, any namespace that runs infrastructure that is incompatible with sidecar injection). The injection scope at launch defines the initial mesh coverage boundary; document it explicitly so that the coverage boundary is not a discovered limitation but a known architectural decision. Connect to the build artifact provenance decision record: the service mesh data plane components — sidecar proxy containers injected into every pod — are themselves software artifacts whose supply chain provenance should be verified; specify in this section that the sidecar image source (the Istio or Linkerd release artifact) is verified against the project's SLSA attestation or signature before deployment, and that the sidecar image version is pinned in the admission webhook configuration to prevent automatic upgrade to an unverified version.

Section 2: mTLS enforcement mode and the peer authentication migration path. Specify the initial mTLS enforcement mode, the migration path to STRICT enforcement, and the timeline and owner for each migration milestone. The initial mode decision: if all services in the target namespaces can be fully meshed at deployment time, deploy directly in STRICT mode and skip the PERMISSIVE migration phase entirely; PERMISSIVE mode is a migration aid, not a required initial state, and teams that can avoid it should. If some services cannot be meshed at deployment (legacy applications with incompatible thread models, external services, batch jobs with resource constraints), deploy in PERMISSIVE mode with a per-namespace migration plan: identify which namespaces contain only fully-meshed services, apply STRICT PeerAuthentication policies to those namespaces immediately after the initial deployment is verified healthy, and leave the namespaces with unmeshed services in PERMISSIVE mode with a defined migration deadline. Specify per-namespace migration deadlines: assign a named owner to each namespace in PERMISSIVE mode; define a specific deadline (a calendar date, not a conditional like "after the legacy service is migrated") for each namespace to reach STRICT mode; configure an alert that fires 30 days before each deadline to prompt the migration if it has not started. Specify the verification procedure before switching a namespace to STRICT: run istioctl experimental authz check or the equivalent to identify services in the namespace that receive plaintext connections, and enumerate the sources of plaintext connections before applying the STRICT policy. After applying STRICT, watch the Envoy access logs for connection rejections with PEER_CERT_NOT_PROVIDED — each rejection identifies a plaintext caller that was previously accepted and must either be added to the mesh or given an explicit exclusion. Connect to the secrets management decision record: the SPIFFE X.509 certificates issued by the mesh's certificate authority are short-lived secrets managed by the mesh infrastructure — the leaf certificates rotate every 24 hours by default, the intermediate CA rotates on a longer schedule, and the root CA is a long-lived secret that should be stored in an HSM or a hardware-backed key management service; specify in the secrets management decision record that the mesh's root CA private key is a tier-1 secret, that its storage and access controls are documented, and that its compromise would require a full root CA rotation that re-issues all service identity certificates in the mesh — a high-impact operation whose procedure should be tested before an incident requires it.

Section 3: Traffic policy scope and per-service retry, circuit breaker, and timeout configuration. Specify the traffic policy model — global or per-service — and the criteria that determine which VirtualService retry and DestinationRule circuit breaker settings apply to each service. The global-policy-is-wrong principle: a global retry policy applied to all services in a namespace treats all services as having the same failure model (transient, idempotent, stateless), which is false for any system with write endpoints, resource-bounded concurrency, or downstream dependencies. The correct model is per-service traffic policies configured from the service's failure characteristics, not from a getting-started template. Define a classification for each service: services whose endpoints are all idempotent and stateless (GET-only services, read-through caches, static asset services) may use a retry policy with retryOn: "connect-failure,reset,retriable-4xx" and attempts: 2; services with mixed idempotent and non-idempotent endpoints require per-route retry policies that apply retries only to the safe routes; services with resource-bounded capacity (services backed by a connection pool, a thread pool, or a rate-limited external API) require a circuit breaker with outlierDetection configured to open after a threshold of consecutive 5xx responses from a host, and should not include 5xx in retryOn for endpoints that return 503 due to resource exhaustion. Specify the DestinationRule outlierDetection parameters: consecutiveGatewayErrors (the number of consecutive 502/503/504 responses before the host is ejected from the load balancing pool), interval (the evaluation window for ejection decisions), and baseEjectionTime (the minimum ejection duration before the host is eligible for re-admission). Configure timeouts: every VirtualService should specify a timeout for the overall request (not per-attempt), and this timeout should be shorter than the upstream caller's timeout to ensure that the call graph's timeout budget propagates inward. Connect to the timeout and deadline propagation decision record: the VirtualService timeout is the mesh's contribution to the deadline propagation model — it is the enforcement point that prevents zombie requests from accumulating in services that have already exceeded the caller's deadline; if the VirtualService timeout is longer than the caller's deadline, the service will continue processing a request after the caller has abandoned it and moved on; the service mesh timeout should be configured as part of the overall deadline propagation architecture, not independently.

Section 4: Mesh coverage boundary and non-HTTP traffic annotation model. Specify which traffic in the system is covered by the mesh's protocol-aware telemetry and which requires explicit application-level instrumentation. Begin with a traffic inventory: list every category of connection in the system (service-to-service HTTP calls, service-to-database connections, service-to-queue connections, service-to-cache connections, background job executions, external API calls) and classify each as: HTTP/gRPC (covered by mesh golden-signal telemetry), opaque TCP within the mesh (proxied but not decoded — requires application-level instrumentation), or external (not proxied at all — requires application-level instrumentation). For opaque TCP connections in the mesh (database connections, Kafka consumers, Redis connections), specify the application-level instrumentation required to provide equivalent observability: per-operation error counters and latency histograms instrumented via the OpenTelemetry SDK in the application code; consumer group lag metrics exported from the message broker; per-query error rates from the database server via a Prometheus exporter. Specify the opaque port annotations for non-HTTP ports that the mesh's protocol detection should not attempt to parse: annotate pods with config.linkerd.io/opaque-ports (Linkerd) or configure the Istio traffic interception port exclusions for known database and queue ports. This prevents the proxy from attempting HTTP protocol detection on PostgreSQL (port 5432), MySQL (port 3306), Kafka (port 9092), Redis (port 6379), and other non-HTTP protocols — misidentification of these protocols can cause connection failures or unexpected behavior when the proxy forwards bytes that do not match the detected protocol's framing. Specify the SLO coverage audit procedure: after any new service deployment, run an audit that maps each service's SLOs to the observability signals that measure them, verifies that each signal has complete coverage for the operation it claims to measure, and flags SLOs defined against proxy signals for operations that are not covered by the proxy. Connect to the observability strategy decision record: the mesh coverage boundary is an input to the observability strategy — the observability strategy should specify the primary signal for each SLO, and the mesh coverage boundary determines which SLOs can use mesh telemetry as the primary signal and which require dedicated application-level instrumentation; the audit procedure should be run both at service deployment time and during the periodic observability review, because a service that was correctly instrumented at deployment may accumulate new non-HTTP dependencies over time whose observability gaps are not caught until a silent failure exposes them.

Section 5: Service identity model and the certificate rotation and incident revocation procedure. Specify the SPIFFE service identity model — which SPIFFE IDs are issued to which services, the trust domain (the cluster-level trust root that all service identities are scoped to), the certificate rotation interval for leaf certificates, and the certificate rotation interval for the intermediate and root CAs. The service identity model defines the authorization boundary for the mesh: the identity of a service is the claim it makes when initiating a mTLS connection, and authorization policies (Istio AuthorizationPolicy, Linkerd Server and ServerAuthorization resources) use service identities to specify which services may call which services. Specify the leaf certificate rotation interval: 24 hours is the correct default for most deployments; do not extend this interval to reduce rotation overhead because the sidecar agent handles rotation automatically without service disruption; the rotation interval is the window during which a compromised private key remains valid without explicit revocation. Specify the AuthorizationPolicy model for each service: for services that receive calls only from known callers within the mesh, configure an explicit allow-list (principal-based AuthorizationPolicy) that allows connections only from the specific SPIFFE identities of the authorized callers; for services that receive calls from external sources (ingress gateways, external load balancers), configure a policy that allows the ingress gateway's identity and denies all other mesh identities that are not on the allow-list. Specify the revocation procedure for a compromised service identity: (1) terminate the affected pod immediately — pod termination destroys the private key and the certificate is no longer usable from that pod; (2) if the private key was exfiltrated and may be in use by an attacker, rotate the intermediate CA to invalidate all certificates signed by the current intermediate CA; test the CA rotation procedure in staging before an incident requires it in production; (3) update the AuthorizationPolicy to deny the compromised SPIFFE identity explicitly, as a belt-and-suspenders measure while the CA rotation completes; (4) review the Envoy access logs for any connections made using the compromised identity between the time of the compromise and the pod termination — these logs are the incident timeline for establishing what the compromised identity accessed. Specify the log retention requirement for the Envoy access logs: retain for at least 90 days and index by source SPIFFE identity and destination service, so that an incident response team can reconstruct the access history for a compromised identity across the retention window.

FAQ

How do you migrate from PERMISSIVE to STRICT mTLS mode without breaking existing services?

Migrate namespace by namespace, not cluster-wide. A namespace-scoped PeerAuthentication with mtls.mode: STRICT overrides the cluster-wide PERMISSIVE default for pods in that namespace, allowing you to enable STRICT enforcement for fully-meshed namespaces while leaving PERMISSIVE in place for namespaces with legacy services. The migration procedure for each namespace: (1) verify all pods have sidecars injected — kubectl get pods -n <namespace> and check the READY column; a pod with 2/2 containers has a sidecar; a pod with 1/1 does not; (2) enumerate known plaintext callers — services in non-meshed namespaces, Prometheus scraping on pod metrics ports, Kubernetes liveness probes in older Istio versions, external health checkers; (3) apply the STRICT PeerAuthentication; (4) watch the Envoy access logs for PEER_CERT_NOT_PROVIDED connection rejections — each rejection identifies a plaintext caller that was previously accepted; (5) for each rejected source, either inject a sidecar into the source namespace or add a namespace-scoped PeerAuthentication exemption for that specific service. Common plaintext callers that break STRICT mode: Prometheus scraping pod metrics ports directly (fix by configuring Istio to expose metrics on the Envoy sidecar's port instead of the application port); external monitoring agents that health-check services via TCP (configure them to use the Kubernetes liveness probe endpoint via the kubelet instead of direct pod access); services in the kube-system namespace that interact with application services (these should be routed through an ingress gateway rather than directly). The migration from PERMISSIVE to STRICT for a well-meshed namespace takes 2–4 hours including the verification period. Do not defer the migration to "after the full cluster is meshed" — apply STRICT per-namespace as each namespace reaches full mesh coverage, rather than holding all enforcement until the last legacy service is migrated.

How do you configure retry policies that don't amplify resource exhaustion failures?

Configure retry policies per-service via VirtualService, not globally, and exclude 503 from retryOn for any service that returns 503 due to resource exhaustion rather than transient fault. The classification test: if the service returns 503 within 100–200 milliseconds (indicating a fast rejection from a connection pool, a semaphore, or a rate limiter), the 503 is a resource exhaustion signal and should not be retried. If the service returns 503 after a significant latency (indicating a slow upstream dependency that timed out), the 503 may be a transient fault appropriate for retry. Configure a DestinationRule with outlierDetection for every service that has resource-bounded capacity: set consecutiveGatewayErrors: 5 to open the circuit after 5 consecutive 503 responses from a host, interval: 10s as the evaluation window, and baseEjectionTime: 30s as the minimum ejection duration. Envoy will not retry requests to an ejected host, which prevents retry amplification against an already-exhausted service. For non-idempotent endpoints: add per-route retry policy with attempts: 0 for every VirtualService route that matches a non-idempotent endpoint (POST, PATCH, DELETE, or any endpoint that records a write). Use retryOn: "connect-failure,reset" only — not 5xx — for services that have a mix of idempotent and non-idempotent endpoints but where per-route policies are not yet configured; this limits retries to network-level failures where the request provably did not reach the destination, which is always safe to retry. The combination of per-service circuit breakers and conservative retryOn conditions eliminates the resource exhaustion amplification surface while preserving the transient failure recovery that motivated the retry policy adoption.

How do you get observability for database and message queue connections in a service mesh?

The service mesh provides bytes-transferred metrics for opaque TCP connections but not per-operation golden signals. Instrument the non-HTTP layer explicitly using three complementary mechanisms. For database connections: use OpenTelemetry SDK instrumentation for the database client library — OTEL instrumentation is available for most database drivers (pgx and sqlx for Go, SQLAlchemy for Python, pg and mysql2 for Node.js); the instrumentation records per-query latency, error type (constraint violation, type mismatch, timeout), and affected row count as OTEL spans; configure the spans to export to the same collector as the mesh telemetry so that a single dashboard shows both HTTP-layer and database-layer signals. Complement with server-side metrics via the postgres_exporter Prometheus exporter: alert on pg_stat_activity connections in state idle in transaction (indicates connection leaks), and on pg_stat_statements rows with high calls and increasing mean_exec_time (indicates accumulating slow query performance degradation). For Kafka consumers: expose consumer group lag as a Prometheus metric using the Kafka JMX exporter's kafka_consumer_group_lag metric, Strimzi's built-in consumer group lag metrics (if using the Strimzi operator), or the Burrow consumer lag monitoring tool. Consumer group lag — the number of unconsumed messages per partition relative to the producer offset — is the primary signal for consumer health; a lag that is growing indicates the consumer is falling behind or has stopped. Alert on consumer group lag exceeding a threshold corresponding to your data staleness SLO; alert on consumer group lag that is growing monotonically for more than 10 minutes (which indicates the consumer is stopped, not just slow). For Linkerd-specific configuration: annotate pods whose non-HTTP ports carry meaningful data flow with config.linkerd.io/opaque-ports listing those port numbers; this prevents Linkerd's protocol detection from interfering with non-HTTP protocols that the proxy would otherwise misidentify.

How do you handle service identity revocation when a pod is compromised?

For most threat models, immediate pod termination is sufficient: terminating the pod destroys the pod's private key, and the SPIFFE certificate associated with that key expires within the leaf rotation interval (24 hours by default). The attacker cannot use the private key after the pod is terminated because the key no longer exists on any running system. The replacement pod started by the Deployment controller receives a fresh key and certificate from Istiod within seconds of starting. For the threat model where the private key was exfiltrated before pod termination and may be in use by an attacker from a different location: (1) terminate the compromised pod immediately; (2) deploy an Istio AuthorizationPolicy that explicitly denies the compromised SPIFFE identity as a belt-and-suspenders measure while the intermediate CA rotation completes; (3) rotate the intermediate CA — issuing a new intermediate CA causes Istiod to re-issue all leaf certificates signed by the new CA; services already running will receive new certificates within one rotation cycle; connections carrying the old intermediate CA's certificates will fail verification against the new trust root within minutes as the new certificates are distributed; test the intermediate CA rotation in staging before executing in production, as the rotation causes a brief disruption to in-flight connections while services pick up new certificates; (4) after the CA rotation, review the Envoy access logs for connections made by the compromised SPIFFE identity between the compromise event and the pod termination — these logs are the incident timeline for establishing which services the compromised identity accessed and which operations it performed; retain Envoy access logs for at least 90 days indexed by source SPIFFE identity to support this retrospective analysis. Document the intermediate CA rotation procedure in the incident response playbook — its operational complexity means it should not be executed under incident pressure without prior practice.

Further reading

  • Observability strategy decision record — the instrumentation and SLO signal model that determines which operations have golden-signal coverage and which require dedicated application-level metrics; the service mesh provides HTTP/gRPC request-level telemetry automatically, but the observability strategy must specify which operations are not covered by the mesh (database writes, message queue consumption, background job execution) and require explicit OpenTelemetry instrumentation; an SLO defined against mesh telemetry is only as complete as the fraction of the system's value delivered through HTTP connections.
  • Incident response playbook decision record — the incident tier classification and evidence collection procedures that determine how a service mesh security event is handled; the Envoy access log (recording every connection with source SPIFFE identity, destination service, result, and latency) and the mTLS rejection log (recording every PEER_CERT_NOT_PROVIDED rejection) are the primary incident timeline artifacts for lateral movement detection; both must be retained with defined retention windows and indexed by SPIFFE identity and destination service before an incident requires them, not collected ad hoc during an active response.
  • Secrets management decision record — the secret classification and storage model for the mesh's root CA private key and any intermediate CA keys managed outside the mesh control plane; the root CA compromise would require a full root CA rotation that re-issues all service identity certificates in the mesh, which is a cluster-wide operation whose scope and procedure must be documented and tested before an incident; the mesh's short-lived leaf certificates eliminate the long-lived service-specific key management problem, but introduce a certificate authority management problem at the control plane level that is not automatically covered by existing secrets policies.
  • Timeout and deadline propagation decision record — the timeout layer model that determines whether the service mesh's per-route timeout is shorter than the upstream caller's deadline (as required for correct deadline propagation) or longer (which creates zombie requests that continue processing after the caller has abandoned them); the mesh VirtualService timeout is part of the deadline propagation architecture, not an independent configuration; the retry policy's perTryTimeout interacts with resource exhaustion error latency to determine the effective retry amplification rate, which can be much higher than the perTryTimeout analysis suggests when pool exhaustion errors return in milliseconds rather than in seconds.
  • Build artifact provenance decision record — the SLSA attestation and supply chain security model for the sidecar proxy containers that are injected into every pod in the mesh; the sidecar is a privileged network interceptor with access to all pod traffic — its supply chain integrity is as important as the application container's; specify the attestation verification procedure for the sidecar image source (the Istio or Linkerd release artifact) and the process for responding to a disclosed vulnerability in a sidecar version that is deployed at scale across the cluster.
  • Open-source extractor — find the service mesh adoption decisions buried in your AI chat history: the session where PERMISSIVE mode was set as a temporary migration aid without a deadline, the session where the global retry policy was pasted from a getting-started guide without per-service customization, and the session where the mesh deployment scope was defined as "HTTP service-to-service traffic" without asking whether the system's critical data flows were HTTP — each of these is a decision that determined the mesh's security boundary, reliability profile, and observability coverage, and whose reasoning is recoverable only from the conversation history where it was made.