Tracing a few services manually is easy. Doing it across dozens of microservices, languages, and teams quickly becomes repetitive and inconsistent.
OpenTelemetry auto-instrumentation solves that by generating traces for supported frameworks and libraries without adding tracing code to every service. On Kubernetes, Middleware packages the OpenTelemetry Operator with automatic language detection, so you can turn on tracing per workload and follow a request across every service in one place.
TL;DR
- OpenTelemetry auto-instrumentation creates spans for supported frameworks and libraries without tracing code in your application.
- Java, Python, Node.js, .NET, and Go use different instrumentation methods, so coverage varies by runtime and library.
- The Middleware Kubernetes Agent installs the OpenTelemetry Operator, a language detector, and an injection webhook when you set one Helm flag.
- You enable tracing per workload from the Middleware UI or with pod annotations, then restart the pods.
- Plan sampling, filtering, service naming, and context propagation before a large production rollout.
What is OpenTelemetry auto-instrumentation?
OpenTelemetry auto-instrumentation is a way to generate traces from an application without writing tracing code. A language agent or runtime library hooks into supported frameworks, such as HTTP servers, database clients, and message queues, and creates a span for each operation automatically. OpenTelemetry calls this zero-code instrumentation.
Zero-code doesn’t mean zero configuration. You still need to configure service names, exporters, sampling, and where the telemetry goes.
This approach is especially useful for existing microservice environments. Teams can establish broad distributed tracing coverage first, then add custom instrumentation only where they need more context.
Auto-instrumentation vs manual instrumentation
| Factor | Auto-instrumentation | Manual instrumentation |
|---|---|---|
| Code changes | None for supported libraries | Import the SDK and create spans in code |
| Coverage | Supported frameworks and libraries | Any code path you instrument |
| Business context | Limited | Full, such as order IDs or tenant |
| Rollout effort | One install, then enable per workload | Engineering work in every service |
| Best for | Baseline coverage across a fleet | Critical business operations |
Most teams use both. Auto-instrumentation provides the baseline, and manual spans add context where it matters.
How does auto-instrumentation work in Middleware?
Middleware’s Kubernetes auto-instrumentation is built on the OpenTelemetry Operator. It adds a language detector and a mutating admission webhook, so you don’t have to work out each workload’s runtime yourself. Here’s what happens from install to first trace:
- Detection. The language detector scans your Deployments, StatefulSets, and DaemonSets and lists them in the Middleware UI with their inferred language.
- Selection. You enable tracing for a workload with a UI toggle or a pod annotation.
- Injection. When the pod restarts, the Middleware auto-injector webhook adds the matching OpenTelemetry instrumentation to the new pod.
- Span creation. The application starts with instrumentation attached, and supported HTTP, database, and messaging operations generate spans.
- Export. Spans are sent over OTLP to the Middleware Kubernetes Agent on port 9319 (gRPC) or 9320 (HTTP).
- Analysis. Traces appear in Middleware APM with service maps, latency, and error rates.
Your application source code stays unchanged. What changes is the runtime environment: depending on the language, injection adds an init container, instrumentation files, environment variables, or startup arguments.
Which languages can be auto-instrumented?
OpenTelemetry doesn’t instrument every runtime in the same way, and some runtimes need an extra setting in Middleware.
| Language | How it works | Common coverage | Middleware setup note |
|---|---|---|---|
| Java | Java agent and bytecode instrumentation | Spring, HTTP clients, JDBC, Kafka | Enabled by default |
| Python | Runtime and library instrumentation | Flask, Django, FastAPI, database clients | Enabled by default; add otel-python-platform: "musl" for Alpine images |
| Node.js | Module instrumentation | HTTP, Express, supported database libraries | Enabled by default |
| .NET | CLR profiling and runtime instrumentation | ASP.NET Core, HttpClient, SQL clients | Enabled by default; set otel-dotnet-auto-runtime to linux-x64 or linux-musl-x64 |
| Go | A helper process attaches to the compiled binary | Varies by library | Off by default in the Operator; needs the binary path, a privileged container, and a single-container pod |
Java has one of the most mature auto-instrumentation implementations. Python and Node.js rely on runtime and library instrumentation, while .NET uses runtime profiling.
Middleware can also instrument Apache HTTPD and Nginx web servers through their own annotations. Nginx, like Go, is turned off by default in the Operator’s feature gates.
Because the mechanism varies, so does coverage. A supported framework may be instrumented automatically while an internal library or custom function stays invisible.
What does auto-instrumentation capture, and where does it stop?
Consider a single checkout request and the operations it runs:
| Operation | Captured automatically? | Why |
|---|---|---|
POST /checkout | Yes | Incoming HTTP request to a supported framework |
validateCart() | No | Internal application function |
calculateDiscount() | No | Internal application function |
SELECT on inventory | Yes | Query through a supported database client |
authorizePayment() | No | Business logic; needs a manual span |
POST to payment-service | Yes | Outgoing HTTP call through a supported client |
That already gives developers more than knowing the checkout service is slow. The trace shows which downstream operation consumed most of the request time.
But auto-instrumentation doesn’t understand your application logic. OpenTelemetry knows how supported HTTP and database libraries behave. It doesn’t know what calculateDiscount() means to your business. The practical approach is:
- Auto-instrument your services.
- Inspect real traces from real traffic.
- Find the blind spots where business logic runs without a span.
- Add manual spans only where they help debugging.
After that, the same checkout trace combines both kinds of spans:
| Span | Source |
|---|---|
POST /checkout | Automatic |
checkout.validate | Manual |
PostgreSQL SELECT | Automatic |
payment.authorize | Manual |
POST /payments | Automatic |
Automatic spans provide framework and dependency visibility. Manual spans add context around operations such as payment authorization, inventory reservation, fraud checks, and order processing.
How do traces stay connected across microservices?
Generating spans in every service isn’t enough. Those spans need to belong to the same trace. When a request moves from the browser to the API gateway, then to checkout, inventory, and payment, each service must pass trace context to the next one.
Middleware’s default instrumentation configuration propagates W3C Trace Context, W3C Baggage, and B3 headers. Here’s what a connected trace looks like:
| Service | Trace ID | Span ID | Parent span |
|---|---|---|---|
| Checkout | abc123 | 01 | None (root) |
| Inventory | abc123 | 02 | 01 |
| Payment | abc123 | 03 | 02 |
Each operation gets its own span ID, but all of them share one trace ID. The parent-child links let Middleware rebuild the request path in the trace view and service map.
If context is lost at a service boundary, the downstream service starts a new trace instead of continuing the existing one.
How to set up Kubernetes auto-instrumentation in Middleware
These steps follow the Middleware Kubernetes Agent docs and the Kubernetes auto-instrumentation docs. The running example is a checkout, inventory, and payment application.
Prerequisites
- Kubernetes 1.21 or later, with
kubectlpointed at the target cluster - Helm 3.5 or later
- A Middleware account and API key (the Installation page in Middleware pre-fills your key and account URL)
Step 1: Install cert-manager
cert-manager issues the TLS certificates the injection webhook needs, and it’s the recommended option. The chart can generate self-signed certificates instead, but then you manage renewal yourself.
helm repo add jetstack https://charts.jetstack.io --force-update
helm install cert-manager jetstack/cert-manager \
--namespace cert-manager \
--create-namespace \
--version v1.14.5 \
--set installCRDs=trueStep 2: Install the Kubernetes Agent with auto-instrumentation enabled
One Helm release installs the Middleware Agent. With mw-autoinstrumentation.enabled=true, it also deploys the OpenTelemetry Operator, the language detector, the auto-injector webhook, and a default Instrumentation resource named mw-autoinstrumentation. Setting opsai.enabled=true adds the OpsAI component for AI-powered cluster insights and fixes.
helm repo add middleware-labs https://helm.middleware.io
helm install mw-agent middleware-labs/mw-kube-agent-v3 \
--set global.mw.apiKey=<MW_API_KEY> \
--set global.mw.target=https://<MW_UID>.middleware.io:443 \
--set global.clusterMetadata.name=<your-cluster-name> \
--set mw-autoinstrumentation.enabled=true \
--set opsai.enabled=true \
-n mw-agent-ns --create-namespaceIf your cluster already has OpenTelemetry Operator CRDs, the install can fail with an ownership error. Either remove the existing CRDs and let the Middleware chart manage them, or add --set mw-autoinstrumentation.opentelemetry-operator.crds.create=false.
Step 3: Verify the components
kubectl get daemonset mw-kube-agent -n mw-agent-ns
kubectl get deployment mw-agent-opentelemetry-operator -n mw-agent-ns
kubectl get deployment mw-auto-injector -n mw-agent-ns
kubectl get deployment mw-lang-aggregator -n mw-agent-ns
kubectl get daemonset mw-lang-detector -n mw-agent-nsEvery component should show as available before you enable any workload.
Step 4: Scope the first rollout
By default, auto-instrumentation applies to all namespaces except system namespaces. For the first rollout, limit it to the namespaces on your test request path by adding one of these flags to the install command:
# Instrument only these namespaces
--set mw-autoinstrumentation.webhook.includedNamespaces="{checkout,catalog,inventory}"
# Or instrument everything except these
--set mw-autoinstrumentation.webhook.userExcludedNamespaces="{payments}"Keeping a namespace such as payments out of the first rollout lets you validate trace quality, propagation, and telemetry volume before you touch compliance-sensitive services.
Step 5: Choose which workloads to instrument
From the Middleware UI (default auto mode). Open the Kubernetes Agent section of the Installation page. Middleware lists detected Deployments, StatefulSets, and DaemonSets with their inferred language and status. Toggle the workloads you want, or select all, and save.
With pod annotations (manual mode). Use annotations when language detection misses a workload, you need custom settings, or you run on EKS Fargate or GKE Autopilot, where UI-driven detection isn’t available. Disable the workload in the UI first to avoid conflicts. Then add the annotation to the Pod template, not the Deployment’s top-level metadata:
spec:
template:
metadata:
annotations:
instrumentation.opentelemetry.io/inject-python: "mw-agent-ns/mw-autoinstrumentation"Java, Node.js, .NET, and Go use inject-java, inject-nodejs, inject-dotnet, and inject-go with the same value. For clusters that run entirely in manual mode, install with --set mw-autoinstrumentation.mode=manual.
In pods with sidecars such as Istio or Linkerd, name the application container explicitly. Otherwise the first container in the pod is instrumented, and that may be the sidecar:
instrumentation.opentelemetry.io/inject-java: "mw-agent-ns/mw-autoinstrumentation"
instrumentation.opentelemetry.io/container-names: "checkout"Step 6: Set sampling and resource attributes
The default Instrumentation resource uses a parent-based sampler with a ratio of 1.0, so it keeps every trace. That’s useful while you validate a few services. Before a wide rollout, check it and decide what rate you need:
kubectl get otelinst mw-autoinstrumentation -n mw-agent-ns -o yamlTo use a different rate for some workloads, create your own Instrumentation resource and point their annotations at it, for example "mw-agent-ns/prod-sampled". To tag traces with ownership or environment, add resource attribute annotations to the Pod template:
resource.opentelemetry.io/team: "checkout"
resource.opentelemetry.io/environment: "production"Step 7: Restart, send real traffic, and open APM
kubectl rollout restart deployment/checkout -n checkoutDon’t verify the setup from configuration alone. Send a real request through the application, then open APM in Middleware and select the service. You’ll see the service map, trace details, latency charts, and error rates. Learn more about distributed tracing in Middleware.
Example: following a polyglot request in Middleware
Consider an e-commerce application built from services in three languages. The storefront calls checkout, and checkout calls inventory and payment:
| Service | Language | Called by | Dependency |
|---|---|---|---|
| Checkout | Node.js | Storefront | Inventory, Payment |
| Inventory | Java | Checkout | PostgreSQL |
| Payment | Python | Checkout | Payments database |
With all three services auto-instrumented, one checkout request produces a single trace with these spans:
| Span | Service | Duration |
|---|---|---|
POST /checkout | Checkout | 1.8 s |
| Inventory request | Inventory | 120 ms |
| PostgreSQL query | Inventory | 45 ms |
| Payment request | Payment | 1.3 s |
| Payments database query | Payment | 1.1 s |
The checkout request took 1.8 seconds, and 1.1 seconds of that came from one database query in the payment service. Instead of investigating every service, you know where to look.
Investigating beyond the trace
The slowest span tells you where to look. It doesn’t always tell you why the operation became slow. Here’s how that investigation works in practice:
- Spot the symptom. Checkout latency rises on the service dashboard or triggers an alert.
- Open the distributed trace. Pick a slow checkout request in Middleware APM to see every service it touched.
- Find the slowest path. The payment service accounts for 1.3 seconds of the 1.8-second request.
- Narrow it to one operation. Inside the payment service, a single database query takes 1.1 seconds.
- Check the surrounding signals. Review the payment service’s logs, database metrics, and pod resource usage for the same time window.
- Identify the likely cause. For example, connection pool exhaustion, CPU throttling on the database pod, or a burst of application errors.
Middleware shows the trace, related logs, Kubernetes metrics, and database telemetry on one timeline, so you don’t have to switch tools to line them up. OpsAI, Middleware’s AI SRE agent, correlates these signals and surfaces the likely root cause automatically.
Generation Esports found troubleshooting harder as its Kubernetes-based microservices grew. After moving to Middleware, the team improved MTTR by 75%.
Why do auto-instrumented traces break?
Auto-instrumentation removes repetitive setup, but configuration problems can still produce incomplete or misleading traces. These are the most common symptoms and what to check:
| Symptom | What to check |
|---|---|
| Services appear as separate traces | Context propagation across proxies, gateways, HTTP clients, and messaging systems |
| No traces appear | Pods restarted after enabling, namespace in scope, annotation under spec.template.metadata.annotations, and webhook logs from kubectl logs deploy/mw-auto-injector -n mw-agent-ns |
| Spans are created but not received | Connectivity from the pod: kubectl exec -it <pod> -n <namespace> -- curl http://mw-service.mw-agent-ns:9320 |
| Workload missing from the UI | Supported language, pod in Running state, and detector logs from kubectl logs daemonset/mw-lang-detector -n mw-agent-ns |
| Webhook timeouts on GKE private clusters | A firewall rule that lets the control-plane CIDR reach nodes on TCP 9443 |
Services named unknown_service | Consistent service.name values set early |
| Database spans missing | Whether your database client and version are supported by the language instrumentation |
| Duplicate spans | Whether the app already has an SDK, another APM agent, or a second instrumentation library attached |
How to roll out auto-instrumentation in production
Begin with an entry point such as an API gateway and two or three downstream services. Choose services that represent the languages and frameworks you use most.
Check that service names are correct, expected spans appear, traces stay connected, captured attributes are useful, and telemetry volume is reasonable.
Once the request path looks right, expand to another service group or namespace. Expanding in stages makes problems easier to isolate. If 100 services are instrumented at once and traces start breaking, it’s much harder to tell whether the cause is one runtime, one environment, or the global configuration.
Control trace volume before scaling
Auto-instrumentation can generate a lot of telemetry because common framework operations are captured automatically. A single request may produce spans for the incoming HTTP request, several downstream calls, database operations, and messages. Across dozens of high-traffic services, volume grows quickly. Three sampling approaches help:
- Head sampling makes the sampling decision near the start of a request.
- Parent-based sampling keeps downstream services aligned with the parent’s decision. Middleware’s default Instrumentation resource uses this.
- Tail sampling decides after more of the trace is available, which helps keep traces based on latency or errors.
In Middleware, you control volume in two places. The first is the sampler in the Instrumentation resource. The second is Middleware Pipelines, which can filter data at the agent, sample or transform it at ingestion, and drop it before storage. Set your strategy before a broad rollout, not after volume becomes hard to manage.
Watch for cardinality and sensitive data
Auto-instrumentation may collect attributes developers didn’t deliberately add, including URLs, routes, database metadata, network details, messaging destinations, and request attributes. Dynamic values such as /users/182739 or /orders/89d71c8 deserve particular attention, because every unique value adds cardinality.
Sensitive data is an even bigger concern. Tokens, secrets, personal information, and confidential request values should be removed before they’re stored. Review telemetry from representative workloads before expanding the rollout, and use Middleware Pipelines to remove fields that shouldn’t be kept.
Where telemetry is collected and processed
You don’t need to design a collector topology from scratch. With Middleware, telemetry moves through four stages:
- Instrumented pods send spans over OTLP to the
mw-serviceService inmw-agent-ns, on port 9319 (gRPC) or 9320 (HTTP). - The Middleware Kubernetes Agent receives them. It runs as a DaemonSet on every node for node-level telemetry, plus a Deployment for cluster-wide telemetry.
- Middleware Pipelines filter, sample, or transform the data based on your rules.
- Middleware APM, logs, and infrastructure views store and display the result, correlated by service and time.
Run only one Middleware Agent per cluster. Multiple agents cause unexpected behavior.
Auto-instrumentation best practices
- Start with a representative request path. Validate a few services before expanding across the cluster.
- Standardize service names. Consistent
service.namevalues make traces easier to search and navigate. - Verify context propagation. Spans from every service aren’t useful if they aren’t connected.
- Configure sampling early. Measure trace volume before expanding coverage.
- Review captured attributes. Remove sensitive or high-cardinality values.
- Avoid duplicate instrumentation. Check whether services already have an SDK or another APM agent attached.
- Test overhead. Compare CPU, memory, and latency on representative workloads before and after instrumentation.
- Add manual spans selectively. Focus on business operations that automatic instrumentation can’t explain.
- Version your configuration. Keep Helm values and Instrumentation resources in source control like other infrastructure code.
When auto-instrumentation isn’t enough
Consider a payment flow with four business stages:
payment.validatechecks the request.payment.fraud_checkscores the risk.payment.authorizereserves the funds.payment.capturecompletes the charge.
An automatically instrumented HTTP client shows the requests made during these stages, but it doesn’t know which business stage is running.
Manual spans add that missing context. A hybrid approach also makes sense for unsupported frameworks, complex asynchronous workflows, compliance-sensitive services, or applications where domain-specific attributes matter during debugging. Manual spans exported to Middleware appear in the same trace as the automatic ones.
Conclusion
OpenTelemetry auto-instrumentation removes much of the repetitive work involved in tracing microservices. Teams can capture supported HTTP requests, database operations, messaging activity, and service-to-service calls without manually instrumenting every application.
Auto-instrumentation provides the baseline, and manual spans fill the gaps where business context matters. With Middleware, you enable that baseline with one Helm install and a UI toggle, then investigate each trace alongside logs, infrastructure, database, and application telemetry to move from a slow request to the dependency behind it.
FAQs
How does OpenTelemetry auto-instrumentation work for microservices?
OpenTelemetry attaches language-specific instrumentation to supported frameworks and libraries. It creates spans for HTTP requests, database calls, and messaging operations. Trace context is passed between services, so the spans form one distributed trace.
How do I enable OpenTelemetry auto-instrumentation in Kubernetes with Middleware?
Install the Middleware Kubernetes Agent with mw-autoinstrumentation.enabled=true. Enable workloads from the Middleware UI or add a language-specific annotation to each pod template. Then restart the pods.
Does OpenTelemetry auto-instrumentation require code changes?
No code changes are needed for supported frameworks and libraries. You still configure service names, exporters, sampling, and the telemetry destination.
Do I need to restart pods after enabling auto-instrumentation?
Yes. Instrumentation is injected when a pod starts, so existing pods need a rollout restart.
What is the difference between auto-instrumentation and manual instrumentation?
Auto-instrumentation captures supported framework and dependency operations without code. Manual instrumentation adds spans for custom business logic that automatic instrumentation can’t see.
Can auto-instrumentation and manual instrumentation be used together?
Yes. Automatic and manual spans can belong to the same distributed trace. You get broad framework coverage plus application-specific context.
Which languages does Middleware auto-instrument on Kubernetes?
Middleware auto-instruments Java, Python, Node.js, .NET, and Go, plus Apache HTTPD and Nginx. Go and Nginx are off by default in the Operator. Go also needs the target binary path, a privileged container, and a single-container pod.
Does Middleware auto-instrumentation work on EKS Fargate or GKE Autopilot?
Yes, in manual mode. Automatic language detection and the UI toggle aren’t available there. Install with mw-autoinstrumentation.mode=manual and add annotations to your workload manifests.
Why are spans from my microservices showing as separate traces?
Broken context propagation is the most common cause. Check propagator settings, HTTP headers, gateways, proxies, messaging systems, and unsupported libraries to find where trace context is lost.
How do I reduce trace volume from auto-instrumentation?
Lower the sampler ratio in your Instrumentation resource. Use Middleware Pipelines to filter, sample, or drop telemetry you don’t need. Configure both before a fleet-wide rollout.
Does OpenTelemetry auto-instrumentation affect application performance?
It adds some runtime and export overhead. Test representative workloads before broad deployment, and use sampling and filtering to keep telemetry volume under control.

