OpenTelemetry for LLMs means tracing every model call, token, and retrieval step with the same OTel SDK and collector you already run, using the GenAI semantic conventions so each call lands as a standard gen_ai.* span. That matters because LLM calls break the assumptions of ordinary APM: prompts run to kilobytes, responses are nondeterministic, and the token bill swings between identical requests.

This guide shows what a compliant LLM span looks like, how to send LLM traces to Middleware in a few minutes, how to track token usage and latency, and how to trace a full RAG pipeline so a wrong answer points to one stage instead of one opaque span.

TL;DR

  • The OpenTelemetry GenAI semantic conventions define a vendor-neutral gen_ai.* schema for LLM spans, metrics, and events. They are in Development status as of September 2026.
  • Every LLM call becomes a CLIENT span named {operation} {model}, for example chat gpt-4o-mini.
  • The OpenTelemetry standard for token usage and latency is two histograms: gen_ai.client.token.usage and gen_ai.client.operation.duration.
  • A RAG pipeline needs one parent span per query and a child span for each stage: routing, retrieval, and generation.
  • Middleware’s Python SDK, middleware-llmobs, traces LLM calls with one register() call and turns on GenAI semantic conventions automatically.
  • Traceloop and OpenLIT cover Node.js, Next.js, Go, Ruby, and TypeScript, and export to the same Middleware LLM Observability view.

See every LLM call in one trace

Send gen_ai spans to Middleware and view prompts, token usage, cost, and latency next to your APM data.

What is OpenTelemetry for LLMs?

OpenTelemetry for LLMs is the practice of instrumenting model calls, embeddings, retrieval, and agent steps as OpenTelemetry spans and metrics that follow the GenAI semantic conventions. Any OTLP-compatible backend can then read provider, model, token, and latency data the same way.

Distributed tracing already answers “which service was slow.” An LLM call raises different questions: which model responded, how many tokens it used, and whether a bad answer came from a retrieval miss. Traditional APM has no vocabulary for any of them. The GenAI conventions add that vocabulary.

An LLM span is still an ordinary SpanKind.CLIENT span. There is no separate AI tracing protocol, which is why any OTLP-compatible backend can receive one. If you are new to OTel, start with what OpenTelemetry is and how its components fit together.

Spec status (last verified September 2026): In June 2026, the GenAI conventions moved out of the main semantic-conventions repository into open-telemetry/semantic-conventions-genai. They are marked Development and have no tagged release yet. Attribute names can still change, so pin your instrumentation versions.

What does a gen_ai span look like?

Each GenAI operation produces a span. gen_ai.operation.name identifies the call type, and the span name follows the pattern {gen_ai.operation.name} {gen_ai.request.model}. These are the attributes to set on every model call:

AttributeExampleWhat it tells you
gen_ai.operation.namechat, embeddingsThe type of call
gen_ai.provider.nameopenai, anthropicWhich provider served it (replaces the deprecated gen_ai.system)
gen_ai.request.modelgpt-4o-miniThe model you asked for
gen_ai.response.modelgpt-4o-mini-2024-07-18The exact model version that answered
gen_ai.usage.input_tokens412Prompt tokens billed
gen_ai.usage.output_tokens96Completion tokens billed
gen_ai.conversation.idsess-8f2cTies multi-turn and agent steps together

This is what that looks like as a manual span in Python:

from opentelemetry import trace

tracer = trace.get_tracer("llm-chat-service")

def call_model(prompt: str, model: str = "gpt-4o-mini") -> str:
    with tracer.start_as_current_span(
        f"chat {model}", kind=trace.SpanKind.CLIENT
    ) as span:
        span.set_attribute("gen_ai.operation.name", "chat")
        span.set_attribute("gen_ai.provider.name", "openai")
        span.set_attribute("gen_ai.request.model", model)

        response = client.chat.completions.create(
            model=model, messages=[{"role": "user", "content": prompt}]
        )

        span.set_attribute("gen_ai.response.model", response.model)
        span.set_attribute("gen_ai.usage.input_tokens", response.usage.prompt_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", response.usage.completion_tokens)
        return response.choices[0].message.content

In practice you rarely write this by hand. Instrumentation libraries patch the provider client and set these attributes for you, which is the approach the next section uses.

How to trace LLM calls with OpenTelemetry in Middleware

To trace LLM calls with OpenTelemetry, install an instrumentation library for your provider, point its OTLP exporter at your backend, and run one real request. With Middleware’s Python SDK, middleware-llmobs, that takes five steps. The steps follow the middleware-llmobs Python SDK docs.

Prerequisites

  • A Python LLM app on Python 3.10 or later (below 3.15)
  • Your Middleware UID for the ingest endpoint and your Middleware API key

Step 1: Install the SDK

pip install middleware-llmobs

Step 2: Install instrumentation for each provider and framework

The SDK has no LLM-provider dependencies. It traces a library only when that library’s OpenInference instrumentation package is installed, so install one per provider, vector store, or framework you use:

pip install openinference-instrumentation-openai
pip install openinference-instrumentation-langchain
# Also available: anthropic, bedrock, llama-index, crewai, and more

Step 3: Configure your Middleware credentials

export OTEL_EXPORTER_OTLP_ENDPOINT="https://<MW_UID>.middleware.io:443"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=<MW_API_KEY>,X-Trace-Source=openinference"
export OTEL_SERVICE_NAME="my-llm-app"

X-Trace-Source=openinference tells Middleware which SDK produced the spans so it routes them correctly. Pass only the base URL. The SDK appends /v1/traces, /v1/logs, and /v1/metrics automatically.

Step 4: Register and auto-instrument

Call register() once at startup, before you import or call your LLM libraries:

from middleware.llmobs import register

providers = register(auto_instrument=True)

With auto_instrument=True, the SDK attaches every installed instrumentor with GenAI semantic conventions enabled, so your spans carry gen_ai.* attributes. It exports over OTLP/HTTP with a batch processor, so tracing doesn’t block your requests.

Step 5: Send a request and open LLM Observability

Run your app and trigger at least one model call. In Middleware, open LLM Observability and select the trace. You’ll see the chat span with its model, prompt and response, total tokens, and per-call cost.

If your app sends LLM traffic through a fleet of services, run an OpenTelemetry Collector in front of Middleware so batching and filtering happen in one place.

Which instrumentation path should you use?

Middleware accepts LLM traces from three OpenTelemetry-compatible SDKs. Pick the one that matches your language and providers.

PathLanguagesSetupBest for
Middleware middleware-llmobsPythonOne register(auto_instrument=True) callPython apps that want tracing and evaluations in one package
Traceloop (OpenLLMetry)Python, Node.js, Next.js, Go, RubySDK init at startup with your Middleware endpoint and keyMulti-language teams
OpenLITPython, TypeScriptOne openlit.init() call with your Middleware endpoint and keyTeams already using OpenLIT integrations
Manual spansAny language with an OTel SDKYou set every attributeInternal model gateways and unsupported providers

All paths export over OTLP. Attribute coverage can still differ by SDK and version, so if you mix SDKs, standardize on one per service or normalize attribute names in the Collector. The Traceloop and OpenLIT setup guides cover each language.

Is there an OpenTelemetry standard for token usage and latency?

Yes. The GenAI semantic conventions define two histogram metrics: gen_ai.client.token.usage for tokens and gen_ai.client.operation.duration for latency. Both can be filtered by model and provider, and token usage is split into input and output with gen_ai.token.type.

Span attributes help you debug one request. Metrics tell you whether token spend is climbing week over week. If you emit them yourself, record both token types as OpenTelemetry histograms:

import time
from opentelemetry import metrics

meter = metrics.get_meter("llm-chat-service")
token_usage = meter.create_histogram(
    "gen_ai.client.token.usage", unit="{token}",
    description="Tokens used per GenAI operation",
)
duration = meter.create_histogram(
    "gen_ai.client.operation.duration", unit="s",
    description="GenAI operation duration",
)

base = {
    "gen_ai.operation.name": "chat",
    "gen_ai.provider.name": "openai",
    "gen_ai.request.model": model,
}

start = time.monotonic()
response = client.chat.completions.create(model=model, messages=messages)
duration.record(time.monotonic() - start, base)

token_usage.record(response.usage.prompt_tokens, {**base, "gen_ai.token.type": "input"})
token_usage.record(response.usage.completion_tokens, {**base, "gen_ai.token.type": "output"})

In Middleware, you don’t need to build these charts from scratch. Each trace shows total tokens, per-call cost, and model details. Pre-built LLM dashboards chart token usage, latency percentiles, and cost trends, and you can alert when any of them crosses a threshold.

How to instrument a RAG pipeline with OpenTelemetry

To instrument a RAG pipeline, create one parent span per user query and a child span for each stage: routing, retrieval, and generation. A slow or wrong answer can then be traced to the stage that caused it, instead of one multi-second span with no detail.

Middleware’s SDK ships decorators for this. @task marks a pipeline step, @retriever marks a retrieval step, and annotate_rag() records the query and returned documents on the retrieval span. The provider’s chat and embedding calls are traced automatically.

import openai
from opentelemetry import trace
from middleware.llmobs import register, task, retriever, annotate_rag, using_session

providers = register(service_name="docs-assistant", auto_instrument=True)
tracer = trace.get_tracer(__name__)
client = openai.OpenAI()

@task
def needs_retrieval(question: str) -> bool:
    return router.classify(question) == "knowledge"

@retriever
def retrieve(question: str):
    vector = client.embeddings.create(
        model="text-embedding-3-small", input=question
    ).data[0].embedding
    docs = [
        {"id": d.id, "score": d.score, "text": d.text}
        for d in vector_db.search(vector, k=5)
    ]
    annotate_rag(query=question, documents=docs)
    return docs

def answer(question: str, session_id: str) -> str:
    with using_session(session_id=session_id):
        with tracer.start_as_current_span("rag.query"):
            docs = retrieve(question) if needs_retrieval(question) else []
            context = "\n\n".join(d["text"] for d in docs)
            response = client.chat.completions.create(
                model="gpt-4o-mini",
                messages=[{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"}],
            )
            return response.choices[0].message.content

Each request produces one trace with this structure:

SpanStageSource
rag.queryParent for the whole requestManual
needs_retrievalRouting@task
retrieveRetrieval, with the query and documents attached@retriever
Embeddings callQuery embeddingAutomatic
Chat callGeneration, with tokens and costAutomatic

using_session adds the session ID to every span inside it, so a multi-turn conversation groups together in Middleware. The RAG tracing cookbook has a full, runnable version.

Find the slow stage in your RAG pipeline

Middleware shows routing, retrieval, and generation spans on one timeline, so a stale index and a hallucinated answer stop looking the same.

How to trace AI agents and tool calls

Agentic systems add span types that traditional APM has no equivalent for: create_agent, invoke_agent, and execute_tool. Each should carry gen_ai.agent.name and a shared gen_ai.conversation.id. A multi-step run (plan, call a tool, re-plan, call the model again) then reconstructs as one trace instead of a pile of disconnected spans.

With Middleware’s SDK, install the OpenInference package for your agent framework, such as LangChain, CrewAI, or LlamaIndex, and register(auto_instrument=True) turns each agent step, tool call, and model call into a span. See tracing an AI agent end to end for a working recipe, and AI agent monitoring for the metrics to watch once agents are traced.

Example: debugging a wrong answer in Middleware

Consider an internal docs assistant where users report outdated answers about a feature that changed last month. With the RAG pipeline above instrumented, the investigation works like this:

  1. Spot the symptom. A user flags a wrong answer, or an evaluation score on the generation span drops on the LLM dashboard.
  2. Open the trace. In LLM Observability, filter traces by the session ID or by a failed evaluation and open the request.
  3. Check the routing span. needs_retrieval returned true, so the router did its job.
  4. Inspect the retrieval span. The documents recorded by annotate_rag() are all from the old version of the feature page. The answer was grounded, but in stale context.
  5. Confirm the generation span. The prompt shows the model faithfully used what it was given, so this is not a model or prompt problem.
  6. Fix the right layer. Re-index the updated docs, then watch evaluation scores on new traffic to confirm the fix.

Without the retrieval span, this looks identical to a hallucination, and the team might have spent a day rewriting prompts. Middleware also correlates LLM traces with logs, so an error thrown by the vector store shows up on the same timeline.

How to score answer quality with evaluations

GenAI semantic conventions standardize how you capture model attributes, tokens, and latency. They don’t tell you whether an answer was correct. That takes evaluations.

Middleware supports two ways to run them. Server-side evaluations are configured in the Middleware UI and run on incoming traces. Client-side evaluations run in your code with the SDK’s LLMJudge, @evaluator decorator, or submit_evaluation(). Either way, results attach to the span that produced the answer and export as gen_ai.evaluations.* metrics.

from middleware.llmobs import submit_evaluation

with tracer.start_as_current_span("chat"):
    response = client.chat.completions.create(...)

    submit_evaluation(
        label="answer_grounded",
        value=True,
        assessment="pass",
        reasoning="Every claim is supported by the retrieved context.",
    )

Because scores are metrics, you can chart pass rates over time, filter traces by failed evaluations, and alert when groundedness drops. Teams running their own inference also get GPU monitoring in the same view, so a slow trace and a saturated GPU show up together.

How to control prompt and completion capture

Prompts and completions are exactly the data that ends up in a data-leak ticket. Under the GenAI conventions, gen_ai.input.messages and gen_ai.output.messages are opt-in because they may contain sensitive data. Instrumentation libraries differ in their defaults, so check yours before production:

  • OpenTelemetry contrib instrumentations capture content only after you enable a flag such as OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true.
  • OpenInference-based instrumentation, which Middleware’s SDK uses, lets you pass a custom TraceConfig to control what is recorded when you wire instrumentors manually.
  • Middleware Pipelines can filter, transform, or drop fields at ingestion, before they are stored.

Keep full content capture for environments where redaction and access controls are already in place.

Why LLM traces go missing

SymptomWhat to check
Nothing shows up in LLM ObservabilityOTEL_EXPORTER_OTLP_ENDPOINT and OTEL_EXPORTER_OTLP_HEADERS are correct, register() runs before your LLM calls, and a real request ran after startup
No spans from one provider or frameworkThe matching openinference-instrumentation-* package is installed
Spans missing from a short script or serverless functionCall providers.tracer.force_flush() before the process exits
gRPC export errorsmiddleware-llmobs exports over OTLP/HTTP only; remove gRPC settings
Traceloop or OpenLIT sends no dataThe service can reach https://<MW_UID>.middleware.io:443 and the Authorization header holds your Middleware API key
OpenLIT sends traces but no metricsdisableMetrics or OPENLIT_DISABLE_METRICS is not set to true
Token counts missing on streaming responsesOpenAI returns usage for streams only with stream_options={"include_usage": True}
Dashboards miss some spansDifferent SDKs or versions using different attribute names for the same field

Conclusion

The OpenTelemetry GenAI semantic conventions give LLM calls the same standard shape as the rest of your telemetry: one span per call, token and latency metrics you can alert on, and per-stage spans that turn a RAG pipeline from a black box into a diagnosable trace.

With Middleware LLM Observability, you get there with one register() call in Python, or Traceloop and OpenLIT in other languages. Traces, token costs, evaluations, logs, and GPU metrics land in one view, next to the APM and infrastructure data you already monitor. For a broader overview of the practice, see the LLM observability guide.

Score answers where they were generated

Attach LLM-as-judge evaluators to any span and see correctness and groundedness next to latency and token cost.

FAQs

How do I trace LLM calls using OpenTelemetry?

Wrap each model call in a CLIENT span named {operation} {model} and set gen_ai.operation.name, gen_ai.provider.name, and gen_ai.request.model. Instrumentation libraries do this automatically. In Middleware, install middleware-llmobs and call register(auto_instrument=True) to trace Python LLM calls.

Is there an OpenTelemetry standard for tracking token usage and latency in AI apps?

Yes. The GenAI semantic conventions define gen_ai.client.token.usage and gen_ai.client.operation.duration as histogram metrics. Token usage is split into input and output by gen_ai.token.type. Both are in Development status as of 2026.

How to instrument a RAG pipeline with OpenTelemetry?

Create a parent span for each query and child spans for routing, retrieval, and generation. Record the query and returned documents on the retrieval span. With Middleware’s SDK, use @task, @retriever, and annotate_rag(), and let auto-instrumentation trace the model calls.

Do I need a separate tool to trace LLM calls?

No. GenAI spans are ordinary OpenTelemetry spans, so your existing SDK, Collector, and exporter work unchanged. You add an instrumentation library for your model provider.

Is prompt and completion content captured by default?

It depends on the instrumentation. The GenAI conventions treat message content as opt-in because it may be sensitive. Check your library’s defaults, and use Middleware Pipelines to drop or transform fields before storage.

Is gen_ai.system still valid?

No. gen_ai.system is deprecated. Use gen_ai.provider.name instead.

Which languages does Middleware LLM Observability support?

Middleware’s own SDK supports Python. Traceloop adds Node.js, Next.js, Go, and Ruby, and OpenLIT adds TypeScript. All three send traces to the same LLM Observability view.

Can Middleware evaluate LLM response quality?

Yes. You can run evaluations server-side from the Middleware UI or client-side with the SDK. Scores attach to the span that produced the answer and export as metrics you can chart and alert on.