ChangeGamer

← All guides · Agent observability and evaluation

How to Choose an LLM Observability Platform for AI Agents

Part 10 of Agent observability and evaluation · 1,124 words · ~5 min read · published 2026-10-06 · updated 2026-10-06 · Markdown variant

How to choose an LLM observability platform: decide on OTel-native ingestion, self-host versus cloud, export portability, redaction hooks and retention support before comparing vendors.

In short

  • Choosing an LLM observability platform is best done by criteria rather than by brand: OTel-native ingestion, self-host versus cloud, export portability, a redaction hook and support for your sampling and retention rules.
  • As of October 2026 the OpenTelemetry GenAI semantic conventions still carry Development status, so a backend that accepts plain OTLP lets you change vendors without re-instrumenting your agent code.
  • OpenLLMetry is instrumentation only and routes standard OTel data to any compatible backend, which makes it a way to postpone the backend decision rather than a backend itself.
  • The reference corpus contains no benchmark, pricing or scale data for any tracing platform, so the criteria in this article are reasoned inference and not a ranking of vendors.
  • A platform that cannot keep the fields your debugging and evaluation need, such as the full call tree, token cost per span and redacted tool payloads, is the wrong choice regardless of its other features.

Part of the AI Agent Observability and the Production Evaluation Playbook guide.


To choose an LLM observability platform, decide your criteria first (OTel-native ingestion, self-host versus cloud, export portability, a redaction hook, and support for your sampling and retention rules) and only then compare vendors. The corpus lists tools but, as of October 2026, holds no benchmark, pricing or scale data, so everything below is inference rather than a ranking. The topic sits under the pillar on observability and evaluation.

Why decide on criteria before vendors?

Criteria come first because the tools overlap heavily and the corpus gives no measured basis for preferring one. The agent-observability resource splits the landscape into framework-agnostic OTel platforms (Langfuse, Phoenix, OpenLLMetry, Logfire) and tracing built into a framework (LangSmith, the OpenAI Agents SDK). It describes what each does, not how well. Picking by popularity or by whatever your framework suggests skips the questions that are expensive to reverse later.

The cost of reversing is mostly in instrumentation and in stored history, not in the dashboard. So the useful question is how hard it would be to leave, and the criteria below are ordered with that in mind.

Which criteria matter most?

Five criteria matter most, and they are ordered by how costly a wrong answer is to undo. The ordering is the author's judgement, not a published standard.

Criterion The question to ask Why it is costly to get wrong
OTel-native ingestion Does it accept standard OTLP traces? Otherwise you instrument against a proprietary SDK
Self-host versus cloud Where may trace data live? Traces contain prompts and tool payloads
Export portability Can you get spans out again? History is stranded if you switch
Redaction hook Can you scrub before storage? Sensitive data is hard to delete afterwards
Sampling and retention support Can it honor your keep rules? Cost and coverage both depend on it

The first three decide how locked in you are. The last two decide whether you can run the platform responsibly.

Does it need to be OpenTelemetry-native?

Prefer an OpenTelemetry-native backend, because the GenAI semantic conventions still have Development status as of October 2026 and attribute names may change. The resource notes that the spec has moved into a dedicated open-telemetry/semantic-conventions-genai repository and that major vendors already support the gen_ai.* naming despite the label.

That combination argues for a layer you can swap. If your code emits standard OTel data, a naming change is handled in one place, and the backend can be replaced without touching agent code. Per the corpus, Langfuse exposes an OTLP ingestion endpoint and aims to comply with the GenAI conventions, and Phoenix is OTel-native and accepts traces over OTLP. For which fields to emit, see agent observability for reliability; those fields are deliberately not listed again here.

Should you self-host or use a cloud service?

Self-host if trace content cannot leave your environment, and use cloud if you would rather not operate a store. The corpus supports only the availability, not a comparison: Langfuse is described as self-host or cloud, Phoenix as running fully local with no API key required, and LangSmith as a paid product with a free tier.

Traces can hold prompts, completions and tool inputs, so the data-location question is a privacy and contract question before it is an engineering one. No operating cost or effort for self-hosting is stated in the corpus, and none is claimed here. Estimate it yourself against your own volume.

Where does OpenLLMetry fit?

OpenLLMetry fits as an instrumentation layer that postpones the backend decision. The corpus describes it as OTel instrumentations for LLM providers and vector databases, licensed Apache 2.0, with a Traceloop SDK wrapper that emits standard OTel data to any compatible backend such as Langfuse, Datadog or Grafana Tempo.

It stores and displays nothing, so it is not a platform to compare against the others. Its value is portability: instrument once, then point the exporter wherever the other criteria lead. Logfire is the contrasting case, described as an OTel-based platform with first-class Python and Pydantic AI integration, so it suits a stack already built around those.

What about framework-native tracing?

Framework-native tracing is the lowest-effort route if you already use that framework, with coupling as the price. LangSmith is described as paired tracing and evaluation for LangChain and LangGraph, with end-to-end OTel ingestion so spans can go to LangSmith and other backends simultaneously. The OpenAI Agents SDK has a built-in trace processor that exports to the OpenAI Traces dashboard by default, and custom TracingProcessor implementations can redirect spans to any OTel-compatible backend.

Both therefore have an exit route to OTel in the corpus. Whether that route preserves every field you rely on is not stated, so test it with a real run before you rely on it.

What must the trace keep, whatever you pick?

The trace must keep the full call tree, token usage and derived cost per span, latency, redacted tool inputs and outputs, and errors with retries on the span that failed. These come from the resource's list of key signals, and a platform that drops any of them fails the job whatever else it offers. Check by running a representative multi-step agent and reading the stored result.

Then check the platform against your own rules. How you decide which runs to keep, and for how long, is in sampling and retention of agent traces. Confirm the platform can express those rules and that redaction happens before storage. And confirm what happens to runs that die mid-flight, since that is a property of your pipeline as much as the tool, covered in crash-safe run-start records.

What does ChangeGamer use?

ChangeGamer does not use any of them in production, because it runs no agent trace store. It has not evaluated or benchmarked any vendor named here, and every tool description in this article is taken from the reference corpus, not from operating experience.

That is why the advice is framed as criteria and not as a recommendation. The corpus entry's sources, including the Langfuse, Phoenix, OpenLLMetry, Logfire, LangSmith and OpenAI Agents SDK documentation, are the place to verify current behaviour before you commit.

What the corpus does not tell you

No ranking, price, performance figure, scale limit or retention feature for these tools exists in the corpus as of October 2026. It supports the tool descriptions and the Development status of the GenAI conventions. The criteria and their order are reasoned inference and should be treated as a hypothesis. Test any shortlisted backend by sending it a real agent run, exporting the result, and checking that redaction and your keep rules behave as you expect.

Frequently asked questions

How do I choose an LLM observability platform?
Choose an LLM observability platform by testing it against fixed criteria: whether it ingests OpenTelemetry data natively, whether you can self-host, whether you can export your traces again, whether it lets you redact before storage, and whether it supports your sampling and retention rules. Compare brands only after those answers narrow the field.
Do I need an OpenTelemetry-native LLM observability tool?
An OpenTelemetry-native tool is the safer default because the GenAI semantic conventions are still in Development status as of October 2026 and may change. Instrumenting against standard OTLP keeps the backend replaceable, whereas a proprietary SDK ties your agent code to one vendor.
What is the difference between OpenLLMetry and Langfuse?
OpenLLMetry emits OpenTelemetry data from LLM providers and vector databases and is licensed Apache 2.0, but it stores nothing. Langfuse is a platform that receives OTel traces over an OTLP endpoint and can be self-hosted or used as cloud. They are often used together.
Is it safe to use the observability built into an agent framework?
Framework-native tracing such as LangSmith or the OpenAI Agents SDK is a reasonable choice if you accept coupling to that framework. The OpenAI Agents SDK lets a custom TracingProcessor redirect spans to an OTel-compatible backend, and LangSmith accepts OTel ingestion, so lock-in can be reduced.

#agents #observability #tracing #opentelemetry #tooling #decision

Put this corpus inside your own agents

This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.

Agents: this page as Markdown · JSON · offers at /api/pricing.json · payment methods at /api/payment.json · single-resource access via HTTP 402 (how that works)