Overview
For years, distributed tracing was a Datadog or New Relic feature, priced per host and locked to a vendor SDK. OpenTelemetry changed that. The instrumentation is open-source, the protocol is standardized, and you can point the output at whatever backend you want — self-hosted, cloud, or nothing at all during development.
I set it up on a small API last month. Twenty minutes of work, and I now have per-request traces across three services with no vendor lock-in.
What OpenTelemetry actually is
| Component | Purpose |
|---|---|
| API | The interface your code calls |
| SDK | The implementation that processes data |
| Instrumentation | Libraries that auto-instrument frameworks and Databases |
| OTLP | The wire protocol between your app and the backend |
| Collector | Optional middleware that receives, processes, and exports |
The important shift from older systems: OpenTelemetry separates instrumentation from the backend. You instrument once, and you can change backends without changing code. In practice this rarely happens, but it means the instrumentation is vendor-neutral and you're not paying per-host for the privilege of collecting data.
The three signals
OpenTelemetry handles three types of telemetry:
| Signal | What it is | Typical backend |
|---|---|---|
| Traces | A request's path through services | Jaeger, Tempo |
| Metrics | Counters, gauges, histograms | Prometheus, Mimir |
| Logs | Structured log lines | Loki, Elasticsearch |
Traces are what most people start with, because they answer the question that nothing else does: "this request took 800ms, where did the time go?"
Setting up a Node.js app
npm install @opentelemetry/api \
@opentelemetry/sdk-node \
@opentelemetry/auto-instrumentations-node \
@opentelemetry/exporter-trace-otlp-http \
@opentelemetry/resources \
@opentelemetry/semantic-conventions
Create a file that runs before anything else:
// tracing.js
const { NodeSDK } = require("@opentelemetry/sdk-node");
const { getNodeAutoInstrumentations } = require("@opentelemetry/auto-instrumentations-node");
const { OTLPTraceExporter } = require("@opentelemetry/exporter-trace-otlp-http");
const { Resource } = require("@opentelemetry/resources");
const { ATTR_SERVICE_NAME, ATTR_SERVICE_VERSION } = require("@opentelemetry/semantic-conventions");
const sdk = new NodeSDK({
resource: new Resource({
[ATTR_SERVICE_NAME]: "my-api",
[ATTR_SERVICE_VERSION]: "1.2.0",
"deployment.environment": process.env.NODE_ENV || "development",
}),
traceExporter: new OTLPTraceExporter({
url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT || "http://localhost:4318/v1/traces",
}),
instrumentations: [
getNodeAutoInstrumentations({
"@opentelemetry/instrumentation-fs": { enabled: false },
}),
],
});
sdk.start();
process.on("SIGTERM", () => {
sdk.shutdown().then(() => process.exit(0));
});
Load it first:
node --require ./tracing.js server.js
That's it. HTTP requests, database queries, and outgoing fetch calls are now traced automatically. Every request gets a trace ID, every span gets a duration, and errors get attached to the span where they occurred.
The instrumentation-fs disable is worth noting. The file system instrumentation is noisy — it traces every read and write — and for most applications it produces more spans than everything else combined. Turn it off unless you're specifically debugging file I/O.
What auto-instrumentation gets you
With no code changes, you get spans for:
- Incoming HTTP requests (with method, path, status)
- Outgoing HTTP requests (with target URL, status)
- Database queries (Postgres, MySQL, MongoDB, Redis)
- Message queue operations (Kafka, RabbitMQ, SQS)
- GraphQL operations
For a typical API, this covers 90% of what you'd want to see. The remaining 10% is custom business logic that auto-instrumentation can't know about.
Adding custom spans
const { trace, SpanStatusCode } = require("@opentelemetry/api");
const tracer = trace.getTracer("my-api");
async function processOrder(orderId) {
return tracer.startActiveSpan("processOrder", async (span) => {
try {
span.setAttribute("order.id", orderId);
const order = await fetchOrder(orderId);
span.setAttribute("order.total", order.total);
span.setAttribute("order.items", order.items.length);
await validateInventory(order);
await chargePayment(order);
await sendConfirmation(order);
span.setStatus({ code: SpanStatusCode.OK });
return order;
} catch (err) {
span.setStatus({
code: SpanStatusCode.ERROR,
message: err.message,
});
span.recordException(err);
throw err;
} finally {
span.end();
}
});
}
Attributes on spans become searchable fields in the backend. Setting order.id means you can find every trace for a specific order, which is the fastest debugging tool I've used.
The recordException call attaches the full stack trace to the span. In Jaeger or Tempo, you see exactly where the error occurred, not just that an error happened.
Context propagation
The magic that makes tracing work across services is context propagation. The trace ID gets passed from service to service via HTTP headers, and every span becomes part of the same trace.
With auto-instrumentation, this happens without code changes for HTTP. Service A makes a request to Service B, and the OTel SDK injects a traceparent header. Service B reads it and continues the trace.
The header looks like this:
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
Trace ID, span ID, and flags. That's the whole W3C Trace Context spec, and it's why tracing works across languages — a Go service can propagate context to a Python service with no integration code.
For non-HTTP communication, you need to propagate manually:
const { propagation, context } = require("@opentelemetry/api");
// Producer: inject into message headers
const headers = {};
propagation.inject(context.active(), headers);
await kafka.send({ topic: "orders", headers, value: payload });
// Consumer: extract from message headers
const parentContext = propagation.extract(context.active(), message.headers);
context.with(parentContext, () => {
// any spans created here are children of the producer's span
});
Running Jaeger locally
docker run -d --name jaeger \
-p 16686:16686 \
-p 4317:4317 \
-p 4318:4318 \
jaegertracing/all-in-one:1.62
Point your OTLP exporter at http://localhost:4318 and open http://localhost:16686. Every trace appears in the UI with a waterfall view showing exactly how long each span took and where they overlap.
Jaeger's all-in-one image uses in-memory storage by default, which means traces disappear when the container restarts. For development that's fine. For anything persistent, configure it with a backend or switch to Grafana Tempo, which uses object storage and handles volume better.
The Collector: when you need it
For a single service, your app can export directly to Jaeger or Tempo. The Collector becomes useful when:
- You want to sample traces before sending (drop 99% of successful requests, keep all errors)
- You're exporting to multiple backends
- You want to scrub sensitive data from spans before export
- Your services can't be reconfigured easily and you need a stable endpoint
# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 1s
send_batch_size: 1024
tail_sampling:
policies:
- name: errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: slow
type: latency
latency: { threshold_ms: 500 }
- name: sample-rest
type: probabilistic
probabilistic: { sampling_percentage: 5 }
attributes:
actions:
- key: http.request.header.authorization
action: delete
exporters:
otlp:
endpoint: tempo:4317
tls:
insecure: true
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch, tail_sampling, attributes]
exporters: [otlp]
Tail sampling is the interesting one. Unlike head sampling (deciding at the start of a trace whether to keep it), tail sampling waits until the trace is complete and then decides. This means you can keep 100% of errors and slow requests, and sample heavily from the boring ones. Head sampling can't do this because you don't know at the start whether a request will fail.
Cost considerations
Tracing produces a lot of data. A service handling 100 requests per second with 20 spans per request generates 172 million spans per day. At typical SaaS pricing, that's thousands of dollars a month.
Three ways to control it:
- Head sampling — sample at the SDK level. Keep 10% of traces. Cheap and simple, but you might drop the one that failed.
- Tail sampling in the Collector — keep all errors and slow traces, sample the rest. Best quality-to-cost ratio, but requires running the Collector.
- Self-hosted backend — Tempo and Jaeger are free. Storage costs are the only expense, and object storage is cheap.
I run the Collector with tail sampling and ship to a self-hosted Tempo. Storage cost for a month of traces from a low-traffic API is a few dollars. The setup took an afternoon, and it's not an ongoing expense.
When tracing is worth it
- Microservices. The moment you have more than one service, tracing is the only way to see the full request path.
- Performance debugging. "Why is this endpoint slow?" becomes a one-look answer.
- Production incident response. Trace IDs in logs let you find the exact request that failed and follow it through every service.
For a single monolith with good logs, tracing is a nice-to-have. For anything distributed, it's the difference between guessing and knowing.
One last recommendation: always include the trace ID in your log lines. It's two lines of setup and it connects your two observability sources. When a customer reports an error and gives you the trace ID, you can jump straight to the trace and skip the log search entirely.
