OTLP
OTLP in production
Exports are plain OTLP/HTTP JSON (spec-stable for traces, metrics and logs) — Grafana Tempo/Mimir/Loki, Honeycomb, Jaeger, Datadog, an OTel collector: anything with an OTLP HTTP receiver works, on port 4318 by default.
TELEMETRY_EXPORTERS=otlp
TELEMETRY_OTLP_ENDPOINT=https://otlp.example.com:4318
Authenticated backends:
'otlp' => [
'endpoint' => env('OTEL_EXPORTER_OTLP_ENDPOINT'),
'headers' => ['Authorization' => 'Bearer '.env('OTLP_TOKEN')],
],
Schedule the metrics flush
Spans and events push themselves at terminate. Metrics need the scheduler:
Schedule::command('telemetry:flush')->everyMinute()->onOneServer();
onOneServer() matters in multi-node setups: the store is cluster-wide,
so one flusher is enough (and avoids duplicate datapoints).
Metrics are exported with cumulative temporality — backends see monotonic series regardless of how many PHP processes contributed.
When the backend rejects a batch
telemetry:flush answers for delivery, not just for collection. A batch
the endpoint refused is printed with the exporter, the HTTP status and
the backend's own error body, written to the log at error level (under
cron nobody reads stdout), and the command exits non-zero:
$ php artisan telemetry:flush
ERROR Flushed 57 metric families: 0 of 1 exporter accepted the batch.
otlp ...... rejected: HTTP 400: {"code":3,"message":"unknown metric type"}
$ echo $?
1
Partial delivery is reported as partial — 2 of 3 exporters accepted the batch — and a backend that took the batch but refused some data points
(OTLP partial success) is reported too. One exporter failing never stops
the others from being tried.
Because the exit code is real, the scheduled flush can be monitored like any other cron job:
Schedule::command('telemetry:flush')
->everyMinute()
->onOneServer()
->emailOutputOnFailure('[email protected]');
Spool mode counts entries held back for the next tick as a failure of this run: nothing is lost — they stay queued — but the endpoint is not taking data and the exit code says so.
The request path is unchanged: a rejection at terminate is counted in the
telemetry.export.count{outcome=...} self-metric and never surfaced to
the user's request.
High traffic: the spool + flush daemon
At scale, two costs bite: per-request OTLP POSTs at terminate, and a one-minute metrics cadence that is too coarse. The spool solves both — an agent-plus-daemon model, with Redis standing in for a local socket:
TELEMETRY_OTLP_SPOOL=true
php artisan telemetry:flush --daemon --interval=1 --metrics-interval=15
With the spool enabled, requests serialize their spans/events and push
them to a capped Redis list — two commands (one RPUSH for everything
the request carries, and the LTRIM that caps the list), microseconds,
no HTTP in the request lifecycle. The daemon (one process, under
supervisor) drains the list every --interval seconds, merges up to
--max-batch entries into a single OTLP request, and flushes metrics
every --metrics-interval seconds — sub-second span delivery,
sub-minute metrics.
Delivery semantics
Stated precisely, because the useful thing to know about a buffer is what it does not promise.
- Endpoint down → the chunk is requeued at the front and retried
next tick; nothing is lost to a collector hiccup. A
Retry-Afteris honoured, so an overloaded collector is not hammered. - Endpoint returns 4xx → that chunk is dropped. A payload the collector will always reject must not wedge the queue behind it, and the drop is counted and reported.
- Daemon down → the list caps at
otlp.spool.max_items(20 000 by default) with drop-oldest semantics; app memory and Redis stay bounded. - SIGTERM → the daemon drains for up to five seconds and exits. Whatever does not fit stays in Redis and the next start picks it up; the spool survives restarts. It is deliberately not "drain whatever remains", because a supervisor sends SIGKILL about ten seconds later and a drain that outlasts that is a kill with extra steps.
- Redis connection lost mid-drain → entries are claimed with a
single atomic
LPOP key count, so two daemons never take the same ones. But if that command executes and its reply is lost, those entries are gone: the spool is at-most-once for the drain itself. It is a buffer in front of an endpoint, not a durable queue, and telemetry is lossy by design at several points before it (sampling, the span-buffer cap, the circuit breaker). If you need acknowledged delivery of telemetry, the thing to run is a local collector with its own on-disk queue, and point this package at it. - Collector accepts a POST and its response is lost → that chunk is requeued and re-sent, so the other edge is at-least-once. Duplicate spans share a span id and a well-behaved backend deduplicates them.
Supervisor program:
[program:telemetry-flush]
command=php /var/www/artisan telemetry:flush --daemon --interval=1
autorestart=true
stopwaitsecs=10
Cron mode still works with the spool — telemetry:flush (no flags)
drains it once per run. Without the spool, spans export directly at
terminate and only metrics need the scheduler, as above.
Watch the drain, not just the daemon process. The spool is drained
exclusively by telemetry:flush — nothing else touches it. If the
daemon dies (or cron was never scheduled) the list just grows until it
hits max_items and starts silently dropping its oldest entries; there
is no other warning. php artisan telemetry:doctor reports current
depth as a fraction of max_items and fails the check above 90% full
(warns above 50%) — run it from your deploy pipeline or an uptime check,
not just once at setup.
Latency budget
Trace export happens after the response is sent (terminable middleware), but still occupies the FPM worker. The transport uses tight timeouts (3 s total / 1 s connect by default) and never retries in-request; 429/503 responses are classified retryable and simply dropped for that batch — telemetry is best-effort by design.
If your OTLP backend is slow or far away, enable the spool above — it is exactly that fast local buffer, without the extra binary. A local OTel collector or Grafana Alloy works too — supported, just never required.
No collector? No problem
The whole point: a bare Laravel app + Redis exports production-grade telemetry with zero extra infrastructure. Add infrastructure only when you need buffering, tail sampling or fan-out.