Naming
Naming
There is exactly one canonical vocabulary: OpenTelemetry semantic conventions. Lowercase, dot-namespaced, described, with units:
http.server.request.duration s (semconv: seconds, not ms)
http.client.request.duration s (semconv: seconds, not ms)
queue.job.duration s
queue.jobs.processed
system.memory.usage By
system.cpu.utilization 1 (fraction 0-1)
Where a name is a stable OpenTelemetry metric, its unit is fixed by
the spec and this package follows it — http.server.request.duration and
http.client.request.duration are in seconds, with the semconv advisory
buckets. Reusing a semconv name with your own unit is the worst of both
worlds: no stock dashboard matches it, and a collector fed the same name
from another SDK sees one metric arrive with two units.
Names the spec does not define — queue.job.duration, command.duration,
schedule.task.duration — are free choices, but their unit is not:
"when instruments are measuring durations, seconds (i.e. s) SHOULD be
used". Every duration this package emits is in seconds, so one scale reads
across the stock HTTP metrics and our own, and a dashboard never has to
know which it is looking at.
A dimensionless 1 means a ratio. The Prometheus translation suffixes a
gauge with that unit _ratio, so a count wants a braced annotation
instead ({run_queue_item}, {entry}), which carries no suffix.
Prometheus names are derived automatically — dots become underscores and
counters get _total:
http_server_request_duration_seconds_bucket{le="0.1"}
queue_jobs_processed_total
Shape is part of the name
A metric's type travels with it and decides what arithmetic a backend is allowed to do. Getting it wrong produces a number that is individually correct and collectively a lie.
Counter monotonic; may be rate()'d system.network.io
UpDownCounter a sum that can go down system.memory.usage
Gauge a level, not a sum system.cpu.utilization
Histogram a distribution http.server.request.duration
Bytes of memory in use are the clearest case. They are a sum: add them
across ten hosts and you have the fleet's memory, which is the question you
actually have. Recorded as a gauge, a backend is entitled to average them
instead, and the answer is a tenth of the truth. Conversely, only a counter
may be rate()'d, and only a counter's reset is understood as a reboot
rather than a cliff.
Declare it with Telemetry::observable() for a reading taken at scrape
time, or Telemetry::pushed() for one written into the shared store:
Telemetry::observable('system.memory.usage', $read, MetricType::UpDownCounter, unit: 'By');
Telemetry::pushed('system.network.io', MetricType::Counter, unit: 'By');
Prometheus has only four types and no up-down counter, so it receives a
gauge — the same lossy mapping the OTLP-to-Prometheus translation makes.
The distinction survives where it can: in the OTLP payload, as
isMonotonic: false.
Your own metrics
Namespace by domain, most-general first:
orders.created — not created_orders
checkout.duration — not time_of_checkout
billing.invoices.overdue — hierarchy reads left to right
Package authors: prefix with the package domain (queue_autoscale.workers.desired)
so dashboards group naturally.
Rules
- Names match
[a-z][a-z0-9._]*— invalid names throw at registration. - A name is one instrument type forever; re-registering
orders.createdas a histogram after it was a counter throwsInstrumentTypeMismatch. - Declare units in the instrument (
unit: 's','By','1') rather than in the name; exporters surface them appropriately. - Label keys follow the same conventions (
http.route,tenant.id). Non-conforming characters are sanitized to_for Prometheus.