Skip to main content

Metrics Reference

Every service exports Prometheus metrics on its metrics port, and the monitoring stack scrapes them without per-service setup (see Observability & Metrics). This page lists the service-specific metrics that are not described with a feature elsewhere. The ones that belong to a feature are on that feature's page: ingest backpressure, maintenance passes, caches, the detection loop and the edge services.

Names are devicechain_, then the service's name with its dashes removed, then the name in the first column, so batch_refusals_total in command delivery is devicechain_commanddelivery_batch_refusals_total. A name in braces after a metric lists its labels. No metric on this page is labelled by tenant or by device, so none of them is a cardinality risk to scrape. Counters end in _total and only ever rise, so read them as a rate. Gauges are read as they are.

Most of what follows is diagnostic: it lets you tell what a service did without reading its log. Where a metric is one an operator should act on, its row says what to do, and the alert that watches it, if there is one, is named.

Command delivery​

Prefix: devicechain_commanddelivery_.

The three command delivery alerts read command_delivery_claims_stranded_total, command_delivery_presence_read_errors_total and batch_refusals_total{bound="reserve"}. A Command Delivery dashboard charts the rest. Several of these are expected to be busy and are not faults: a steady not_sole share of declined nudges, claims lost to the sweep racing a nudge, and holds placed while devices are offline. command_delivery_responses_refused_total is expected to read zero. A rising rate on command_delivery_stranded_observed_total or command_response_lost_unsettled_total means commands are timing out against devices that did nothing wrong. (The doubled command_delivery_ in the older names is historical and stays, because renaming a series breaks every dashboard that reads it.)

MetricMeans
batch_cancel_commands_total{disposition}Commands examined by a batch cancel, by what the cancel was able to do with each. A standing already_sent share is a brake that keeps arriving late.
batch_devices_total{disposition}Devices a command batch resolved to, by whether the batch enqueued to them.
batch_enqueues_total{target_kind, outcome}Command batches decided, by how the target was named and the outcome.
batch_refusals_total{code, bound}Per-device batch refusals, by code and — for a ceiling refusal — which bound caused it. bound=reserve means the device would have fitted against the tenant's own ceiling and was refused only by the delivery reserve.
command_dead_letters_not_ours_totalDead letters read from the shared stream that describe some other kind of work. Acked and ignored — every producer writes to one stream.
command_dead_letters_unreadable_totalDead letters this consumer could not act on — no parseable tenant, a body that is not an envelope, or an envelope naming no command. Acked and counted rather than retried, because no redelivery makes a malformed message parse.
command_delivery_claims_lost_total{path}Dispatches abandoned because another dispatcher claimed the command first, by the dispatch path that lost. "sweep" is the periodic pass, "nudge" is the dispatch issued when a command is enqueued; the two racing for one row is expected and safe (the claim is a compare-and-set), so read a rate on one path with none on the other rather than the total.
command_delivery_claims_stranded_totalCommands left reading SENT because their publish failed and the release failed too. On LwM2M the stranded reconciler re-arms these; on MQTT they still expire as TIMEOUT, wrongly blaming the device.
command_delivery_dispatches_exhausted_totalCommands failed because the platform could not publish them to their device as many times as the configured bound allows, so it stopped retrying. Each row records FAILED with the platform named as the cause, rather than retrying until its TTL and then recording TIMEOUT against a device it never reached.
command_delivery_holds_placed_totalCommands withheld from dispatch because the device is authoritatively absent.
command_delivery_holds_released_totalWithheld commands returned to the dispatch queue because their device came back.
command_delivery_nudges_applied_totalDispatch nudges that found exactly one queued command and put it through the delivery gates. NOT a count of publishes: the presence gate may still hold or fail the command, exactly as it would on a sweep tick.
command_delivery_nudges_declined_total{reason}Dispatch nudges the drain refused to act on, by reason. A high and steady reason="not_sole" rate is expected — the nudge stands down whenever a device has more than one queued command — and is not a fault.
command_delivery_nudges_dropped_totalDispatch nudges discarded because the queue was full. A LATENCY signal, not an error rate: the command still goes out on the delivery sweep, which is the net under every nudge.
command_delivery_nudges_requested_totalDispatch nudges accepted onto the enqueue-time dispatch queue, one per command created through createCommand (a fleet batch issues none).
command_delivery_presence_read_errors_totalSweep passes that could not read the presence projection; the gate fails OPEN, so a standing rate here means commands are being dispatched ungated.
command_delivery_responses_dead_lettered_totalDevice command responses written to the dead-letter stream after every attempt to record them failed, so an answer the device did give can be seen rather than leaving its command looking unanswered.
command_delivery_responses_refused_totalDevice responses rejected because the publishing device does not own the command they name. Expected to be zero: either a device is answering for another device, or dispatch addressed a command to the wrong one.
command_delivery_stranded_observed_totalCommands found sitting in SENT with no outcome for longer than the platform could still have been retrying them.
command_response_lost_not_actionable_totalDead-lettered command responses this consumer left alone because their reason says the platform declined to write the answer rather than tried and failed. Nothing is settled: there is no lost outcome to record, and the command named may still be live. A rising rate on a reason nobody expected is worth looking at, because this gate is deliberately closed by default.
command_response_lost_not_answerable_totalDead-lettered responses whose command was not in a state a response could settle: it had already reached a terminal outcome some other way, or it has gone back to being live (re-dispatched or held) since the answer was lost. Nothing is written in either case — a late dead letter must not overwrite an outcome that really happened, nor fail a command the platform still intends to deliver.
command_response_lost_settled_totalCommands driven to a terminal state because the device's answer to them was dead-lettered, so a command whose response the platform lost stops reading as though it were still in flight.
command_response_lost_unsettled_totalDead-lettered responses that exhausted every delivery attempt without their command's disposition being written. Those commands read as in flight until their TTL and then lapse to TIMEOUT, blaming a device that did answer.

Device management​

Prefix: devicechain_devicemanagement_.

geofence_set_publish_failures_total feeds the GeoFenceSetPublishFailing alert. A rise in credential_misconfigured_total means a database was restored next to a root key other than the one that sealed it, or a device credential was never stored; wrong passwords are not counted here. The three dead-letter counters mean a state change reached the database but not the bus: read them with dcctl dead-letters.

MetricMeans
alarm_event_dead_lettered_totalAlarm state-change events that could not be published to the alarm-events stream and were written to the dead-letter stream instead. Each one is an alarm transition that reached the database but not the bus, so nobody was paged about it — it is visible to an operator rather than only logged.
callout_in_flightDevice auth-callout requests being authorized right now.
callout_refused_busy_totalDevice auth-callout requests refused with the generic denial because the in-flight bound was reached.
credential_misconfigured_total{path}Device credential checks refused because the stored secret can never match: none is stored, or its digest was made under a different key than this instance's root key derives (most likely a database restored next to the wrong root key). path="connect" is the MQTT auth callout, path="event" the per-event check. Wrong passwords are not counted here.
geofence_set_publish_failures_totalGeofence-set manifests that could not be published — a marshal error, a broker refusal, or a transport fault. Each one means event-processing was not told about a fence edit, so containment for that tenant holds its previous fence set until a reconcile sweep repairs it. A sustained non-zero rate means fence edits are not reaching the detection engine.
raise_alarm_dead_lettered_totalRaise-alarm edges written to the dead-letter stream after every attempt to apply them failed, so an alarm that should have been raised or cleared is visible rather than only logged.
resolve_event_time_bounded_totalReported event times refused for leading the server clock by more than the configured tolerance, and replaced with the latest time the tolerance allows.

Event processing: detection​

Prefix: devicechain_eventprocessing_.

The metrics that alerts read are detect_is_leader, detect_live, detect_checkpoints_total (the DetectLeaderIsNotConsuming alert reads all three) and the detect_fence_* counters that back the geofence alerts. detect_consumer_pending is the lag signal that DetectConsumerBacklogHigh watches, and detect_watermark_lag_seconds is the one behind DetectWatermarkLagHigh. detect_loop_heartbeat_timestamp_seconds is the engine's own liveness: if it is stale while detect_is_leader and detect_live both read 1, the loop is hung inside a call.

MetricMeans
detect_applied_stream_seqHighest JetStream stream sequence captured in the committed snapshot.
detect_checkpoint_failures_total{stage}Scheduled DETECT checkpoints that did not commit, by stage (publish, serialize, save). A checkpoint call that outlives its deadline counts against the stage it was stuck in. A stale-writer refusal is not counted here: that ends the process.
detect_checkpoints_totalCommitted DETECT snapshot checkpoints.
detect_consumer_ack_pendingDelivered-but-unacked messages on the resolved-events durable consumer (in-flight work).
detect_consumer_pendingUndelivered messages waiting on the resolved-events durable consumer (the primary DETECT lag signal).
detect_derived_events_published_totalDerived signal events published.
detect_derived_events_rejected_total{reason}Detections dropped before publish, by reason (bounded enum).
detect_events_applied_totalResolved events fed into the DETECT engine.
detect_fact_persist_retries_totalRetries of a fact projection write (rules, roster, attributes, deletions) that failed with a non-terminal error. The fact stays unacked and this consumer is blocked behind it until a retry commits.
detect_fanout_events_totalPer-rule core events produced by the resolved-event fan-out.
detect_fence_archive_skew_totalGeofence archive reads that failed because device-management does not serve the manifest doors — it is running a build from before manifest delivery. Repairs itself when that service rolls forward.
detect_fence_geometry_cache_evictions_totalCompiled-geometry cache entries dropped to stay inside the cache's vertex bound.
detect_fence_geometry_cache_hits_totalCompiled-geometry cache lookups served from cache, avoiding both a cross-service read and a recompile.
detect_fence_geometry_cache_misses_totalCompiled-geometry cache lookups that had to fetch and compile the document.
detect_fence_geometry_cache_verticesTotal vertices held in the compiled-geometry cache — the quantity its bound is counted in.
detect_fence_geometry_hash_mismatch_totalGeofence geometry documents that did not hash to the content address they were requested under. Always a bug, never transient: the peer served the wrong row, something re-encoded the document in transit, or the archive is corrupt.
detect_fence_geometry_unresolved_totalGeofence manifest entries whose geometry could not be obtained from device-management's archive. Each one leaves that fence reporting unresolvable rather than answering, repaired by the next reconcile sweep if the cause was transient.
detect_idle_advances_totalWall-clock idle advances that produced at least one detection.
detect_idle_detections_totalDetections produced by wall-clock idle advance (absence/duration/session firing on silence).
detect_is_leader1 while this replica holds the DETECT partition lease, from acquisition rather than from the end of the term build.
detect_live1 while this replica is consuming inside a held leadership term; 0 while standing by OR while building a term it has already acquired.
detect_live_gap_fills_total{outcome}Times live consumption met a message whose stream sequence was more than one past the engine's and read the missing range from the stream, by outcome: filled (the range was read and held messages), absent_only (the stream held none of it: purged or evicted), failed (an attempt to read the range did not complete; live consumption parks and the attempt repeats every tick, so a sustained outage counts every attempt, not once).
detect_live_gap_sequences_total{outcome}Stream sequences inside live gap fills, by outcome: applied (a message the broker had counted as delivered that never reached the loop, now applied), absent (not in the stream: purged or evicted), skipped (read back unprocessable, and recorded as handled without applying anything). A non-zero applied rate means deliveries are being lost between the broker and this loop.
detect_loop_heartbeat_timestamp_secondsUnix time of the last pass of the DETECT single-writer loop, refreshed at least once per tick while a term is live. Stale while detect_is_leader and detect_live read 1 means the loop is hung inside a call.
detect_restore_secondsTime to restore engine state from the snapshot store at startup.
detect_rules_activeRules loaded into the DETECT engine.
detect_stale_absence_dropped_totalAbsence detections dropped at publish because the device left the rule's scope: it was deleted, re-typed, or the rule version was superseded.
detect_superseded_frontier_dropped_totalDetections of the duration, session and aggregate kinds that fire from the passage of time, dropped at publish because their profile version has been superseded.
detect_tenants_over_live_key_budgetTenants currently exceeding the per-tenant live-key budget.
detect_tenants_over_retained_sample_budgetTenants currently exceeding the per-tenant retained-sample budget.
detect_tenants_over_rule_budgetTenants currently exceeding the per-tenant rule-count budget.
detect_watermark_lag_secondsWall-clock time minus the engine watermark at the last checkpoint.

Event processing: reactions​

Prefix: devicechain_eventprocessing_.

react_events_poison_dropped_total is the series the ReactPoisonDropping alert fires on. react_actions_dropped_total{reason="unknown_kind"} should always be zero: a non-zero value is a defect to investigate, not load. react_actions_not_enabled_total should never show raiseAlarm or clearAlarm, because that sink is always wired.

MetricMeans
react_actions_dispatched_total{action}REACT actions handed to their sink, by action type (includes idempotent replays).
react_actions_dropped_total{reason}REACT actions skipped without being attempted, by reason. unknown_kind: the rule definition carried an action type this build cannot dispatch (a forged or hand-edited definition; unreachable through the supported authoring path), so a non-zero value is a defect to investigate, not load.
react_actions_not_enabled_total{action}REACT actions recognized but dropped because this deployment has no sink for them, by action type: sendCommand without command-delivery configured, or httpCall/publish without outbound connectors enabled. The alarm sink is always wired, so raiseAlarm/clearAlarm should never appear here.
react_connector_egress_shed_total{action}Connector dispatch attempts (httpCall, publish) shed at the source for being over the tenant's outbound rate, by action type. Counted per attempt: a redelivery may shed and later admit the same action, so this is not a count of permanently dropped actions.
react_connector_shed_dead_lettered_total{action}Connector actions (httpCall/publish) shed at the source and recorded as an individual dead letter with reason shed, by action type.
react_events_dead_lettered_totalDerived events written to the dead-letter stream after the redelivery cap, so their actions can be inspected rather than vanishing.
react_events_orphaned_totalDerived events whose rule was gone from the projection (nothing dispatched).
react_events_poison_dropped_totalDerived events dropped after the redelivery cap, because their dispatch kept failing. They are also written to the dead-letter stream, so this counts the same events as react_events_dead_lettered_total; it is kept because the ReactPoisonDropping alert fires on it.

Event sources​

Prefix: devicechain_eventsources_.

The presence metrics belong to the broker-asserted MQTT presence that this service reads from the broker. presence_tap_off is the one to alert on: it is 1 when that presence is not running on a replica, labelled by why, and a long-lived MQTT fleet emits no advisories to show it. The canary pair tells you whether presence is being read at all: presence_canary_missed_total rising means it is not. presence_reconcile_regressed_sessions standing above zero means the repairs are not converging. The total_msg_* counters are per source.

MetricMeans
command_wake_dropped_totalCommand wakes discarded because the queue was full; their commands are released later by command-delivery's reconcile pass, so this is a latency signal rather than a loss.
command_wake_failed_totalCommand wakes that could not reach command-delivery; the reconcile pass covers them.
command_wake_released_totalWithheld commands returned to the delivery queue because their device reconnected.
command_wake_requested_totalReturning devices queued for a command wake.
presence_advisories_skipped_total{reason}Broker connection advisories that produced no presence event, by reason.
presence_canary_missed_totalCanary probes whose own broker connection was NOT observed — presence is not being read.
presence_canary_observed_totalCanary probes whose own broker connection was observed end to end.
presence_events_failed_totalPresence transitions that could not be written to the inbound stream.
presence_events_refused_totalPresence transitions refused by the ingest admission gate (deleted tenant or tenant ceiling).
presence_reconcile_regressed_sessionsDevices found LIVE on a session id lower than the one the projection holds, as of the last reconciliation pass. A standing non-zero value means the repairs are not converging.
presence_reconcile_repaired_total{direction}Presence transitions emitted by reconciliation because an advisory was missed, by direction.
presence_reconcile_runs_total{outcome}Presence reconciliation passes, by outcome.
presence_reconcile_withheld_disconnects_totalDevices that would have been marked offline had the broker inventory been provably complete.
presence_released_totalDevices handed back from asserted to inferred presence because this source stopped reading the broker.
presence_sessions_regressed_totalPresence transitions observed with a session id lower than this replica's high-water mark (a broker node's clock may be trailing its peers). Diagnostic only: whether such a transition is applied is decided downstream by the projection.
presence_still_assertedDevices this source still had asserted at the start of the last release pass. The work empties itself, so a healthy drain walks this to zero and stays there.
total_http_connections_closed_before_request{source}Count of connections to an HTTP ingest listener that closed before delivering a request — a header-timeout close, but equally a port scan, a TCP health check or a client that hung up.
total_msg_decode_successful{source}Count of total messages successfully decoded.
total_msg_failed_decode{source}Count of total messages that failed to decode.
total_msg_tenant_deleted{source}Count of inbound messages refused because their tenant has been deleted and its data is being reclaimed.

LwM2M ingest​

Prefix: devicechain_lwm2mingest_.

The leader and serving gauges are described with the edge services, as are the metrics that page. The ones below are the rest. The command counters tell the story of a command through the adapter: attempted, then succeeded or failed, with the offline, parked and drained counters showing commands that waited for a device to wake. command_park_errors_total and command_response_publish_failures_total are the two that mean a command can end as a TIMEOUT although the device was not at fault.

MetricMeans
bad_requests_totalMalformed /rd requests (e.g. a registration-item request with no location).
coap_requests_total{code}Transport-level CoAP requests handled by the L0 health probe, by response code. The /rd registration outcomes are metered by the registrations/updates/deregistrations/expiries counters.
command_drain_claim_errors_totalBacklogged commands NOT dispatched because ownership could not be established with command-delivery (fail-closed; retried shortly, in order).
command_drain_claims_lost_totalBacklogged commands another dispatcher or the delivery sweep claimed first, so this drain did not dispatch them — a duplicate actuation avoided, not a fault.
command_drain_errors_totalDrain fetches that failed; retried after a short delay while the device stays live.
command_park_errors_totalUndeliverable commands that could NOT be handed back to command-delivery; left unacked to retry on redelivery, and stuck in SENT until one succeeds.
command_park_settled_totalHand-backs that moved no row because the command had already been answered, cancelled, expired or re-claimed — a settled outcome, not a fault.
command_park_skipped_totalUndeliverable commands NOT handed back because no parker is wired or the delivery envelope carried no dispatch nonce; they stay SENT and ride their TTL.
command_response_publish_failures_totalCommand outcomes that could not be published to command-responses after local retries (the op already ran, so the command is not redelivered; it will TIMEOUT).
commands_attempted_total{op}LwM2M commands dispatched to a live device (= succeeded + failed), by operation.
commands_drained_totalBacklogged (held or parked) commands dispatched to a live device by a drain, oldest first.
commands_poison_totalCommands dropped as unprocessable (no parseable tenant in the subject, or an undecodable envelope).
commands_served_offline_totalCommands for a device this adapter serves that had no live connection — parked in command-delivery and drained on the device's next wake, or EXPIRED at their TTL if it never wakes.
commands_succeeded_total{op}LwM2M commands the device acknowledged with a 2.xx, by operation.
commands_tenant_deleted_totalCommands ack-dropped because their tenant has been deleted and its data is being reclaimed — the platform declining to actuate an offboarded customer's hardware.
deregistrations_totalLwM2M explicit deregistrations (DELETE /rd/{id}).
devices_registered_totalDevices auto-created on their first LwM2M registration.
handshakes_totalCompleted DTLS handshakes (each is a new authenticated LwM2M session).
measurements_emitted_totalIndividual measurement samples durably written from LwM2M Observe/Notify.
notifies_received_totalLwM2M Notify messages received on an observed object instance.
observe_terminal_notifications_totalNotifications that terminated an observation (RFC 7641, e.g. 4.04 after the observed instance was deleted).
presence_dropped_totalPresence transitions dropped: an unregistered device on a no-auto-register credential, or a durable-emit budget exhausted.
sessions_rejected_totalNew sessions refused because the live-session ceiling (maxSessions) was reached.
telemetry_devices_registered_totalDevices auto-created on a first telemetry sample (rare: LwM2M devices are created at /rd registration).
telemetry_tenant_deleted_dropped_totalTelemetry samples dropped because the tenant has been deleted and its data is being reclaimed.
telemetry_unknown_dropped_totalTelemetry samples dropped for an unregistered device (auto-registration off for the credential).

Sparkplug ingest​

Prefix: devicechain_sparkplugingest_.

The rest of the Sparkplug metrics, including the leader gauge, are described with the edge services. samples_skipped_total exists because a node that publishes only values with no numeric sample otherwise looks identical to an idle one.

MetricMeans
devices_registered_totalDevices auto-registered on first sight of their Sparkplug identity.
samples_skipped_total{reason}Metrics of accepted Sparkplug messages that produced no sample, by reason: non_numeric (a boolean, string, bytes, dataset or template value), null (is_null) or unnamed (no name and no resolvable alias). A node publishing only such metrics otherwise looks identical to an idle one.

Notification management​

Prefix: devicechain_notificationmanagement_.

Notifications that fail every delivery attempt are written to the dead-letter stream, where dcctl dead-letters lists them. This counter tells you that pages were never sent.

MetricMeans
notifications_dead_lettered_totalAlarms written to the dead-letter stream after every delivery attempt failed, so an operator can see which pages were never sent.

User management​

Prefix: devicechain_usermanagement_.

These count the dead-letter consumer, which stores a failure a service gave up on in a queryable store so it outlives the stream's own seven-day window. dead_letters_unstored_total is the one that matters: a failure the store could not take is recorded nowhere, and DeadLetterStoreLosing fires on it.

MetricMeans
dead_letters_stored_totalDead letters written to the queryable store, so a failure a consumer gave up on outlives the stream's own seven-day window.
dead_letters_unstorable_totalDead-letter messages this consumer could not make sense of — no parseable tenant, or a body that is not an envelope. They are ACKED and counted rather than retried, because no redelivery makes a malformed message parse.
dead_letters_unstored_totalDead letters that exhausted every delivery attempt without being stored — the store was unreachable for longer than the retries last. The failure they described is now recorded nowhere, which is the one thing this consumer exists to prevent.

Every service​

Prefix: devicechain_<area>_, where <area> is the service's name without dashes (devicechain_devicemanagement_ready). Every service exports these.

safety_gate_enabled is what the SafetyGateDisabled alert reads. The JetStream gauges report each stream's size against its configured limits and its replica health; the replication alerts under Replication are built on the latter.

MetricMeans
auth_gate_attempts_totalBackground auth-gate JWKS fetch attempts.
auth_gate_failures_totalBackground auth-gate JWKS fetch failures.
ready1 when the data plane is ready (auth live), else 0.
safety_gate_enabled{gate}1 when an optional safety gate is wired in this service, 0 when it is OFF because its configuration is absent. An OFF gate lets work through that the gate exists to refuse (tenant_lifecycle: work for a deleted tenant; presence: commands to absent devices; rule_validation: uncompilable detection rules at profile publish).
jetstream_peers_current{stream}RAFT peers currently caught up and online for this stream or KV bucket, including the leader. Below jetstream_replicas_actual means the stream is labelled replicated but is not.
jetstream_replicas_actual{stream}Replica count JetStream reports for this stream or KV bucket.
jetstream_stream_limit_bytes{stream}Configured MaxBytes ceiling for a JetStream stream.
jetstream_stream_limit_messages{stream}Configured MaxMsgs ceiling for a JetStream stream.
jetstream_stream_used_bytes{stream}Current on-disk bytes stored in a JetStream stream.
jetstream_stream_used_messages{stream}Current message count in a JetStream stream.