Tenant Deletion
Deleting a tenant is a lifecycle, not a single action. The delete cuts access immediately. The tenant's data is reclaimed afterwards, in the background, and the tenant's token does not become available for reuse until that finishes.
This page covers what happens between those two moments: what to expect, what can hold a deletion open, and what the platform deliberately keeps.
What happens, and when
| Immediately | Sign-in is revoked. The tenant is disabled and marked as deleting, and the moment of the cut is recorded. New device connections, new ingest and new command dispatch are refused within about a minute. |
| Within a minute or two | A background coordinator begins reclaiming the tenant's data from every storage system that holds it, and keeps going until each one reports that it holds nothing. |
| After at least 12 hours | The tenant's record is removed and its token becomes available again. |
The delete itself is idempotent. Deleting a tenant that does not exist, or one that is already being deleted, succeeds and changes nothing, so you can re-run a teardown script that half failed.
A tenant that still has user memberships is refused, because removing those is what actually revokes people's access to it. Remove the memberships, then delete the tenant.
Why the token is held for 12 hours
The tenant's data is typically gone within the first minute or two. The token is what stays reserved, and two separate waits must pass before the platform reports the deletion finished.
Every storage system must report clean, and keep reporting clean. One clean answer is not enough: a sweep can finish just as a write that was already in flight lands behind it. The platform therefore requires a sustained quiet period (five minutes by default), and restarts that clock whenever a pass still finds something to erase.
No connection from before the delete may still be able to write. A device that was already connected keeps its broker credential until that credential expires, up to 12 hours. Its credentials have been erased, so it cannot reconnect, but the existing connection survives. Releasing the token before then would let a straggler's writes land under a name that a new tenant could then be created at, and inherit. Preventing that is the point of the whole process, so the wait is measured from the moment of the delete and matches the credential lifetime.
Until both waits have passed, the token stays taken. Creating a tenant at it returns an error saying so.
What is erased
A tenant's data does not live in one place, so the reclamation asks each storage system in turn:
- the main database, across every functional area's schema in it;
- the telemetry database, where the event history lives;
- the message broker: the tenant's streams, plus the MQTT gateway's own session state, queued messages and subscriptions;
- the key-value store: cached lookups and resolutions;
- the live event-processing engine: open detection windows, running timers and armed absence checks, which exist in memory and are asked to evict the tenant rather than queried;
- the object store: uploaded assets such as a tenant's branding logo.
A deletion cannot be reported as finished until all of them report clean. A storage system that is unreachable is retried on the next pass, not skipped.
What is deliberately kept
Some records survive on purpose, and none of them contains the tenant's own data:
- People. A user account is instance-wide, not owned by a tenant. What is removed is the membership that linked them to it.
- Instance-level definitions: roles, tiers, OAuth clients, signing keys and system settings. These exist once per installation and outlive any tenant.
- The deletion record. Each deletion writes a durable record of what was erased and when, plus a line per storage system. It deliberately does not keep the tenant's name or contact details; otherwise the erasure's own evidence would be the last place the customer's details lived.
- The audit journal, with the identities in it destroyed. See below.
Backups are the exception to that sentence. A backup is a copy of a whole database, so a
deleted tenant's data stays in the backup archive until it ages out of that database's
recovery window: its core data for 30 days by default
(backup_retention_rdb), and its event data for 7 (backup_retention_tsdb). Until then, a restore
to a point before the deletion brings that data back. Backups sent to an object store you supplied
are also subject to that store's own lifecycle rules. With
volume-snapshot base backups, the snapshots at your cloud
provider are pruned to the same window while the cluster runs; if the cluster is deleted before
it has deleted them, they stay at the provider, with that data, until you delete them there.
The audit journal
Every change to an entity is recorded in an audit journal: when it happened, which table, which operation, how many rows, and who did it. The journal is kept through a deletion, because it is the record that the deletion happened. Sweeping it away would erase the evidence of the erasure.
What does not survive is anybody's identity. Two fields can name a person:
- the acting user, which for a human sign-in is their email address;
- the label of the affected row, which can be an email or a name you chose yourself for a customer, device or asset.
Both are emptied for a deleted tenant, permanently. They are written blank rather than scrambled or hashed. An email address is short and guessable, so anyone holding a list of addresses could match a hash back to one, which would be erasure in name only.
After a deletion, a journal entry for that tenant still shows the shape of what happened ("three devices deleted at 14:02") with nobody named. That is the retention working, not data missing.
A former member's sign-in history survives a tenant deletion. If your obligations extend to sign-in records, handle them through database and log retention rather than through tenant deletion.
Signing in happens in two steps: you authenticate as a person, then you choose a tenant. The first step is recorded, successes and failures alike, with the email address and no tenant attached, because no tenant has been chosen yet. Records with no tenant cannot be reached by any tenant's deletion. What a deletion does cover is everything a member did inside the tenant.
Two smaller limits, for completeness:
- A former member who tries to sign in to the tenant after its deletion completes creates a new record naming them, under a token that may by then belong to someone else.
- A journal entry about a person's own profile still records which account was changed. Accounts are not owned by a tenant and are not deleted with it, so an operator with access to both can still connect the two. What is destroyed is the name in the journal, not the existence of the account it referred to.
What can hold a deletion open
A deletion that is not finishing is almost always held by one of these, and each names itself in the platform's logs:
- Event processing is not running. The live engine's state can only be erased by the process that holds it. If that service is scaled to zero, the deletion waits, correctly, because the data really is still there. Starting the service resolves it on the next pass.
- No object storage is configured, but the tenant uploaded something. The reference to the object is known, but the object itself cannot be reached. Configure the storage backend that holds it, or remove the object out of band.
- A storage system is unreachable. This is treated as "try again", not as a failure. The deletion resumes by itself once the system is back.
None of these loses the deletion. The work list is the tenant's own record, so a coordinator that was stopped, a replica that was rescheduled, and a system that was down for a week all converge on the next pass.
How you are told
A stalled deletion makes nothing else look broken, which is why it needs its own alert. The coordinator visits the tenant on every pass, finds nothing it can report as an error, and the pass succeeds. The scheduled-task metrics that would tell you the coordinator had stopped therefore stay healthy throughout; they answer a different question.
TenantPurgeStalled answers this one. It fires when the oldest open deletion has been open for more
than twice the configured token hold, well past the point where every mandatory wait has elapsed.
The admin API's tenantDeletions query then names the tenant and the storage system still
outstanding. It is paged like the admin API's other lists; to read the open deletions:
query {
tenantDeletions(criteria: {pageNumber: 1, pageSize: 50, completed: false}) {
results { token epoch awaiting blockedBy stores { store complete retaining } }
pagination { totalRecords }
}
}
Its counterweight, TenantPurgeVisibilityLost, fires when those figures stop being collected at all.
Without it, a user-management service that was not being scraped would look exactly like one with no
deletions open.
Configuration
These settings live under tenantPurge in the user-management service. For all three, 0 means "use
the default".
| Setting | Default | Meaning |
|---|---|---|
intervalSeconds | 60 | How often the coordinator runs. A negative value disables it entirely. |
settleSeconds | 300 | How long every storage system must keep reporting clean. Must be greater than 140. |
tokenHoldSeconds | 43200 | How long a deleted token stays reserved, measured from the delete. |
Disabling the coordinator is a supported operational lever, for example during a maintenance window in which nothing should be deleting rows. It is safe: pending deletions stay pending rather than being lost, and no token is released while it is off.
The two windows cannot be disabled. A negative value for either is rejected when the
configuration loads, rather than taking effect. Turning off the settle window would mean declaring data erased without ever
having observed it gone. Turning off the token hold would release a name while a pre-existing session
could still write under it. Lowering tokenHoldSeconds is a real choice, since a deleted tenant's
name is unavailable for that long, but it is a choice about correctness, not tidiness.
During a deletion
Device traffic stops within about a minute. New connections, ingest on every transport, and command dispatch are all refused once the deletion is picked up. A refused device is answered exactly as one that is over its rate limit, and backs off and retries.
Already-connected devices are not disconnected. They keep their existing broker credential until it expires, up to 12 hours, and cannot reconnect once it does. This is the reason for the token hold.
If the user-management service is unreachable, these refusals stop applying until it is back. That is deliberate: otherwise user-management would be a hard dependency of device connectivity for every tenant on the instance. The refusals stop traffic early so the reclamation is not chasing data that is still arriving. The erasure itself does not depend on them.
Database write refusal
Writing to the databases is refused separately, and that refusal does not depend on any other service. From the moment the reclamation first reaches the main database and the telemetry database, each refuses every write for the deleted tenant, whether it arrives through the API, through a stream a service consumes, or through a background job of its own. The check runs inside the same database transaction as the write it refuses.
This is what makes the guarantee an erasure rather than a cleanup: reclaimed rows cannot come back to stay, even if everything that was supposed to stop the traffic earlier failed. A late write is swept again, and the deletion does not complete until the writes have stopped. The refusal is lifted only when the deletion completes, which is also when the token is released, so a new tenant created at that token writes normally from its first request.
The refusal is checked at a transaction's first write for a tenant, so a transaction that had
already written for the tenant when the deletion began can still commit, including writes it makes
afterwards. The deletion catches such writes, provided the transaction commits before the
settle window ends. The databases guarantee that for a transaction whose service stops talking to
them, for example because its pod froze or lost the network: a transaction left idle for 60 seconds
is ended and rolled back, well inside the shortest settle window allowed. The database logs
terminating connection due to idle-in-transaction timeout and nothing the transaction wrote is
kept. The request it belonged to fails with an error, and work a service takes from a stream is
delivered to it again.
The refusal covers the two databases, which is where a tenant's records live. The broker, the key-value store and the object store are reclaimed by the same passes but have no equivalent refusal. A message or object arriving late there is collected on a later pass rather than turned away, which is also why the deletion is not reported as finished until every one of them has stayed clean for the settle window.
Outbound connectors and notifications
Outbound connectors are covered by the same refusal, and it matters most there. Everywhere else, a message admitted a moment too late is data that stays on the platform until the reclamation reaches it. An outbound connector sends a tenant's data to a system you own but the platform does not (a webhook endpoint, an MQTT broker, a Kafka topic, an SNS or SQS queue), and once it has been sent, no later pass can take it back. So a connector dispatch for a deleted tenant is refused and discarded rather than held for inspection.
Notifications stop too. A deleted tenant's alarms no longer send email or fire notification webhooks, and open alarms stop escalating. This matters because it needs no device traffic to happen: escalation re-pages open, unacknowledged alarms on a timer. Without this refusal, a deleted tenant's on-call recipients would keep being paged about alarms belonging to a tenant that no longer exists.
An action already in flight when the delete lands will complete, and refusals begin within a minute of the delete rather than instantly. If a deletion must guarantee that nothing further reaches an external system or recipient, disable the tenant's connectors and notification policies before deleting it.
What you can see today
A tenant that is being deleted has a Deletion tab on its detail page in the admin console. It answers the question an operator actually has, is it done, and if not, why not, with:
- a plain-language status line;
- anything that is blocking;
- a row per storage system showing whether that system is clean, still holding data, or retrying after a failure.
Those three row states are kept apart on purpose. Data still held will not clear until someone changes something, while a failure clears itself on the next sweep.
Some systems add a short note to a clean line, saying what they declined to look at: a telemetry database that does not exist on this instance, for example, or the internal caches a deletion deliberately exempts. A note is a footnote to "clean", not a problem to act on.
Completed deletions live at Admin → Deletions, not on a tenant page. Finishing a deletion removes the tenant, so a finished deletion has no tenant page left to appear on. That instance-wide list is the durable record: token, when it was requested, when it completed, how much was erased, and what each storage system reported. Open it when someone asks you to show that a customer's data was erased.
Two actions are deliberately missing from it:
- No retry. Every sweep already retries, so a button would imply the absence of one.
- No force-complete. It is the one action that could record an erasure that did not happen.