Releases & Upgrades
DeviceChain ships as a set of prebuilt, versioned container images plus a Helm chart. You do not need to build anything to run it — pull a released version, install the chart, and upgrade in place with zero downtime.
v0.9.0 replaced every service's database migration chain with a single frozen baseline, so
an existing v0.8.x instance cannot be moved onto it with helm upgrade — the migration
fails on already exists. Recreate the instance instead. See
The v0.9.0 baseline squash below before you do anything else.
Versioning model
Every release is a single semantic-version git tag (vX.Y.Z). That one version covers
everything together — each service image, the operator, the Helm chart, and the
dcctl CLI are all published at the same version. There is no per-service version skew to
reason about: a deployment is one coherent number.
- Stable releases are
vX.Y.Z(e.g.v1.2.0). The:latesttag tracks the most recent stable release. - Pre-releases are
vX.Y.Z-rc.N(e.g.v1.2.0-rc.1). These never move:latest.
Pre-1.0 stability
Until v1.0.0, any release — including a patch release — may change APIs, schemas, or behavior without a compatibility shim. This is deliberate: while the data model is still settling, we prefer a clean cutover to carrying a shim we would have to support forever.
Every breaking change is called out at the top of that release's notes. Read them before upgrading. They are the authoritative list; the version number alone does not tell you whether a release is safe for your deployment.
Concretely, before v1.0.0 you should expect that a release may:
- tighten validation, so a request that previously succeeded is now rejected — usually because it was being silently accepted or silently discarded
- change or remove a GraphQL field, rather than deprecating it for a cycle
- alter database schema in ways that a downgrade will not undo
- replace the migration baseline outright, which removes the upgrade path entirely rather
than merely making it one-way. When that happens the release notes say so at the top, and
the only route forward is to recreate the instance.
v0.9.0is such a release
The "upgrade in place with zero downtime" property above describes the mechanics of a rolling upgrade. It is not a promise that your existing API calls keep the same meaning across a pre-1.0 version bump.
Once v1.0.0 ships, this section is replaced by a normal semantic-versioning compatibility promise: breaking changes only in a major version.
Because releases are frequent before GA, the minor version marks a milestone (a significant feature or subsystem landing) and the patch version carries the ongoing cadence of fixes and hardening. A patch release is not automatically a low-risk upgrade during this period — again, the release notes are what tell you.
Images
Images are published to the public GitHub Container Registry under
ghcr.io/devicechain-io — for example ghcr.io/devicechain-io/device-management. They are
multi-arch (linux/amd64 and linux/arm64) and built on a distroless nonroot base, so
they run as an unprivileged user with no shell and a minimal attack surface.
Because the registry is public, no credentials are required to pull released images.
Installing a specific version
Pin the image tag to the release you want:
helm install dc deploy/helm/devicechain \
--set instance.id=devicechain \
--set image.tag=v1.2.0
The Helm chart itself is also published as an OCI artifact, so you can install it without a checkout of the repository:
helm install dc oci://ghcr.io/devicechain-io/charts/devicechain \
--version 1.2.0 \
--set instance.id=devicechain \
--set image.tag=v1.2.0
Zero-downtime upgrades
Upgrading to a new version is normally a plain helm upgrade, and the chart and services are
built to roll customers forward without dropping traffic. Two releases so far are exceptions,
both documented below: the durable-ingest cutover, which is still a helm upgrade but has a
visible side effect, and v0.9.0, which cannot be upgraded into at all. Check the release
notes for the version you are moving to before running the command:
helm upgrade dc deploy/helm/devicechain \
--set instance.id=devicechain \
--set image.tag=v1.3.0
What makes the rollout safe:
- Surge-before-terminate. Each Deployment uses a
RollingUpdatestrategy withmaxUnavailable: 0andmaxSurge: 1, so a new pod must pass its/readyzreadiness probe before an old pod is removed. Capacity never dips during the rollout. - Graceful shutdown / connection draining. When a pod is asked to terminate it first
reports "not ready" (so the Service stops routing new requests to it), waits a short
drain window for that change to propagate, and only then finishes in-flight work and
shuts down. Configure the window with
shutdownDrainSeconds(default5), kept safely underterminationGracePeriodSeconds(default30). - Coordinated schema migrations. Services run database migrations under a database-level lock, so when several replicas start at once exactly one applies migrations and the rest wait — no races, no duplicate DDL.
For true zero-downtime, run replicas: 2 (or more) for each area so the rollout always has
a live pod serving traffic. A single replica still has a brief gap while its one pod is
replaced. Set it globally with --set replicas=2, or per area under
functionalAreas.<area>.replicas. A PodDisruptionBudget is rendered automatically for any
area with more than one replica, so node drains can't evict every replica at once.
The v0.9.0 baseline squash
v0.9.0 is the one release that cannot be reached with helm upgrade.
Before it, each service's schema was built by a chain of migrations applied in order. v0.9.0
replaces every one of those chains with a single frozen baseline — one migration per
service that creates the whole schema as it stands. A database created by v0.8.x has already
applied the old chain, so when it meets the baseline it tries to create tables that are
already there and fails with already exists. The failure is loud and happens at startup; it
does not corrupt anything.
There is no migration path, and before v1.0.0 there will not be one. Carrying a compatibility
shim for a schema shape that is still moving is exactly the cost this project has chosen not to
take on while every install is still an early one.
To move to v0.9.0, recreate the instance:
# Export anything you need first — this discards the databases.
dcctl destroy local devicechain
dcctl bootstrap local devicechain
The destroy guard protects the databases from an ordinary helm operation,
not from a deliberate dcctl destroy. If the instance holds telemetry, device definitions or
dashboards you care about, dump them before you start. There is no in-place path that preserves
them across this release.
This is a one-time cost, taken deliberately while the platform's only installs are early ones.
From v0.9.0 onward a schema change appends a new migration to the baseline, which is a
normal helm upgrade again — so this section describes a single release, not a new policy.
The one-time durable-ingest cutover
The release that introduces durable MQTT ingest changes how event-sources receives
device telemetry: instead of subscribing to the broker as an MQTT client, it consumes a
durable capture stream that the broker writes to before it acknowledges the device. This
is what stops telemetry being lost when event-sources is down.
Crossing that release once is a normal helm upgrade — but expect a brief window of
duplicated telemetry, and plan for it:
- During the rollout the outgoing pod is still ingesting over MQTT while the incoming pod has already begun consuming the capture stream, so messages published in that overlap are ingested by both. The window is bounded by how long the two pods coexist — the incoming pod's startup plus the outgoing pod's drain.
- Events that carry both an
altIdand a device-suppliedoccurredTimeare unaffected: the write-side dedup key is(tenant, altId, occurredTime), so those duplicates collapse. An event with analtIdbut nooccurredTimedoes not collapse — the decoder stamps the current time when the device omits one, and the two copies are decoded in different pods at different instants, so they get different timestamps and land as two rows. Telemetry with noaltIdis not deduplicated at all. - The overlap is preferred deliberately. The alternative ordering — stopping the old pod before the capture stream exists — loses every message the broker acknowledges in the gap, and that loss is silent: the device is told the message was accepted and it is never stored. A duplicate reading is visible and correctable; a missing one is neither.
event-sources to Recreatestrategy: Recreate on event-sources produces exactly the lossy ordering above, because
it terminates the old pod before the new one creates the capture stream. The chart refuses
to render this configuration rather than let it drop telemetry silently. event-sources
is not a single-writer service and gains nothing from Recreate — once cut over it can run
multiple replicas, which the MQTT-client path it replaces could not.
Data durability
The database tier is intentionally lifecycle-independent from the application. Both databases are provisioned as separate infrastructure with a destroy guard, so upgrading, reinstalling or uninstalling the application never touches them. That is the common case and it is safe.
The guard protects each database while it is in the infrastructure configuration. It does not protect one that has been taken out of it: a resource removed from the configuration is no longer covered by rules the configuration declares, and the removal plan will succeed. The database clusters also own their volumes, so removing one takes its data with it rather than leaving an unattached volume behind.
Do not edit the database out of the infrastructure configuration as a way of replacing it.
Upgrading an instance created before the databases moved onto the operator is the one case
where this comes up, and it is refused at plan time rather than left to chance. Dump both
databases first, then re-run the bootstrap with --allow-legacy-db-removal — which asserts
you have handled the data, and verifies nothing. For a local instance, dcctl destroy
followed by a fresh bootstrap is simpler and discards the data deliberately.
This is durability of the running volumes — it is not a substitute for scheduled backups and point-in-time recovery, which are provisioned with the production infrastructure. See Deployment & Operator for how the infrastructure and application layers are separated.