- AI Workspace
- next
- Production Deployment
Operate the deployment¶
Once AI Workspace is serving traffic, three things keep it that way: logs you can search, backups that restore, and an upgrade you can undo.
Step 1: Emit structured logs¶
Both services log to standard output, so whatever collects container logs on your platform already collects theirs. Switch the format to JavaScript Object Notation (JSON) so a log aggregator can index fields instead of matching text.
Keep the level at info in production. debug is verbose enough to affect throughput and to put request detail in your logs. Leave browser_debug off, because it turns on verbose console logging in every user's browser.
Step 2: Watch the right signals¶
Both services answer a health endpoint that doesn't require authentication: /healthz on AI Workspace and /health on the Platform API. Everything else comes from logs and from your platform's own metrics.
Alert on these:
| Signal | Why it matters |
|---|---|
| Either health endpoint failing | The service is down or can't reach a dependency |
| A restart loop on either service | A configuration or secret stopped resolving; the startup message names which |
| Database connection errors | The pool is exhausted or the database is unreachable |
| Certificate expiry inside 14 days | Renewal stopped, and an expired certificate takes the workspace offline |
| A gateway moving to inactive | The gateway lost its connection to the control plane, so it stops receiving configuration updates |
| Sustained CPU or memory at the limit | The deployment is undersized for its traffic |
| Authentication failures rising sharply | An identity provider problem, or credential stuffing |
Gateway connection log lines¶
The Platform API's WebSocket manager logs each gateway transition at info level, so these three messages are available without changing the log level:
| Message | Attributes | What it tells you |
|---|---|---|
Gateway connected |
gatewayID, connectionID |
A gateway established its control plane connection |
Gateway disconnected |
gatewayID, connectionID |
A gateway dropped. One at a time is normal during a gateway restart; several at once points at the control plane |
Heartbeat timeout detected |
gatewayID |
The gateway stopped answering pings, and the Platform API is about to close the connection. This precedes a gateway going inactive |
Alert on Heartbeat timeout detected rather than waiting for the console to show a gateway as inactive, since the timeout is the earlier signal.
WebSocket metrics¶
The Platform API also logs connection metrics on an interval, which is the closest signal to how many gateways are connected:
The Platform API writes this line at debug level, so it needs level = "debug" as well as metrics_log_enabled. At info the setting produces no output. Debug logging on the Platform API is verbose, so turn it on to investigate a connection problem rather than leaving it on permanently.
The line's message is WS Metrics, and its payload attribute holds a JSON object with these fields:
| Field | Meaning |
|---|---|
from, to |
The interval the counters cover, in Request for Comments (RFC) 3339 format |
totalActiveConnections |
Gateways connected at the end of the interval |
totalActiveOrgs |
Organizations with at least one connected gateway |
successfulConnections |
Connections established during the interval |
failedConnections |
Connection attempts that failed during the interval |
disconnections |
Connections lost during the interval |
eventsSent |
Configuration events pushed to gateways during the interval |
The counters reset each interval, so they're rates rather than running totals. totalActiveConnections is the one to graph: a drop with no matching deployment change means gateways are losing the control plane. A failedConnections count that stays above zero points at a registration token or network path problem rather than a gateway fault.
Step 3: Back up what you can't regenerate¶
Two things matter, and they're only useful together. A database backup restored without its encryption key gives you rows whose secret columns can't be decrypted.
| What | How often | Where |
|---|---|---|
| The database | On your organization's schedule for a system of record | Your database backup tooling |
| The at-rest encryption key | Once, when created, and after any change | Your secret manager, separate from the database backup |
Back up the database with your standard tooling, and treat the key files as secrets rather than as part of a host backup. This example uses PostgreSQL; use the equivalent for whichever database you run:
pg_dump --host postgres.example.com --username platform_api \
--format=custom --file platform_api-$(date +%F).dump platform_api
Keep configs/config.toml and docker-compose.yaml in version control. Keep api-platform.env, resources/keys/, and secrets/ out of it.
If you still run SQLite on a single host, back up the database file with the stack stopped, or through SQLite's own backup command. Copying a file that's being written produces a backup that may not restore.
Back up the database with your standard tooling. Back up the Secrets separately, into your secret manager rather than into a cluster snapshot.
The Platform API's persistent volume claim carries a helm.sh/resource-policy: keep annotation, so it survives helm uninstall and a later install with the same release name re-adopts it. That protects you from an accidental uninstall; it isn't a backup, because it doesn't survive losing the cluster.
To remove that data deliberately:
Test a restore into a non-production environment on a schedule. Only a tested restore proves the backup works.
Step 4: Upgrade¶
Read the release notes first, then back up the database, then upgrade. Both procedures below upgrade the whole stack in one step, so plan the upgrade as a change window that covers both services.
- Back up the database, and confirm the backup completed.
- Pin the new image tags in
docker-compose.yaml. -
Pull the images before you stop anything, so the pull isn't part of your downtime:
-
Recreate the services:
-
Confirm both services are healthy and check the logs for startup errors.
Across several hosts, upgrade one at a time. Remove a host from the load balancer, upgrade it, wait for its health check, then return it to rotation before moving to the next.
- Back up the database, and confirm the backup completed.
-
Resolve the component chart versions:
-
Preview what changes before you apply it:
umask 077 helm template <release-name> ./ai-workspace -n <namespace> \ -f values-secrets.yaml -f my_values.yaml > ./next.yamlThe rendered output contains Secret manifests in the clear. Review it, then remove it:
-
Upgrade:
-
Confirm the rollout finished and the pods are ready:
With two or more replicas and a PodDisruptionBudget in place, the rollout replaces pods gradually and keeps serving. Workspace sessions held by a replaced pod end, so users on that pod sign in again.
Step 5: Roll back¶
Roll back the application, then decide separately about the database. The previous version can't always read a schema the upgrade changed, which is why the backup comes first.
Set the image tags back to the previous version and recreate:
Restore the database only if the upgrade changed the schema in a way the previous version can't read. Restoring loses everything written since the backup.
Step 6: Uninstall¶
Volumes survive. docker compose down -v deletes them, including the database when you still run SQLite. Check what a volume holds before you remove it.
The Platform API's persistent volume claim survives, because of its helm.sh/resource-policy: keep annotation. Secrets you created outside the chart, including the ones generate-secrets.sh made, also survive. Remove them deliberately once you're certain no data depends on them.
Housekeeping settings¶
Two background tasks have configurable bounds. The defaults suit most deployments; review them if your database grows unexpectedly or deployments appear stuck.
| Setting | Default | What it controls |
|---|---|---|
event_hub.poll_interval |
3s |
How often each replica checks for events from the others |
event_hub.cleanup_interval |
10m |
How often the Platform API purges delivered events |
event_hub.retention_period |
1h |
How long the Platform API keeps delivered events |
deployments.timeout_duration |
60 seconds |
How long a stuck deployment waits before the Platform API marks it failed |
Related¶
- Run in high availability: the replica setup that makes a rolling upgrade possible
- Provision secrets and keys: what to back up alongside the database
- Connect a database to the Platform API: the database this page backs up