Skip to main content
CID222 Docs

Guided runbooks

The full catalogue of symptom-driven runbooks, grouped by lifecycle phase and executable against the appliance's own diagnostics.

  • Version: 0.4
  • Role: admin_user
  • Type: troubleshooting

Each runbook starts from a symptom you can recognise, checks the appliance's own state, and names both the cause and the fix. Work one from the top: the checks are ordered so the cheapest and most likely cause is eliminated first.

Tip

Checks that read the diagnostics snapshot run themselves — in the appliance's own copy of this documentation you get a live verdict beside each one, with no model configured and no internet connection. Checks marked manual are the ones no tool can see; you answer those.

If a finding in the diagnostic snapshot sent you here, it named a runbook id: use the heading that carries it. If you are starting from a symptom in the install, first-boot, licensing, TLS or update phases, the install and activation hub is a shorter route.

Install

Provisioning stops at "6/8 container images"

The machine has no working path to the registry. Read the indented line above the failure: it separates a blocked port from a broken resolver, stale credentials and an inspected TLS connection.

Runbook install-provisioning-stops-at-images is not in this bundle. Run npm run docs:sync to pull the current corpus.

This network has no DHCP

The unattended boot entry expects an address to be offered. A segment that offers none needs the boot entry that asks for one, or a static address set from the appliance console.

Runbookinstall

This network has no DHCP, and the appliance comes up with no address

You might see: no IPv4 network (no default route) · the VM summary shows only an fe80:: address · the appliance has no IP address after the install · this network has no DHCP

Before you start

Checks

  1. 1

    Check which boot menu entry this appliance was installed from

    expected · The install was started, or will be started, from the "set the network by hand — no DHCP" entry.

  2. 2

    Check whether a static address was set from the appliance console instead

    expected · An IPv4 address, gateway and name servers are set, and a default route exists.

  3. 3

    Verify the name servers you configured actually answer

    expected · Both queries answer. A configured resolver that times out is worse than none, because the failure looks like a network outage.

The resource check fails before anything downloads

Memory and free disk are checked before the first pull. The installer's own thresholds predate the current supported minimum, so a run that only warned can still be undersized.

Runbookinstall

The installer stops on memory or free disk before anything downloads

You might see: this machine has 4096 MB of RAM — the appliance cannot run below 6 GB · only 12 GB free on / — the release images need about 40 GB · memory is 8192 MB; 16 GB is the designed size · the installer stopped at 0/8 this machine

Before you start

Checks

  1. 1

    Read the memory the installer measured on this machine

    expected · The reported memory is at least 32 GB, the supported minimum for the full stack.

  2. 2

    Read the free disk the installer measured on the root filesystem

    expected · At least 100 GB free, the supported minimum. The installer refuses below 20 GB and warns below 45 GB.

  3. 3

    Check whether the kernel stopped the previous attempt

    expected · No line reports "KILLED BY THE KERNEL".

The appliance cannot get out

Name resolution, a TCP connection on 443 and an authenticated request through the site proxy fail independently. Each passes routinely while the others do not.

Runbookinstall

The appliance cannot get out — one or more egress tests failed

You might see: egress tests failed · the appliance cannot reach the registry · ETIMEDOUT after 5000ms · proxy returned 407 Proxy Authentication Required

Before you start

Checks

  1. 1

    Verify the appliance actually ran its outbound probes

    diagnostics · appliance.egress

  2. 2

    Read the name-resolution results for the hosts this appliance must reach

    diagnostics · appliance.dns

  3. 3

    Read whether an outbound proxy is configured on the appliance

    diagnostics · appliance.proxy.configured = true

  4. 4

    Read whether the release channel answered on the last attempt

    diagnostics · appliance.update.channelReachable = true

  5. 5

    Check whether a TLS-inspection device is presenting its own certificate to the appliance

    expected · The issuer is a public certificate authority, not your organisation's inspection CA.

The machine has less RAM than the package needs

The full stack was measured at about 30 GiB in use. Below the package minimum the kernel evicts whichever detection service is largest when pressure peaks.

Runbookinstall

The machine has less RAM than the licensed package needs

You might see: This box has 16.0 GB of RAM; the package needs 32 GB · ML services will be OOM-killed under load · memory is 8192 MB; 16 GB is the designed size · the appliance runs but services keep being evicted

Before you start

Checks

  1. 1

    Compare total memory against the 32 GB minimum for the full stack

    diagnostics · host.memTotalMb = 32768

  2. 2

    For a deployment that processes documents and images, compare against the 64 GB recommendation

    diagnostics · host.memTotalMb = 65536

  3. 3

    Check whether the undersizing has already cost a service

    expected · No container has been killed for memory.

A filesystem is nearly full

An update needs room for a second copy of every image, and the state directory fails silently when the root filesystem fills.

Runbookinstall

A filesystem is nearly full

You might see: is 94% full (6 GB free of 100 GB) · an update needs room for a second copy of every image · no space left on device · the update fails part way through

Before you start

Checks

  1. 1

    Read the free space on the root filesystem

    diagnostics · host.disks[/].freeGb = 20

  2. 2

    Read the free space where Docker keeps its images and volumes

    diagnostics · host.disks[/var/lib/docker].freeGb = 40

  3. 3

    Read how large the CID database has grown

    diagnostics · datastores.postgres.dbSizeMb = 50000

The host clock is not synchronised

Licence expiry, certificate validity, token lifetimes and detection timestamps all come from this clock. A drifting one produces three failures that each look like something else.

Runbookinstall

The host clock is not synchronised

You might see: The host clock is not synchronised · clock is not synchronised — if the registry login fails on a certificate error, this is why · a valid licence is reported as expired · tokens are rejected as expired immediately after sign-in

Before you start

Checks

  1. 1

    Read whether the host clock is disciplined by a time service

    diagnostics · host.timeSynced = true

  2. 2

    Verify the appliance can reach a time source at all

    expected · The status reports the clock as synchronised, or names a server the network permits.

  3. 3

    Check whether the skew has already aged the licence out

    diagnostics · appliance.license.state = "active|trial"

Sample accounts with published passwords still exist

Appliance images ship with sample seeding off. Where it was on, the demo tenants are working sign-ins for anyone who has read the documentation.

Runbookinstall

Sample accounts with published passwords exist on this appliance

You might see: Sample accounts are seeded · sarah_smith exists on this box · demo users are present on a customer appliance · the documented demo password works

Before you start

Checks

  1. 1

    Read whether sample-account seeding is switched off for this deployment

    diagnostics · posture.seedSampleAccounts = false

  2. 2

    Check whether the sample tenants still exist, since switching seeding off does not delete what was already created

    expected · No sample tenant is present.

The OCR service has no models on an air-gapped appliance

An image built before the model check shipped can carry the engine and not its models: healthy on a networked machine, which fetches them on demand, and fatal air-gapped.

Runbookinstall

The OCR service has no models on an air-gapped appliance

You might see: HTTP 503: det/rec models not found · ocr-service answered but is not healthy · image analysis returns no text on the appliance and works on the dev box · OCR works when the machine has internet and fails when it does not

Before you start

Checks

  1. 1

    Read whether the OCR service reports itself healthy

    diagnostics · services[ocr-service].healthy = true

  2. 2

    Rule out an out-of-memory kill, which produces an unhealthy OCR service for a different reason

    expected · oomKilled is false.

A first-boot provisioning step fails on the console

Provisioning aborts on the steps the appliance cannot work without and warns on the ones it can, so a warned step leaves a usable appliance with something missing and nothing further happening on its own.

Runbookinstall

A first-boot provisioning step fails on the console

You might see: docker install failed · docker daemon will not start · secret generation failed · console UI install FAILED

Before you start

Checks

  1. 1

    Read the first-boot log rather than the console

    expected · The log names the step and carries the failing command's output.

  2. 2

    Separate a step that aborted provisioning from one that warned and continued

    expected · You can say whether provisioning aborted.

  3. 3

    Check whether Docker installed and is running

    expected · The daemon answers.

  4. 4

    Check that this box generated its own per-unit secrets

    diagnostics · posture.jwtSecretPlaceholder = false

  5. 5

    For the HTTPS step, check whether the front-end container exists yet

    diagnostics · appliance.tls.servedBy

First boot and the setup wizard

The install finished but the dashboard does not load

An ISO-installed appliance is HTTPS-only, and the front end refuses to start with a certificate it cannot use. A blank page, a refused connection and a brief upstream error have three causes.

Runbookfirst-boot

The install finished but the dashboard does not load

You might see: the browser shows a blank white page · https://<ip>/ refuses the connection · HTTPS sidecar FAILED · 3 of 16 containers are running; these are not:

Before you start

Checks

  1. 1

    Read the console's own count of running containers

    expected · Every container is running, and no line reports "N of M containers are running".

  2. 2

    Verify something holds port 443 on the appliance host

    diagnostics · appliance.tls.servedBy

  3. 3

    Verify the gateway itself is reachable

    diagnostics · services[nestjs-core].reachable = true

  4. 4

    Verify the gateway container is not restarting repeatedly

    diagnostics · containers[nestjs-core].restarts = 3

The connectivity step returns 500

Almost always a permissions problem on the state directory, not a network problem — the connectivity probe itself may well have succeeded before the write failed.

Runbookfirst-boot

The setup wizard's connectivity step returns 500

You might see: POST /setup/connectivity → 500 · setup wizard connectivity test fails with a server error · Internal server error on the network step of the wizard · wizard will not advance past connectivity

Before you start

Checks

  1. 1

    Verify the gateway container is running and is not restarting repeatedly

    diagnostics · containers[nestjs-core].restarts = 3

  2. 2

    Verify the gateway can write its state directory, where the wizard persists each step

    expected · The command exits 0 and prints nothing. A permission-denied error confirms this cause.

  3. 3

    Verify the hostnames the connectivity step probes actually resolve

    diagnostics · appliance.dns

  4. 4

    Read the outbound connectivity probe results

    diagnostics · appliance.egress

The wizard finishes and the appliance asks for setup again

The completion write is wrapped in a handler that logs and continues, so an unwritable directory produces a wizard that completes and a product that locks itself again at the next restart.

Runbookfirst-boot

The setup wizard finishes and the appliance asks for setup again

You might see: 423 SETUP_REQUIRED · Appliance setup is not complete. Finish first-boot setup to use the product. · the wizard completes and the dashboard sends me back to it · every page returns 423 after setup

Before you start

Checks

  1. 1

    Read whether the gateway considers setup complete

    diagnostics · appliance.setupComplete = true

  2. 2

    Verify the gateway can write the state directory the wizard persists into

    expected · The command exits 0 and prints nothing.

  3. 3

    Verify the gateway is not restarting between wizard steps

    diagnostics · containers[nestjs-core].restarts = 3

Settings the appliance saves do not survive a restart

The same root cause seen from the other side: setup state, the licence, update intents and the diagnostics hand-off share one directory and fail silently together.

Runbookfirst-boot

Settings the appliance saves do not survive a restart

You might see: the setup state directory is not writable · setup-state write failed · the licence I uploaded is gone after a restart · the wizard completes and comes back

Before you start

Checks

  1. 1

    Write and remove a probe file in the shared state directory as the gateway's own user

    expected · The command exits 0 and prints nothing. Permission denied or a read-only filesystem confirms this runbook.

  2. 2

    Repeat the probe in each subdirectory the product writes

    expected · All four exit 0.

  3. 3

    Check that the filesystem holding the state directory has space

    diagnostics · host.disks[/].freeGb = 2

The providers step shows an empty list

Provider and model rows come from the database seeder, not from code. Nothing is broken; the data was never inserted.

Runbookfirst-boot

The setup wizard's providers step shows an empty list

You might see: GET /setup/providers returns an empty list · no providers to choose from in the wizard · provider dropdown is empty during setup · cannot select OpenAI or Anthropic in the wizard

Before you start

Checks

  1. 1

    Verify the gateway can reach PostgreSQL — the provider list is read from it, not hard-coded

    diagnostics · datastores.postgres.reachable = true

  2. 2

    Verify no migrations are pending, so the provider and model tables exist in their current shape

    diagnostics · datastores.postgres.migrationsPending

  3. 3

    Ask the setup API for the provider list

    GET /setup/providers

The token signing secret is a placeholder

A published placeholder lets anyone with the source mint an administrator token. An unset secret is different and not safe either: every restart invalidates every session.

Runbookfirst-boot

The JWT signing secret is a placeholder or unset

You might see: JWT_SECRET is a known placeholder · the gateway refuses to boot with a message about the signing secret · everyone is signed out after every restart · tokens stop working when the container restarts

Before you start

Checks

  1. 1

    Read whether the effective signing secret is a placeholder or unset

    diagnostics · posture.jwtSecretPlaceholder = false

  2. 2

    Verify the gateway actually started, since a known placeholder refuses boot by design

    diagnostics · services[nestjs-core].reachable = true

A wizard step is refused

Some refusals are the wizard protecting a running appliance — it will not re-provision one whose setup is already complete — and some are a step that cannot be completed as configured.

Runbookfirst-boot

A first-boot wizard step is refused

You might see: setup already complete · an account already exists · cannot complete: required steps are missing · default tenant-group 'All Users' not found

Before you start

Checks

  1. 1

    Check whether setup has in fact already been completed

    diagnostics · appliance.setupComplete = true

  2. 2

    For a refusal to finish, read which steps are still outstanding

    expected · Every critical step is done.

  3. 3

    For the provider step, check that the default tenant group exists

    expected · The group exists under **Tenant Groups**.

  4. 4

    For the connectivity step, check the proxy fields together

    expected · Either the toggle is off, or a URL is present.

  5. 5

    Check whether a previous host-side change is still being applied

    expected · No reconcile is pending.

Sign-in and access

A user cannot sign in

Wrong credentials, an inactive account, an account the directory owns and a directory that could not be reached are four deliberately distinguishable refusals, because each needs a different person to act.

Runbookaccess

A user cannot sign in to the dashboard

You might see: Invalid credentials · Account is inactive · Account not found · User not found

Before you start

Checks

  1. 1

    Establish whether this account is local or comes from the directory

    expected · You can say which of the two it is.

  2. 2

    For a directory account, check that the directory answered

    expected · The directory test succeeds.

  3. 3

    Read whether the account is active

    expected · The account is active.

  4. 4

    For a user who signs in and is immediately signed out again, check the token signing secret

    diagnostics · posture.jwtSecretPlaceholder = false

The product refuses an action with 402, 403 or 423

Licence state, role, read-only demo accounts and incomplete setup all refuse here, and each carries a machine-readable code. The error codes page reads them one by one.

Runbookaccess

The product refuses an action with 402, 403 or 423

You might see: READ_ONLY_ROLE · This is a read-only demo account (viewer role) — actions and changes are disabled. · ROLE_NOT_FOR_CHAT · FEATURE_NOT_LICENSED

Checks

  1. 1

    Read the machine-readable code in the error body, not the sentence

    expected · You can name the code.

  2. 2

    For a 423, check whether first-boot setup ever completed

    diagnostics · appliance.setupComplete = true

  3. 3

    For a 402, read the licence state

    diagnostics · appliance.license.state = "active|trial"

  4. 4

    For FEATURE_NOT_LICENSED, read which tier is installed and whether the feature is in it

    diagnostics · appliance.license.tier

  5. 5

    For a 403 naming a role, read what the account's role may do

    expected · The role holds the page or capability the action needs.

Licensing

Two failures dominate activation, and they share one root cause: an appliance that demands a licence it is structurally unable to verify.

License upload fails with 400 "rejected: signature"

The licence file is usually fine. The appliance has no signing public key to check it against, so every licence looks forged.

Runbooklicensing

License upload fails with 400 "rejected: signature"

You might see: 400 rejected: signature · 400 rejected: ENOENT · License upload → 400 · licence file will not upload

Before you start

Checks

  1. 1

    Verify the licence trust anchor exists on the host

    diagnostics · appliance.license.trustAnchorPresent = true

  2. 2

    Verify the appliance reports an installation id

    diagnostics · appliance.license.installationId

  3. 3

    Verify the host clock is disciplined by NTP

    diagnostics · host.timeSynced = true

  4. 4

    Read the licence state the gateway reports

    GET /admin/license/status → {"state":"active|trial"}

The appliance has no licence trust anchor

The same missing key before anyone uploads anything. With enforcement off the appliance quietly runs its trial, so the gap surfaces only when the trial ends.

Runbooklicensing

This appliance has no licence trust anchor and can only run the trial

You might see: No license trust anchor (/etc/cid/license-pubkey.pem) · no licence can be verified on this box · the appliance only ever runs the built-in trial · every licence file is rejected, whichever one we upload

Before you start

Checks

  1. 1

    Read whether the licence-signing public key exists on the host

    diagnostics · appliance.license.trustAnchorPresent = true

  2. 2

    Read whether licence enforcement is switched on, which decides how bad the missing anchor is

    diagnostics · appliance.license.requireLicense = false

Every request returns 402 after the wizard completes

Licence enforcement is on and no licence resolves as valid. Administration stays reachable so you can fix it without a rescue procedure.

Runbooklicensing

Every request returns 402 LICENSE_EXPIRED after the wizard completes

You might see: 402 LICENSE_EXPIRED · chat returns 402 after finishing setup · the dashboard loads but every action fails with payment required · product blocked immediately after setup

Before you start

Checks

  1. 1

    Read whether licence enforcement is switched on for this deployment

    diagnostics · appliance.license.requireLicense = true

  2. 2

    Read the licence state the gateway resolved at boot

    diagnostics · appliance.license.state = "active|trial"

  3. 3

    Verify the licence trust anchor exists, since without it no licence can ever resolve as active

    diagnostics · appliance.license.trustAnchorPresent = true

  4. 4

    Check how long the installed licence has left

    diagnostics · appliance.license.expiresAt

The licence has expired and the product is locked

An appliance that worked and now blocks its product routes. Check the clock before requesting a renewal.

Runbooklicensing

The licence has expired and licensed endpoints answer 402

You might see: License expired / product locked · 402 LICENSE_EXPIRED · License state is "expired" · the trial ran out

Before you start

Checks

  1. 1

    Read the licence state the gateway resolved

    diagnostics · appliance.license.state = "active|trial"

  2. 2

    Verify the host clock is disciplined, because a skewed clock ages a valid licence out early

    diagnostics · host.timeSynced = true

  3. 3

    Verify the trust anchor exists, since without it no renewal can ever verify either

    diagnostics · appliance.license.trustAnchorPresent = true

  4. 4

    Read the licence status the API reports

    GET /admin/license/status → {"state":"active|trial"}

TLS and certificates

The certificate has expired, or nothing is serving 443

The umbrella runbook: a refused connection and an expired certificate look different to the user and share a cause more often than you would expect.

Runbooktls

The certificate has expired, or nothing is serving 443

You might see: ERR_CERT_DATE_INVALID · NET::ERR_CERT_AUTHORITY_INVALID · connection refused on 443 · the dashboard cannot be reached over HTTPS

Before you start

Checks

  1. 1

    Verify something is actually holding port 443 on the appliance host

    diagnostics · appliance.tls.servedBy

  2. 2

    Verify the appliance has a certificate to serve

    diagnostics · appliance.tls.mode = "none"

  3. 3

    Read how many days the served certificate has left

    diagnostics · appliance.tls.daysLeft = 0

  4. 4

    Check whether the certificate is inside the renewal window — 30 days, the `CID_TLS_RENEW_DAYS` default `appliance/tls/ensure-tls.sh` renews an appliance-issued leaf at and warns about a customer certificate at

    diagnostics · appliance.tls.daysLeft = 30

  5. 5

    Verify the appliance root CA is present, so clients can be made to trust an appliance-issued certificate

    diagnostics · appliance.tls.caPresent = true

Nothing is listening on 443

Connections are refused before any handshake, so the browser reports a network error rather than a certificate error.

Runbooktls

Nothing is listening on 443

You might see: Nothing is listening on 443 · connection refused on 443 · https://<ip>/ refuses the connection · the dashboard cannot be reached over HTTPS

Before you start

Checks

  1. 1

    Read what holds port 443 on the appliance host

    diagnostics · appliance.tls.servedBy

  2. 2

    Verify the appliance has a certificate to serve at all

    diagnostics · appliance.tls.mode = "none"

  3. 3

    Verify the dashboard container the listener proxies to is running

    diagnostics · containers[frontend].state = "running"

The certificate is expiring or has expired

Clients configured to trust this appliance fail closed rather than degrading, so enforcement stops at the moment the dashboard becomes unreachable.

Runbooktls

The appliance HTTPS certificate is expiring or has expired

You might see: The appliance HTTPS certificate expires in 6 day(s) · ERR_CERT_DATE_INVALID · certificate expired · your connection is not private

Before you start

Checks

  1. 1

    Read how many days the served certificate has left

    diagnostics · appliance.tls.daysLeft = 0

  2. 2

    Check whether the certificate has entered the renewal window — 30 days, the `CID_TLS_RENEW_DAYS` default in `appliance/tls/ensure-tls.sh`

    diagnostics · appliance.tls.daysLeft = 30

  3. 3

    Read which kind of certificate the appliance serves, because the renewal path differs

    diagnostics · appliance.tls.mode = "unknown"

The appliance refuses the certificate you uploaded

Certificate, key and chain are validated in sequence and each has its own vocabulary. Fixing the wrong file leaves the same message in place.

Runbooktls

The appliance refuses the certificate you uploaded

You might see: cert is not a PEM X.509 certificate · cert is a CA certificate; upload the server (leaf) certificate, with the CA in chain · the private key does not match the certificate · key is not a PEM private key (encrypted keys are not accepted — decrypt it first)

Before you start

Checks

  1. 1

    Read which of the three files the refusal is about — the certificate, the key or the chain

    expected · You can say which file is being refused.

  2. 2

    Confirm you uploaded the server certificate, not the CA

    expected · CA:FALSE on the certificate field; the CA belongs in the chain field.

  3. 3

    Confirm the key matches the certificate

    expected · The moduli match and the key is not passphrase-protected.

  4. 4

    Confirm the certificate carries a subjectAltName

    expected · The hostname the appliance is reached by is listed as a SAN.

  5. 5

    Confirm the chain is ordered leaf-side first

    expected · Each link issues the one before it.

  6. 6

    Check whether a previous change is still being applied on the host

    expected · No reconcile is pending.

Directory (AD / LDAP)

The directory cannot be reached or bound

Local accounts keep working throughout, and content policy never consults the directory — what stops is directory sign-in and membership updates.

Runbookldap

The directory cannot be reached or bound

You might see: LDAP url, bind DN and base DN must be configured · LDAP bind password file unreadable · Directory authentication is unavailable · The directory server could not be reached, so directory sign-in is temporarily unavailable.

Before you start

Checks

  1. 1

    Read the three fields the connection cannot be attempted without

    expected · All three are non-empty.

  2. 2

    Check where the bind password is coming from

    expected · One of the three resolves to a password the directory accepts.

  3. 3

    If you are testing a URL you just typed, check whether you also re-entered the password

    expected · Either the URL is unchanged, or you typed the bind password into the test form.

  4. 4

    Check that the directory host answers from the appliance

    expected · The port answers and, for LDAPS, the certificate chain validates.

A sync scope will not run, or imports the wrong people

On a partial directory answer the sync imports and updates what it saw and removes nobody, because removing on incomplete data deprovisions people whose entries were simply not returned.

Runbookldap

A directory sync scope will not run, or imports the wrong people

You might see: Sync scope not found · This sync scope is disabled. Re-enable it before syncing. · Sync scope is mapped to a CID tenant group that no longer exists · Sync scope would assign a role which cannot be granted from a directory scope

Before you start

Checks

  1. 1

    Read the scope's state before anything else

    expected · The scope is enabled.

  2. 2

    Check the CID tenant group the scope maps into

    expected · The target group exists.

  3. 3

    Read the role the scope assigns

    expected · The role is one of those the message lists as allowed.

  4. 4

    Read the result counts of the last run

    expected · The run completed rather than reporting a partial answer.

Providers

No providers are in the catalogue

Chat cannot resolve a model, so every completion fails. The catalogue is data, not code.

Runbookproviders

No LLM providers are in the catalogue, so every completion fails

You might see: No LLM providers are in the catalogue · chat cannot resolve a model · the model dropdown is empty · every completion fails with a model error

Before you start

Checks

  1. 1

    Verify the gateway can reach the database, since the catalogue is read from it rather than hard-coded

    diagnostics · datastores.postgres.reachable = true

  2. 2

    Verify no migrations are pending, so the provider and model tables have their current shape

    diagnostics · datastores.postgres.migrationsPending

  3. 3

    Ask the gateway for the model catalogue

    GET /models

A provider credential will not save, or its test fails

Three unrelated refusals share this screen: the endpoint validator, the assignment rule, and the provider refusing the key itself. Only the third is about the key.

Runbookproviders

A provider credential will not save, or its test fails

You might see: Invalid Anthropic API key · Test failed to run · Failed to create credential · Failed to update credential

Before you start

Checks

  1. 1

    For a credential that carries an endpoint (Azure OpenAI, or any self-hosted endpoint), read the exact rejection text

    expected · The endpoint is an absolute http(s) URL that does not resolve into link-local space.

  2. 2

    Check what the credential is assigned to

    expected · Exactly one of them is set.

  3. 3

    Check whether the appliance can reach the provider at all before blaming the key

    diagnostics · appliance.egress

  4. 4

    Read what the credential test actually reported

    expected · The test succeeds.

A chat request fails before the model is reached

Model row, then credential, then provider call — a failure at each step reads the same to the user and needs a different fix.

Runbookchat

A chat request fails before the model is reached

You might see: Provider 'openai' not found · Model 'gpt-4o' not found · Model 'gpt-4o' is not active · Model not found for provider

Before you start

Checks

  1. 1

    Check that the provider and model catalogue is populated at all

    expected · Providers and models are listed.

  2. 2

    Read whether the model the caller named exists and is active

    expected · The model is listed and active.

  3. 3

    Check that a credential resolves for this caller and provider

    expected · One active credential resolves.

  4. 4

    Test the credential that resolves

    expected · The test succeeds.

A request is refused for being too frequent, or a quota is exhausted

A gateway throttle, an API-key quota and the provider's own refusal look alike and have completely different remedies.

Runbookproviders

A request is refused for being too frequent, or a quota is exhausted

You might see: Too many requests, please try again later. · Too many help requests, please try again in a few minutes. · Too many unlock requests, please try again later. · Too many attestation requests

Checks

  1. 1

    Establish which limit refused the call — the gateway's own throttle, an API-key quota, or the provider's

    expected · You can say which of the three it is.

  2. 2

    For a quota message, read the key's configured request and token quotas

    expected · The key has headroom left in the current window.

  3. 3

    Check that Redis is reachable, because the throttles keep their counters there

    diagnostics · datastores.redis.reachable = true

A gateway API key is refused

Keys are hashed and cannot be recovered, only replaced. A group key borrows the identity of the group's earliest-added member, so emptying the group stops it resolving.

Runbookaccess

A gateway API key is refused

You might see: Invalid API key · Invalid or expired API key · API key is not active · API key has expired

Before you start

Checks

  1. 1

    Find the key in the estate

    expected · The key exists.

  2. 2

    Read the key's status and expiry

    expected · The key is active and either has no expiry or expires in the future.

  3. 3

    Check what identity the key resolves to

    expected · The key is assigned to a tenant, or to a group that has at least one member.

Chat and detections

A message is rejected, masked or flagged by the content policy

A decision is not an error. Every decision writes a detection naming the engine, the label and the action — start from that record rather than the user's paraphrase.

Runbookchat

A message is rejected, masked or flagged by the content policy

You might see: Document rejected due to policy violation · my prompt came back with names replaced by placeholders · the assistant refused an ordinary business question · a customer record was masked and the model could not answer

Before you start

Checks

  1. 1

    Find the request in the detection record

    expected · One detection explains the outcome.

  2. 2

    Read which engine produced the match

    expected · You can name the engine.

  3. 3

    Read the filter that fired and the action it carries

    expected · The action matches what the user experienced.

  4. 4

    Decide whether the match was correct

    expected · The matched span really is what the rule is for.

A user's AI access is locked

Repeated detections fill a sliding window that opens one per-user review. A lock the analyst could not tie back to detections is downgraded to a warning, so a lock that survived has evidence behind it.

Runbookaccess

A user's AI access is locked and they cannot chat

You might see: user_locked · the chat stream ends immediately with a lock notice · Your AI access is not locked · An unlock request is already pending for this lock

Before you start

Checks

  1. 1

    Read whether the user actually holds an active lock

    expected · One active lock exists for that user.

  2. 2

    Read the review that produced the lock

    expected · The review explains the lock in terms of specific detections.

  3. 3

    For a user who cannot file an unlock request, read the state of the existing one

    expected · Either no request is pending, or the pending one is waiting for a reviewer.

  4. 4

    After an unlock, confirm the lock is gone rather than only appearing gone

    expected · The message goes through.

A file is refused, or its analysis fails

Format, size and a document parser behind its circuit breaker all refuse in the same place. A chunked upload is refused rather than passed through uninspected.

Runbookperformance

A file is refused, or its analysis fails

You might see: No file uploaded (expected multipart field "file") · Image size exceeds maximum allowed size of 10MB · Unsupported file type. Upload a PDF, DOCX or TXT. · legacy and macro-enabled spreadsheet formats cannot be safely redacted

Before you start

Checks

  1. 1

    Read whether the format is one the product accepts at all

    expected · The format is accepted on that surface.

  2. 2

    Read the size limit for that surface

    expected · The file is under the limit named in the message.

  3. 3

    Check whether the client was doing a chunked or resumable upload

    diagnostics · posture.resumableUploadPolicy = "block"

  4. 4

    Read whether document analysis is switched on for this deployment

    expected · The service is enabled.

  5. 5

    Read the document parser's own health

    diagnostics · services[document-parser].healthy = true

Updates

The System Updates page shows 0.0.0, or the buttons return 500

A missing version stamp, an unset channel, an unreachable channel and a stalled updater all surface on the same page and look alike.

Runbookupdates

The System Updates page shows 0.0.0, or the update buttons return 500

You might see: System Updates shows version 0.0.0 · current version 0.0.0 · update check returns 500 · install update → 500

Before you start

Checks

  1. 1

    Read the product version the appliance reports

    diagnostics · appliance.version = "0.0.0"

  2. 2

    Verify an update channel is configured

    diagnostics · appliance.channel

  3. 3

    Verify the release channel answered on the last attempt

    diagnostics · appliance.update.channelReachable = true

  4. 4

    Check whether an update or host-repair intent is stuck waiting for the host updater

    diagnostics · appliance.update.pendingIntent

  5. 5

    Read the host-repair status, which is what the page renders alongside the version

    GET /admin/system-update/host-repair

The appliance cannot say which release it runs

Every comparison against the channel manifest is meaningless until the version stamp is restored, and so is every support answer.

Runbookupdates

The appliance cannot say which release it is running

You might see: This box cannot say which release it is running · System Updates shows version 0.0.0 · current version 0.1.0 on a box that is not 0.1.0 · the update page offers nothing and shows no version

Before you start

Checks

  1. 1

    Read the product version the appliance reports

    diagnostics · appliance.version = "0.1.0"

  2. 2

    Read the host-repair status, which is what rewrites the version stamp

    GET /admin/system-update/host-repair

The update channel is unreachable

Expected on a deliberately air-gapped appliance, which updates from a signed offline bundle. On a connected one it is an egress problem wearing an update-shaped mask.

Runbookupdates

The update channel is not reachable from the appliance

You might see: The update channel host is not reachable from inside the gateway · Check for updates does nothing · update check returns 500 · the appliance never finds a new version

Before you start

Checks

  1. 1

    Verify an update channel is configured at all

    diagnostics · appliance.channel

  2. 2

    Read whether the channel answered on the last attempt

    diagnostics · appliance.update.channelReachable = true

  3. 3

    Read the name-resolution results for the hosts the appliance must reach

    diagnostics · appliance.dns

  4. 4

    Read whether an outbound proxy is configured, on a network that requires one

    diagnostics · appliance.proxy.configured = true

An update bundle will not upload or install

Resumable uploads fail on offsets and expiry rather than on the release, and an online install is refused outright unless the update host is pinned.

Runbookupdates

An update bundle will not upload or install from the dashboard

You might see: unknown or expired upload · offset mismatch · incomplete upload: have N of M bytes · bundle filename must end with .cidupd

Before you start

Checks

  1. 1

    Read the bundle's file name

    expected · The file ends in `.cidupd`.

  2. 2

    For a resumed upload, check that the server still holds the partial file

    expected · The status call returns an offset.

  3. 3

    For an offset mismatch, compare what the client thinks it sent with what the server holds

    expected · The client resumes from the offset the server reports.

  4. 4

    For "incomplete upload", compare the assembled size with the declared size

    expected · The two sizes agree.

  5. 5

    For an online install, read the update host allowlist

    expected · The allowlist names the host in the manifest URL, over https.

The updater refuses the release, or rolls it back

Every refusal before the smoke test happens with nothing changed. The one that surprises people is the rollback check: pruned images mean there is no way back, so the updater will not go forward.

Runbookupdates

The updater refuses the release, or stops part-way through installing it

You might see: signature missing · bundle not found · compose file missing from release · cannot jump from X to Y: this release requires at least

Before you start

Checks

  1. 1

    Read whether the failure is the release's signature or the appliance's trust anchor

    diagnostics · appliance.license.trustAnchorPresent = true

  2. 2

    Read the version this appliance runs against the release's minimum

    diagnostics · appliance.version

  3. 3

    Check that the bundle carries everything its manifest references

    expected · No file is reported missing or mismatched.

  4. 4

    Check that the images of the release currently installed are still present

    expected · Every current-release image inspects successfully.

  5. 5

    Read the database pre-flight probes

    expected · Every probe reports clean.

  6. 6

    For an update that installed and then rolled back, read the smoke result

    expected · The smoke test passed.

A host-side repair or reconfigure step fails

The failure mode that costs the most is the quiet one: the host write fails, the handler logs and continues, and the dashboard reports success.

Runbookupdates

A host-side repair or reconfigure step fails

You might see: host command failed · host write failed · cannot determine own image · a certificate was applied in the dashboard and the host never changed

Before you start

Checks

  1. 1

    Check whether the agent can identify its own image

    expected · The agent resolves its own image, or `CID_HOST_EXEC_IMAGE` pins one.

  2. 2

    For "host write failed", read which path it was writing

    expected · The path is writable and the filesystem has room.

  3. 3

    Read free disk on the host

    diagnostics · host

  4. 4

    Check the shared state directory, which several features fail through together

    expected · The directory exists and the gateway can write it.

Database and schema

Migrations are pending

The endpoints reading the new columns answer 500 while the rest of the product looks healthy.

Runbookdata

Migrations are pending, and the pages that read the new columns fail

You might see: compiled migration(s) have never been applied to this database · one page returns 500 while the rest of the product works · column does not exist · relation does not exist

Before you start

Checks

  1. 1

    Read the list of compiled migrations that are not recorded in the database

    diagnostics · datastores.postgres.migrationsPending

  2. 2

    Verify auto-synchronise is off, since it hides an unapplied chain behind a schema that looks right

    diagnostics · posture.dbSynchronize = false

  3. 3

    Verify the gateway is not restarting, which is what a failing boot migration usually produces

    diagnostics · containers[nestjs-core].restarts = 3

PostgreSQL or Redis is unreachable

A database that cannot be reached empties every list rather than erroring; a Redis that cannot be reached degrades rate limiting open and stops scheduled work.

Runbookdata

PostgreSQL or Redis is unreachable from the gateway

You might see: Postgres is unreachable · Redis is unreachable · ECONNREFUSED on the database port · every list in the dashboard is empty

Before you start

Checks

  1. 1

    Read whether the gateway can reach PostgreSQL

    diagnostics · datastores.postgres.reachable = true

  2. 2

    Read the state of the database container

    diagnostics · containers[postgres].state = "running"

  3. 3

    Read whether the gateway can reach Redis

    diagnostics · datastores.redis.reachable = true

  4. 4

    Check that the filesystem holding the database has space, since PostgreSQL stops accepting writes when it does not

    diagnostics · host.disks[/].freeGb = 2

The schema is auto-synchronised on every boot

Auto-synchronise drops every object the entities do not declare, including the indexes the migrations created. Nothing errors; the appliance just gets slower after each restart.

Runbookdata

The database schema is being auto-synchronised on every boot

You might see: DB_SYNCHRONIZE=true · indexes disappear after a restart · search is slow again after every reboot · the schema changes without a migration

Before you start

Checks

  1. 1

    Read the effective auto-synchronise setting

    diagnostics · posture.dbSynchronize = false

  2. 2

    Verify the migration chain is fully applied, since auto-sync hides an unapplied chain

    diagnostics · datastores.postgres.migrationsPending

Services and capacity

One service is unreachable or reports itself unhealthy

Reachable and unhealthy is the interesting case, and usually means a model that did not load.

Runbookperformance

One service is unreachable or reports itself unhealthy

You might see: hap-guard-v2 is unreachable: no response · ocr-service answered but is not healthy · ml-detector is unreachable · one detector is down and the others are fine

Before you start

Checks

  1. 1

    Read whether the named service answered its health endpoint at all

    expected · reachable is true. When it is false the failure is network or container level, not model level.

  2. 2

    Read the container state for the named service

    expected · state is running, restarts are low and oomKilled is false.

  3. 3

    For a service that answers and reports itself unhealthy, read the error it returns

    expected · The error names what did not initialise — a model file, a device, a dependency.

  4. 4

    For the analyst and the MCP server, check what they depend on before blaming them

    expected · The dependency is healthy, so the unhealthy verdict is about this service.

A container keeps restarting

Separate an out-of-memory kill from a configuration the service refuses, before restarting it again.

Runbookperformance

A container keeps restarting

You might see: has restarted 12 times — it is in a crash loop · container restarting · a service comes back and dies again · the dashboard works for a minute and then errors

Before you start

Checks

  1. 1

    Check whether the last termination was an out-of-memory kill rather than a crash

    expected · oomKilled is false. When it is true, this is a sizing problem and not a crash.

  2. 2

    Read how much memory the host has left

    diagnostics · host.memFreeMb = 1024

  3. 3

    Verify the schema is current, since a failing boot migration restarts the gateway forever

    diagnostics · datastores.postgres.migrationsPending

  4. 4

    Check whether the container refuses its own configuration at boot

    expected · The service reports that it started and is listening, rather than refusing a setting.

A container was killed for running out of memory

Exit code 137 is the kernel, not the service. Restarting it restores service and changes nothing.

Runbookperformance

A container was killed for running out of memory

You might see: was killed for running out of memory (exit 137) · exit code 137 · an OOM kill (exit 137) means this machine needs more RAM · the detector dies under load

Before you start

Checks

  1. 1

    Compare the host's total memory against the minimum for the licensed package

    diagnostics · host.memTotalMb = 32768

  2. 2

    Read how much memory the host has available right now

    diagnostics · host.memFreeMb = 2048

  3. 3

    Compare the killed container's declared limit with what it was using

    expected · Usage sits well below the limit. Usage at the limit means the limit is the constraint, not the host.

  4. 4

    Check whether every service running on this box is actually in use

    expected · Everything running is something this deployment needs.

A PDF or report will not render

Rendering is a separate service. With it absent every PDF export fails and everything else on the page keeps working, which is why it looks specific to one button.

Runbookperformance

A PDF or report will not render

You might see: Failed to render compliance scorecard PDF · Failed to render detections PDF · Failed to render executive summary PDF · Failed to render regulations PDF

Before you start

Checks

  1. 1

    Read whether the report renderer answered

    diagnostics · services[report-renderer].reachable = true

  2. 2

    Read whether the renderer reports itself healthy

    diagnostics · services[report-renderer].healthy = true

  3. 3

    Check whether only the large exports fail

    expected · The small export succeeds.

  4. 4

    For a compliance dossier, read whether the render was reconciled against the record

    expected · The renderer reported its rendered ids.

The LLM risk analyst is disabled, unreachable, or times out

The analyst and its MCP server are opt-in. When enabled they are the largest thing on the box and the first thing memory pressure kills.

Runbookperformance

The LLM risk analyst is disabled, unreachable, or times out

You might see: The risk analyst service is not enabled on this deployment. · The risk analyst could not be reached. · The analyst is already running an analysis. · The analysis did not start.

Before you start

Checks

  1. 1

    Read whether the gateway believes the analyst is enabled

    expected · The flag is set.

  2. 2

    Read whether the analyst service answered

    diagnostics · services[risk-analyst].reachable = true

  3. 3

    Read the MCP server, which the analyst reads the estate through

    diagnostics · services[mcp-server].reachable = true

  4. 4

    Check whether an analysis is already running

    expected · No analysis is in flight.

  5. 5

    For timeouts, compare the configured deadline against how long a reading actually takes

    expected · The deadline is longer than a typical completed review took.

A dashboard action fails with "Failed to …" and no reason

The dashboard's failure toasts are generic by design. The HTTP status separates five unrelated causes in one step.

Runbookperformance

A dashboard action fails with "Failed to …" and no reason

You might see: Failed to create · Failed to update · Failed to delete · Error loading data. Please try again.

Before you start

Checks

  1. 1

    Read the HTTP status the request actually returned

    expected · You have a status code and a response body.

  2. 2

    For 401, check whether the session simply expired

    expected · The action succeeds after signing in again.

  3. 3

    For 402, 403 or 423, read the error code

    expected · The request is not being refused by role, licence or setup state.

  4. 4

    For 400 or 409, read the message the server sent

    expected · The message does not describe something you can correct in the form.

  5. 5

    For 5xx or no response at all, check that the gateway is serving

    diagnostics · services[nestjs-core].healthy = true

Inline proxy and inspection

Every ICAP caller arrives as the same docker bridge address

A gateway running as a virtual machine beside CID222 does not reach the listener with its own address, so a source-IP allowlist there admits all callers or none.

Runbookinline-proxy

Every ICAP caller arrives as the same docker bridge address

You might see: every SWG appears as 172.19.0.1 · the ICAP source-IP allowlist does not distinguish our proxies · allowed_source_ips has no effect · the allowlist blocked every caller after we set it

Before you start

Checks

  1. 1

    Read how the ICAP listener admits callers

    diagnostics · posture.inspectionSourceIpMode = "unrestricted"

  2. 2

    Check whether callers arrive with their own addresses or with the container bridge address

    expected · The recorded source address is the gateway's own address.

  3. 3

    When an explicit allowlist is set, verify it lists every caller that must be admitted

    expected · Every caller is listed.

The inspection endpoint refuses the proxy calling it

Admission is evaluated in a fixed order — an mTLS requirement, then a bearer key, then network position — and a key that resolves to no tenant is refused outright rather than treated as no key at all.

Runbookinline-proxy

The inspection endpoint refuses the proxy or gateway calling it

You might see: Client certificate required · API key required · Invalid API key · Internal API only

Before you start

Checks

  1. 1

    Read whether this deployment mandates a client certificate

    expected · Either no DN is required, or the caller presents a certificate matching it.

  2. 2

    Read whether the caller presents a gateway API key

    expected · The caller sends `Bearer cid_key_…` and the key resolves.

  3. 3

    Check whether the call arrives through a reverse proxy

    expected · Either the caller presents a key, or it reaches the listener directly from an internal address.

  4. 4

    Before trusting any source-IP control, check what address the caller actually arrives as

    expected · Callers arrive with their own addresses.

The proxy CA or a helper download is not available

Recreating the proxy with a new volume generates a new CA. Every client that trusted the old one must be given the new one; old trust does not transfer.

Runbookinline-proxy

The inline proxy's CA or helper download is not available

You might see: CA not available. Is cid-inline-proxy running? · Unsupported format · dc-agent.ps1 not bundled in this build · the CA download button returns an error

Before you start

Checks

  1. 1

    Read the container state for the inline proxy

    expected · state is running.

  2. 2

    Check whether the proxy has generated its CA yet

    expected · The proxy has completed a first start and holds a CA.

  3. 3

    For an "Unsupported format" refusal, read which format was asked for

    expected · The format is one of the offered ones.

  4. 4

    For the domain-controller helper script, check whether this build ships it

    expected · The build carries the helper.

Endpoint agent and browser extension

An endpoint agent will not enrol, or stops receiving policy

A fleet that is not active refuses new enrolments with the same message as an unknown token, and an agent adopts a bundle only when its version is strictly higher than the one it applied.

Runbookendpoint-agent

An endpoint agent will not enrol, or stops receiving policy

You might see: Invalid enrollment token · Missing agent API key · Invalid agent API key · Agent key does not match device

Before you start

Checks

  1. 1

    Check the enrollment token against the fleet

    expected · The token belongs to a fleet whose status is active.

  2. 2

    Read the Authorization header the agent sends on policy and heartbeat calls

    expected · The agent sends a key the gateway recognises.

  3. 3

    Check that the key and the device id agree

    expected · The key resolves to the device id the agent claims.

  4. 4

    Read the fleet's policy for the two fields the bundle cannot be signed without

    expected · Both are present.

  5. 5

    For a policy change the fleet never adopted, check that the bundle version advanced

    expected · The version is higher than the one the agent reports.

The browser extension cannot attest, sign in, or fetch its policy

A device never enrolled, a device enrolled to another tenant, key material of the wrong shape and a device throttled for re-attesting in a loop are four different failures.

Runbookextension

The browser extension cannot attest, sign in, or fetch its policy

You might see: Attestation rejected · Device is not enrolled · Device is enrolled to another tenant · Device not found

Before you start

Checks

  1. 1

    Read the device's enrolment

    expected · The device is enrolled to the tenant the user belongs to.

  2. 2

    For a rejected enrolment, read what key material the extension sent

    expected · The extension sends a P-256 public JWK.

  3. 3

    Check whether the device is being refused for frequency rather than identity

    expected · The device attests at its normal interval.

  4. 4

    Read the role of the account signing in

    expected · The role is one that uses the extension.

  5. 5

    For the firewall EDL feed, read the token the firewall presents

    expected · The firewall presents a current EDL token.

Integrations and repositories

Events stop arriving at the SIEM or webhook collector

Nothing is dropped silently — the exporter advances its cursor only after a successful send — so a stalled export has a reason attached and a zero drop count is structural rather than reassuring.

Runbookdata

Events stop arriving at the SIEM or webhook collector

You might see: a backlog stopped moving · the collector kept refusing delivery · Configuration store unavailable; settings cannot be saved · ITSM handoff is unavailable in this process; cannot send a test event

Before you start

Checks

  1. 1

    Read why the exporter stopped

    expected · A reason is recorded.

  2. 2

    Send a test event to the collector

    expected · The collector accepts it.

  3. 3

    For syslog, check host, port and protocol

    expected · The collector is reachable on the configured transport.

  4. 4

    For "settings cannot be saved", check the configuration store

    diagnostics · datastores.postgres.reachable = true

A repository connector cannot read the repository

Credentials are encrypted and cannot be read back, so a connector saved with a missing field looks complete on screen and fails at first use.

Runbookdata

A repository connector cannot read the repository

You might see: generic_git connector needs base_url set to the clone URL · GitHub App connector needs both app_id and private_key_pem · connector needs an access_token · Assignment has no GitHub installation id

Before you start

Checks

  1. 1

    Read the connector's credential fields for its type

    expected · Every field the type needs is set.

  2. 2

    Read the clone URL

    expected · The URL is one of the accepted forms.

  3. 3

    Check that the credential type and the URL scheme can work together

    expected · The URL scheme matches the credential the connector holds.

  4. 4

    For GitHub, read the host URL field

    expected · Blank for github.com, or the bare GHES host.

  5. 5

    Before reading findings, check that a baseline scan completed

    expected · A completed baseline exists for the repository.

  6. 6

    Read whether the last scan ran with every analyzer available

    expected · The scan did not run degraded.

If none of these match

  • Read the diagnostic snapshot end to end. Its findings section names a runbook for each rule that fired.
  • Check the health matrix — a service that is reachable but unhealthy is the air-gapped failure mode.
  • Collect the support bundle and escalate. Include the installation id and the appliance version.

Last updated on

On this page

Install
Provisioning stops at "6/8 container images"
This network has no DHCP
The resource check fails before anything downloads
The appliance cannot get out
The machine has less RAM than the package needs
A filesystem is nearly full
The host clock is not synchronised
Sample accounts with published passwords still exist
The OCR service has no models on an air-gapped appliance
A first-boot provisioning step fails on the console
First boot and the setup wizard
The install finished but the dashboard does not load
The connectivity step returns 500
The wizard finishes and the appliance asks for setup again
Settings the appliance saves do not survive a restart
The providers step shows an empty list
The token signing secret is a placeholder
A wizard step is refused
Sign-in and access
A user cannot sign in
The product refuses an action with 402, 403 or 423
Licensing
License upload fails with 400 "rejected: signature"
The appliance has no licence trust anchor
Every request returns 402 after the wizard completes
The licence has expired and the product is locked
TLS and certificates
The certificate has expired, or nothing is serving 443
Nothing is listening on 443
The certificate is expiring or has expired
The appliance refuses the certificate you uploaded
Directory (AD / LDAP)
The directory cannot be reached or bound
A sync scope will not run, or imports the wrong people
Providers
No providers are in the catalogue
A provider credential will not save, or its test fails
A chat request fails before the model is reached
A request is refused for being too frequent, or a quota is exhausted
A gateway API key is refused
Chat and detections
A message is rejected, masked or flagged by the content policy
A user's AI access is locked
A file is refused, or its analysis fails
Updates
The System Updates page shows 0.0.0, or the buttons return 500
The appliance cannot say which release it runs
The update channel is unreachable
An update bundle will not upload or install
The updater refuses the release, or rolls it back
A host-side repair or reconfigure step fails
Database and schema
Migrations are pending
PostgreSQL or Redis is unreachable
The schema is auto-synchronised on every boot
Services and capacity
One service is unreachable or reports itself unhealthy
A container keeps restarting
A container was killed for running out of memory
A PDF or report will not render
The LLM risk analyst is disabled, unreachable, or times out
A dashboard action fails with "Failed to …" and no reason
Inline proxy and inspection
Every ICAP caller arrives as the same docker bridge address
The inspection endpoint refuses the proxy calling it
The proxy CA or a helper download is not available
Endpoint agent and browser extension
An endpoint agent will not enrol, or stops receiving policy
The browser extension cannot attest, sign in, or fetch its policy
Integrations and repositories
Events stop arriving at the SIEM or webhook collector
A repository connector cannot read the repository
If none of these match
Download PDF