Skip to content

Troubleshooting

Find your symptom in the table, jump to the fix. Every runbook page follows the same shape: Symptom → Confirm → Fix → Prevent.

If you see…Go to
The Python Inference card shows CPU (fallback) instead of CUDARunning on CPU when GPU expected
Log line starting TensorRT EP not usedExpected on the TX2, no action — see Running on CPU when GPU expected
The stack won’t start because Docker doesn’t know the nvidia runtime, or the inference log shows no CUDA-capable device is detectedGPU not used — CUDA errors
With STRICT_EP=1, the inference container exits with a STRICT_EP: error and keeps restartingGPU not used — CUDA errors
Enabling a flow in Task Manager bounces back to off with “Could not connect to the datasource” (adapter_connect_failed) or a shape mismatchDatasource down or faulted
A streaming OPC-UA / MQTT / CSV source stops; a fault indicator or red banner appearsDatasource down or faulted
A flow is enabled but no predictions appearDatasource down or faulted
MQTT input faults after a topic goes quiet, or the log shows tag stale / ChannelReadExceptionMQTT input stops or faults
All predictions read NaN, or a topic publishes a NaN/Infinity sentinelMQTT input stops or faults
MQTT sink logs transport: not connected, or the log shows SessionTakenOverMQTT input stops or faults
Tag mapping rejected with tag_outside_subscription_*MQTT input stops or faults
Settings → About highlights the frontend / backend version in amberFrontend / backend version mismatch
Version reads 0.0.0-dev… or ends in -dirtyFrontend / backend version mismatch
Disk filling up; database file far larger than the data it holds; deleting rows doesn’t shrink itDisk fills / database grows unbounded
A container shows unhealthy in docker ps but the app serves fineContainer marked unhealthy but service works