Troubleshooting
Find your symptom in the table, jump to the fix. Every runbook page follows the same shape: Symptom → Confirm → Fix → Prevent.
Find your symptom
Section titled “Find your symptom”| If you see… | Go to |
|---|---|
| The Python Inference card shows CPU (fallback) instead of CUDA | Running on CPU when GPU expected |
Log line starting TensorRT EP not used | Expected on the TX2, no action — see Running on CPU when GPU expected |
The stack won’t start because Docker doesn’t know the nvidia runtime, or the inference log shows no CUDA-capable device is detected | GPU not used — CUDA errors |
With STRICT_EP=1, the inference container exits with a STRICT_EP: error and keeps restarting | GPU not used — CUDA errors |
Enabling a flow in Task Manager bounces back to off with “Could not connect to the datasource” (adapter_connect_failed) or a shape mismatch | Datasource down or faulted |
| A streaming OPC-UA / MQTT / CSV source stops; a fault indicator or red banner appears | Datasource down or faulted |
| A flow is enabled but no predictions appear | Datasource down or faulted |
MQTT input faults after a topic goes quiet, or the log shows tag stale / ChannelReadException | MQTT input stops or faults |
All predictions read NaN, or a topic publishes a NaN/Infinity sentinel | MQTT input stops or faults |
MQTT sink logs transport: not connected, or the log shows SessionTakenOver | MQTT input stops or faults |
Tag mapping rejected with tag_outside_subscription_* | MQTT input stops or faults |
| Settings → About highlights the frontend / backend version in amber | Frontend / backend version mismatch |
Version reads 0.0.0-dev… or ends in -dirty | Frontend / backend version mismatch |
| Disk filling up; database file far larger than the data it holds; deleting rows doesn’t shrink it | Disk fills / database grows unbounded |
A container shows unhealthy in docker ps but the app serves fine | Container marked unhealthy but service works |