Running on CPU when GPU expected
Symptom
Section titled “Symptom”Inference is slower than expected, and the Python Inference card on the dashboard shows CPU (fallback) instead of CUDA.
On the TX2 the execution provider chain is CUDA → CPU. The default execution mode, auto, tries CUDA first. If CUDA cannot start, the service keeps running on the CPU. That safety net also hides the problem: a box that fell back still predicts, just slower.
A plain CPU badge, without “(fallback)”, means the box was set to CPU on purpose (EXECUTION_MODE=cpu or FORCE_CPU=1). That is not a fallback.
Confirm
Section titled “Confirm”The badge is the ground truth. It shows the provider that the loaded model actually runs on, as the inference service reports it.
The inference container’s console shows errors only. Warnings go to a log file inside the container. Search it for the fallback warning:
docker exec aiboard-inference-real grep -E "resolved to CPU only|STRICT_EP" /data/logs/inference.logEach line starts with a timestamp. Look at the lines from the latest start.
resolved to CPU only→ ONNX Runtime in this container offers no CUDA provider at all.- No match, but the badge still shows CPU (fallback) → CUDA was offered but could not start on the GPU.
To see which execution mode the container started with:
docker logs aiboard-inference-real 2>&1 | grep "Starting REAL inference"A box on the default prints Starting REAL inference service (Jetson GPU, EXECUTION_MODE=auto)....
The dashboard’s Logs page also streams the inference service’s informational lines. When it has the start-up lines of a healthy box, you see CUDA EP enabled, then a Providers: line that lists CUDAExecutionProvider first.
-
Check that the inference container runs with the NVIDIA runtime:
Terminal window docker inspect -f '{{.HostConfig.Runtime}}' aiboard-inference-realIt must print
nvidia. The NVIDIA runtime is what makes the TX2’s CUDA libraries visible inside the container. The bundle’s compose file sets it. If it prints anything else, the compose file was changed. Restore it by runningsudo ./update.shfrom a release bundle, of the same version or newer. The update copies the bundle’s compose files back into/opt/aiboard/compose/. -
Check that the inference image is the Jetson GPU image:
Terminal window docker inspect -f '{{.Config.Image}}' aiboard-inference-realThe tag must end in
-jetson-gpu. Any other image is not built for the TX2’s GPU. Install thejetson-gpurelease bundle. See Offline Bundle Install. -
If both checks pass and the box still falls back, the host’s CUDA stack is the problem. Follow GPU not used — CUDA errors.
Make the fallback loud
Section titled “Make the fallback loud”On a box that must use the GPU, make the service refuse to start instead of falling back. Pin the execution mode to cuda and turn on strict enforcement.
The release compose file sets EXECUTION_MODE=auto for the inference service and does not set STRICT_EP. The .env file does not control either one. Change them in the installed compose file:
-
Open
/opt/aiboard/compose/docker-compose.release.yml. In theenvironment:list of theinferenceservice, changeEXECUTION_MODE=autotoEXECUTION_MODE=cudaand addSTRICT_EP=1:- EXECUTION_MODE=cuda- STRICT_EP=1 -
Recreate the inference container so it picks up the change:
Terminal window sudo docker compose --env-file /opt/aiboard/.env \-f /opt/aiboard/compose/docker-compose.release.yml up -d inference
With this setting, a session that does not start on CUDA stops the service with a STRICT_EP: error, and the container restarts until you fix the cause. STRICT_EP=1 has no effect while EXECUTION_MODE is auto.
Prevent
Section titled “Prevent”- Check the provider after every install, update, or host change. The dashboard shows CUDA on a healthy TX2. Treat CPU (fallback) as a defect, not normal variance.
- Keep the bundle’s compose file intact. It carries
runtime: nvidia, which the GPU needs. - Pin and enforce on GPU-mandatory boxes.
EXECUTION_MODE=cudawithSTRICT_EP=1turns a silent slowdown into a startup failure you can see.
Related
Section titled “Related”- GPU not used — CUDA errors — when CUDA itself does not work on the host.
- Hardware Setup — the TX2’s execution providers.
- Environment Variables —
EXECUTION_MODE,STRICT_EP, andFORCE_CPU. - Monitoring — reading the active provider and latency.