GPU not used — CUDA errors
Symptom
Section titled “Symptom”You see one of these on a Jetson TX2:
- The Python Inference card shows CPU (fallback), and Running on CPU when GPU expected found no problem with the compose file or the image.
- The stack does not start, and Docker reports that it does not know the
nvidiaruntime. - The inference container logs
no CUDA-capable device is detected. - Strict GPU enforcement is on (
STRICT_EP=1), and the inference container exits with aSTRICT_EP:error and keeps restarting.
Confirm
Section titled “Confirm”Run these on the TX2.
-
Check the JetPack release:
Terminal window head -n 1 /etc/nv_tegra_releaseJetPack 4.5 prints a line that starts with
# R32 (release), REVISION: 5.. The release bundle is built for JetPack 4.5. The container gets its CUDA libraries from the host, so on another JetPack release they may not match what the image expects. -
Check that Docker knows the NVIDIA runtime:
Terminal window docker info 2>/dev/null | grep -i runtimeThe
Runtimes:line must includenvidia. -
Check how the inference container runs:
Terminal window docker inspect -f '{{.HostConfig.Runtime}} {{.Config.User}}' aiboard-inference-realIt must print
nvidia 0:0. The TX2’s GPU device nodes need the NVIDIA runtime and root. Without root, CUDA reportsno CUDA-capable device is detectedon a healthy board. -
Look for CUDA errors from the latest start:
Terminal window docker logs --tail 200 aiboard-inference-real 2>&1 | grep -iE "cuda|STRICT_EP"
Fix the check that failed, then restart inference.
-
No
nvidiaruntime in Docker. Install the NVIDIA container runtime that ships with JetPack 4.5, from the JetPack 4.5 apt repository or with NVIDIA SDK Manager:Terminal window sudo apt-get install -y nvidia-container-runtimeThen register it in
/etc/docker/daemon.jsonand make it the default:{"runtimes": {"nvidia": { "path": "nvidia-container-runtime", "runtimeArgs": [] }},"default-runtime": "nvidia"}Restart Docker with
sudo systemctl restart docker. -
Container not on
nvidia 0:0. The installed compose file was changed. Restore it by runningsudo ./update.shfrom a release bundle, of the same version or newer. -
Not JetPack 4.5. Run the bundle on a TX2 flashed with JetPack 4.5.
-
Restart inference and check the dashboard:
Terminal window docker restart aiboard-inference-realThe service allows itself 90 seconds to start. The card then shows CUDA.
While the host is still broken
Section titled “While the host is still broken”Run on the CPU on purpose instead of fighting the GPU. In the installed compose file, set EXECUTION_MODE=cpu for the inference service, and remove STRICT_EP=1 if you added it. Running on CPU when GPU expected shows where the file is and how to apply the change. The card then shows CPU, and the box keeps predicting until you fix the host. Set EXECUTION_MODE back to auto afterwards.
Do not combine STRICT_EP=1 with EXECUTION_MODE=cuda on a broken host. The container then refuses to start, which is the opposite of what you want during the workaround.
Prevent
Section titled “Prevent”- Re-run the checks after any host change. After a JetPack, Docker, or NVIDIA package change, run the checks above and confirm the card shows CUDA.
- Make a silent fallback loud where the GPU is mandatory. Pin
EXECUTION_MODE=cudaand setSTRICT_EP=1, so a broken CUDA stack stops the container at startup instead of quietly running on the CPU. - First start is slow, not stuck. The inference health check shows
startingfor up to 90 seconds while CUDA initialises. TensorRT is not used on the TX2, so there is no engine compile.
Related
Section titled “Related”- Running on CPU when GPU expected — when the GPU works but the service still falls back.
- Jetson Deployment — the TX2 settings the bundle’s compose file carries.
- Hardware Setup — the TX2’s execution providers.
- Monitoring — where the active provider is shown.