Skip to content

GPU not used — CUDA errors

You see one of these on a Jetson TX2:

  • The Python Inference card shows CPU (fallback), and Running on CPU when GPU expected found no problem with the compose file or the image.
  • The stack does not start, and Docker reports that it does not know the nvidia runtime.
  • The inference container logs no CUDA-capable device is detected.
  • Strict GPU enforcement is on (STRICT_EP=1), and the inference container exits with a STRICT_EP: error and keeps restarting.

Run these on the TX2.

  1. Check the JetPack release:

    Terminal window
    head -n 1 /etc/nv_tegra_release

    JetPack 4.5 prints a line that starts with # R32 (release), REVISION: 5.. The release bundle is built for JetPack 4.5. The container gets its CUDA libraries from the host, so on another JetPack release they may not match what the image expects.

  2. Check that Docker knows the NVIDIA runtime:

    Terminal window
    docker info 2>/dev/null | grep -i runtime

    The Runtimes: line must include nvidia.

  3. Check how the inference container runs:

    Terminal window
    docker inspect -f '{{.HostConfig.Runtime}} {{.Config.User}}' aiboard-inference-real

    It must print nvidia 0:0. The TX2’s GPU device nodes need the NVIDIA runtime and root. Without root, CUDA reports no CUDA-capable device is detected on a healthy board.

  4. Look for CUDA errors from the latest start:

    Terminal window
    docker logs --tail 200 aiboard-inference-real 2>&1 | grep -iE "cuda|STRICT_EP"

Fix the check that failed, then restart inference.

  1. No nvidia runtime in Docker. Install the NVIDIA container runtime that ships with JetPack 4.5, from the JetPack 4.5 apt repository or with NVIDIA SDK Manager:

    Terminal window
    sudo apt-get install -y nvidia-container-runtime

    Then register it in /etc/docker/daemon.json and make it the default:

    {
    "runtimes": {
    "nvidia": { "path": "nvidia-container-runtime", "runtimeArgs": [] }
    },
    "default-runtime": "nvidia"
    }

    Restart Docker with sudo systemctl restart docker.

  2. Container not on nvidia 0:0. The installed compose file was changed. Restore it by running sudo ./update.sh from a release bundle, of the same version or newer.

  3. Not JetPack 4.5. Run the bundle on a TX2 flashed with JetPack 4.5.

  4. Restart inference and check the dashboard:

    Terminal window
    docker restart aiboard-inference-real

    The service allows itself 90 seconds to start. The card then shows CUDA.

Run on the CPU on purpose instead of fighting the GPU. In the installed compose file, set EXECUTION_MODE=cpu for the inference service, and remove STRICT_EP=1 if you added it. Running on CPU when GPU expected shows where the file is and how to apply the change. The card then shows CPU, and the box keeps predicting until you fix the host. Set EXECUTION_MODE back to auto afterwards.

Do not combine STRICT_EP=1 with EXECUTION_MODE=cuda on a broken host. The container then refuses to start, which is the opposite of what you want during the workaround.

  • Re-run the checks after any host change. After a JetPack, Docker, or NVIDIA package change, run the checks above and confirm the card shows CUDA.
  • Make a silent fallback loud where the GPU is mandatory. Pin EXECUTION_MODE=cuda and set STRICT_EP=1, so a broken CUDA stack stops the container at startup instead of quietly running on the CPU.
  • First start is slow, not stuck. The inference health check shows starting for up to 90 seconds while CUDA initialises. TensorRT is not used on the TX2, so there is no engine compile.