Skip to content

Running on CPU when GPU expected

Inference is slower than expected, and the Python Inference card on the dashboard shows CPU (fallback) instead of CUDA.

On the TX2 the execution provider chain is CUDA → CPU. The default execution mode, auto, tries CUDA first. If CUDA cannot start, the service keeps running on the CPU. That safety net also hides the problem: a box that fell back still predicts, just slower.

A plain CPU badge, without “(fallback)”, means the box was set to CPU on purpose (EXECUTION_MODE=cpu or FORCE_CPU=1). That is not a fallback.

The badge is the ground truth. It shows the provider that the loaded model actually runs on, as the inference service reports it.

The inference container’s console shows errors only. Warnings go to a log file inside the container. Search it for the fallback warning:

Terminal window
docker exec aiboard-inference-real grep -E "resolved to CPU only|STRICT_EP" /data/logs/inference.log

Each line starts with a timestamp. Look at the lines from the latest start.

  • resolved to CPU only → ONNX Runtime in this container offers no CUDA provider at all.
  • No match, but the badge still shows CPU (fallback) → CUDA was offered but could not start on the GPU.

To see which execution mode the container started with:

Terminal window
docker logs aiboard-inference-real 2>&1 | grep "Starting REAL inference"

A box on the default prints Starting REAL inference service (Jetson GPU, EXECUTION_MODE=auto)....

The dashboard’s Logs page also streams the inference service’s informational lines. When it has the start-up lines of a healthy box, you see CUDA EP enabled, then a Providers: line that lists CUDAExecutionProvider first.

  1. Check that the inference container runs with the NVIDIA runtime:

    Terminal window
    docker inspect -f '{{.HostConfig.Runtime}}' aiboard-inference-real

    It must print nvidia. The NVIDIA runtime is what makes the TX2’s CUDA libraries visible inside the container. The bundle’s compose file sets it. If it prints anything else, the compose file was changed. Restore it by running sudo ./update.sh from a release bundle, of the same version or newer. The update copies the bundle’s compose files back into /opt/aiboard/compose/.

  2. Check that the inference image is the Jetson GPU image:

    Terminal window
    docker inspect -f '{{.Config.Image}}' aiboard-inference-real

    The tag must end in -jetson-gpu. Any other image is not built for the TX2’s GPU. Install the jetson-gpu release bundle. See Offline Bundle Install.

  3. If both checks pass and the box still falls back, the host’s CUDA stack is the problem. Follow GPU not used — CUDA errors.

On a box that must use the GPU, make the service refuse to start instead of falling back. Pin the execution mode to cuda and turn on strict enforcement.

The release compose file sets EXECUTION_MODE=auto for the inference service and does not set STRICT_EP. The .env file does not control either one. Change them in the installed compose file:

  1. Open /opt/aiboard/compose/docker-compose.release.yml. In the environment: list of the inference service, change EXECUTION_MODE=auto to EXECUTION_MODE=cuda and add STRICT_EP=1:

    - EXECUTION_MODE=cuda
    - STRICT_EP=1
  2. Recreate the inference container so it picks up the change:

    Terminal window
    sudo docker compose --env-file /opt/aiboard/.env \
    -f /opt/aiboard/compose/docker-compose.release.yml up -d inference

With this setting, a session that does not start on CUDA stops the service with a STRICT_EP: error, and the container restarts until you fix the cause. STRICT_EP=1 has no effect while EXECUTION_MODE is auto.

  • Check the provider after every install, update, or host change. The dashboard shows CUDA on a healthy TX2. Treat CPU (fallback) as a defect, not normal variance.
  • Keep the bundle’s compose file intact. It carries runtime: nvidia, which the GPU needs.
  • Pin and enforce on GPU-mandatory boxes. EXECUTION_MODE=cuda with STRICT_EP=1 turns a silent slowdown into a startup failure you can see.