Skip to content

Jetson Deployment

The Xisom Edge AI Box runs on the NVIDIA Jetson TX2 with JetPack 4.5. The stack has three services: frontend, backend, and the Python inference service. The inference image is the Jetson-specific part. It is built against the CUDA libraries that JetPack provides, with an ONNX Runtime wheel made for JetPack 4.5.

ProfileDevice / JetPackBase imageONNX RuntimeExecution provider
jetson-gpuTX2, JetPack 4.5 (L4T r32.5)l4t-base:r32.5.0onnxruntime-gpu 1.10.0 (cp36 wheel)CUDA, with CPU fallback

TensorRT is not used on the TX2. The ONNX Runtime build for JetPack 4.5 needs a newer TensorRT than the platform ships, so inference runs on CUDA.

Two consequences of this profile to know before you start:

  • ONNX opset ceiling. ONNX Runtime 1.10 supports up to opset 15. Export models for the TX2 with --opset 15. The bundle seeds an opset-15 model for this reason.
  • Frozen Python 3.6 dependency set. JetPack 4.5 is locked to Python 3.6, so the jetson-gpu image uses a pinned set of older dependencies.

You build images natively on a Jetson. Cross-building L4T images under QEMU is unreliable and unsupported.

Terminal window
docker version # >= 20.10
docker compose version # Compose v2 plugin required
sudo apt-get install -y zstd openssl
df -h ~ # keep a few GB free (TX2 eMMC is small)
docker info | grep -i nvidia # the NVIDIA container runtime must be wired in

Pick the JetPack 4.5, Python 3.6 ONNX Runtime GPU wheel from the Jetson Zoo. The generic onnxruntime-gpu wheel from PyPI does not work on Jetson. A “missing library” error at startup is the classic symptom.

flowchart LR
  A["build-release-images.sh\n(build + tag 3 images)"] --> B["build-release-bundle.sh\n(docker save → dist/<profile>-<version>/)"]
  B --> C["install.sh on the target\n(load images → compose up → seed admin)"]
  A -. "build host: Jetson, internet once" .-> B
  C -. "target host: air-gapped OK" .-> C
  1. Build the images (on the build Jetson, internet required once):

    Terminal window
    ORT_WHEEL_URL="https://<jetson-zoo-wheel-for-jetpack-4.5>.whl" \
    ./scripts/build-release-images.sh <version> --profile jetson-gpu
  2. Package the offline bundle:

    Terminal window
    ./scripts/build-release-bundle.sh --profile jetson-gpu --version v<version>

    The output dist/jetson-gpu-v<version>/ contains the compressed image archive with an integrity manifest, the Tegra compose file, the seed model, and the install.sh / update.sh / uninstall.sh runbook scripts. The layout is described in Offline Bundle Install.

  3. Install on the target device (copy the bundle directory via USB/scp):

    Terminal window
    cd dist/jetson-gpu-v<version>
    sudo ./install.sh # detects the Jetson and accepts the jetson-gpu bundle
    # sudo ./install.sh --with-systemd # optional: start on boot

    The installer verifies the archive checksum, loads images, stages the running stack in /opt/aiboard, generates a unique JWT secret and admin password, starts the stack, and waits for health. It prints the dashboard URL and admin password once. Store it immediately.

  4. Verify GPU inference is real:

    Terminal window
    docker ps --format '{{.Names}} {{.Status}}' # inference and backend (healthy), frontend Up
    docker inspect aiboard-inference-real --format '{{.HostConfig.Runtime}}' # expect: nvidia
    docker exec aiboard-inference-real grep -E "resolved to CPU only|STRICT_EP" /data/logs/inference.log # expect: no output

    The stack runs from /opt/aiboard and reads its settings from /opt/aiboard/.env, so check it with docker ps rather than a docker compose command from the bundle directory. Then sign in to the dashboard and confirm the Python Inference card shows CUDA as the execution provider, not CPU (fallback). The inference console shows errors only, so the provider lines are not in docker logs; the warning above goes to /data/logs/inference.log inside the container. See Running on CPU when GPU expected.

The bundle’s compose file carries settings the TX2 needs:

  • runtime: nvidia and root inside the inference container — required for GPU device access on Tegra.
  • /sys mount — GPU load and temperatures come from Tegra sysfs (there is no nvidia-smi on Jetson), feeding the dashboard’s GPU panel.
  • OPENBLAS_CORETYPE pin — avoids an OpenBLAS illegal-instruction crash on Tegra cores.
  • Per-field N/A metrics — dashboard health fields a Jetson cannot report show as N/A instead of fake zeros.

The inference service uses the CUDA execution provider. If CUDA cannot start, it falls back to the CPU, logs a warning, and the dashboard shows CPU (fallback). To make the box refuse to start instead of falling back, see Execution provider fell back to CPU.

Jetson releases are versioned by the VERSION file, not by Git tags. The dashboard’s About panel shows the release version of the installed bundle. If it differs from what you expect, compare it against the bundle’s images/manifest.txt. See Versions & Updates.

SymptomFix
unknown shorthand flag: 'f' in -fCompose v2 plugin missing — sudo apt-get install -y docker-compose-plugin
Cannot autolaunch D-Bus during buildHeadless credential-helper issue — the build script isolates it automatically; for ad-hoc docker commands use an empty DOCKER_CONFIG dir
GPU container: no CUDA-capable deviceHost is missing the NVIDIA container runtime, or the compose lacks runtime: nvidia (the bundle compose already sets it)
Inference exits with a missing-library errorWrong ONNX Runtime wheel — pick the JetPack 4.5 wheel from the Jetson Zoo and rebuild
install.sh: “does not match this bundle”The machine is not a Jetson. Copy the jetson-gpu bundle to the TX2 and install there
Model fails validation at uploadModel exported with opset > 15 — re-export with --opset 15 for ONNX Runtime 1.10