Skip to content

gRPC Inference API

The backend and the Python inference runtime communicate over gRPC on port 50051 (protos/inference.proto, service InferenceService). Operators rarely call this API directly — the backend is the client — but the contract matters when you debug the pipeline, integrate a custom client on the box, or read inference logs.

RPCTypePurpose
StreamInferencebidirectional streamSensor windows in, predictions out — the real-time hot path
LoadModelunaryValidate + register + hot-swap an ONNX model (zero downtime)
ValidateModelunaryNon-mutating inspection: shapes + SHA-256, nothing loaded
HealthCheckunaryService status + CPU/GPU/memory metrics
StreamLogsserver streamTail inference logs filtered by minimum level

Each InferenceRequest carries request_id, a timestamp (Unix epoch ms, stamped once by the backend at sensor-sample time and echoed back unchanged — so there is no clock-sync problem), a flat features array, and a metadata map.

Each InferenceResponse returns predictions and confidence_scores, the serving model_id/model_version, inference_time_ms, and three timestamps:

FieldMeaning
window_start_timestampOldest input sample in the prediction’s window
window_end_timestampNewest input sample — the “as-of” time of the prediction
emitted_atWall-clock at emission; emitted_at − window_end_timestamp approximates end-to-end latency

During model warm-up (the sliding window is not yet full) responses carry metadata["warming_up"] = "true"; the backend suppresses these from the dashboard.

Registers a model version and hot-swaps it into the running session with double-buffering — the new session is loaded and warmed outside the lock, then the reference is swapped atomically, so streaming never pauses.

  • Idempotent by content hash: re-sending the same (model_id, version) with an identical SHA-256 is a success no-op. The same version with a different SHA-256 is rejected as version_conflict — versions are immutable.
  • The response echoes input_shape, output_shape, registered_path, and the computed sha256.
  • If the new model’s (window_size, n_features) differ from the previous one, the windowing pipeline is rebuilt and warm-up restarts.

Pure inspection used by the backend during model upload: opens a throwaway runtime session, extracts input/output shapes and SHA-256, and returns {valid, message, input_shape, output_shape, sha256} without registering anything.

Validation enforces the model contract:

  • ONNX file ≤ 500 MB, loadable, and passing the ONNX checker
  • Opset 13–18
  • Input tensor input shaped (batch, window_size, n_features) — 3-D required; the batch dimension may be dynamic
  • Output tensor output shaped (batch, 3) = [health_score, failure_probability, rul_normalized]

HealthCheck powers the dashboard health panel (GPU load, memory, uptime). Besides status, CPU, host memory and uptime, the reply carries these fields (the full message is HealthCheckResponse in protos/inference.proto):

FieldMeaning
swap_used_mb, swap_total_mbThe host’s swap; -1 when the host has none
container_memory_used_mb, container_memory_limit_mbThe inference container against the memory limit it is stopped at (RAM plus swap); -1 when the container has no limit or it cannot be read
gpu_metrics.memory_kindMEMORY_KIND_SHARED (Jetson: the GPU uses system RAM, so the GPU memory fields repeat the host reading), MEMORY_KIND_DEDICATED (discrete VRAM), or MEMORY_KIND_UNKNOWN
model_window_size, model_feature_countWindow length and feature count of the model session the engine actually has loaded; 0 when no model is loaded
model_loaded, stub_modeWhether a usable model session is loaded, and whether the engine is emitting synthetic output

These fields are additive: an older inference engine does not send them, so they arrive as 0 (or MEMORY_KIND_UNKNOWN) and the dashboard shows the value as not reported.

StreamLogs forwards the Python service’s log stream to the backend, which persists it for the LogViewer — filter with min_level.

Environment variables the inference container honors (set in the compose file):

VariableDefaultEffect
EXECUTION_MODEautoExecution provider: auto (CUDA, CPU as fallback), cuda, cpu. TensorRT is not used on the TX2.
STRICT_EPoff1 = refuse to start instead of silently degrading a pinned GPU mode to CPU
MODEL_PATH/app/models/predictive_maintenance_op15.onnxSeed model loaded at startup
MODEL_STORAGE_PATH/data/modelsWhere registered model versions are stored
LOG_DIRunsetEnables the rotating on-disk log file (inference.log); unset = console only
FORCE_CPUoffLegacy alias for EXECUTION_MODE=cpu