Triton AI Inferenceby Inseyab connecting…
Range

Pipelines

load / unload here — each drives its underlying model(s) –
Pipeline State VRAM Inferences Capacity Req/s Avg ms Feeds Action VRAM MB Instances Batch Max feeds Info
waiting for data…

Hardware live GPU telemetry via NVML

Pipelines Loaded
– / –
Live video feeds
– streams attached
leases held across all pipelines
Disk — physical volumes
–
▶

Add a model to the pool

HuggingFace · git · YOLO .pt · ONNX URL

Model from HuggingFace or git

Paste a repo id — the server reads its config.json and works out what kind of model it is before downloading anything. Chat/vision, embeddings, speech-to-text and text-to-speech all come from here.

Type a repo id to inspect it.

YOLO model (.pt upload)

Upload an ultralytics .pt; it is exported to ONNX on the server (streamed, large files OK).

Export at the size the model was trained/validated at. A mismatch gives plausible boxes with near-zero confidence, not an error.

ONNX model by URL

Direct link to a .onnx file (InsightFace, custom detectors, …).

⚠GPU memory pressure