Pipelines
load / unload here — each drives its underlying model(s) –| Pipeline | State | VRAM | Inferences | Capacity | Req/s | Avg ms | Feeds | Action | VRAM MB | Instances | Batch | Max feeds | Info |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| waiting for data… | |||||||||||||
Hardware live GPU telemetry via NVML
Add a model to the pool
HuggingFace · git · YOLO .pt · ONNX URLModel from HuggingFace or git
Paste a repo id — the server reads its config.json and works out what kind of model it is before downloading anything. Chat/vision, embeddings, speech-to-text and text-to-speech all come from here.
YOLO model (.pt upload)
Upload an ultralytics .pt; it is exported to ONNX on the server (streamed, large files OK).
ONNX model by URL
Direct link to a .onnx file (InsightFace, custom detectors, …).