Operate
Manage model capacity and GPU policy
Choose processing policy and distinguish document slots, model requests, and section concurrency.
Choose the workload and node
Generation and embedding are separate workloads. Inspect the settings of the selected execution node. CPU, GPU first, and GPU only policies have different fallback behaviour.
GPU first can fall back to CPU when the selected node lacks a usable accelerator. GPU only requires an eligible GPU. A saved setting is not proof of native execution.
Understand the limits
| Setting | Default | Meaning |
|---|---|---|
| Files prepared in parallel | 3 per processing service | 1 to 64 document slots. A restart is required for this setting. |
| GPU utilisation target | 90% | Admission target, not a hard hardware percentage cap |
| Maximum parallel GPU requests | 8 | Safety ceiling, not a promise that eight requests run |
| Maximum parallel sections per document | Automatic, value 0 | Or choose 1 to 8 sections |
Apply a change
- Open the relevant Inference or model settings and choose the intended node.
- Review generation and embedding choices separately.
- Save one controlled change.
- Allow healthy node synchronisation to apply it. In progress calls keep their acquired execution settings.
- Inspect a new call's observed backend and profile.
Read capacity waits
Unavailable or stale GPU measurements reduce automatic admission. Background tests and model fingerprinting can wait while document processing and foreground inference use capacity.
Large models can take time to copy and verify before loading. Storage failure is not fixed by switching the same inaccessible source to CPU.
Next step
Use document processing controls for a specific job. Keep hardware acceptance separate from saved policy.
Was this page helpful?
Your answer helps us improve the documentation.
Do not include personal information, customer information, passwords, or keys.