Operate

Manage model capacity and GPU policy

Choose processing policy and distinguish document slots, model requests, and section concurrency.

Choose the workload and node

Generation and embedding are separate workloads. Inspect the settings of the selected execution node. CPU, GPU first, and GPU only policies have different fallback behaviour.

GPU first can fall back to CPU when the selected node lacks a usable accelerator. GPU only requires an eligible GPU. A saved setting is not proof of native execution.

Understand the limits

SettingDefaultMeaning
Files prepared in parallel3 per processing service1 to 64 document slots. A restart is required for this setting.
GPU utilisation target90%Admission target, not a hard hardware percentage cap
Maximum parallel GPU requests8Safety ceiling, not a promise that eight requests run
Maximum parallel sections per documentAutomatic, value 0Or choose 1 to 8 sections

Apply a change

  1. Open the relevant Inference or model settings and choose the intended node.
  2. Review generation and embedding choices separately.
  3. Save one controlled change.
  4. Allow healthy node synchronisation to apply it. In progress calls keep their acquired execution settings.
  5. Inspect a new call's observed backend and profile.

Read capacity waits

Unavailable or stale GPU measurements reduce automatic admission. Background tests and model fingerprinting can wait while document processing and foreground inference use capacity.

Large models can take time to copy and verify before loading. Storage failure is not fixed by switching the same inaccessible source to CPU.

Next step

Use document processing controls for a specific job. Keep hardware acceptance separate from saved policy.

Pūnaha Docs

Search the guides

Enter at least two characters.

    Product screen

    View the full screenshot