Operate
Run local AI models
Install the local AI runtime, add GGUF models, and manage their memory use.
Open Models
Open AI, then Models. Check the selected server, bundled runtime, model files, Connections, and deployments.
Add a local model
- Use the runtime supplied with the matching package.
- Obtain a permitted GGUF model and review its licence and hardware requirements.
- Upload or select the model file. A file may be up to 32 GB, subject to available storage.
- Create or select the Embedded llama.cpp Connection.
- Discover the model and enable the intended generation or embedding deployment.
- Test the intended workload with safe input.
Load now and load at startup
Load now requests memory loading on the selected server. Load at startup records a preference for a later start. Neither setting guarantees that memory or storage will be available.
Choose processor policy
Generation and embedding settings are separate. CPU, GPU first, and GPU only policies depend on the selected node's available runtime and hardware.
Inspect the observed backend and fallback reason for an actual call. Saving a GPU preference does not prove GPU execution.
Understand waits
Copying and verifying a large shared model can precede loading. Background tests may yield to production work. A verified local model cache still depends on the authoritative source.
Next step
Read Manage model capacity and GPU policy and Control document processing.
Was this page helpful?
Your answer helps us improve the documentation.
Do not include personal information, customer information, passwords, or keys.