Operate

Run local AI models

Install the local AI runtime, add GGUF models, and manage their memory use.

Open Models

Open AI, then Models. Check the selected server, bundled runtime, model files, Connections, and deployments.

Add a local model

  1. Use the runtime supplied with the matching package.
  2. Obtain a permitted GGUF model and review its licence and hardware requirements.
  3. Upload or select the model file. A file may be up to 32 GB, subject to available storage.
  4. Create or select the Embedded llama.cpp Connection.
  5. Discover the model and enable the intended generation or embedding deployment.
  6. Test the intended workload with safe input.

Load now and load at startup

Load now requests memory loading on the selected server. Load at startup records a preference for a later start. Neither setting guarantees that memory or storage will be available.

Choose processor policy

Generation and embedding settings are separate. CPU, GPU first, and GPU only policies depend on the selected node's available runtime and hardware.

Inspect the observed backend and fallback reason for an actual call. Saving a GPU preference does not prove GPU execution.

Understand waits

Copying and verifying a large shared model can precede loading. Background tests may yield to production work. A verified local model cache still depends on the authoritative source.

Next step

Read Manage model capacity and GPU policy and Control document processing.

Pūnaha Docs

Search the guides

Enter at least two characters.

    Product screen

    View the full screenshot