Get started

Download Gemma and create your first chat

Download Gemma 4 E4B from Hugging Face, upload the GGUF model to Pūnaha, and publish a workflow you can use in Chat.

Your progress

Follow this journey

0 steps complete
0%

Progress is saved only in this browser.

Recorded RC1 frontend walkthrough with synthetic narration and English captions. Pauses are shortened. The upload uses a real still from the earlier session. The existing model is reused, so the download is shown without downloading it again. Read the video transcript.

What you will make

You will download a language model, make it available in Pūnaha, and build a workflow that answers your messages.

Chat input → Prompt → Language model → Chat output

This example uses Gemma 4 E4B Instruct for text chat. It does not need a knowledge database or an embedding model.

Download the narrated frontend walkthrough or read its transcript. These video links need an internet connection. The written guide and screenshots are included in the offline documentation.

Before you start

  • Install RC1, finish first setup, and open a test tenant. Use harmless practice messages.
  • Use an account allowed to manage models, create and publish workflows, and open Chat.
  • Check that the AI and Workflow services are ready. For this first example, use one server with the required services.
  • Allow about 5 GB for the model download and another copy on the Pūnaha server. Leave additional storage and memory for loading, context, and other services.
  • Use the runtime supplied with your RC1 package. A downloaded model still needs a successful load and response test on your hardware.

Download the RC1 package and current Windows repair instructions. The recorded installation had the published Windows runtime repair applied.

RC1 registration is unavailable. Export anything you need before its 720 powered on hour allowance ends. Read RC1 availability and limits.

1. Download the correct Gemma file

A GGUF file contains model weights in the format used by Pūnaha's local runtime. Upload the model file, not a Python package or a whole repository.

  1. Read Google's Gemma 4 E4B Instruct model card and the linked model licence.
  2. Open the GGUF files published by Unsloth on Hugging Face. This example links to a fixed revision reviewed on 7 October 2026.
  3. Choose gemma-4-E4B-it-Q4_K_M.gguf, then choose Download → Download file. The file is about 4.98 GB.
  4. Wait for the download to finish. Keep the .gguf extension and check that the browser has not saved an incomplete download.
  5. Do not choose a file beginning mmproj or mtp. Those are companion files, not the language model for this text chat example. Files ending .safetensors are also not this upload format.

Q4_K_M identifies this model's quantization. Its disk size is not the memory needed to run it. The 4.98 GB download is about 4.64 GiB. Some screens display this as 4.6 GB.

Already have this file? Verify it and continue to upload. The walkthrough reuses an existing download.

Hugging Face download menu for the Gemma Q4_K_M file
The verified file page, captured 7 October 2026. Open the full image.
Optional: verify the downloaded file

For the exact linked revision, the file has 4,977,171,584 bytes and this SHA256 value:

85a896a047553e842f25297ee5b031d64ff30147d9c4af17b1e4b394cd1fab87

On Windows, open PowerShell in the download folder and run:

Get-FileHash .\gemma-4-E4B-it-Q4_K_M.gguf -Algorithm SHA256

On Linux, run:

sha256sum gemma-4-E4B-it-Q4_K_M.gguf

Compare the complete result. A different model revision can have a different hash. Do not continue with a partial or mismatched file.

2. Upload the model to Pūnaha

  1. Open the target Pūnaha server's frontend and select your test tenant.
  2. Open AI → Models. The page is titled Model runtime and connections.
  3. Check the selected server and runtime status. If it is missing, follow your package's installation or repair instructions before continuing.
  4. Choose Add GGUF model and select gemma-4-E4B-it-Q4_K_M.gguf from your computer.
  5. Keep the page open while the upload finishes. Confirm the completed model appears in the local model list.

The model file is copied to the Pūnaha server. It runs through the local runtime. You do not need a Hugging Face API key in the local model Connection.

If the file is already listed, reuse it. Uploading the same filename again is rejected.

The Add GGUF model control showing the Gemma upload in progress
Real upload progress from 6 October 2026. The uploaded file was reused for the final test on 7 October. Open the full image.

If the upload control is missing, open the target server's own frontend. Selecting a remote server does not enable uploading to it from another node.

3. Enable the model for workflows

  1. Look for the local LLM Model Connection and the uploaded model. Local discovery may have created them already.
  2. If there is no local Connection, choose Create connection. Give it a clear name and select LLM Model as its Provider. This is the embedded llama.cpp option.
  3. On that Connection, choose Discover models. Select the uploaded Gemma file and choose Enable selected models if it is not already enabled.
  4. Check that the Connection and model are enabled and that the model supports Chat / response generation.
  5. Use Change tenant label to name the deployment Gemma E4B chat. Keep its underlying model ID unchanged.
  6. Choose Load into memory now and check the result. Initial loading can take time. A memory or compatibility error must be resolved before continuing.

A deployment is the model entry that a workflow can select. Uploading a file and enabling its deployment are separate checks. Load at startup is optional.

If the model already says In memory, continue without loading it again.

Enabled Gemma E4B chat deployment with In memory status
The enabled deployment is loaded in the installed RC1 frontend. Open the full image.

4. Create the chat workflow

  1. Open Workflows, then choose Create workflow.
  2. Select Blank workflow. Name it Gemma chat tutorial and describe it as Practice text chat with a local Gemma model.
  3. Create the draft, then open it. The address is generated from the name. Use a different name if that address is already taken.
  4. Use Add module to add Chat input, Prompt, Language model, and Chat output.
  5. Check the connections. They must form one path in that order, with no disconnected module or extra branch. Keep After the previous stage selected and leave the required Protection gateway enabled.
  6. Open the Prompt module and enter the settings below.
Blank workflow creation named Gemma chat tutorial
Create a small practice workflow before adding other capabilities. Open the full image.

System instructions

You are a helpful assistant. Answer clearly and briefly.
If you do not know, say so. Do not invent facts.

Paste this complete User prompt template. It includes the instructions in the text passed to the next module:

You are a helpful assistant. Answer clearly and briefly.
If you do not know, say so. Do not invent facts.

Conversation so far:
{{input.messages}}

Reply to the latest user message:
{{inputText}}

{{input.messages}} supplies the conversation messages. {{inputText}} supplies the latest user message. Keep both pairs of braces.

Using only {{input}} supplies the latest message. It does not pass earlier replies as context. Use the complete template above for this short conversation example.

5. Choose the model and test the draft

  1. Open the Language model module.
  2. Set Model deployment to Gemma E4B chat. Confirm that it refers to the uploaded Gemma file.
  3. For a short first test, use the example settings below. They are a starting point, not a model benchmark.
  4. Choose Save draft, then Check readiness. Resolve each reported error.
  5. Choose Test draft. Send In two sentences, explain what a language model does.
  6. Check the run result and the Language model step. You should receive a relevant answer without a model or runtime error.
SettingExample value
Model deploymentGemma E4B chat
Temperature0.2
Response detailBalanced
Maximum response tokens256, under Advanced response settings
Gemma deployment with temperature 0.2 and Balanced response detail
Choose the enabled local deployment and check the response limit. Open the full image.

Draft testing runs a saved draft snapshot. It does not publish the workflow. A readiness check alone does not prove that the model can generate an answer.

Keep this practice conversation short. Longer conversations need separate context and capacity testing.

6. Publish and start chatting

  1. After the draft test succeeds, return to Workflow design. Choose Publish version and confirm publication.
  2. Open Chat in the same tenant and choose Gemma chat tutorial.
  3. Start a new chat and send Suggest three names for a fictional garden club.
  4. Read the answer. Choose one returned name and ask for a change, quoting that name explicitly.
  5. Check the answer and saved run. Responses vary, so check relevance and successful completion rather than exact wording.

For example: Make the name "The Verdant Patch Club" more playful. Give one name only. Replace the quoted name with one from your own answer.

Our test returned The Patchwork Petals Club. An earlier request to change "the second name" changed a different name. Quoting the intended name made the request clearer.

Chat lists published workflows. Saving or testing a draft is not enough. After changing the workflow, publish it again and start a new chat to use that version.

Gemma Chat with the named follow-up and The Patchwork Petals Club reply
A real conversation using published workflow version 2, captured 7 October 2026. Open the full image.

If something does not work

What you seeWhat to check
Upload rejected or interruptedCheck the GGUF file and free server storage. RC1 accepts one file up to 32 GB. Follow the upload error before retrying.
Gemma is missing from Model deploymentConfirm the same tenant, enabled Connection, enabled chat deployment, and your access to it.
Model cannot loadCheck the complete file, memory, and the package runtime's support for this Gemma architecture. Keep the error for your administrator.
Follow up misses earlier detailsUse the complete conversation template in step 4. Publish the changed draft and start a new chat. Quote the detail you want changed.
Raw model data or an empty answerDo not count this as success. Check the model step and prompt context. RC1 can display raw result data when the model returns no text.
Slow answer or timeoutStart with a short message and modest response limit. Check CPU or GPU use and other jobs. Ask an administrator to review capacity and timeout settings.
An answer stops earlyInspect the run result. A response token limit can truncate an answer. Adjust it carefully after the first short test succeeds.
The workflow is missing from ChatPublish it, use the same tenant, and check Chat access and workflow Run access.
Normal use is restrictedCheck the RC1 powered hour allowance. Registration is unavailable in RC1 and cannot restore access after its cutoff.

Check your result

You should have one complete Gemma file, an enabled local chat deployment, a successful draft test, and a published workflow that answers in Chat.

This installed example was checked on 7 October 2026 using 0.1 RC1 (0.1.0-rc.1), Product revision decfd63165046095b38c2062725a89edd268ab35, and Windows runtime repair2.

The site header identifies the broader documentation source baseline. The installed build above identifies this tutorial's frontend test. This was an existing repaired installation, not a fresh installation acceptance test.

The video combines actual frontend screen captures with synthetic narration and captions. Pauses are shortened. Upload progress is a still from the original session. No Product screen or model answer is simulated.

Next steps

Read Run local AI models, Check and publish workflows, and Use Chat. To add document search later, read Prepare a model and knowledge database.

Sources

Model information: Google model card, Unsloth GGUF files, and Google's llama.cpp guidance. Follow the linked model terms separately from the Pūnaha licence.

Pūnaha Docs

Search the guides

Enter at least two characters.

    Product screen

    View the full screenshot