Gemma: from model file to Chat Real frontend recordings from 7 October 2026. Upload progress is a still from 6 October. Pauses are shortened; synthetic narration. Existing model reused; no repeat download. 00:00:00.000 Download the right model This walkthrough uses the real Pūnaha RC1 frontend. On Hugging Face, choose the Gemma 4 E4B instruct Q4_K_M file. Open Download, then Download file. We already have the verified file, so we do not download it again. 00:00:19.727 Open AI, then Models In your test tenant, open AI, then Models. Check that the local runtime is installed. This Windows installation had the published runtime repair applied before the successful tests. 00:00:33.659 Upload the GGUF file Choose Add GGUF model and select the downloaded file. This real still shows the original upload in progress. Wait until the model appears in the file list. Reuse an existing upload instead of uploading the same filename twice. 00:00:50.876 Enable the deployment Find the local LLM Model connection. If discovery has not enabled the uploaded model, choose Discover models and Enable selected models. The local connection does not need a Hugging Face API key. 00:01:05.758 Name and load Gemma Set the tenant label to Gemma E4B chat. If the model is not already loaded, choose Load into memory now. In memory confirms it is loaded. Our model was already in memory, so we reuse it. 00:01:21.365 Create a blank workflow Create a workflow and choose Blank workflow. Name it Gemma chat tutorial, add a short description, and open the saved draft from the workflow list. 00:01:43.317 Build one connected path Use Add module to build one path: Chat input, Prompt, Language model, then Chat output. Keep After the previous stage selected. Leave the required protection gateway enabled. 00:02:08.501 Pass the conversation Copy the full prompt template from the guide. It includes the conversation messages and latest user message. Using input alone passes only the latest message, so earlier replies are missing from the model context. 00:02:24.102 Select Gemma Select Gemma E4B chat in the Language model module. Use temperature zero point two, Balanced response detail, and a maximum of two hundred and fifty six response tokens. These are example settings, not a performance guarantee. 00:02:41.889 Return the answer Add Chat output as the final module. Check the connections, save the draft and check readiness. Resolve any reported errors before running your first test. 00:02:53.961 Test the saved draft Choose Test draft. Ask: In two sentences, explain what a language model does. Start the test and inspect the result. A readiness check alone does not prove the model can generate an answer. 00:03:12.931 Check the result The revised draft returned an answer and completed all four steps. Read the answer as well as the status. Model output always needs a human check. 00:03:24.349 Publish the workflow Return to Workflow design and choose Publish version. Confirm the new version. This example publishes version two after correcting the prompt. Your first publication normally shows version one. 00:03:38.846 Use the workflow in Chat Open Chat in the same tenant and select Gemma chat tutorial. Check the published version. Send: Suggest three names for a fictional garden club. The local model returns three suggestions. 00:04:01.914 Review the model's answer We asked for the second name to be more playful. The model changed a different name. This is why checking relevance matters, even when the workflow completes successfully. 00:04:20.709 Make the request precise Now quote the intended name explicitly: Make the name The Verdant Patch Club more playful. Give one name only. The model responds with The Patchwork Petals Club. 00:04:35.623 Your first local AI chat You now have a local model, a tested published workflow, and a working chat. Use the written guide for the exact template and troubleshooting. RC1 registration is unavailable. Export what you need before its seven hundred and twenty powered on hour cutoff.