Gemma: chat with your knowledge Actual frontend captures, with pauses shortened and synthetic narration. The existing verified model download is reused. Static captures are held while their instructions are explained. 00:00:00.000 Download EmbeddingGemma This walkthrough uses two models. EmbeddingGemma turns documents and questions into searchable vectors. Gemma E4B writes the answer. Open the linked EmbeddingGemma 300M Q8_0 GGUF file on Hugging Face. We already have the verified file, so we reuse it. The recorded installation has the signed R C one embedding repair and uses C P U only for answer generation. See the guide for the tested build and processing setting. 00:00:34.499 Upload the embedding file In your test tenant, open AI, then Models. Choose Add GGUF model and select the embedding file. Wait for it to finish. This uploads the weights to the Pūnaha server. You do not need a Hugging Face API key in the local connection. 00:00:55.371 Enable and load the model Find the uploaded model under the local LLM Model connection. Enable its deployment if necessary, and label it Gemma embeddings. Choose Load into memory now. In memory confirms loading, but successful document processing must still be tested. 00:01:15.507 Create a knowledge database Open AI, then Knowledge, and create a database. Name it Garden club knowledge. Select Semantic Chunking and the uploaded EmbeddingGemma model with seven hundred and sixty eight dimensions. Leave the processing model as None for this small plain text example. 00:01:35.269 Add harmless practice facts Download the practice text file from the guide. It describes a fictional garden club, with meeting times, tool borrowing and seed exchange rules. Choose Upload documents, then Choose documents. Select the file and start processing. Use the supplied fictional facts, not private customer material. 00:02:00.380 Check indexing and search Wait for document processing to finish successfully. Then search for when and where the club meets. Check that the returned passage says every Sunday at 9:30 am in the Glasshouse Room. Stop and resolve any processing or search failure before continuing. 00:02:19.646 Use the knowledge template Create a workflow using Answer from approved knowledge. Name it Gemma knowledge chat. This creates a connected retrieval, prompt and language model path. Keep the required protection gateway enabled. 00:02:35.868 Connect the database Open Find approved knowledge. Select Garden club knowledge and Focused, with three matches. Retrieval uses the latest question. Similar passages can still be returned when the requested fact is missing, so the next step controls when to refuse. 00:02:55.695 Require a supported answer Open Prepare grounded answer and copy the complete prompt from the guide. It passes the latest question and the knowledge step's results. It asks for a brief supported answer, a source and an exact quote. If the passages do not support the full answer, it requests the exact refusal sentence. 00:03:17.781 Choose the answer model In Generate answer, select Gemma E4B chat. Use temperature zero, Balanced response detail and a maximum of two hundred and fifty six response tokens. Save the draft and check readiness. Readiness alone does not prove the answers are correct. 00:03:38.458 Detect a refusal After Generate answer, add a Condition named Refusal detected. Inspect steps dot model dot content. Choose contains, select Text, and enter the exact refusal sentence. This detects a refusal even when the model adds a partial answer before it. 00:03:59.239 Return the exact refusal Add two Prompt modules after the condition, alongside each other. Return exact refusal contains only the fixed refusal sentence. Return supported answer uses steps dot model dot content inside double braces. Edit their connections to use the true and false condition branches respectively. These modules format text without another model call. This check does not independently verify the truth of a model answer. 00:04:30.733 Test a supported question Test the draft with a question answered by the document: How long can a Patchwork Garden Club member borrow a hand tool? The recorded answer says seven days and labels the source Document. Compare the answer with the retrieved passage, even if the model omits a separate quote. A completed run alone does not establish accuracy. 00:04:54.774 Test missing information Now ask the club's annual membership fee. The document does not contain a fee. The required reply is: I can't find an answer to your question. Also test an unrelated question and a question that mixes a known fact with a missing fact. Do not publish if these tests fail. 00:05:16.920 Publish after the tests pass After reviewing the tests, publish the workflow version. Open Chat in the same tenant and select Gemma knowledge chat. Start a new chat so it uses the version you just published. 00:05:31.092 Check both outcomes in Chat Repeat a supported question and a missing-information question in Chat. Check both results and the saved runs. The prompt guides the model but cannot guarantee every future answer. Retest whenever you change the model, prompt or source documents. Use the written guide for the full template and troubleshooting.