RAG in ISMS

RAG stands for Retrieval-Augmented Generation. RAG is a form of generative AI like ChatGPT— it's an architecture/technique that combines a generative model (typically a large language model) with a retrieval step.

In our proof-of-concept project we use RAG to answer questions about information stored in our own security documentation. That means that our local AI will only process and make its reasoning based on our own resources, not something found in internet.

This video shows how RAG formulates an answer to "Can I send a classified document via email to a third party?"

Note: the reason the search takes 11 seconds (analyzing about 18 documents stored in our ISMS) is that in our PoC we use a simple Intel® Core™ i5-9400F Processor, 6 cores, 6 threats, 9M Cache, up to 4.10 GHz from 2019 with 16GB RAM, simple GPU. So no real beast.

Theory behind

At its core, RAG still uses an LLM to produce the final output — text generated token by token based on learned patterns, just like "plain" generative AI (e.g., a chatbot without retrieval).

Before generating a response, the system searches an internal knowledge source (a document database, vector store, the web, etc.) for relevant information, and feeds that retrieved content into the model's context alongside the user's query. The model then generates its answer grounded in that retrieved material.

Why this matters
- reduces hallucination: the model can reference actual retrieved facts rather than relying purely on what it "remembers" from training.
- access to current/private data: it can pull in information the model was never trained on (recent events, internal company documents, etc.).
- traceability: outputs can often be traced back to specific source documents.