Local GraphRAG setup with Ollama and Chaos Cypher explained
A new tutorial demonstrates how to run a local knowledge graph using Ollama and Chaos Cypher, enabling private document analysis without cloud APIs.
A recent guide published on September 28, 2026, outlines a method for running Graph Retrieval-Augmented Generation (GraphRAG) entirely on local hardware. The tutorial, written by Denis MacPherson for Chaos Cypher, demonstrates how to combine the Ollama local large language model runner with the Chaos Cypher platform to process documents privately. This approach allows developers and system administrators to build knowledge graphs and query them with citations without sending data to external cloud services or incurring API costs.
What happened
The article provides a step-by-step walkthrough for setting up a local AI pipeline that extracts entities and relationships from uploaded documents. The process begins with installing Ollama and pulling a specific model, qwen3:30b-instruct, which requires approximately 18 to 20 GB of storage. While this model download runs in the background, users are instructed to launch Chaos Cypher via Docker. The container exposes ports 80 and 443 and mounts a volume for data persistence. On Linux systems using Docker Engine rather than Docker Desktop, an additional host gateway configuration is required to allow the container to communicate with Ollama running on the host machine.
Once the services are online, users upload a PDF, DOCX, or text file to the Chaos Cypher interface. The system immediately begins indexing the document by chunking the text and creating embeddings, a process that takes about 30 seconds per 100 pages and does not require the large language model. After indexing, the system detects the document domain, such as legal or technical, and queues entity extraction. This extraction phase relies on the previously downloaded Ollama model and can take five to ten minutes for a 100-page document. During this time, the system builds a graph of nodes representing people, organizations, and concepts, connected by edges that define their relationships.
The final stage involves querying the generated graph through a chat interface. When a user asks a question, Chaos Cypher retrieves relevant text chunks and graph context, passing them to the local LLM. The model generates an answer with inline citations that link directly to the specific paragraph in the source document where the information was found. This verification step ensures that answers are grounded in the provided data rather than hallucinated. The author notes that while each upload currently builds its own graph, cross-source entity resolution is planned for future updates.
Key details
- The setup requires downloading the qwen3:30b-instruct model, which is approximately 18-20 GB in size.
- Chaos Cypher runs in a Docker container and typically takes 30-60 seconds to start all internal services.
- Indexing and embedding occur locally without the large language model, processing at a rate of roughly 30 seconds per 100 pages.
- Entity extraction uses the local LLM and takes about 5-10 minutes for a 100-page document on a 30B-class model.
- Citations in the chat interface link to exact text chunks, allowing users to verify the source of each claim.
- The system supports mixing local processing for privacy with cloud providers for extraction if higher quality is needed for complex documents.
Background
GraphRAG extends traditional Retrieval-Augmented Generation by incorporating knowledge graphs into the retrieval process. Standard RAG systems split documents into chunks and retrieve them based on semantic similarity, which can miss broader contextual relationships between distant parts of a document. GraphRAG identifies entities like people, places, and events, then maps the relationships between them. This structure allows the AI to reason over connected concepts rather than just isolated text fragments, leading to more comprehensive answers for complex queries.
Running these models locally addresses growing concerns about data privacy and operational costs. By using tools like Ollama, organizations can keep sensitive documents within their own infrastructure. This eliminates the risk of data leakage to third-party API providers and removes variable usage fees. However, local deployment requires sufficient hardware resources, particularly GPU memory, to handle the computational load of both embedding creation and large language model inference.
Why it matters
For teams managing sensitive intellectual property or regulated data, the ability to analyze documents without external exposure is critical. Cloud-based AI services often require data to leave the corporate network, creating compliance hurdles for industries like healthcare, finance, and legal services. A local-first approach ensures that proprietary information remains on-premises or within a private cloud environment. This setup gives IT leaders full control over data retention, access logs, and security policies, aligning AI adoption with strict internal governance standards.
Additionally, local execution offers predictable performance and cost structures. Unlike cloud APIs that charge per token or request, local models have a fixed upfront cost in hardware and electricity. Once the model is downloaded, there are no recurring fees for querying the knowledge graph. This makes it easier for small and mid-sized companies to budget for AI capabilities without worrying about unexpected spikes in usage bills. It also reduces dependency on external vendors, ensuring that business operations continue even if third-party services experience outages.
What you can do
- Evaluate your current hardware to ensure it has enough VRAM to run a 30B-parameter model locally, or plan to use a smaller model for initial testing.
- Install Docker and Ollama on a test machine to replicate the workflow described, starting with a small non-sensitive document.
- Configure network settings carefully, especially if using Linux Docker Engine, to ensure secure communication between containers and host services.
- Establish a protocol for verifying AI-generated citations by randomly sampling answers and checking the linked source chunks for accuracy.
- Consider hybrid approaches where sensitive data stays local while less critical tasks might leverage cloud resources for speed or quality, if policy allows.
- Monitor resource usage during the extraction phase to optimize VRAM presets and adjust batch sizes for better performance on your specific hardware.



