Most self-hosted AI apps run on CPU, new data shows
A review of 161 open-source AI projects reveals that 113 do not require a GPU, changing how teams plan infrastructure.
A recent analysis of 161 open-source artificial intelligence applications reveals that the majority can operate without dedicated graphics hardware. Published on October 9, 2026, by Archestack on DEV Community, the data challenges the common assumption that self-hosting AI requires expensive GPU servers. This finding is particularly relevant for small businesses and homelab enthusiasts managing their own infrastructure.
What happened
Archestack maintains a ranked list of open-source AI tools suitable for self-hosting. For each of the 161 entries, the team examined the project documentation to determine hardware requirements. The results show that 113 of these applications run entirely without a GPU. Another 31 projects list a GPU as optional, meaning they can function on standard central processing units but may perform better with acceleration. Only eight projects strictly require a GPU, while nine README files do not specify any hardware preference.
The analysis highlights a clear architectural pattern in modern AI deployment. The heavy computational lifting is typically isolated to the model server layer. Among the 19 model serving projects reviewed, such as Ollama, LocalAI, llama.cpp, and vLLM, fourteen treat GPU usage as optional. Only five serving projects mandate a graphics card. Outside of model serving, most components like chat user interfaces, retrieval-augmented generation systems, gateways, and observability tools communicate via application programming interfaces. These components are lightweight and can run on virtually any server hardware.
Key details
- 113 out of 161 reviewed open-source AI apps run without a GPU.
- 31 apps list a GPU as optional, while only 8 require one.
- Model serving is the primary component where GPU acceleration matters, with 14 of 19 servers offering optional GPU support.
- Image and video training tools are the other main category requiring GPUs, with three trainers needing a card.
- Voice processing tools are split, with 6 of 11 able to use a GPU if available.
- RAM requirements remain largely undocumented, with only 11 of 161 projects specifying minimum memory needs.
Background
In traditional machine learning workflows, graphics processing units were essential for training models due to their ability to handle parallel mathematical operations efficiently. However, the landscape has shifted toward inference, which is the process of using a trained model to generate predictions or answers. Modern inference engines have become highly optimized for central processing units, allowing smaller organizations to run capable models on existing hardware. This decoupling of the model engine from the user interface allows for flexible deployment strategies.
Self-hosting refers to running software on your own servers rather than relying on third-party cloud providers. This approach gives organizations full control over their data and costs but requires them to manage the underlying infrastructure. Understanding the specific hardware needs of each component in an AI stack is crucial for cost-effective planning. By identifying which parts of the stack actually need specialized hardware, teams can avoid overspending on unnecessary equipment.
Why it matters
For IT leads and developers at small and mid-sized companies, this data significantly lowers the barrier to entry for self-hosted AI. Many teams hesitate to adopt local AI solutions because they believe they need to purchase expensive servers with high-end graphics cards. Knowing that the majority of tools, including user interfaces and management layers, run on standard CPUs allows these teams to utilize existing hardware or cheaper virtual private servers. This reduces both capital expenditure and operational complexity.
The separation of concerns also simplifies scaling. Teams can deploy the heavy model server on a machine with a GPU if needed, while running the rest of the application stack on lightweight, cost-effective instances. This modular approach means that if a team chooses a CPU-only model server, they can still run a full-featured AI platform with chat interfaces, document processing, and agent workflows without performance bottlenecks in the non-computational layers. It empowers organizations to start small and upgrade only the specific components that require more power.
What you can do
- Audit your current AI stack to identify which components are model servers versus interface tools.
- Prioritize model serving engines that list GPU support as optional, such as Ollama or llama.cpp, for CPU-based deployments.
- Allocate budget for RAM upgrades instead of GPUs, as memory is often the limiting factor for CPU inference but is rarely documented.
- Test your chosen applications on existing hardware before purchasing new servers, since 113 of 161 apps should run without modification.
- Monitor voice processing tools closely, as they are more likely to benefit from GPU acceleration than text-based tools.
- Keep an eye on project documentation for updates on RAM requirements, as this data is currently missing for most projects.



