New open-source tool turns local books into structured AI courses
A new self-hosted application called Tutor converts EPUB and PDF files into interactive study plans using large language models.
A developer has released an open-source project named Tutor that transforms personal digital libraries into structured, interactive courses. Published on GitHub by Demetrius Edelin on October 10, 2026, the tool runs locally on a user’s computer and leverages external large language model APIs to teach concepts from uploaded books.
What happened
The project, currently at version 0.1.0, allows users to create subjects of study, such as SQL or Rust, and populate them with EPUB files, tagged PDFs, or online documentation. Unlike standard chatbots that answer ad-hoc questions, this application analyzes the source material to extract key concepts, bold terms, and italicized definitions. It then organizes these elements into a concept map, creating a linear learning path rather than a free-form conversation.
The system operates through a command-line interface for setup and data ingestion, followed by a browser-based interface for the actual studying. Users must provide their own API keys for supported providers, including Anthropic, OpenAI, or OpenRouter. The application does not include a built-in model, ensuring that all processing costs and data privacy controls remain with the user. The developer notes that the tool is designed for single-user, single-computer deployment, though container support is planned for future releases.
Key details
- Supported formats: The parser handles EPUB files optimally, utilizing publisher CSS for structure. It also supports tagged PDFs but will reject untagged or scanned PDFs unless the LLM can process images.
- Model compatibility: Users can select models from Anthropic, OpenAI, or OpenRouter. Recent testing indicates Claude Haiku 5.5 and GPT-6 Luna perform well, with reasoning levels adjustable via configuration.
- Ingestion process: The
ingestcommand scans chapters and images, caching results to avoid repeated API charges. Users can ingest entire books or specific chapters using flags like--chapters. - Study workflow: The app features a diagnosis phase to test existing knowledge, a study queue for new concepts, and lessons that reference specific book sections. It includes a dispute mechanism for incorrect AI grading.
- Technical stack: Built with TypeScript on Node.js 22, using Fastify for the server, React 19 for the frontend, and SQLite for local data storage.
- Privacy controls: Book text and API keys are stored in a local
data/folder excluded from Git. The online fetcher respectsrobots.txtand adds delays between requests.
Background
This tool represents a shift from retrieval-augmented generation (RAG) systems that simply search documents toward agentic systems that structure pedagogy. Traditional RAG implementations allow users to ask questions about a document corpus, returning relevant snippets. In contrast, this application performs semantic analysis to identify educational units—concepts—and tracks the user’s mastery status over time. It relies on the structural metadata of digital books, such as heading hierarchies and typographic emphasis, to determine what constitutes a learnable unit. By running locally, it avoids sending entire book contents to third-party servers continuously, sending only specific chunks during the lesson generation phase.
Why it matters
For teams and individuals managing proprietary documentation or specialized technical libraries, this approach offers a way to institutionalize knowledge without relying on public cloud training data. Self-hosting ensures that sensitive internal manuals or purchased technical references remain under local control. The ability to convert static PDFs and EPUBs into interactive curricula reduces the friction of onboarding new team members to complex codebases or regulatory frameworks. It transforms passive reading into active verification, ensuring that learners actually understand the material rather than just skimming it.
Furthermore, the bring-your-own-key model aligns with cost-conscious engineering practices. Organizations can monitor API usage precisely, caching ingested content to prevent redundant processing fees. This contrasts with subscription-based learning platforms where content is fixed and pricing is opaque. For IT leads, the requirement for Node.js 22 and local SQLite databases means the tool can be deployed on existing development workstations without heavy infrastructure overhead, although the lack of multi-user support currently limits its use in collaborative team environments.
What you can do
- Install Node.js 22 or later and clone the repository from GitHub to test the ingestion pipeline with a sample EPUB file.
- Configure a
.envfile with API keys from Anthropic, OpenAI, or OpenRouter, selecting a model optimized for reasoning tasks. - Use the
npm run parsecommand to preview chapter structures and concept extraction before committing to full database ingestion. - Start with tagged PDFs or native EPUBs to ensure accurate concept mapping, avoiding scanned documents that lack text layers.
- Monitor API costs by using the
--previewflag during initial setup to estimate token usage for large textbooks or documentation sets. - Provide feedback on the GitHub repository regarding the planned container image feature if server-based deployment is required for your workflow.



