AnythingLLM
Chat with your own documents: AnythingLLM, from Mintplex Labs, wraps retrieval-augmented generation (RAG) in an open-source application anyone can run. You organize content into workspaces, each an isolated namespace with its own documents, vector embeddings, chat history, and settings, so one instance can hold several separate knowledge bases. Upload PDFs, DOCX, TXT, and other formats, or scrape web pages; the built-in collector parses and chunks them into a vector database (LanceDB by default, with Pinecone, Chroma, Qdrant, and others supported). Answers cite their source documents. It works with both cloud LLMs (OpenAI, Anthropic, Gemini) and local ones via Ollama or LM Studio, and the embedding model is separately configurable. Beyond RAG chat, it includes AI agents that can browse the web and run tools, an embeddable chat widget for your website, a developer API, and multi-user mode with admin, manager, and default roles plus per-workspace access control. Context assembly is smarter than naive RAG: pinned documents, attached files, vector search hits, and recent chat history are combined under a token budget so the model's context window is filled efficiently, and each workspace supports multiple independent conversation threads against the same knowledge base. Because the embedding model, vector store, and chat LLM are all independently swappable, you can move between providers without re-ingesting a single document. The stack is Node.js with a React frontend, MIT-licensed.
Open Canvas
Open-source alternative to OpenAI's Canvas — a collaborative writing and coding environment where AI agents help you draft, edit, and refine documents through an agentic architecture built on LangGraph. The dual-mode editor combines a BlockNote rich text editor for live-rendered markdown with a CodeMirror-based code editor supporting syntax highlighting across multiple programming languages, letting you switch between prose and code artifacts within the same session. The built-in reflection agent automatically generates style rules and user insights from your chat history, storing them in a shared LangGraph memory store that persists across sessions for increasingly personalized assistance. Pre-built quick actions provide one-click access to common writing transformations including summarize, expand, simplify, and translate, while coding actions offer explain, refactor, add comments, and convert between languages. The monorepo architecture separates the Next.js 14 frontend from the LangGraph agent backend, connecting via HTTP and WebSocket protocols through the @langchain/langgraph-sdk client. Seven LLM providers are supported out of the box — OpenAI, Anthropic Claude, Google Gemini, Fireworks AI, Groq, Azure OpenAI, and local Ollama models — with Supabase handling authentication and data persistence. Deploy via Docker or build from source. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
HolaOS
With over 6,500 GitHub stars, HolaOS bills itself as an "open agent computer" that reimagines the traditional operating system as a shared workspace where humans and AI agents collaborate across files, browsers, and 100+ integrated tools simultaneously. Unlike chat-only interfaces, HolaOS places live application UIs—Notion-style editors, browsers, custom workspace apps—side by side with the agent conversation, so operators always see what agents are doing and can intervene at any moment. The persistent memory system stores workspace knowledge locally as Markdown files and embedded vectors via SQLite vec, enabling RAG-powered recall that survives session boundaries without the typical context window bloat. Safe Session Compaction reserves roughly 70% of the model context window for fresh reasoning while folding older history into structured checkpoints that retain goals, constraints, progress, and decisions. Agents connect to Linear, GitHub, Slack, Jira, HubSpot, Gmail, and dozens more through one-click OAuth, automatically fetching relevant signals and converting scattered app data into working memory. BYOK support for Claude, GPT, and Gemini models lets operators use their own API keys at zero markup, while built-in Kimi K3 and GLM-5.2 models provide ready-to-use alternatives. Skills package reusable workflows that any agent can invoke on demand, and scheduled triggers enable autonomous digests, monitors, and reports. The runtime supports independent server deployment alongside the desktop client. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Modified Apache 2.0 licensed.
WeKnora
WeKnora turns scattered corporate documents into a searchable, reasoning-capable knowledge asset that your team can query in plain language and receive cited, sourced answers. Upload PDFs, Word files, web pages, Feishu wikis, Notion databases, Yuque docs, GitLab repositories, or RSS feeds into structured knowledge bases, and three distinct modes make the content actionable: RAG Quick Q&A retrieves relevant chunks and generates answers with source citations; the ReAct Agent autonomously orchestrates multi-step reasoning across knowledge retrieval, MCP tool calls, web search, and sandboxed code execution to produce comprehensive research reports; and Wiki Mode deploys LLM agents to distill raw documents into an interlinked markdown knowledge base with an interactive knowledge graph, revision history, and one-click rollback. Connect 20+ LLM providers including OpenAI, DeepSeek, Qwen, Claude, and local Ollama models without vendor lock-in, and choose from seven vector database backends (Qdrant, Milvus, Weaviate, and more) for embedding storage. Enterprise features include four-tier RBAC with per-resource ownership and per-workspace audit logs, AES-256-GCM credential encryption, scoped API keys, Langfuse observability tracing for every agent loop and tool call, and a runtime task-queue dashboard for worker-pool governance. Cross-session long-term memory preserves conversational context across interactions. The Agent Skills catalog lets teams install and share sandboxed scripts executed in Docker or E2B containers. A Chrome Extension captures web content directly into knowledge bases. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Arkon
With 1,200+ GitHub stars since its April 2026 launch, Arkon provides an enterprise-grade knowledge management layer that turns scattered organizational documentation into AI-accessible structured context. The platform runs as a centralized MCP server, compiling your SOPs, policies, technical docs, and institutional knowledge into a versioned wiki with draft-approval workflows, then serving that wiki to Claude Desktop, Claude.ai, Cursor, and any MCP-compatible client through a single permission-scoped endpoint. OAuth 2.1 with PKCE authentication eliminates manual token management — employees authenticate through a browser login while the system discovers endpoints automatically via RFC 8414. The RBAC v2 system supports custom roles with granular permissions, department-scoped AI Skills, workspace isolation, and comprehensive audit logging so every query and access event is traceable. RAG retrieval powered by pgvector embeddings enables AI clients to search across all organizational documents with source attribution, while the AI Skills system lets teams define reusable instruction sets scoped to specific departments or roles. The architecture runs seven Docker containers coordinated by Compose: PostgreSQL with pgvector for embeddings and metadata, Redis for caching, MinIO for document storage, a FastAPI backend, two ARQ async workers for embedding generation and document processing, and a Next.js frontend portal accessible on port 3119. API keys are encrypted at rest with Fernet, and no telemetry leaves the deployment. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. PolyForm Internal Use licensed.
Eclaire
Built to consolidate fragmented digital life into a private personal cloud, Eclaire centralizes tasks, notes, documents, photos, and web bookmarks into a unified workspace powered by local artificial intelligence. Users can save web articles to generate clean Markdown copies and searchable PDF archives, automatically bypassing clutter and dead links. The integrated document ingestion engine extracts tabular data from spreadsheets, performs optical character recognition across uploaded receipts and handwritten images, and indexes multi-page PDF manuals for rapid full-text search. Conversational assistant panels allow operators to interrogate their accumulated library using natural language, asking for cross-document summaries, research timelines, or action items extracted from meeting transcripts. Background workers prioritize asynchronous indexing queues, categorizing image libraries with visual object tags and managing recurring task deadlines with automated Telegram alert notifications. External automation flows can interact directly with the unified storage layer via an OpenAI-compatible REST API, allowing mobile shortcuts and custom scripts to ingest bookmarks and dictate notes remotely. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Utopia
The first open-source substrate for enterprise knowledge engineering that learns passively and governs itself. The Rust-built backend paired with PostgreSQL and pgvector delivers a bitemporal knowledge graph where every fact carries two timelines: when it held in the real world and when the system came to believe it — enabling full audit trail replay of how understanding evolved. Document ingestion handles PDF, DOCX, PPTX, XLSX, CSV, Markdown, HTML, and plain text with legacy encoding detection, while scheduled syncing pulls from web pages, RSS feeds, GitHub, Jira, Notion, WebDAV, and S3-compatible buckets. Search fuses Tantivy full-text indexing with pgvector semantic vectors using Reciprocal Rank Fusion, streaming answers with inline citations that link directly to source passages. The built-in agent harness drives agentic RAG through conversation — searching documents, walking the knowledge graph at any historical date, and querying mounted databases via Ontology2SQL which achieves state-of-the-art results on BIRD Mini-Dev benchmarks. Five ontology packs ship inside the binary (schema.org, W3C Org, PROV-O, FOAF, IOF Core) with forward-chaining reasoning for transitivity, symmetry, inverses, and relation hierarchy. Entity resolution operates in three stages: exact name matching, embedding similarity, then model-based judgment with every merge reversible. Any OpenAI-compatible endpoint works including DeepSeek, Qwen, Ollama, and vLLM for fully air-gapped deployment. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache-2.0 licensed.
Khoj
A self-hosted "second brain": Khoj indexes your own files and answers questions from them, parsing Markdown (whole Obsidian vaults included), org-mode, PDF, Word, plain text, Notion pages, GitHub repositories, and images described by a vision model, then embedding everything with sentence-transformers into a vector index for semantic search and RAG with cited sources. Any LLM backend works: local models like Llama, Qwen, or Mistral via Ollama, or cloud models like GPT, Claude, and Gemini. You can build custom agents, each with its own persona, scoped knowledge base, chat model, and tools such as web search and code execution. Scheduled automations run recurring research and deliver newsletters or notifications to your inbox, and research mode performs multi-hop web searches with inline citations. Access it from a browser, the Obsidian plugin, Emacs, desktop, or WhatsApp - all clients connect to the same self-hosted instance, making Khoj one of the few AI assistants Emacs users can point at decades of org files. Semantic search means recall works without exact keywords: "that paper about forecasting with transformers" surfaces the right PDF even when you cannot remember its title. Switching LLM backends never requires re-indexing your documents, and with a local model via Ollama, even inference stays on hardware you control - journals, research, and private notes are never sent anywhere. Python/FastAPI stack, AGPL-licensed, with PostgreSQL storage.
Open Notebook
The most feature-complete open-source alternative to Google's NotebookLM — a self-hosted research platform where you upload PDFs, videos, audio files, and web pages into organized notebooks, then chat with your content, generate multi-speaker podcasts, and run semantic search across everything without sending a single byte to Google's servers. The podcast engine supports 1-4 fully customizable speakers with backstories, personalities, and expertise profiles, generating professional audio dialogue through OpenAI, ElevenLabs, Google TTS, or completely local text-to-speech via Kokoro for maximum privacy. Content processing uses token-based chunking with RAG-powered retrieval grounded in your uploaded sources, while both full-text keyword search and semantic vector search via SurrealDB enable conceptual discovery across all notebooks. The 18+ supported AI providers include OpenAI, Anthropic, Google Gemini, Groq, Ollama, LM Studio, and more — configurable per task so you can route cheap models to summarization and powerful models to analysis. Content transformations extract insights, generate summaries, create study guides, and produce structured outputs from any source material. The MCP integration connects Open Notebook to Claude Desktop, VS Code, and other MCP clients for seamless workflow integration. A full REST API on port 5055 enables complete automation of notebook management, source upload, and podcast generation. Deploy via Docker Compose with the application container, SurrealDB v2 on RocksDB, and optional TTS containers. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.