Onyx
Formerly known as Danswer and now backed by over 31,000 GitHub stars with 253 releases, Onyx delivers a production-ready AI platform that turns any LLM into a context-aware enterprise assistant connected to your organization's actual knowledge. The agentic RAG pipeline combines BM-25 keyword search with prefix-aware embedding models in a hybrid index, then deploys AI agents to retrieve, verify, and synthesize answers with source citations from over 40 connected workplace tools including Google Drive, Confluence, Slack, Notion, Jira, SharePoint, GitHub, and Linear. Custom AI assistants with configurable prompts, backing knowledge sets, and document-level access control enable specialized agents for engineering, sales, support, and research workflows. The platform supports every major LLM provider — Anthropic Claude, OpenAI, Google Gemini, plus self-hosted options via Ollama, LiteLLM, and vLLM for fully air-gapped deployments. Beyond chat, Onyx provides web search with Serper, Google PSE, Brave, and SearXNG integration, an in-house web crawler, code execution, file creation, and multi-step deep research with report generation. Enterprise features include SSO via Google OAuth, OIDC, or SAML with SCIM provisioning, role-based access control, usage analytics by team and agent, query history auditing, PII removal through custom code hooks, and full whitelabeling. Deploy via Docker Compose on any infrastructure. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed (Community Edition).
Fess
Eliminate indexing blind spots across fragmented company networks with Fess, a distributed enterprise search engine that crawls internal document repositories, corporate intranets, and cloud services to provide employees with instant, permission-aware file discovery. Organizations can schedule automated crawlers across web pages, network shares, databases, and third-party storage platforms to extract text and generate visual thumbnails from PDFs, spreadsheets, presentations, and compressed archives. Users can search indexed records using natural language queries, filtering results by file metadata, creation dates, categories, and geographic coordinates while viewing contextual snippets with highlighted keyword matches. Granular role-based permissions mirror directory security structures from Active Directory and LDAP so staff only see search results for documents they have explicit authorization to inspect. The browser-based administrative console lets operations teams define path mapping rules, manage virtual hosts, monitor active crawl jobs, inspect error logs, and configure scheduled re-indexing workflows without modifying server configuration files. Security teams can establish single sign-on across the organization using SAML or OpenID Connect authentication providers to secure administrative workflows and search portals. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
PipesHub
Unify fragmented corporate knowledge across chat channels, documents, and cloud drives with PipesHub, an open-source workplace context platform providing permission-aware search and verifiable answers for teams and autonomous agents. Employees can query enterprise knowledge bases through a conversational interface that retrieves relevant information across Google Workspace, Microsoft 365, Slack, Confluence, Jira, and GitHub while strictly respecting individual user access permissions. Knowledge workers can verify every answer through clickable citation blocks that trace claims directly back to source documents, spreadsheet rows, or chat messages. Teams can assemble custom AI agents visually using a drag-and-drop builder, combining retrieval collections with interactive toolsets to automate multi-step operations like drafting customer responses or creating issue tickets. Autonomous agents connect via Model Context Protocol to inspect shared company knowledge and perform external actions across connected SaaS tools. Organizations can index scanned PDFs, slide decks, markdown pages, and spreadsheets with OCR and document parsing, maintaining synchronized records through automated schedules. Administrators can inspect synchronized records, monitor query history, and manage connector credentials across all integrated business applications. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache License 2.0 licensed.
Omni
Omni is an open-source workplace AI agent and enterprise search platform that unites your company's scattered tools into a single context-aware workspace where team members can find answers and automate cross-application workflows. Employees can query connected knowledge sources including Google Drive, Gmail, Slack, Jira, GitHub, Confluence, and HubSpot to locate project specifications, customer records, and operational documents without jumping across tabs. The integrated AI assistant answers complex multi-step inquiries, summarizes lengthy email threads, and cites exact source passages so workers can verify factual accuracy before making business decisions. Granular permission inheritance mirrors existing access controls across every connected service, guaranteeing that sensitive personnel files, financial statements, and confidential repositories remain hidden from unauthorized staff. Engineers can construct bespoke data connectors and define automated agent skills through the extensible Python and TypeScript software development kits, triggering multi-service actions directly from interactive chat dialogues. Administrators can monitor background synchronization jobs, inspect audit trails, manage single sign-on access, and configure custom extraction pipelines using centralized operational dashboards. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.