LiteLLM
Backed by 56,000+ GitHub stars and over 240 million Docker pulls, LiteLLM delivers the open-source AI gateway trusted by Netflix, Lemonade, Rocket Money, and thousands of engineering teams to route every LLM request through one unified API. The Rust-core gateway adds sub-millisecond overhead per request with 8ms P95 latency at 1,000 RPS, 15x throughput improvement and 11x lower memory footprint compared to Python-only proxies. A single OpenAI-compatible endpoint connects to 100+ providers and 1,800+ models spanning OpenAI, Anthropic, Google Gemini, AWS Bedrock, Azure OpenAI, Vertex AI, Hugging Face, vLLM, Nvidia NIM, Ollama, and Mistral with day-zero support for new model releases. The Auto Router V2 classifies request complexity across four tiers using rule-based scoring, semantic keyword matching, and adaptive Thompson sampling to route each request to the most cost-effective model without API calls or training data. Virtual API keys enable multi-tenant governance with per-team, per-user, and per-project cost tracking, budget caps with automatic fallback rerouting, and role-based access control. Built-in guardrails provide PII masking, prompt injection detection, and model-graded evaluation before requests reach providers. The Agent Gateway extends routing from model calls to agent workflows with MCP server integration. Observability integrates with Langfuse, Arize Phoenix, OpenTelemetry, and MLflow for complete request tracing. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Open Connector
With over 5,000 GitHub stars since its June 2026 launch, OOMOL OpenConnector bridges the gap between AI agents and the real world by handling the authentication nightmare that stops LLMs from calling external APIs safely. The runtime connects to more than 1,000 SaaS providers — GitHub, Gmail, Notion, Slack, Microsoft, HubSpot, Google Workspace, and hundreds more — through 10,000+ prebuilt typed Actions that agents can discover and execute without ever touching raw credentials. OAuth2 flows, API key rotation, custom credentials, and no-auth providers are all managed centrally with AES-encrypted storage, scoped runtime tokens, and configurable action allowlists and blocklists that enforce least-privilege access. Agents interact through five access surfaces: the Model Context Protocol endpoint at /mcp for Claude and other MCP-capable hosts, a full REST API at /v1 for programmatic control, an auto-generated OpenAPI specification for code generation, a TypeScript SDK for application integration, and the oo CLI for local agent relay. The built-in Web Console provides browser-based administration for configuring OAuth apps, managing connections, inspecting action schemas, and reviewing execution logs with redacted inputs. Deploy via Docker Compose with SQLite for single-server setups, run from source on Node.js 22+, or push to Cloudflare Workers with D1 and R2 for edge deployment. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
Rivet
Stateful serverless actors that run indefinitely, sleep when idle, and persist state across restarts without external database round trips. Rivet provides a vendor-neutral alternative to Cloudflare Durable Objects, delivering long-running processes with co-located in-memory state and per-actor SQLite databases on any infrastructure you choose. The Rust engine comprises four packages: Pegboard for actor orchestration, Gasoline for durable execution, Guard for traffic routing via Envoy, and Epoxy implementing multi-region KV storage through EPaxos consensus. Each actor maintains instant-access in-memory state plus a dedicated SQLite database for relational queries, backed by RocksDB on single nodes or PostgreSQL with NATS pub/sub for multi-node clusters. RivetKit SDKs in TypeScript, Rust, and Python support built-in WebSocket connections, task queues, scheduling, and CRDT-based real-time collaboration. The v2.3 rewrite moved the core runtime from JavaScript to native Rust via WebAssembly, eliminating main-thread blocking across Node.js, Bun, Deno, and Cloudflare Workers. A built-in dashboard provides actor inspection with real-time Prometheus metrics. Deploys as a single Docker container on port 6420 with optional PostgreSQL for persistence. Available on RepoCloud with a dedicated VPS under the Apache 2.0 license.
Knowhere
With 2,600+ GitHub stars since its May 2026 open-source launch, Knowhere solves the last-mile problem of document intelligence for AI systems — transforming complex unstructured PDFs, reports, and multi-page documents into structured JSON chunks that LLMs can consume without hallucination. The platform processes documents through an AI-native parsing pipeline that handles 20+ page documents with deep hierarchies, intricate tables, and multimodal content including images with OCR, achieving 95% precision in information extraction while reducing token costs by 50% compared to raw document ingestion. The knowledge tree architecture maintains historical context across multiple documents, enabling cross-document graph navigation for agentic retrieval that goes beyond simple chunk-based RAG. Built on Python 3.11+ with MinerU as the default PDF parser, the backend API runs alongside async workers that process document ingestion, graph construction, and embedding generation. The self-hosted Docker Compose stack packages the API server, processing workers, and Next.js dashboard for managing API keys, webhooks, and document-processing jobs, backed by PostgreSQL and Redis. Both Python and Node.js SDKs provide programmatic access for integration into existing AI pipelines and agent frameworks. LLM providers include DeepSeek and Alibaba Cloud DashScope with configurable key rotation for rate-limit management. Deploy on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
MetaMCP
With 2,600+ GitHub stars, MetaMCP solves the MCP server sprawl problem by aggregating any number of upstream servers into a single authenticated endpoint that any MCP client connects to once. Group servers into namespaces — development tools in one, data sources in another — then publish each namespace as its own SSE, Streamable HTTP, or OpenAPI endpoint with API-key authentication in headers or query parameters, or full OAuth per the MCP Spec 2025-06-18 standard. The aggregation engine discovers tools, resources, and prompts from all active servers in parallel, prefixes tool names with server identifiers to prevent collisions, and applies configurable middleware including tool filtering to reduce context-window bloat and description overrides to improve LLM comprehension. The web management UI lets you configure MCP servers with stdio, SSE, or Streamable HTTP transports, toggle servers active or inactive per namespace, create and revoke API keys per endpoint, and inspect discovered tools with their schemas. Nested MetaMCP support enables hierarchical architectures where one MetaMCP instance consumes another, creating multi-level tool organization with automatic name resolution. Compatible with Claude Desktop, Claude Code, Cursor, Open WebUI, and any MCP-compatible client through a single connection URL. The Docker container packages the TypeScript backend with PostgreSQL for configuration persistence, exposing the management UI on port 12005 and MCP endpoints on configurable ports. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
InsForge
With 12,600 GitHub stars and 52 releases in under a year of development, InsForge is the fastest-growing open-source backend platform purpose-built for AI coding agents — giving Claude, Cursor, and any MCP-compatible tool direct access to database, authentication, storage, compute, and AI model infrastructure through a single self-hosted stack. The native MCP server exposes every InsForge operation as callable tools, letting coding agents autonomously create database tables, manage user authentication, upload files, deploy edge functions, and ship complete full-stack applications without human intervention. The Model Gateway provides an OpenAI-compatible API that routes requests across multiple LLM providers (OpenAI, Anthropic, Google, and open-source models) with unified billing, rate limiting, and fallback logic. PostgreSQL with pgvector handles both relational data and vector embeddings for RAG pipelines, while S3-compatible storage manages file uploads and static assets. Edge Functions run serverless TypeScript code on Deno with sub-millisecond cold starts for API endpoints, webhooks, and scheduled tasks. The authentication system provides user management, OAuth2 flows, sessions, and magic links with JWT token handling built in. Site Deployment builds and serves frontend applications with automatic SSL and custom domain configuration. The CLI paired with Agent Skills enables terminal-based workflows where agents invoke InsForge operations directly from the command line. Deploy via Docker with PostgreSQL as the only required external dependency. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache-2.0 licensed.
Arkon
With 1,200+ GitHub stars since its April 2026 launch, Arkon provides an enterprise-grade knowledge management layer that turns scattered organizational documentation into AI-accessible structured context. The platform runs as a centralized MCP server, compiling your SOPs, policies, technical docs, and institutional knowledge into a versioned wiki with draft-approval workflows, then serving that wiki to Claude Desktop, Claude.ai, Cursor, and any MCP-compatible client through a single permission-scoped endpoint. OAuth 2.1 with PKCE authentication eliminates manual token management — employees authenticate through a browser login while the system discovers endpoints automatically via RFC 8414. The RBAC v2 system supports custom roles with granular permissions, department-scoped AI Skills, workspace isolation, and comprehensive audit logging so every query and access event is traceable. RAG retrieval powered by pgvector embeddings enables AI clients to search across all organizational documents with source attribution, while the AI Skills system lets teams define reusable instruction sets scoped to specific departments or roles. The architecture runs seven Docker containers coordinated by Compose: PostgreSQL with pgvector for embeddings and metadata, Redis for caching, MinIO for document storage, a FastAPI backend, two ARQ async workers for embedding generation and document processing, and a Next.js frontend portal accessible on port 3119. API keys are encrypted at rest with Fernet, and no telemetry leaves the deployment. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. PolyForm Internal Use licensed.
OpenLLM
OpenLLM serves any large language model as an OpenAI-compatible API endpoint from a single CLI command, handling model download, backend selection, quantization, and port binding automatically. It supports the full spectrum of popular models including Llama 3.3, Qwen2.5, DeepSeek, Mistral, and Phi3, choosing between vLLM and PyTorch inference backends based on hardware capabilities. When vLLM is available, continuous batching with PagedAttention achieves up to 23x throughput improvement over naive serving, while GPTQ and bitsandbytes quantization reduces memory requirements for GPU-constrained deployments. The server exposes a RESTful API on port 3000 with full OpenAI client library compatibility, enabling drop-in replacement for commercial providers in any application using the standard chat completions format. A built-in web chat UI at the /chat endpoint provides immediate interactive testing without external clients. Custom model repositories allow teams to maintain private catalogs of fine-tuned models alongside the default repository that tracks the latest releases. Deployment workflows generate production-ready Docker images automatically, with Kubernetes manifest support for orchestrated scaling. Native integration with LangChain and LlamaIndex supports RAG pipelines, Transformers Agents enables tool-calling workflows, and HuggingFace Hub handles model discovery. Server-Sent Events enable real-time token streaming across all API endpoints. Backed by BentoML's production ML infrastructure. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.