LiteLLM screenshot thumbnail

LiteLLM

Backed by 56,000+ GitHub stars and over 240 million Docker pulls, LiteLLM delivers the open-source AI gateway trusted by Netflix, Lemonade, Rocket Money, and thousands of engineering teams to route every LLM request through one unified API. The Rust-core gateway adds sub-millisecond overhead per request with 8ms P95 latency at 1,000 RPS, 15x throughput improvement and 11x lower memory footprint compared to Python-only proxies. A single OpenAI-compatible endpoint connects to 100+ providers and 1,800+ models spanning OpenAI, Anthropic, Google Gemini, AWS Bedrock, Azure OpenAI, Vertex AI, Hugging Face, vLLM, Nvidia NIM, Ollama, and Mistral with day-zero support for new model releases. The Auto Router V2 classifies request complexity across four tiers using rule-based scoring, semantic keyword matching, and adaptive Thompson sampling to route each request to the most cost-effective model without API calls or training data. Virtual API keys enable multi-tenant governance with per-team, per-user, and per-project cost tracking, budget caps with automatic fallback rerouting, and role-based access control. Built-in guardrails provide PII masking, prompt injection detection, and model-graded evaluation before requests reach providers. The Agent Gateway extends routing from model calls to agent workflows with MCP server integration. Observability integrates with Langfuse, Arize Phoenix, OpenTelemetry, and MLflow for complete request tracing. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
Agent Gateway screenshot thumbnail

Agent Gateway

Backed by the Linux Foundation with contributions from AWS, Cisco, IBM, Microsoft, Red Hat, and Shell, Agentgateway is the first data plane built from the ground up for AI agent workloads — providing a unified Rust-based proxy that handles conventional HTTP and gRPC traffic alongside MCP tool servers, A2A agent communication, and LLM inference endpoints through a single deployment. The LLM gateway routes requests to OpenAI, Anthropic, Gemini, AWS Bedrock, and other providers through an OpenAI-compatible unified API with per-tenant budget controls, spend tracking, prompt enrichment, load balancing across multiple model endpoints, and automatic failover when providers experience outages. The MCP gateway federates multiple tool servers behind one endpoint, supporting stdio, HTTP/SSE, and Streamable HTTP transports with built-in OAuth authentication compliant with the MCP auth specification, integrating Auth0 and Keycloak out of the box. OpenAPI integration exposes existing REST APIs as MCP-native tools without code changes, enabling legacy services to participate in agent workflows. Policy-based RBAC controls which agents access which tools, while OpenTelemetry integration provides distributed tracing across agent communication chains. Deploy as a standalone binary with flat YAML configuration or on Kubernetes using the built-in controller with Gateway API support for declarative infrastructure-as-code management. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
BitRouter screenshot thumbnail

BitRouter

BitRouter is a context-aware LLM router that learns which model delivers the cheapest successful outcome per workflow step, cutting agent costs by up to 80% while maintaining 96% quality versus all-frontier baselines. Point any agent runtime at http://localhost:4356 with a one-line OPENAI_BASE_URL change and BitRouter routes to OpenAI, Anthropic, Google, Groq, DeepSeek, Mistral, Moonshot, MiniMax, Nvidia, and any OpenAI-compatible endpoint simultaneously, normalizing authentication, streaming, and cross-protocol translation between wire formats. The act-observe-evaluate-learn loop traces every hop with cost, tokens, and latency attribution, scores each decision against a versioned policy-lock.yaml, then tightens routes automatically with no LLM judge in the path. Native MCP gateway auto-discovers tools from connected servers and makes them routable and governed alongside model calls. Agent Client Protocol integration enables the TUI to manage Claude Code, Codex, OpenCode, OpenClaw, Gemini, and Copilot sessions in real time with inline tool-call approval and live streaming. Built-in guardrails inspect, redact, or block risky content before requests leave your network. Virtual keys scope API access per agent or user without exposing upstream credentials. Per-agent spend caps and loop guards contain runaway cost automatically. Multi-account failover reroutes mid-run so rate limits never re-pay completed work. Ships as a single Rust binary via npm or Cargo. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
CPA Manager Plus screenshot thumbnail

CPA Manager Plus

CPA Manager Plus is a self-hosted observability dashboard and management panel that tracks every AI request flowing through your CLI Proxy API gateway, breaking down failures, costs, and account health across providers like OpenAI, Anthropic, xAI, and Codex in one interface. When a request fails, drill into the persistent history to see status codes, affected models, latency, and redacted failure evidence without exposing raw response bodies. The cost analytics engine breaks down token consumption and estimated spend by model, provider, account, API key, project, channel, and time range while tracking input, output, reasoning, cache, and service-tier pricing semantics separately. Model prices sync automatically from models.dev with LiteLLM and OpenRouter fallbacks, and you can add local overrides for aliases or internal models. For teams running Codex or xAI accounts, the health inspector reads quota windows, reset evidence, credential state, and workspace status on a configurable schedule, routing credential failures into an action queue for review rather than letting them silently degrade throughput. Deploy the Lightweight Panel to replace your existing CPA management UI without adding another service, or run Full Mode as a single Docker container that adds the Manager Server with persistent SQLite storage for request history, historical analytics, and automated account inspections. Export or import request history as JSONL for external analysis, and back up the SQLite files alongside your encrypted management keys. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
Bifrost screenshot thumbnail

Bifrost

Bifrost is an open-source AI gateway that unifies 23+ LLM providers into a single OpenAI-compatible endpoint with automatic failover, semantic caching, and built-in cost governance, so one provider going down never takes your production AI application with it. Point your existing OpenAI or Anthropic SDK at Bifrost's local endpoint and gain access to OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, Groq, Mistral, and Ollama without changing application code. Define fallback chains that automatically switch providers when one returns errors or exceeds latency thresholds, keeping response times stable during outages. The built-in web dashboard at port 8080 lets you configure providers, create virtual API keys, monitor live request traffic, and review analytics without editing configuration files. Semantic caching combines exact hash matching with vector similarity search via Weaviate, serving cached responses for identical or paraphrased prompts in sub-millisecond time to cut costs on repetitive workloads. The MCP gateway connects AI agents to external tools like filesystems, databases, and web APIs, exposing them to clients such as Claude Desktop and Cursor with per-key allow-lists. Four-tier budget hierarchy at customer, team, virtual key, and provider levels enforces spend caps, rate limits, and model restrictions across your organization. Extend functionality through custom Go plugins for analytics, monitoring, or security middleware. Native Prometheus metrics and OpenTelemetry distributed tracing give operations teams full production observability. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
Apache APISIX screenshot thumbnail

Apache APISIX

With 17,000 GitHub stars, 460+ contributors, and deployments across telecommunications, automotive, and financial services running on over 10,000 CPU cores at the largest known installations, Apache APISIX delivers a fully dynamic API gateway achieving 140,000 QPS on eight cores with sub-millisecond latency through NGINX's event-driven architecture and LuaJIT-compiled plugin execution. The 100+ open-source plugins cover authentication (JWT, OAuth 2.0, OIDC, Keycloak, LDAP), observability (Prometheus, Datadog, SkyWalking, OpenTelemetry), traffic management (rate limiting, circuit breaking, canary releases, traffic splitting), and security (CORS, IP restriction, CSRF protection) — all hot-reloadable without process restarts via etcd-based real-time configuration synchronization. Multi-protocol support handles HTTP, gRPC, MQTT, TCP, UDP, and WebSocket traffic for both north-south API access and east-west service mesh communication. AI gateway capabilities proxy requests to 20+ LLM providers with semantic caching, token-aware rate limiting, provider failover routing, and content moderation. Custom plugins extend the gateway in Lua, Go, Java, Python, or WebAssembly. Radixtree route matching handles 100,000+ routes without performance degradation. Functions as a Kubernetes ingress controller with native service discovery for Consul, Nacos, and Eureka. Deploy via Docker or Helm charts with horizontal scaling through etcd cluster coordination. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
Inference Gateway screenshot thumbnail

Inference Gateway

Inference Gateway puts a single OpenAI-compatible API endpoint in front of OpenAI, Anthropic, Groq, Cohere, Ollama, DeepSeek, Google, Mistral, MiniMax, Moonshot, Nvidia, and llama.cpp, so your application code never changes when you switch models or providers. The Go binary starts on port 8080 and normalizes authentication, streaming protocols, and response formats across all backends transparently. Native Model Context Protocol support auto-discovers tools from connected MCP servers and injects them into LLM requests without client-side management, enabling server-side tool execution across any provider that supports function calling. Agent-to-Agent protocol integration allows distributed agent communication through a declarative Agent Definition Language that generates production-ready Go or Rust servers from a single YAML manifest. The dedicated Kubernetes Operator manages Gateway, Agent, MCP, and Orchestrator custom resources with automatic HPA scaling, OIDC authentication, and service discovery that rebuilds MCP configurations when the discovered server set changes. Prometheus metrics and OpenTelemetry tracing provide full request-level observability across the entire inference pipeline. Middleware controls enable per-request provider selection, model routing, and fallback strategies. Official SDKs in Go, Python, TypeScript, and Rust provide typed client interfaces with streaming support. Docker Compose deployment requires only environment variables for API keys. A CNCF Sandbox applicant. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
Kong screenshot thumbnail

Kong

With over 43,000 GitHub stars and adoption by companies including Nasdaq, Samsung, and Expedia, Kong Gateway is the world's most deployed open-source API gateway, processing billions of API requests daily across hybrid-cloud and multi-cloud architectures. Built on the battle-tested NGINX engine with OpenResty's LuaJIT runtime, Kong delivers sub-millisecond proxy latency while supporting REST, gRPC, GraphQL, WebSocket, SOAP, and Kafka protocols. The plugin architecture includes authentication via JWT, Basic Auth, HMAC, key authentication, OAuth 2.0, and LDAP, alongside rate limiting with configurable windows per consumer, IP address, or API key. The AI Proxy plugin provides a universal LLM API that routes across OpenAI, Anthropic, Google Gemini, AWS Bedrock, Azure AI, Databricks, Mistral, and Hugging Face through a single standardized interface, while MCP proxy capabilities convert REST APIs into MCP tools and provide traffic governance for AI agents. Kong supports declarative configuration via YAML for GitOps workflows, a RESTful Admin API for dynamic configuration, and decK CLI for version-controlled infrastructure-as-code management. Upstream health checking with active and passive probes enables automatic failover, and the ring balancer distributes traffic across upstream targets with consistent hashing, round-robin, or least-connections algorithms. The Kong Plugin Hub hosts over 100 community and official plugins covering logging, monitoring, transformation, security, and traffic control. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy