BitRouter
BitRouter is a context-aware LLM router that learns which model delivers the cheapest successful outcome per workflow step, cutting agent costs by up to 80% while maintaining 96% quality versus all-frontier baselines. Point any agent runtime at http://localhost:4356 with a one-line OPENAI_BASE_URL change and BitRouter routes to OpenAI, Anthropic, Google, Groq, DeepSeek, Mistral, Moonshot, MiniMax, Nvidia, and any OpenAI-compatible endpoint simultaneously, normalizing authentication, streaming, and cross-protocol translation between wire formats. The act-observe-evaluate-learn loop traces every hop with cost, tokens, and latency attribution, scores each decision against a versioned policy-lock.yaml, then tightens routes automatically with no LLM judge in the path. Native MCP gateway auto-discovers tools from connected servers and makes them routable and governed alongside model calls. Agent Client Protocol integration enables the TUI to manage Claude Code, Codex, OpenCode, OpenClaw, Gemini, and Copilot sessions in real time with inline tool-call approval and live streaming. Built-in guardrails inspect, redact, or block risky content before requests leave your network. Virtual keys scope API access per agent or user without exposing upstream credentials. Per-agent spend caps and loop guards contain runaway cost automatically. Multi-account failover reroutes mid-run so rate limits never re-pay completed work. Ships as a single Rust binary via npm or Cargo. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
KrakenD
KrakenD processes over 18,000 requests per second on a single instance while consuming under 50MB of RAM at 1,000 concurrent connections, operating as a stateless API gateway that requires no database whatsoever. The Community Edition has earned over 2,600 GitHub stars by outperforming database-dependent alternatives like Kong and Tyk in independent benchmarks. Written entirely in Go, it uses declarative JSON or YAML configuration files that integrate directly into GitOps workflows for version-controlled infrastructure management. The gateway aggregates responses from multiple backend services into a single API call, transforms request and response payloads with field filtering, grouping, and mapping, and applies zero-trust security policies including JWT validation, OAuth 2.0, CORS, HSTS, clickjacking protection, and XSS prevention. Traffic management features include multi-layer rate limiting at both the router and proxy levels, circuit breakers for backend fault isolation, spike arrest policies, and concurrent call support that requests the same data from multiple backends in parallel for improved response times. Telemetry integrates with OpenTelemetry, Prometheus, Grafana dashboards, Datadog, Zipkin, and Jaeger for distributed tracing and metrics collection. The gateway extends through Go plugins, Lua scripting, Martian modifiers, and Google CEL expressions for custom request processing logic. AI workload routing supports OpenAI, Anthropic, Gemini, and other model endpoints with built-in fallback, retries, and load balancing. Deploy via Docker with the devopsfaith/krakend image as a single binary. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
TrailBase
With 5,400+ GitHub stars and sub-millisecond response times that eliminate the need for dedicated caches entirely, TrailBase consolidates your database, API server, auth service, realtime subscriptions, and admin interface into a single Rust binary weighing under 40MB. The platform generates type-safe CRUD APIs automatically from your SQLite schema with configurable access control rules using SQL expressions, while realtime subscriptions push data changes to connected clients via Server-Sent Events. The embedded Wasmtime runtime executes custom server-side logic as WebAssembly components compiled from Rust, JavaScript, Python, or any language targeting WASI, enabling complex business logic without external services. First-class geospatial support through the in-house LiteGIS GEOS extension provides GeoJSON integration, spatial indexing via R-Trees, and query operators including @within, @intersects, and @contains for location-based applications. Client SDKs span JavaScript/TypeScript, Dart/Flutter, Rust, C#/.NET, Swift, Kotlin, Go, and Python — covering mobile, web, desktop, and IoT platforms. The admin dashboard offers visual schema editing, a data browser, Record API configuration, OAuth provider setup, user management, SQL query editor, ERD visualization, and server logs. Experimental PostgreSQL support (v0.28+) allows connecting to existing Postgres instances via connection string. Deploy via a single binary, Docker container, or the one-line install script across Linux, macOS, and Windows. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. OSL-3.0 licensed.
OpenLLM
OpenLLM serves any large language model as an OpenAI-compatible API endpoint from a single CLI command, handling model download, backend selection, quantization, and port binding automatically. It supports the full spectrum of popular models including Llama 3.3, Qwen2.5, DeepSeek, Mistral, and Phi3, choosing between vLLM and PyTorch inference backends based on hardware capabilities. When vLLM is available, continuous batching with PagedAttention achieves up to 23x throughput improvement over naive serving, while GPTQ and bitsandbytes quantization reduces memory requirements for GPU-constrained deployments. The server exposes a RESTful API on port 3000 with full OpenAI client library compatibility, enabling drop-in replacement for commercial providers in any application using the standard chat completions format. A built-in web chat UI at the /chat endpoint provides immediate interactive testing without external clients. Custom model repositories allow teams to maintain private catalogs of fine-tuned models alongside the default repository that tracks the latest releases. Deployment workflows generate production-ready Docker images automatically, with Kubernetes manifest support for orchestrated scaling. Native integration with LangChain and LlamaIndex supports RAG pipelines, Transformers Agents enables tool-calling workflows, and HuggingFace Hub handles model discovery. Server-Sent Events enable real-time token streaming across all API endpoints. Backed by BentoML's production ML infrastructure. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
Bifrost
Bifrost is an open-source AI gateway that unifies 23+ LLM providers into a single OpenAI-compatible endpoint with automatic failover, semantic caching, and built-in cost governance, so one provider going down never takes your production AI application with it. Point your existing OpenAI or Anthropic SDK at Bifrost's local endpoint and gain access to OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, Groq, Mistral, and Ollama without changing application code. Define fallback chains that automatically switch providers when one returns errors or exceeds latency thresholds, keeping response times stable during outages. The built-in web dashboard at port 8080 lets you configure providers, create virtual API keys, monitor live request traffic, and review analytics without editing configuration files. Semantic caching combines exact hash matching with vector similarity search via Weaviate, serving cached responses for identical or paraphrased prompts in sub-millisecond time to cut costs on repetitive workloads. The MCP gateway connects AI agents to external tools like filesystems, databases, and web APIs, exposing them to clients such as Claude Desktop and Cursor with per-key allow-lists. Four-tier budget hierarchy at customer, team, virtual key, and provider levels enforces spend caps, rate limits, and model restrictions across your organization. Extend functionality through custom Go plugins for analytics, monitoring, or security middleware. Native Prometheus metrics and OpenTelemetry distributed tracing give operations teams full production observability. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
Anakin
Backed by Y Combinator and powering scraping infrastructure across 195 countries, Anakin delivers a production-grade web scraping API purpose-built for AI agents and RAG pipelines that need clean, structured data from sites that actively block conventional scrapers. The single Go binary server handles JavaScript-heavy SPAs through its Camoufox anti-detect browser service with automatic fingerprint rotation, while the HTTP-first handler chain tries lightweight extraction before escalating to full browser rendering — keeping response times under 2 seconds for static pages. The built-in React 19 dashboard provides visual scraping with live results, job tracking with status filters, domain configuration management with handler chain CRUD, and proxy performance monitoring via Thompson Sampling scoring. Structured JSON extraction leverages Gemini AI to transform raw HTML into typed schemas without manual selector maintenance. SDKs span Python, TypeScript, Go, .NET, Java, and Ruby, while the MCP server exposes all 21 tools directly to Claude, Cursor, Windsurf, and any Model Context Protocol-compatible agent. The hosted platform extends the open-source engine with AI web search returning full page content with citations, multi-source agentic research across 20+ sources per query, Wire pre-built actions covering 944 websites with 5,201 structured endpoints, persistent browser sessions for authenticated scraping, and website change monitoring with scheduled alerts. Deploy via Docker Compose with three containers or run the binary directly with optional PostgreSQL persistence. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
Steel Browser
With over 7,400 GitHub stars and benchmarked at 0.89 seconds average session lifecycle — 1.7x to 9x faster than competing browser automation platforms — Steel Browser delivers production-grade headless Chrome infrastructure purpose-built for AI agents that need to interact with the modern web. The TypeScript-based server exposes a REST API providing on-demand browser sessions with full CDP (Chrome DevTools Protocol) access, allowing connections from Puppeteer, Playwright, or Selenium through standard WebSocket endpoints without framework lock-in. Each session maintains persistent state including cookies, localStorage, IndexedDB, and authentication credentials across requests, enabling stateful multi-step agent workflows that survive session restarts. Built-in anti-detection includes stealth plugins, browser fingerprint randomization, and configurable user-agent rotation, while the proxy chain manager handles IP rotation through residential, datacenter, or custom proxy pools. CAPTCHA solving integrates natively so agents encounter fewer blocking interrupts during autonomous navigation. The Session Viewer provides real-time WebRTC-streamed visual debugging of live sessions and playback of recorded sessions with full network request logging. Browser Tools APIs convert any page to clean Markdown, readability-optimized text, PDF documents, or high-resolution screenshots with a single API call. The MCP Server integration exposes Steel sessions as tools accessible to Claude, Cursor, and other Model Context Protocol-compatible AI agents. Deploy via Docker with a single container or use Docker Compose for production configurations with automatic resource cleanup and session lifecycle management. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
ZincSearch
ZincSearch runs full-text search as a single Go binary that consumes a fraction of the memory and CPU that Elasticsearch demands while staying API-compatible, earning 17,800+ GitHub stars as a lightweight alternative. The bluge-powered indexing library processes documents through analyzers, tokenizers, and token filters while maintaining Elasticsearch-compatible ingestion APIs for single-record and bulk operations, letting existing pipelines connect with minimal configuration changes. Schema-less document ingestion accepts JSON payloads without predefined mappings, allowing different documents within the same index to carry different field structures while the engine automatically detects and indexes field types. An embedded Vue.js web console provides a browser-based interface for creating indexes, querying with full-text syntax, browsing results with hit highlighting, managing users, and monitoring system status. A dual API architecture exposes native ZincSearch endpoints under /api alongside Elasticsearch-compatible endpoints under /es, supporting boolean operators, wildcards, phrase matching, fuzzy search, date ranges, and aggregation pipelines including terms, histogram, date histogram, and range aggregations. Multi-tenancy with user-level access control isolates data across teams. Official SDKs for Go, Python, and Node.js provide typed client libraries for programmatic integration. Deploys via Docker or direct binary download with no external dependencies beyond disk storage. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
Saleor
Backed by 23,000+ GitHub stars and trusted by global brands processing millions of orders, Saleor delivers the open-source headless commerce API that replaces monolithic ecommerce platforms with a composable, GraphQL-native architecture where APIs are the only way to interact with the system. The core engine built on Python and Django handles catalog management, order processing, payment orchestration, inventory tracking, and fulfillment workflows while remaining completely decoupled from any frontend technology. Native multichannel support enables per-channel control of pricing, currencies, warehouses, product availability, and payment methods, managing Instagram, Amazon, regional websites, and retail POS from a single backend. The extensibility layer provides 160+ webhooks spanning synchronous payment callbacks, asynchronous event notifications via Google Cloud Pub/Sub and AWS SQS, and subscription queries that shape webhook payloads to deliver only the data your services need. Dashboard UI Extensions offer 45+ mount points for embedding custom interfaces via iframes without forking, while the Apps system allows building payment gateways, PIM integrations, loyalty programs, and discount logic in any language. The React-based administration dashboard provides product management, order processing, customer segmentation, and analytics with multi-language and multi-currency support. OIDC integration connects existing identity providers for single sign-on across the merchant organization. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. BSD 3-Clause licensed.
Apache APISIX
With 17,000 GitHub stars, 460+ contributors, and deployments across telecommunications, automotive, and financial services running on over 10,000 CPU cores at the largest known installations, Apache APISIX delivers a fully dynamic API gateway achieving 140,000 QPS on eight cores with sub-millisecond latency through NGINX's event-driven architecture and LuaJIT-compiled plugin execution. The 100+ open-source plugins cover authentication (JWT, OAuth 2.0, OIDC, Keycloak, LDAP), observability (Prometheus, Datadog, SkyWalking, OpenTelemetry), traffic management (rate limiting, circuit breaking, canary releases, traffic splitting), and security (CORS, IP restriction, CSRF protection) — all hot-reloadable without process restarts via etcd-based real-time configuration synchronization. Multi-protocol support handles HTTP, gRPC, MQTT, TCP, UDP, and WebSocket traffic for both north-south API access and east-west service mesh communication. AI gateway capabilities proxy requests to 20+ LLM providers with semantic caching, token-aware rate limiting, provider failover routing, and content moderation. Custom plugins extend the gateway in Lua, Go, Java, Python, or WebAssembly. Radixtree route matching handles 100,000+ routes without performance degradation. Functions as a Kubernetes ingress controller with native service discovery for Consul, Nacos, and Eureka. Deploy via Docker or Helm charts with horizontal scaling through etcd cluster coordination. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
CPA Manager Plus
CPA Manager Plus is a self-hosted observability dashboard and management panel that tracks every AI request flowing through your CLI Proxy API gateway, breaking down failures, costs, and account health across providers like OpenAI, Anthropic, xAI, and Codex in one interface. When a request fails, drill into the persistent history to see status codes, affected models, latency, and redacted failure evidence without exposing raw response bodies. The cost analytics engine breaks down token consumption and estimated spend by model, provider, account, API key, project, channel, and time range while tracking input, output, reasoning, cache, and service-tier pricing semantics separately. Model prices sync automatically from models.dev with LiteLLM and OpenRouter fallbacks, and you can add local overrides for aliases or internal models. For teams running Codex or xAI accounts, the health inspector reads quota windows, reset evidence, credential state, and workspace status on a configurable schedule, routing credential failures into an action queue for review rather than letting them silently degrade throughput. Deploy the Lightweight Panel to replace your existing CPA management UI without adding another service, or run Full Mode as a single Docker container that adds the Manager Server with persistent SQLite storage for request history, historical analytics, and automated account inspections. Export or import request history as JSONL for external analysis, and back up the SQLite files alongside your encrypted management keys. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Hasura
A PostgreSQL database becomes a production-grade GraphQL API the moment Hasura GraphQL Engine points at it: track tables and relationships - existing schemas included - and full query, mutation, and subscription types appear with where, order_by, limit, offset, and on_conflict arguments, no resolvers or boilerplate written. Its Haskell core compiles GraphQL to efficient SQL, and any query becomes a real-time live query with a single keyword, powering dashboards and collaborative UIs over standard GraphQL subscriptions. Authorization is where Hasura earns its enterprise reputation: role-based access control with row- and column-level permission policies driven by session variables from JWTs, auth webhooks, or headers - each role effectively sees its own GraphQL schema containing only what it may touch, integrating cleanly with Auth0, Firebase, or homegrown auth. Event triggers fire webhooks on inserts, updates, and deletes for asynchronous business logic; Actions extend the schema with custom REST handlers; remote schema stitching merges external GraphQL services into one endpoint; and auto-generated REST endpoints serve clients that skip GraphQL. A browser console handles data modeling and API exploration, the CLI manages migrations and metadata as code, and deployment is a single stateless Docker container beside Postgres.
FreeLLMAPI
FreeLLMAPI collapses the chaos of 29 free LLM providers — Google AI, Cerebras, Groq, Mistral, OpenRouter, GitHub Models, Cohere, Cloudflare Workers AI, NVIDIA NIM, HuggingFace, SiliconFlow, Reka, Z.ai, and more — into a single /v1 endpoint that speaks both OpenAI and Anthropic protocols. The smart router selects the best available model for each request, automatically fails over to the next provider when rate limits hit, and tracks per-key token consumption so you never exceed a free-tier cap. Keys are stored with AES-256-GCM encryption and clients authenticate using a single unified bearer token, never exposing upstream provider credentials to downstream applications. The catalog tracks 251 model families across 358 provider/model endpoints with approximately 4 billion tokens per month of aggregate free-tier capacity, auto-refreshing from a signed manifest at freellmapi.co twice daily without requiring git pulls. Beyond chat completions, the proxy handles embedding, image generation, and audio/TTS endpoints, plus structured outputs with JSON schema forwarding, JSON healing, and format-ignore failover. An integrated MCP server at /mcp provides gateway introspection for coding agents, while the self-hosted OpenAPI reference at /v1/docs documents every route. Compatible with OpenAI SDKs, LangChain, LlamaIndex, Continue, Claude Code, and Hermes — just change base_url. Deploy via Docker, npm, or build from source. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Crawl4AI
With over 77,000 GitHub stars, Crawl4AI is the most-starred open-source web crawler on GitHub and the go-to tool for converting the web into AI-ready data. Built on Python and Playwright, it transforms any website into clean Markdown with headings, tables, code blocks, and citation hints optimized for LLM ingestion, or extracts structured JSON via CSS selectors, XPath expressions, or direct LLM-based schema extraction through OpenAI, Anthropic, and Ollama providers. The self-hosted Docker server exposes a REST API on port 11235 with endpoints for crawling, streaming results, screenshots, PDF generation, JavaScript execution, and LLM-powered extraction. Version 0.9.x introduced secure-by-default operation with mandatory JWT authentication, strict request validation, declarative hooks replacing inline code, and bounded job queues. Adaptive crawling uses information foraging algorithms to determine when sufficient data has been gathered, while deep crawl mode traverses link graphs intelligently. The async browser pool manages concurrent sessions with stealth plugins, proxy rotation, custom headers, and session persistence for authenticated scraping. A built-in MCP server enables direct integration with Claude, ChatGPT, and Cursor for AI-driven web research workflows. Content filtering applies BM25 and TF-IDF relevance scoring to extract only pertinent sections from noisy pages. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
MetaMCP
With 2,600+ GitHub stars, MetaMCP solves the MCP server sprawl problem by aggregating any number of upstream servers into a single authenticated endpoint that any MCP client connects to once. Group servers into namespaces — development tools in one, data sources in another — then publish each namespace as its own SSE, Streamable HTTP, or OpenAPI endpoint with API-key authentication in headers or query parameters, or full OAuth per the MCP Spec 2025-06-18 standard. The aggregation engine discovers tools, resources, and prompts from all active servers in parallel, prefixes tool names with server identifiers to prevent collisions, and applies configurable middleware including tool filtering to reduce context-window bloat and description overrides to improve LLM comprehension. The web management UI lets you configure MCP servers with stdio, SSE, or Streamable HTTP transports, toggle servers active or inactive per namespace, create and revoke API keys per endpoint, and inspect discovered tools with their schemas. Nested MetaMCP support enables hierarchical architectures where one MetaMCP instance consumes another, creating multi-level tool organization with automatic name resolution. Compatible with Claude Desktop, Claude Code, Cursor, Open WebUI, and any MCP-compatible client through a single connection URL. The Docker container packages the TypeScript backend with PostgreSQL for configuration persistence, exposing the management UI on port 12005 and MCP endpoints on configurable ports. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Languagetool
Grammar, punctuation, and style errors a dictionary lookup can't see: LanguageTool is open-source proofreading powered by a Java rule engine covering English, German, Spanish, French, Portuguese, Dutch, and 25+ other languages. Self-hosting the HTTP server is how you get Grammarly-class checking without sending every sentence you write to a third party - a real concern when the text being proofread is confidential email, legal drafts, or unreleased documentation. Your instance exposes the standard /v2/check API, so the official ecosystem plugs straight in: browser extensions for Chrome and Firefox accept a custom server URL, and integrations exist for VS Code, LibreOffice, Obsidian, Vim, Emacs, and many editors. Notably, self-hosting restores free browser-extension checking that the hosted service moved behind a premium subscription - your server, no character limits, no paywall. Detection quality is tunable: optional n-gram datasets (multi-gigabyte language models for en, de, es, fr, nl) teach the engine word-order and confusion-pair errors like there/their and brakes/breaks, and a fastText model improves automatic language identification. Everything runs offline once models are downloaded. The core is LGPL, the API is documented with Swagger, and rules are community- maintained and constantly expanding.
Spree
Spree Commerce provides a complete headless ecommerce backend where products, orders, payments, and promotions are managed through typed REST APIs with OpenAPI 3.0 specs and TypeScript SDKs providing autocomplete and type safety. Fifteen years of production history and 15,600+ GitHub stars back a mature ecosystem that ships a production-ready Next.js 16 storefront built with React 19 and Tailwind CSS 4, including multi-region URL routing and Stripe payments supporting Apple Pay, Google Pay, Klarna, and Affirm. Sales Channels model distinct contexts from a single instance: DTC storefronts, wholesale portals, mobile apps, and point-of-sale terminals, each with its own catalog, pricing, and checkout flow. The rules-based promotion engine supports coupon codes, multi-condition discounts, gift cards, and digital product fulfillment. Multi-warehouse inventory tracks stock across locations in real time with reservations and advanced order routing that splits shipments across fulfillment centers. B2B capabilities include customer-specific price lists, storefront access gating, wholesale portals with approval workflows, and quick order forms. The admin dashboard built with Tailwind CSS provides role-based permissions, product management, order processing with refunds, and scaffold generators for custom pages. BSD 3-Clause licensed with zero platform or transaction fees. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console.
Inference Gateway
Inference Gateway puts a single OpenAI-compatible API endpoint in front of OpenAI, Anthropic, Groq, Cohere, Ollama, DeepSeek, Google, Mistral, MiniMax, Moonshot, Nvidia, and llama.cpp, so your application code never changes when you switch models or providers. The Go binary starts on port 8080 and normalizes authentication, streaming protocols, and response formats across all backends transparently. Native Model Context Protocol support auto-discovers tools from connected MCP servers and injects them into LLM requests without client-side management, enabling server-side tool execution across any provider that supports function calling. Agent-to-Agent protocol integration allows distributed agent communication through a declarative Agent Definition Language that generates production-ready Go or Rust servers from a single YAML manifest. The dedicated Kubernetes Operator manages Gateway, Agent, MCP, and Orchestrator custom resources with automatic HPA scaling, OIDC authentication, and service discovery that rebuilds MCP configurations when the discovered server set changes. Prometheus metrics and OpenTelemetry tracing provide full request-level observability across the entire inference pipeline. Middleware controls enable per-request provider selection, model routing, and fallback strategies. Official SDKs in Go, Python, TypeScript, and Rust provide typed client interfaces with streaming support. Docker Compose deployment requires only environment variables for API keys. A CNCF Sandbox applicant. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.