3 apps GPU
Airi screenshot thumbnail

Airi

Project AIRI is the most popular open-source AI companion platform — a self-hosted recreation of Neuro-sama that brings AI-powered virtual characters into your world across web, desktop, and mobile. The system renders Live2D, Spine, and VRM 3D character models with auto-blink, eye tracking, and lip-sync driven by real-time voice synthesis, while the xsAI abstraction layer connects to 40+ LLM providers including OpenAI GPT-4, Anthropic Claude, Google Gemini, DeepSeek, and local models via Ollama and OpenRouter. Built from day one on WebGPU, WebAudio, Web Workers, WebAssembly, and WebSocket technologies, the browser version runs entirely client-side with PWA offline support while the server runtime enables persistent memory via PostgreSQL with pgvector embeddings and DuckDB WASM for client-side storage. The Minecraft agent plays autonomously using mineflayer with pathfinding, and a Factorio integration provides cooperative gameplay. Social integrations deploy your companion as a Discord bot joining voice channels, a Telegram bot, and a Twitter/X agent posting and replying autonomously. The desktop Stage Tamagotchi app provides an always-on-screen companion for Windows and macOS, while Stage Pocket brings the experience to mobile. Voice features include client-side speech recognition via VAD, multiple TTS providers including ElevenLabs, and screen vision capabilities. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
LocalAI screenshot thumbnail

LocalAI

With over 48,000 GitHub stars and monthly releases since March 2023, LocalAI is the self-hosted AI engine that replaces every OpenAI endpoint with a single Docker container running on your own infrastructure — serving chat completions, image generation, text-to-speech, speech-to-text, embeddings, vision, video generation, and function calling through identical API schemas that require zero application code changes. The composable backend architecture isolates each inference engine as a separate gRPC service running in its own OCI container, so llama.cpp, vLLM, SGLang, transformers, whisper.cpp, diffusers, MLX, Stable Diffusion, and Flux install on demand without touching the core, can run on separate machines, and a fault in one never affects others. Hardware acceleration spans NVIDIA CUDA 12 and 13, AMD ROCm, Intel oneAPI/SYCL, Apple Silicon Metal, Vulkan, and NVIDIA Jetson L4T — or runs entirely on CPU without any GPU. Built-in AI agents support autonomous tool use, retrieval-augmented generation, Model Context Protocol integration, and skill-based workflows directly in the web interface. The model gallery provides curated YAML configuration files for hundreds of models that install with a single command, while P2P federated inference distributes model shards across multiple machines for running models larger than any single node's memory. Multi-user API key authentication with quotas and role-based access enables team deployments. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
SD WebUI Forge screenshot thumbnail

SD WebUI Forge

With 12,800 GitHub stars and backing from the same developer who created ControlNet, Stable Diffusion WebUI Forge replaces Automatic1111's inference backend with a dynamic GPU memory management system that runs SDXL 30-75% faster while consuming significantly less VRAM — enabling 1024x1024 generation on 6GB cards where A1111 requires 8GB or more. The Gradio 4 interface provides txt2img, img2img, inpainting, and outpainting workflows with a Forge Canvas supporting pressure-sensitive input from Wacom tablets and Microsoft Surface devices. Native Flux.1 model support loads Flux Dev and Schnell checkpoints using BitsandBytes NF4 and FP8 quantization for deployment on consumer GPUs without model splitting. Built-in ControlNet integration includes all preprocessors — Canny, Depth, Normal, OpenPose, MLSD, Scribble, Segmentation, Tile, and IP-Adapter — without requiring separate extension installation. The extension ecosystem maintains full compatibility with popular Automatic1111 extensions including Adetailer for face enhancement, After Detailer, Regional Prompter, and Dynamic Prompts. LoRA loading supports standard, LyCORIS, and DoRA formats with automatic weight detection. The API provides RESTful endpoints for txt2img, img2img, extra single/batch processing, and progress monitoring enabling headless batch generation. Deploy via one-click installer package, Python virtual environment, or Docker with NVIDIA GPU passthrough. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.

Deploy