Apache NiFi
Deployed at thousands of enterprises across financial services, healthcare, government, and telecommunications, Apache NiFi is the industry-standard platform for building automated data pipelines through a visual drag-and-drop browser interface that requires zero coding for common integration patterns. The flow-based programming model connects over 300 built-in processors covering relational databases via ExecuteSQL and PutDatabaseRecord, Apache Kafka with PublishKafka and ConsumeKafka, HTTP endpoints through InvokeHTTP and ListenHTTP, cloud storage for AWS S3, Azure Blob, and Google Cloud Storage, SFTP/FTP file transfers, and JSON, XML, CSV, and Avro transformations. Data provenance tracking logs every routing decision, transformation, and delivery for every FlowFile, creating a searchable lineage graph from source to destination with full content replay capability for auditing and debugging. Guaranteed delivery uses configurable backpressure thresholds, prioritized queuing with latency or throughput optimization, and automatic retry with exponential backoff, ensuring no data loss even during downstream outages. The zero-leader clustering architecture distributes processing across nodes with automatic load balancing, while site-to-site protocol enables secure data transfer between NiFi instances across network boundaries. Security includes OpenID Connect and SAML 2.0 single sign-on, role-based access control with fine-grained policies per component, and TLS encryption for all communication. Custom processors can be written in Java and packaged as NAR bundles, or implemented directly in Python through the native scripting framework. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
Open-Meteo
High-resolution weather forecasts became a free commodity because of Open-Meteo - and this deployment puts the whole open-source engine on your own infrastructure. The public open-meteo.com service aggregates national weather models (NOAA GFS, DWD ICON, ECMWF, Meteo-France, and others) into one consistent JSON interface; self-hosting gives you that same API without rate limits, third-party dependency, or usage metering. The architecture is two cooperating services: the API server exposes forecast endpoints fully compatible with Open-Meteo query parameters - latitude, longitude, hourly and daily variables like temperature, precipitation, wind, and radiation - while a background sync worker downloads fresh weather model data on a configurable interval into a shared persistent volume at /app/data, so forecasts stay current and survive restarts without re-downloading. You control which weather models to mirror, which variables to store, how much historical depth to keep, and how often to refresh - meaning a lean deployment can sync only the model and region you actually query. Responses are plain HTTP/JSON, so integration with dashboards, Home Assistant-style automations, agricultural monitoring, IoT fleets, or any application takes minutes. For anyone making thousands of forecast calls a day, replacing a metered weather API with your own instance turns a recurring bill into a flat infrastructure cost.
xyOps
With 4,500+ GitHub stars and version 1.0.92 released August 2026, xyOps delivers a complete operations platform that unifies workflow automation, job scheduling, server monitoring, alerting, and incident response in one self-hosted system. The platform uses a distributed architecture where a central conductor coordinates lightweight xySat satellite agents running on Linux, macOS, or Windows worker nodes via persistent WebSocket connections. The visual workflow builder lets you chain events, triggers, actions, and monitors into multi-step pipelines with conditional logic, fan-out/fan-in parallelism, multiplex controllers for fleet-wide execution, and configurable resource limits. QuickMon provides per-second CPU, memory, disk, and network visibility streamed live to the web UI, while user-defined monitor plugins sample metrics every minute with time-series storage at hourly, daily, monthly, and yearly resolutions. Alert triggers evaluate expressions against live data and fire notifications via email, webhook, or custom actions, with full server snapshots attached showing every running process, network connection, and resource utilization at the moment of detection. Failed jobs and alerts automatically create tickets with linked logs, metrics history, and context for end-to-end incident tracking. The plugin marketplace supports extensions written in any language, and the Docker plugin enables container-based job execution. Deploy via Docker with persistent volumes on port 5522 for the web UI and 5523 for API access, running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. BSD-3-Clause licensed.
FalkorDB
FalkorDB is the first queryable property graph database to leverage sparse adjacency matrices and linear algebra for graph traversal, replacing traditional pointer-chasing with GraphBLAS-accelerated computation. Originally the RedisGraph engine, it was relaunched as FalkorDB in 2023 and rewritten from C to Rust in 2026 for improved memory safety and performance. The database supports the OpenCypher query language with proprietary extensions, translating queries into linear algebra expressions that exploit AVX hardware acceleration. Indexing options include full-text search, vector similarity for embedding-based retrieval, and range indexing, while connectivity supports both the RESP protocol for Redis clients and the Bolt protocol for Neo4j-compatible tooling. The GraphRAG SDK enables ingestion of documents in text, PDF, and Markdown formats into knowledge graphs, with schema-guided entity extraction, hybrid retrieval combining vector and graph traversal, relationship expansion, and cited answers for LLM applications. Official client libraries cover Python, Node.js, Java, Rust, Go, PHP, and C#. Multi-tenant support handles over 10,000 concurrent graphs with zero overhead and full isolation. Docker deployment runs the falkordb/falkordb image on ports 6379 for the database server and 3000 for the built-in browser UI, with persistent volume storage and optional authentication. A production falkordb-server image excludes the browser for lighter deployments. On RepoCloud, deploy FalkorDB on a dedicated VPS with root SSH access, persistent storage for your graph data, and complete control over authentication, thread count, and memory configuration, all under the SSPLv1 license.
Rivet
Stateful serverless actors that run indefinitely, sleep when idle, and persist state across restarts without external database round trips. Rivet provides a vendor-neutral alternative to Cloudflare Durable Objects, delivering long-running processes with co-located in-memory state and per-actor SQLite databases on any infrastructure you choose. The Rust engine comprises four packages: Pegboard for actor orchestration, Gasoline for durable execution, Guard for traffic routing via Envoy, and Epoxy implementing multi-region KV storage through EPaxos consensus. Each actor maintains instant-access in-memory state plus a dedicated SQLite database for relational queries, backed by RocksDB on single nodes or PostgreSQL with NATS pub/sub for multi-node clusters. RivetKit SDKs in TypeScript, Rust, and Python support built-in WebSocket connections, task queues, scheduling, and CRDT-based real-time collaboration. The v2.3 rewrite moved the core runtime from JavaScript to native Rust via WebAssembly, eliminating main-thread blocking across Node.js, Bun, Deno, and Cloudflare Workers. A built-in dashboard provides actor inspection with real-time Prometheus metrics. Deploys as a single Docker container on port 6420 with optional PostgreSQL for persistence. Available on RepoCloud with a dedicated VPS under the Apache 2.0 license.
LightDash
With 5,600+ GitHub stars and deep dbt integration, Lightdash is the open-source Agentic BI platform that treats analytics like software — defining metrics, dimensions, joins, permissions, and caching in a governed context layer that powers dashboards, AI agents, data apps, embedded analytics, and MCP server endpoints simultaneously. The dbt Write-Back feature lets business users create custom metrics and models in the UI, then automatically generates pull requests in GitHub or GitLab so every change flows through code review and CI validation before reaching production. Context-specific AI analysts automatically select relevant models and metrics, build queries, and present insights in plain English, while row-level security, user attributes, and customer-facing permissions ensure data governance at every layer. The platform connects to BigQuery, Snowflake, Redshift, Databricks, PostgreSQL, Trino, and ClickHouse through warehouse adapters, with the TypeScript monorepo built on React, Mantine, Vite, and TanStack Query on the frontend plus Node.js, Express, Knex, and PostgreSQL on the backend. Data teams build analytics with coding agents, preview changes from the CLI, validate in CI pipelines, and review charts and dashboards in pull requests — making the entire analytics lifecycle version-controlled and reproducible. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Passbolt
Security-conscious IT departments pick Passbolt for its cryptography: every user holds an OpenPGP key pair, and shared credentials are encrypted individually to each recipient's public key - real end-to-end encryption, not a vault password handed around. All crypto runs client-side in the mandatory browser extension (distributed and signed through the Chrome and Firefox stores, deliberately separating the crypto code from the server that stores ciphertext); private keys and passphrases never touch your instance, and the server admin cannot read a single secret. Authentication uses the challenge-based GpgAuth protocol, secrets are digitally signed to verify sender integrity, and metadata encryption extends protection to resource names and URLs. Day to day it behaves like a polished commercial manager: auto-fill and auto-save in forms, strong password generation, anti-phishing protection, TOTP storage, folder hierarchies shared per-user or per-group with fine-grained permissions and instant cryptographic revocation. Native iOS, Android, and desktop apps ship alongside a JSON API, CLI, and SDKs for CI/CD secret retrieval and rotation. The PHP server runs on MariaDB and is AGPL-licensed open source - including the paid tiers' codebase - with published security audits.
Papercups
Companies with privacy and security concerns about piping customer conversations through Intercom or Zendesk run Papercups - open-source live customer chat. The stack is a deliberate strength: an Elixir/Phoenix API over PostgreSQL, with real-time messaging powered by Phoenix Channels and Presence - the same BEAM foundation trusted by Discord and PagerDuty for fault-tolerant, low-latency messaging. Customers see a customizable chat widget that embeds in any site as an HTML snippet, a React component, or even inside React Native apps, with configurable colors, greetings, and away messages. Your team sees a dashboard for managing conversations - close, assign, and prioritize - with Markdown and emoji in replies. The killer workflow is the reply-channel integration: connect Slack or Mattermost and every customer conversation becomes a synced thread your team answers without leaving the tool they already live in, with two-way message syncing handled by webhooks. Email and SMS channels extend intake beyond the widget, an analytics dashboard tracks communication patterns, and the Storytime feature adds real-time screen sharing to watch users navigate while you help them. A documented API supports fully custom chat UIs in Svelte, Flutter, or Vue. MIT-licensed and GDPR-conscious - customer data stays in your PostgreSQL.
Weblate
Over 2,500 open-source projects and companies in more than 165 countries localize with Weblate - the libre continuous localization platform and the standard self-hosted answer to Crowdin and Lokalise. Its defining trait is that translations live in the same version control as your code: Weblate talks directly to Git and Mercurial, pulls new source strings automatically via webhooks, and pushes finished translations back either as direct commits or as pull/merge requests on GitHub, GitLab, Gitea, Bitbucket, Azure DevOps, Gerrit, or Pagure. Every translator is properly credited in the commit history. For translators, it is a full computer-aided translation tool: translation memory, glossaries, customizable quality checks that catch placeholder and formatting mistakes, propagation of identical strings across components, and automatic suggestions from machine translation services - DeepL, Amazon Translate, LibreTranslate, and others, with per-service priorities and support for custom Python engines. It handles the format zoo (gettext PO, JSON, YAML, Android XML, iOS strings, and dozens more) and supports crowdsourced workflows with granular access control, workspaces, two-factor authentication, and reviewer approval steps. A REST API, CLI client, and add-on system automate everything else. Built on Python/Django, GPL-licensed, with no per-string or per-seat pricing when self-hosted.
Redmine
Nearly two decades running engineering organizations: Redmine is the veteran open-source project management and issue tracker, a Ruby on Rails application (GPLv2) still in active development. Its core strength is configurability: define your own trackers (bug, feature, task, or anything else), issue statuses, and role-based workflows that control exactly which transitions each role may perform, then extend records with custom fields of every type. Issues support subtasks, relations (blocks, precedes, duplicates), watchers, categories, and full journaled history, with saved custom queries and cross-project filtering for slicing the backlog any way you need. Around the tracker sit Gantt charts and calendars, a roadmap driven by versions, per-project wikis, forums, news, and document repositories, plus time tracking with estimated versus spent hours and activity-based reporting. Multi-project support runs deep - subprojects, per-project modules, and granular role-based permissions - and repository integration (Git, Subversion, Mercurial) links commits to issues automatically. Email notifications, inbound email-to-issue creation, LDAP authentication, a REST API, and a large plugin and theme ecosystem round it out. Recent 6.x releases brought substantial query and rendering optimizations. Self-hosting keeps your entire project history in your own database, free of per-seat licensing.
GoatCounter
GoatCounter delivers meaningful web traffic insights — pageviews, referrers, browsers, screen sizes, country-level geolocation — without setting a single cookie, without collecting personal data, and without forcing GDPR consent banners on your visitors. Written entirely in Go and distributed as a single compiled binary consuming roughly 25MB of RAM, it adds just 3.5KB to your pages via the tracking script, with a JavaScript-free tracking pixel alternative for sites that avoid scripts entirely, plus backend middleware integration and log file import for server-side collection. The dashboard displays pageview counts per path with hourly resolution, referrer sources grouped by domain with full URL on hover, browser and OS version breakdowns, screen size distributions, and country-level location data derived from IP addresses that are immediately discarded after geolocation. Campaign tracking supports UTM parameters and custom data attributes. A public stats option exposes your dashboard at a shareable URL for build-in-public transparency. SQLite serves as the default database requiring zero administration, while PostgreSQL handles higher-traffic deployments with multi-site setups. Built-in ACME and TLS certificate management eliminates reverse proxy requirements for HTTPS — no Nginx or Caddy needed. The REST API provides programmatic access to all analytics data. Deploy as a single binary, via Docker with the official arp242/goatcounter image, or through native packages. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. EUPL-1.2 licensed.
Baserow
Airtable's spreadsheet-database model, self-hostable and open-source: that is Baserow. It presents data in a spreadsheet-style grid, but underneath each table is a real relational structure with typed fields, links between tables, filters, sorts, and multiple views (grid, gallery, form, kanban, calendar). Beyond the database core, it includes an application builder for composing pages and portals on your data, workflow automations, and dashboards. Everything is API-first: each table exposes a REST endpoint with token auth and webhooks, so it plugs directly into n8n, Zapier, or custom scripts. The stack is Django (Python) on the backend, Vue.js on the frontend, PostgreSQL for storage, with Redis for async tasks. Core features are MIT-licensed; premium features are a paid add-on. The self-hosted version has no row, storage, or API request limits - Airtable's per-base record caps and monthly API quotas simply don't exist here, and capacity is bounded only by your PostgreSQL database and disk. Existing Airtable bases, CSVs, and Excel files import directly with structure preserved, so migration doesn't start from a blank slate, and both the backend and frontend support plugins for custom field types and integrations without forking the core. For non-technical teammates the interface behaves like a spreadsheet; for engineers, the data model is the API.
AgentsView
Software developers running Claude Code, Cursor, Codex, and Aider use AgentsView to trace token expenditures, search full session transcripts, and analyze multi-agent concurrency timelines. Engineers can search through historical developer sessions to locate specific code modifications, prompt chains, or debugging rationales across extensive programming workflows. The real-time activity timeline visualizes concurrent agent operations, revealing peak execution windows, task run durations, and automated tool invocations. The usage analytics engine calculates financial costs and token volumes per model by applying live pricing tables while accounting for prompt caching read discounts. Developers inspect granular turn-by-turn session transcripts complete with raw model payloads, file difference diffs, executed shell commands, and syntax errors. The recall browser extracts key technical context from past programming sessions, organizing facts into an indexed reference knowledge corpus for reuse across future project tasks. Administrators can replicate SQLite session tables into DuckDB analytics files or PostgreSQL warehouses to query longitudinal metrics using SQL analytical tools or custom reporting scripts. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Google Maps Scraper
The leading open-source tool for extracting business leads from Google Maps at production scale. The Go-based engine processes approximately 120 places per minute with optimized concurrency, extracting 33+ data points per listing including business name, address, phone number, website URL, rating, review count, latitude and longitude, opening hours, price level, and optionally crawling business websites for email addresses. Three interfaces serve different workflows: the CLI accepts query files for cron jobs and CI/CD pipelines with output to CSV, JSON, PostgreSQL, S3, or LeadsDB; the Web UI provides a browser-based dashboard with real-time job monitoring, a map view of scraped places, and interactive query submission; and the REST API at /api/v1 enables programmatic integration with full Swagger documentation at /api/docs. Built-in proxy rotation supports SOCKS5, HTTP, and HTTPS with authentication for large-scale runs, while the architecture scales from a laptop to Kubernetes clusters with queue-based worker distribution. The SaaS edition adds multi-user access with API key management, admin UI with 2FA, job queue orchestration, and one-command cloud deployment via an interactive wizard. An AI Agent Skill enables coding agents to run scrapes programmatically. Deploy via Docker or build from source requiring Go 1.26.5+. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Zammad
With 5,700+ GitHub stars and over a decade of active development since 2012, Zammad is the open-source helpdesk platform that unifies every customer communication channel — email, live chat, telephone, WhatsApp, Telegram, Facebook, SMS, and web forms — into a single ticket management interface backed by PostgreSQL, Elasticsearch, and Redis. Version 7.0 introduced native AI features including automated ticket categorization, priority assignment, and title rewriting via AI agents that plug into triggers, macros, and scheduler jobs, plus one-click AI ticket summaries and a writing assistant — all configurable with your choice of LLM provider: OpenAI, Anthropic, Mistral AI, Azure AI, Ollama for local models, or any custom OpenAI-compatible endpoint with full audit logging of every AI action. Define service level agreements with first response, update, and solution time tracking tied to business hours calendars, with automatic escalation notifications. The knowledge base provides multilingual FAQ management with internal and public visibility. Core workflows enable dynamic ticket masks with conditional field dependencies per group. Text modules let agents insert predefined responses via the double-colon shortcut, while macros execute multi-step actions with one click. Security includes two-factor authentication, S/MIME and PGP email encryption, and single sign-on via SAML or OpenID Connect. Integrations connect to GitHub, GitLab, Microsoft 365, LDAP with nested group support, and Exchange. Deploy via Docker Compose, Kubernetes with the official Helm chart, or DEB/RPM packages. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPLv3 licensed.
Countly
Mixpanel, OneSignal, and Crashlytics in one self-hosted stack - Countly is an all-in-one product analytics and engagement platform where every byte of first-party data stays on your server. A Node.js application over MongoDB, it collects through ten battle-tested SDKs spanning iOS, Android, web JavaScript, React Native, Flutter, Unity, and desktop (plus a data write API for anything else), and has powered tens of thousands of mobile, web, and desktop apps since 2012. The analytics core covers sessions, custom events, views, user profiles, and real-time dashboards, with exploration tools built for product managers as much as analysts. What separates Countly from pure analytics tools is acting on the data without third parties: built-in push notifications send automated, transactional, and personalized messages to iOS (APNs), Android (Firebase), and Huawei devices, with the SDK handling token retrieval and permission flows automatically; crash reporting captures symbolicated native crashes on iOS and Android plus JavaScript errors, correlated with the same user and session data. Email reports keep stakeholders updated, and the plugin-based architecture means features load as modules. For GDPR-sensitive products, engagement without piping user data to advertising companies is the entire point. AGPL-licensed server, installable in minutes.
Livebook
Livebook is an interactive notebook for Elixir where you write code alongside rich Markdown prose, execute it cell by cell with reactive dependency tracking, and deploy finished notebooks as standalone web applications with a single click. Nearly 6,000 GitHub stars reflect the Elixir core team's investment in a platform where code cells run on demand alongside Mermaid diagrams and KaTeX mathematical formulas. The Kino visualization library renders Vega-Lite charts, interactive data tables with sorting and pagination, Leaflet maps, and Mermaid diagrams directly within notebook output cells, while custom Kino components enable building interactive controls with sliders, text inputs, and buttons that feed values back into running code. Smart cells abstract high-level tasks into configurable UI widgets: query PostgreSQL, MySQL, SQLite, and BigQuery databases, train machine learning models with Axon, plot charts, and build map visualizations without writing boilerplate code. Real-time collaboration lets multiple users edit the same notebook simultaneously with cursor presence indicators and synchronized cell evaluation. Notebooks are stored as .livemd files, a Markdown-compatible format that renders cleanly on GitHub and integrates with standard version control workflows. Custom runtimes connect Livebook to existing Elixir applications for live introspection and documentation of running systems. The Docker image at ghcr.io/livebook-dev/livebook exposes ports 8080 and 8081 with password or token authentication. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
Mox
Mox offers a complete mail server stack in a single Go binary requiring no external dependencies. The quickstart command configures a working mail server with SMTP, IMAP4, webmail, and full DNS authentication in under ten minutes. SMTP handling includes a delivery server on port 25, a submission server for authenticated clients, and a queue with automatic retries and DKIM signing. IMAP4rev2 implementation provides full mailbox synchronization with CONDSTORE and QRESYNC for efficient offline clients, NOTIFY for multi-mailbox monitoring, MULTISEARCH across mailboxes, and TLS client certificate authentication via the EXTERNAL SASL mechanism. The built-in webmail provides browser-based reading and composing with message threading, attachments, and HTML rendering without requiring a separate web client. Email authentication implements SPF validation, DKIM signing and verification with automatic key rotation, DMARC policy enforcement with aggregate and failure reporting, DANE with DNSSEC-protected TLSA records, and MTA-STS for certificate verification. Junk filtering combines reputation-based sender scoring with Bayesian content analysis trained per account. The web administration interface manages domains, accounts, DNS records, TLS certificates, delivery queue, and real-time log viewing. Internationalized email addresses with EAI and IDNA support handle non-ASCII domains and mailboxes. Account autoconfiguration publishes settings for Thunderbird autoconfig and Outlook autodiscover. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.