Logo
Deploy Now

SaaS Alternative

Claude

Stars

6,630

Forks

519

Watchers

135

Developer links

OpenSquilla

Claiming 60-80% token cost reduction compared to flat single-model deployments and backed by 6,500+ GitHub stars, OpenSquilla delivers an intelligent AI agent runtime where a local ML classifier evaluates every turn on message length, code blocks, keyword patterns, and semantic embeddings before routing it to the optimal model tier from C0 through C3. The pluggable provider layer connects natively to TokenRhythm, OpenRouter, OpenAI, Anthropic, Ollama, DeepSeek, Gemini, DashScope, Moonshot, Mistral, Groq, Zhipu, SiliconFlow, vLLM, LM Studio, and additional compatible backends with primary-plus-fallback selection. The four-tier cognitive memory architecture spans working, episodic, semantic, and raw layers with vector-semantic and BM25 retrieval powered by on-device ONNX embeddings that never leave your infrastructure. Security isolation operates at the syscall level via Bubblewrap on Linux and Seatbelt on macOS, complemented by policy-based execution controls and prompt injection protections. The unified TurnRunner executes identically across the Vue-based control console Web UI, terminal CLI, and chat channel integrations including Slack and Discord, ensuring consistent tool dispatch, retry logic, and decision logging regardless of entry point. Built-in skills cover deep research, multi-search-engine queries, document generation for DOCX, PPTX, XLSX, and PDF formats, GitHub integration, cron scheduling, and bounded subagent delegation. Per-agent workspaces with durable session storage provide transcript replay, context state management, and per-call cost tracking with automatic quota enforcement. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache-2.0 licensed.

OpenSquilla

Benefits

  • 60-80% Token Cost Reduction
  • Local LightGBM classifier routes routine tasks to cheaper models while reserving expensive reasoning-tier models only for complex multi-step problems, cutting API spend dramatically.
  • Syscall-Level Security Isolation
  • Bubblewrap on Linux and Seatbelt on macOS enforce syscall-level sandboxing with policy-based execution controls, preventing tool breakouts and prompt injection attacks at the OS level.
  • Unified Multi-Channel Runtime
  • Single TurnRunner executes identically across Web UI, CLI, Slack, and Discord with consistent tool dispatch, retry logic, decision logging, and streaming artifact support.
  • On-Device Private Embeddings
  • ONNX-based local inference generates embeddings without sending data externally, powering four-tier cognitive memory with vector-semantic and BM25 hybrid retrieval on your hardware.

Features

  • SquillaRouter Model Routing
  • LightGBM plus ONNX classifier scores turns on length, language, code, keywords, and semantic embeddings, routing across four capability tiers to the cheapest adequate model.
  • 20+ LLM Providers
  • Provider registry connects to OpenAI, Anthropic, Ollama, DeepSeek, Gemini, DashScope, Moonshot, Mistral, Groq, vLLM, LM Studio, and more with automatic failover.
  • Four-Tier Cognitive Memory
  • Working, episodic, semantic, and raw memory layers with vector-semantic plus BM25 retrieval enable persistent context across sessions without token waste.
  • Built-In Skills Library
  • Pre-built skills for deep research, document generation (DOCX, PPTX, XLSX, PDF), GitHub integration, cron scheduling, and bounded subagent delegation.
  • Vue Control Console
  • Browser-based Web UI provides session management, agent workspaces, transcript replay, routing diagnostics, cost dashboards, and real-time streaming visualization.