Logo
Deploy Now

Stars

33,514

Forks

3,613

Watchers

107

Developer links

Langfuse

Backed by Y Combinator and trusted by over 2,300 companies processing billions of observations monthly, Langfuse is the most widely adopted open-source platform for building, monitoring, evaluating, and debugging LLM applications. The hierarchical tracing engine captures every LLM call, tool invocation, retrieval step, and agent action as nested spans based on OpenTelemetry, with automatic cost calculation, latency tracking, and token usage attribution across sessions and users. Prompt Management separates prompts from code with versioned artifacts, label-based deployments, one-click rollbacks, and runtime SDK fetching with server-side caching, while linking every generation back to its exact prompt version for attribution analytics. The evaluation system supports LLM-as-a-judge scoring, heuristic code evaluators, user feedback collection, and manual annotation workflows that run automatically on production traces or against curated datasets. The Playground enables interactive prompt testing on real production inputs with side-by-side model comparison across providers. Datasets and Experiments define test cases for systematic benchmarking with comparative result visualization. Native SDKs for Python and TypeScript provide decorator-based instrumentation, while 100+ integrations cover LangChain, LlamaIndex, OpenAI SDK, LiteLLM, Vercel AI SDK, and any OpenTelemetry-instrumented framework. The analytics dashboard surfaces cost breakdowns, quality scores, latency percentiles, and usage trends across models and prompt versions. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Langfuse
Langfuse
Langfuse
Langfuse
Langfuse

Benefits

  • Full LLM Application Observability
  • Hierarchical traces capture every LLM call, tool invocation, retrieval step, and agent action with automatic cost calculation, latency tracking, and token usage attribution.
  • Versioned Prompt Management
  • Separate prompts from code with version control, label-based deployments, one-click rollbacks, and runtime SDK fetching that links each generation to its exact prompt revision.
  • Flexible Evaluation Framework
  • Run LLM-as-a-judge scoring, heuristic code evaluators, user feedback pipelines, and manual annotation workflows automatically on production traces or against curated test datasets.
  • 100+ Framework Integrations
  • Native SDKs for Python and TypeScript with drop-in support for LangChain, LlamaIndex, OpenAI SDK, LiteLLM, Vercel AI SDK, and any OpenTelemetry-instrumented application.

Features

  • Interactive Playground
  • Test and iterate on prompts using real production inputs with side-by-side model comparison across providers directly in the Langfuse interface.
  • Datasets and Experiments
  • Define test sets and benchmarks for systematic LLM evaluation with comparative result visualization and continuous improvement tracking across prompt versions.
  • Cost and Latency Analytics
  • Dashboard surfaces cost breakdowns by model and prompt version, latency percentiles, quality score trends, and usage analytics across all traced LLM operations.
  • OpenTelemetry Foundation
  • Built on the OpenTelemetry standard for reduced vendor lock-in, enabling compatibility with existing OTel-instrumented services and standard observability tooling.
  • Agent Graph Visualization
  • Represent and inspect complex agent workflows as visual graphs showing decision paths, tool calls, and LLM interactions within multi-step autonomous processes.