An LLMOps platform for prompt management, evaluation, and observability to help teams build reliable AI applications.

Agenta screenshot 1

Agenta is an open-source and self-hosted LLMOps platform for building AI applications. It focuses on prompt management, evaluation, and observability to help engineering and product teams create reliable LLM applications through a structured development lifecycle.

The software is deployed as a server and accessed via a web application. It can be self-hosted using Docker Compose or used through a managed cloud version. The platform is specifically designed to bridge the gap between technical developers and subject matter experts, allowing non-technical team members to iterate on prompts and configurations without direct code changes.

Key features

  • Interactive playground to compare prompts side by side against test cases
  • Support for over 50 LLM models including bring-your-own-model options
  • Version control for prompts and configurations with branching and environments
  • Testset creation from production data, playground experiments, or CSV uploads
  • Automated evaluation using LLM-as-judge and 20+ pre-built evaluators
  • OpenTelemetry native tracing compatible with OpenLLMetry and OpenInference
  • Monitoring tools for cost, latency, and usage patterns
  • Integration of human feedback and expert annotations into evaluations

Agenta is built for teams that require a systematic approach to LLM development rather than ad-hoc prompting. It allows subject matter experts to collaborate on complex configuration schemas, ensuring that prompt changes are validated before they reach production. The platform integrates with various models and frameworks, utilizing open standards for tracing to ensure full visibility into complex production workflows and debugging processes.

By combining a UI for subject matter experts and an API for engineers, the platform supports both manual and programmatic evaluation workflows. This dual approach enables teams to run large-scale automated tests while still incorporating qualitative human feedback into the optimization process.

It serves as a centralized hub for the entire LLM lifecycle, from initial prompt experimentation and versioning to production monitoring and cost tracking.

Last Modified
Software TypeWeb App / Server
Platform
Last Activity14 days ago
Repository Age3 years
LicenseMIT
Open Source Alternative to
Open Source Software.io

Join our newsletter to get shiny new open source software delivered to your inbox. Unsubscribe anytime.