Trace, evaluate, and optimize LLM applications with an observability platform featuring prompt management and dataset versioning.

Arize Phoenix is an open-source AI observability platform. It provides a specialized environment for the experimentation, evaluation, and troubleshooting of LLM applications through runtime tracing and performance benchmarking. The platform helps developers identify bottlenecks in AI workflows and refine the accuracy of model responses.
The software is deployed as a server that can run on local machines, within Jupyter notebooks, in containerized environments via Docker or Kubernetes, or in the cloud. It is available as a comprehensive platform package via pip or conda, while also offering lightweight Python and TypeScript sub-packages. These sub-packages allow developers to implement specific functionalities, such as OpenTelemetry instrumentation or API interactions, without deploying the full platform locally.
Phoenix is designed to be vendor and language agnostic, utilizing the OpenInference project for auto-instrumentation. It supports a broad array of LLM providers, including OpenAI, Anthropic, Google GenAI, and AWS Bedrock. The framework integrates with numerous popular AI libraries such as LlamaIndex, CrewAI, and DSPy. To facilitate developer workflows, the architecture includes specialized clients and a CLI for fetching traces and datasets for use with coding agents.
It serves as a technical infrastructure layer for AI engineers to monitor, evaluate, and optimize the performance of generative AI applications throughout the development lifecycle.
Collects and correlates logs, metrics, and traces in a single interface for microservices observability and troubleshooting.
Full stack monitoring platform providing session replay, error tracking, and server side logging for web applications.
Join our newsletter to get shiny new open source software delivered to your inbox. Unsubscribe anytime.