Arize Phoenix
by Arize-aiOpen-source LLM observability — traces, evaluation, datasets, retrieval debugging. OpenTelemetry-native.
This project is no longer available on GitHub
The Arize-ai/Arize Phoenixrepository has been removed, renamed, or made private since we listed it, so the source and download links no longer work. We've kept this page because the details below are still useful for understanding what Arize Phoenix did — and because there are working alternatives.
Try OpenAI Evals, Langfuse or promptfoo instead, or browse all llm evals tools.
About Arize Phoenix
Arize Phoenix is a Python open-source llm evals project by Arize-ai. Open-source LLM observability — traces, evaluation, datasets, retrieval debugging. OpenTelemetry-native. It reached 6,000 GitHub stars and 0 forks before the repository was taken down. The figures below are the last we recorded.
Arize Phoenix — FAQ
What is Arize Phoenix?
Open-source LLM observability — traces, evaluation, datasets, retrieval debugging. OpenTelemetry-native.
Can I still get Arize Phoenix?
Not from the original source. The Arize-ai/Arize Phoenix repository is no longer public on GitHub, so the download and source links no longer work. Forks may still exist elsewhere, and the related llm evals tools listed on this page are active alternatives.
Why is Arize Phoenix no longer available?
The repository was removed, renamed, or made private by its owner after we listed it. We keep this page up so the project's details stay findable and so you can jump straight to working alternatives.
What language is Arize Phoenix written in?
Arize Phoenix is primarily written in Python.
How popular is Arize Phoenix?
Arize Phoenix had 6,000 stars and 0 forks on GitHub when we last recorded it. Those figures are historical — the repository is no longer public.
More LLM Evals skills
See all LLM Evals →OpenAI's framework for benchmarking LLMs and an open-source registry of evals. Industry-standard test harness.
Open-source LLM engineering platform — tracing, prompt management, evaluations, datasets, playground.
CLI and library for evaluating, testing, and red-teaming LLM apps. Side-by-side prompt comparisons in your CI.