OpenAI Evals
by openaiOpenAI's framework for benchmarking LLMs and an open-source registry of evals. Industry-standard test harness.
This project is no longer available on GitHub
The openai/OpenAI Evalsrepository has been removed, renamed, or made private since we listed it, so the source and download links no longer work. We've kept this page because the details below are still useful for understanding what OpenAI Evals did — and because there are working alternatives.
Try Langfuse, promptfoo or Arize Phoenix instead, or browse all llm evals tools.
About OpenAI Evals
OpenAI Evals is a Python open-source llm evals project by openai. OpenAI's framework for benchmarking LLMs and an open-source registry of evals. Industry-standard test harness. It reached 16,000 GitHub stars and 0 forks before the repository was taken down. The figures below are the last we recorded.
OpenAI Evals — FAQ
What is OpenAI Evals?
OpenAI's framework for benchmarking LLMs and an open-source registry of evals. Industry-standard test harness.
Can I still get OpenAI Evals?
Not from the original source. The openai/OpenAI Evals repository is no longer public on GitHub, so the download and source links no longer work. Forks may still exist elsewhere, and the related llm evals tools listed on this page are active alternatives.
Why is OpenAI Evals no longer available?
The repository was removed, renamed, or made private by its owner after we listed it. We keep this page up so the project's details stay findable and so you can jump straight to working alternatives.
What language is OpenAI Evals written in?
OpenAI Evals is primarily written in Python.
How popular is OpenAI Evals?
OpenAI Evals had 16,000 stars and 0 forks on GitHub when we last recorded it. Those figures are historical — the repository is no longer public.
More LLM Evals skills
See all LLM Evals →Open-source LLM engineering platform — tracing, prompt management, evaluations, datasets, playground.
CLI and library for evaluating, testing, and red-teaming LLM apps. Side-by-side prompt comparisons in your CI.
Open-source LLM observability — traces, evaluation, datasets, retrieval debugging. OpenTelemetry-native.