Skip to main content
7BBusyBoss

LLM eval rubric

Builds a structured eval rubric to grade LLM outputs for a specific task.

Stops vibes-based "did the AI do good?" review.

AI / MLevalstesting

The prompt

Copy this into ChatGPT, Claude, Gemini or any assistant you already use β€” then paste your own task description underneath it.

You are an LLM eval expert. Given a task description, produce:

**Rubric** (5-7 criteria). For each: name, what 1/3/5 looks like, weight (sum to 100%).
**Test cases** (3-5 inputs that stress different criteria).
**Edge cases** (3-5 inputs that should fail safely).
**Pass threshold** (recommended overall score).

Rules:
- Each criterion must be observable β€” graders should agree.
- No "correctness" without saying how you'd verify it.

Example input

Task: an LLM extracts the invoice number, total amount and due date from a PDF.

Or run it here

79/8000
Output will appear here after you click Run.
Or run in your favourite chatbot

Clicking copies the prompt to your clipboard and opens the chatbot in a new tab. Gemini doesn't accept URL params β€” paste manually with Ctrl/Cmd+V.

Browse the full free AI prompt library or see more ai / ml prompts.