SEO Eval Builder
SEO Eval Builder tests your prompts systematically: you define test cases with requirements (must contain X, max length Y, valid JSON), run the prompt against all tests, and get pass/fail per test plus a quality score per criterion from an AI judge.
The evaluation is real: requirements are checked deterministically against actual AI output, and scoring is done by a separate judge model against your rubric. The "Suggest tests" button is free and gives you a starting point.
Who this tool is for
Anyone building or maintaining prompts — for their own tools, GPTs or workflows — who wants to measure changes instead of guessing.
Prerequisites
- Signed-in account
- 10 credits per eval run, regardless of test count
- A prompt to test, a rubric (presets exist) and at least one test case
How to use it
- Paste the prompt and pick/adjust the rubric — criterion weights must sum to 1.0.
- Add test cases, or use "Suggest tests" (free) as a starting point.
- Run the eval and read pass/fail per test and score per criterion.
- Adjust the prompt and run again — compare scores manually.
How to read the result
- Pass/fail per test are the hard requirements; the 0–10 criterion score is the judge's quality assessment.
- The output excerpt per test shows what the model actually answered — read it when a test surprises you.
Common errors
- "0 of N passed" with error messages on every test: the AI calls themselves failed (not your prompt) — the charge is not auto-refunded today; retry and contact support if it repeats.
- Weights do not sum to 1.0: adjust to exactly 1.0 before running.
Limitations
- No saving, history or export — note scores and setup yourself before closing; comparing runs is manual.
- Regression alerting does not exist today, even though the UI has room for it.
Related tools
- Code Review, for the code side of prompts
- Content Quality Grader, for content quality
- Brutal Nordic Editor, for the texts your eval approves
Last updated 2026-08-06