Prompt Engineering · 8 min

Prompt Evaluation: Test Before You Trust

Prompt quality should be judged by output quality against a defined standard. A prompt is successful when it reliably produces an acceptable result for its intended use.

Create a test set

Use several representative tasks rather than one easy example. Include normal cases and important edge cases.

Score the output

Evaluate factual accuracy, completeness, relevance, format compliance, tone and risk. Record failures so the prompt can be improved systematically.

Iterate deliberately

Change one important variable at a time when possible. This makes improvement easier to attribute and repeat.

← Back to APEP Knowledge Hub