Decide what a good answer must pass
Turn “looks good” into specific checks that reflect the work and its consequences.
Separate correctness from presentation
A well-written answer can still be wrong. Define the required content, disallowed claims, source requirements and output format before evaluating a workflow. Keep factual checks distinct from tone and readability.
Choose cases that represent normal work, missing inputs and conflicting evidence. Preserve a held-out set when comparing changes so repeatedly tuning to the same examples does not masquerade as general improvement.
A small evaluation for a support-answer workflow
The following cases exercise different failure modes instead of repeating the easiest question.
Test the workflow with a policy-supported question, a question the policy does not answer and two conflicting policy versions. A passing response cites the applicable evidence, avoids unauthorized promises and abstains or requests clarification when needed. Check response format separately. Record both the answer and the reason for each pass or failure; do not average away a critical privacy failure.
A score only covers the checks behind it
Prompt Score reviews prompt quality; it does not independently establish the factual correctness of every generated answer. A small evaluation also cannot prove universal reliability. Report the sample, criteria and unresolved failure modes.
Common questions
Can a model grade its own answer?
It can assist a review, but shared errors and grading bias remain possible. Use deterministic checks where appropriate and human or source-based review for important claims.
Sources and review
Reviewed