Test set

Use your users, language and failure modes

Include common, difficult and adversarial cases. Cover English, Chinese and Singapore-specific terminology where the service requires it.

Scorecard

Separate quality from risk

Score usefulness and factuality separately from privacy leakage, unsafe action, bias and refusal quality.

Open responsible source ↗

Evidence

Keep the prompt, output and reviewer decision

A benchmark score without traceable cases is weak evidence. Preserve the evaluation version and the reason for each critical decision.

Lifecycle

Treat a model change as a system change

Retest after prompt, retrieval, model, tool, policy or data changes. Monitor production failures that the test set missed.

FREQUENTLY ASKED

Two points worth making clear

How many prompts are enough?+

There is no universal number. Coverage of important tasks and failure modes matters more than a large random set.

Can public benchmarks choose a vendor?+

They are useful context, not a procurement decision. Your data, workflow, controls and support requirements can change the result.

NEXT STEP

Apply the framework to your actual task

Desk AI can help structure evaluation questions and controls. It does not make procurement decisions or retain chat data.