Test set
Use your users, language and failure modes
Include common, difficult and adversarial cases. Cover English, Chinese and Singapore-specific terminology where the service requires it.
Scorecard
Separate quality from risk
Score usefulness and factuality separately from privacy leakage, unsafe action, bias and refusal quality.
Open responsible source ↗Evidence
Keep the prompt, output and reviewer decision
A benchmark score without traceable cases is weak evidence. Preserve the evaluation version and the reason for each critical decision.
Lifecycle
Treat a model change as a system change
Retest after prompt, retrieval, model, tool, policy or data changes. Monitor production failures that the test set missed.
FREQUENTLY ASKED
Two points worth making clear
How many prompts are enough?+
There is no universal number. Coverage of important tasks and failure modes matters more than a large random set.
Can public benchmarks choose a vendor?+
They are useful context, not a procurement decision. Your data, workflow, controls and support requirements can change the result.