Golden Set
A fixed set of real tasks whose correct outputs have already been approved by a human, re-run after every prompt edit or model change so the difference between the old and new output shows whether the change actually helped instead of just feeling better.
intermediate