Hermiq skills tutorial series — Part 3: Measuring skills with paired evals
· 9 minuten leestijd
A skill that "feels right" can still add nothing, or make the agent worse. Paired
evals answer the question with two runs of the same cases: one with the skill, one
without. In this part you use the seeded woo-triage-paired-eval dataset to run a
paired baseline against woo-request-triage, read the delta, and watch the L5
evidence appear where Part 2 left a named gap.