Ga naar hoofdinhoud

2 berichten getagd met "Evals"

Laat alle tags zien

Hermiq skills tutorial series — Part 3: Measuring skills with paired evals

· 9 minuten leestijd
Conduction
Open-source workspace stack

A skill that "feels right" can still add nothing, or make the agent worse. Paired evals answer the question with two runs of the same cases: one with the skill, one without. In this part you use the seeded woo-triage-paired-eval dataset to run a paired baseline against woo-request-triage, read the delta, and watch the L5 evidence appear where Part 2 left a named gap.

Claude Skills leerlijn — Deel 3: Skill-evals — meten of je skill werkt

· 13 minuten leestijd
Conduction
Open-source workspace stack

Je hebt nu een werkende skill — maar werkt hij ook écht goed? En blijft hij goed werken als Claude zelf een upgrade krijgt of als je de skill aanpast? Dit derde, optionele deel laat zien hoe je een skill systematisch evalueert: test-scenario's, trigger-tests, een baseline-meting, en de eval-runner die /skill-creator voor je inricht. Dit is de stap van Maturity Level 4 ("voelt goed") naar Level 5 ("gemeten goed").