Skip to main content
AcademytutorialHydra tutorial series — Part 4: Skills

Hydra tutorial series — Part 4: Skills

Which skills run inside the automated Hydra factory, which exist for humans on the CLI, how they get invoked, and when you should add a new skill to the loop yourself. Fourth of seven short modules.

TutorialHydraSkillsOPSXTestingTutorial series
20 min read

This part dives straight into the Hydra-specific skill families. If you'd rather first learn what a Claude Skill even is, how the frontmatter works, and when you'd write one yourself, take the public Claude Skills tutorial series (three short modules, ~40 minutes). From here on we assume you know the basics.

The previous parts were about what Hydra does. This part is about how: the skills that let the personas do their job. By the end you'll know which skills the automated pipeline (the "Hydra factory") runs and which ones you call yourself as a human, you'll know the five families, and you'll be able to judge when a new skill is worth writing.

Skills, in one paragraph

A skill in Claude Code is a folder with a SKILL.md (and optionally scripts, examples/, helpers). The folder sits under .claude/skills/<name>/. The description in the frontmatter tells Claude when the skill is relevant; the contents are the instructions Claude follows once the skill is loaded.

For Hydra, skills are the bundling unit for behaviour: instead of pasting a thousand lines of prompt text straight into a persona's CLAUDE.md, it lives in a skill — reusable and testable.

How a skill gets invoked

A single skill can be triggered in two ways:

  1. Manually — you type /opsx-apply in a Claude session. Direct, predictable, and the skill runs exactly when you want it to.
  2. Automatically — Claude reads the description of every available skill and picks one itself when the current situation matches. Ask "what changed?" and Claude will trigger a skill like summarize-changes on its own if one is available.

Which mode is allowed is set in the skill's own frontmatter:

  • disable-model-invocation: true — only you can call it via /<name>. Use this for skills with side-effects (/opsx-apply modifies code, /create-pr opens a PR). Claude shouldn't just decide to do that on its own.
  • user-invocable: false — only Claude itself may load it. Use this for background knowledge (for example a legacy-system-context that doesn't make sense as an action).
  • Neither set — both are allowed.

So that distinction — manual vs. automatic — is about configuration of one skill, not about two different kinds of files. A pipeline stage has no keyboard: it gets a prompt that says what to do and which brief to follow, and anything else it loads is Claude's own choice. A human on the CLI usually wants explicit control and types /.

Two worlds: the Hydra factory vs. skills for humans

Hydra's hydra/.claude/skills/ holds around 130 skill folders, about 60 of them hydra-gate-* (run ls .claude/skills for today's count). The automated pipeline leans on very few of them. The rest is tooling for you, behind a keyboard. Important to keep separate:

What the Hydra factory actually runs

A real run is driven by the hydra-sequencer flow (part 7). The hermiq app runs each stage as claude -p in a checkout of the target repo, with hydra's own tool tree beside it. The stage's prompt says what to do:

  • Build (Al Gorithm) — the prompt asks for the OpenSpec change named in the issue to be implemented. It does not name a skill. The agent may read and edit files (Read, Edit, Write, Glob, Grep) and is told not to run git; the pipeline commits and pushes. Then the pipeline itself runs scripts/run-hydra-gates.sh --scope-to-diff, allows one fix pass on the gate output, and runs the gates again.
  • Code review, security review, applier (Juan Claude, Clyde Barcode, Axel Pliér) — each reads its own brief from agents/<persona>/ plus the stage brief under images/, runs the gates the way that brief says, and writes a verdict file. A verdict whose checks_run does not name hydra-gates is refused.

The container images are what you start yourself with dev-run.sh, smoke-test.sh or manual-review.sh (part 5). Each has its own, limited set of skills baked in:

ContainerPersonaSkills in imageMain function
BuilderAl GorithmAll of .claude/skills/ (no vendor skills)Implements the change; the image's entrypoint follows opsx-apply. The other opsx skills are on hand for context.
ReviewerJuan Claudehydra-gates + a fixed selection of 16 hydra-gate-* skills + vendor code-reviewMandatory hydra-gates check + deeper code review via the Anthropic community skill.
SecurityClyde Barcodehydra-gates + the same 16 hydra-gate-* skills + vendor trailofbits (its Semgrep plugin) + vendor owaspMandatory hydra-gates + SAST (Semgrep) + OWASP top 10 checklists.
ApplierAxel PliérNo skills; its config allows Read and Bash onlyReads the final diff and both review verdicts and returns one binary answer — pass or fail. It applies no fixes; it's a go/no-go gate, not an editor.

Whichever skill files an image carries, the detection is done by scripts/run-hydra-gates.sh. It is a thin delegator to the conduction/hydra-gates package, which declares about 96 gates today. That number grows with every incident, so don't trust a count written in a doc (this one included): every run prints a COVERAGE line with what it declared and what actually ran. The skill files serve Claude as documentation when fixing; the script does the detection.

What the factory doesn't do — for humans on the CLI

Nearly all the other skills are for you, or for a fellow dev in a Claude Code session on their laptop. Examples:

  • Preparing a change — opsx-new, opsx-ff, opsx-explore, opsx-plan-to-issues are things you do as a human before you throw the work into the pipeline. The pipeline only builds a change that already exists.
  • Taking on a role — team-architect, team-backend, team-po etc. are pure roleplay frames for one-human-one-role sessions. The pipeline doesn't use them.
  • Running tests — /test-counsel (all nine personas) or /test-app (one browser sweep) you start by hand. The pipeline runs no browser at all; the test-* family is yours.
  • PRs and day-to-day work — /create-pr, /review-pr, /report-out are the "three times a day" tools for you, not for the pipeline.

Remember: a pipeline stage follows its prompt, its brief and the gate suite; a human can call every skill in the folder. That's the whole difference.

The five skill families in full

With that split in mind, here are all five families. The tag system in the table: 🤖 factory = used by the pipeline or by a pipeline image's entrypoint, ⌨️ human = you call it yourself. (The builder image copies the whole .claude/skills/ folder, so most ⌨️ skills are present in it; nothing in the pipeline calls them.)

Family 1: OpenSpec workflow (opsx-*, 16 skills)

The OPSX skills implement the Conduction workflow for OpenSpec changes — from proposal to archive.

SkillFor whomDoes
opsx-apply🤖 Builder image + ⌨️Implements tasks from a change. The builder image's entrypoint follows it; the flow's build prompt does not name it.
opsx-verify⌨️Verifies that the implementation covers the change artefacts.
opsx-archive⌨️Archives a completed change, syncing delta into the spec.
opsx-new⌨️Starts a new change (proposal scaffold, schema choice).
opsx-ff⌨️"Fast-forward": creates a change + all artefacts in a single pass.
opsx-continue⌨️Pick up an interrupted change.
opsx-explore⌨️Explore pre-spec what a change should even be.
opsx-onboard⌨️Guided walk through one complete OpenSpec workflow cycle, with narration.
opsx-plan-to-issues⌨️Converts tasks.md into a plan.json and a tracking issue with task checkboxes.
opsx-sync⌨️Sync delta specs back into the main spec.
opsx-bulk-archive⌨️Archive multiple completed changes in one go.
opsx-apply-loop⌨️Runs apply → verify in a loop until verify passes, then archives; per app, in Docker.
opsx-pipeline⌨️Run multiple changes in parallel with subagents.
opsx-coverage-scan⌨️Audit a legacy app for spec ↔ code coverage.
opsx-annotate⌨️Applies @spec PHPDoc tags after a coverage scan.
opsx-reverse-spec⌨️Reverse-engineers a spec from existing code.

Sixteen skills is a lot. In practice you reach for one of a handful, and the choice follows a simple decision aid:

  • Starting fresh and sure of the feature? → opsx-ff (one pass to all artefacts), then opsx-plan-to-issues.
  • Starting fresh but unsure of scope/approach? → opsx-explore first, then opsx-new and opsx-continue step by step.
  • New to the workflow? → opsx-onboard walks you through one full cycle.
  • Legacy code with no specs? → opsx-coverage-scan → opsx-reverse-spec / opsx-annotate.
  • Change is written, want it built locally without the full factory? → opsx-apply-loop (one change) or opsx-pipeline (several in parallel).
  • Inside the factory none of these is called by name: the flow's build prompt implements the change directly. The opsx skills are human-driven preparation and local tooling.

Family 2: Quality + security gates (hydra-gate-*)

The parent skill hydra-gates runs the whole suite through scripts/run-hydra-gates.sh and summarises the failures. The gate logic itself lives in the conduction/hydra-gates package, which that script delegates to; the per-gate skill files serve Claude as reference material when fixing a failure. There are more gates than gate skills, and the numbering (up to gate-116 today) has gaps.

Rather than reproduce a table here, the gates are covered in part 3: Quality gates. For the skill family it's enough to know the gates split into rough groups:

  • Hygiene — spdx, forbidden-patterns, stub-scan, composer-audit, gitignore-then-commit.
  • Security & auth — route-auth, orphan-auth, no-admin-idor, unsafe-auth-resolver, semantic-auth, route-reachability, csrf-cochange.
  • Accessibility (WCAG 2.2 AA) — img-alt, button-name, form-label-association, html-lang, skip-link and more.
  • Conduction conventions & traceability — initial-state, admin-router, nc-input-labels, modal-isolation, dashboard-antipattern, redundant-controller, notification-dialect, manifest-validation, and the two traceability gates spec-coverage (gate-16) and e2e-coverage (gate-19).

In a flow run the pipeline runs the suite after the build and again after the single fix pass, and every review stage is told to run it. When a new incident produces a new gate, it's added here and the count grows — so trust the run's COVERAGE line, not a hard-coded number in any doc.

Family 3: Team roles (team-*, 7 skills) — ⌨️ humans only

Skills that model a specific kind of work via a persona frame. Intended for a human working in a terminal alongside Hydra who wants to step into a single role for a bit.

  • team-po (Product Owner) — writes user stories and acceptance criteria.
  • team-sm (Scrum Master) — manages backlog and sprint planning artefacts.
  • team-architect — makes architecture decisions, writes ADR drafts.
  • team-backend — implements backend work (PHP, services, mappers).
  • team-frontend — implements frontend work (Vue 2, Pinia, NL Design System).
  • team-reviewer — a manual variant of the reviewer work.
  • team-qa — writes test cases and test plans.

The builder image carries them only because it copies the whole folder; nothing in the automated loop calls them.

Family 4: Test suites (test-*) — all ⌨️ for humans

After the gates, the largest family: skills that drive agentic browser and API testing, in three clusters:

Test types (one per test kind):

  • test-app — automated browser test of a whole Nextcloud app (Playwright MCP).
  • test-functional — functional scenarios against implemented features.
  • test-api — API checks (REST, NLGov conventions).
  • test-accessibility, test-performance, test-security, test-regression — specialised variants.

Personas (test-persona-*) — nine Dutch user profiles, each looking at an app from their own angle: annemarie, fatima, henk, janwillem, jasper, mark, noor, priya, sem. Jasper works with a screen reader as his primary way in; Henk reads with large type and looks for simple navigation; Noor hammers on RBAC and audit trails; Annemarie checks NLGov/GEMMA mapping. The persona cards themselves live in hydra/personas/.

Scenario management — test-scenario-create, test-scenario-edit, test-scenario-run write, edit and run reusable TS-NNN-*.md scenarios per app.

The dispatcher for this family is test-counsel: it coordinates all nine personas against one feature and delivers a combined report. Same pattern as hydra-gates for the quality-gates family.

There are three separate "testing" surfaces and they're easy to confuse:

  1. scripts/run-browser-tests.sh — a script that logs into a running app via Playwright MCP, walks the spec's acceptance criteria and returns a verdict JSON. It is runtime proof that the feature works, but only when you run it: no flow calls it, so a pipeline run brings no browser evidence.
  2. gate-19 (e2e-coverage) — a static gate from part 3. It doesn't run a browser at all; it checks that every spec scenario is linked to a Playwright test via an @e2e annotation. Coverage bookkeeping, not execution.
  3. The test-* family — human-driven test suites you run by hand before or after the pipeline.

So: gate-19 checks the links exist and runs in the pipeline; the browser script and the test-* skills actually drive a browser, and both are things you start. Different jobs.

Family 5: Utility & maintenance — all ⌨️ for humans

Everything else. Mostly dev comfort and meta work; a selection:

SkillDoes
create-prCreates a PR from a feature branch — local checks → branch pick → PR body.
review-prReviews a PR (note: manual variant; the factory has its own Juan Claude).
report-outEnd-of-day report: today's commits + GitHub activity → Dutch Slack notification.
clean-envResets the local OpenRegister dev environment (stop, remove volumes, restart, install apps).
local-runRuns builder → reviewer → security locally on a throwaway repo; from Claude Code it does what smoke-test.sh does (part 5).
sync-docsSync {app}/docs/ or .github/docs/claude/ with reality in the repo.
skill-creatorWizard for building a new skill (scaffold, frontmatter, evals).
feature-counselPre-build spec analysis from nine persona perspectives (sibling of test-counsel).
persistence-auditAudits a codebase or architecture for post-compromise persistence vectors: OAuth grants, service accounts, CI/CD secrets, IdP trust, audit trails.
journeydoc-init / journeydoc-add-story / journeydoc-instrumentManual instrumentation + extension of Journey docs.
verify-global-settings-versionCheck whether global-settings/VERSION was bumped after a change.

Beside these sit writing skills (writing, blog-write, tutorial-write), the parity-* skills for capability matrices against competitors, and atom-design for design-system mocks. Nothing in this family runs in the pipeline. They are your daily / commands.

Vendor skills (community)

Next to its own skills, Hydra has vendor skills under hydra/vendor/skills/:

  • code-review — community review skill from Anthropic. → 🤖 Reviewer container.
  • trailofbits — Semgrep-based static-analysis methodology from Trail of Bits. → 🤖 Security container.
  • owasp — OWASP top 10:2025 + ASVS 5.0 checklists. → 🤖 Security container.

So those three do ship inside the reviewer and security images; the builder image carries none of them. They give Juan and Clyde extra coverage on top of our own hydra-gate-*. Maintenance sits with external parties — updates happen by tracking upstream, not by editing them yourself.

When do you write a new skill?

The pragmatic test: write a skill if…

  1. The check / behaviour is repeatable — needed more than once, in more than one place.
  2. It's mechanically describable — you can instruct it in 1-3 paragraphs without it devolving into "it depends on the context".
  3. A persona or a human would benefit from it. Don't write a skill because you can.

For a false-positive gate (part 3): you adjust the existing skill, you don't write a new one. For a new class of mistake that you see come by 3x: yes, that earns its own hydra-gate-* skill.

Test yourself

Four short questions to check whether you've grasped this part. Stuck? Click Hint. Curious about the answer? Click Answer.

1. What are the two ways a skill can be activated, and how do you control that per skill?

Hint

One way requires a human to type something. The other lets Claude decide for itself based on the skill description. Which frontmatter fields determine which mode is allowed?

Answer
  • Manual: you type /<name> in a Claude session. Direct, predictable.
  • Automatic: Claude reads the description of all skills and picks one itself when the current conversation matches.

In the skill's frontmatter you set:

  • disable-model-invocation: true — manual only. Used for skills with side-effects (/opsx-apply, /create-pr).
  • user-invocable: false — automatic only. Used for background knowledge that isn't useful as an action.
  • Both empty → both allowed. The default for most Hydra skills.

In the Hydra pipeline no one types a /: each stage gets a prompt that says what to do and which brief to follow, and any skill it loads beyond that is Claude's own choice. A human on the CLI typically types / explicitly.

2. What does the automated pipeline actually use, which skills do the container images carry, and which ones sit in the repo only for humans?

Hint

Separate the flow-driven run (prompt, brief, gate suite) from the images you start yourself. What does each image carry — and which families sit entirely outside?

Answer

In a flow-driven run (🤖 factory), a stage follows its prompt and brief:

  • Build — the prompt asks for the change to be implemented directly; no skill is named. The pipeline then runs scripts/run-hydra-gates.sh, allows one fix pass, and runs the gates again.
  • Review stages — each persona reads its brief from agents/<persona>/ and images/, runs the gates, and must list hydra-gates in checks_run.

In the container images (what dev-run.sh and smoke-test.sh start):

  • Builder — all of .claude/skills/, no vendor skills. Its entrypoint follows opsx-apply.
  • Reviewer — hydra-gates + 16 individual hydra-gate-* + vendor code-review.
  • Security — hydra-gates + the same 16 hydra-gate-* + vendor trailofbits (Semgrep) + vendor owasp.
  • Applier — no skills, Read and Bash only.

Whatever the image carries, scripts/run-hydra-gates.sh does the detection. It declares far more gates than there are gate skills; the run's COVERAGE line says how many.

Not used by the factory (some are present in the builder image, but no stage calls them), only for humans on the CLI:

  • The entire team-* family.
  • The entire test-* family — the pipeline runs no browser at all.
  • The opsx-* skills (humans prepare changes; the pipeline builds them).
  • The whole utility family (create-pr, review-pr, report-out, clean-env, local-run, sync-docs, skill-creator, feature-counsel, persistence-audit, journeydoc-*, verify-global-settings-version, and the writing and parity skills).

In total: of the roughly 130 skills in the repo, the automated pipeline leans on the gate suite and a brief per persona; the rest is your toolkit.

3. When do you adjust an existing gate skill and when do you write a new one?

Hint

One decision is about "the gate doesn't do what we already wanted it to do". The other is about "we've discovered a new category of mistake".

Answer
  • Adjust existing for a false positive: the gate triggers too broadly or too narrowly on something it was already meant to check. Example: hydra-gate-forbidden-patterns matched $builder->add( incorrectly — you tighten the regex, you don't add a second gate.
  • Write new for a new class of mistake that you see come by 3× and that slips through all existing meshes. Example: hydra-gate-stub-scan was born when a builder shipped a return null; method that PHPCS swallowed and had no test — that was a new category, not a fix on an existing check.

Rule of three: once is chance, twice is coincidence, three times is a pattern that deserves its own gate.

4. What do the "vendor skills" do and why do we keep them separate from our own hydra-gate-*?

Hint

Think about provenance (who wrote them?), which pipeline container they get loaded into, and what happens when you need an external party to update their work.

Answer

Vendor skills (hydra/vendor/skills/) are skills that were not written by us, and that do ship inside the reviewer and security images:

  • code-review (Anthropic community) → Reviewer container (Juan Claude).
  • trailofbits (Trail of Bits, Semgrep methodology) → Security container (Clyde).
  • owasp (OWASP top 10:2025 + ASVS 5.0) → Security container.

We keep them separate because:

  • Maintenance sits with external parties — we update them by tracking upstream (see vendor/skills/VERSIONS.md), not by editing them ourselves. Our own gates we mutate freely; vendor skills we leave alone.
  • Audit trail stays clear: what's ours vs. what's community/external? On a failure you immediately know which camp is responsible for the fix.

Next step

In part 5 we get practical: running the stages on your own machine, and starting a real Hydra run on a real app with a label.