Skip to main content
AcademytutorialHydra tutorial series — Part 3: Quality gates

Hydra tutorial series — Part 3: Quality gates

What are Hydra's mechanical quality gates, why do they deliberately NOT rely on AI judgement, and what do you do when a gate misfires? The third of seven short modules.

TutorialHydraGatesQualityTutorial series
15 min read

In part 2 you saw that the build lane runs the hydra-gates suite on every build, and that a review stage whose verdict doesn't show it ran those gates is not accepted. These are mechanical quality gates — checks that pass or fail deterministically, without AI in the loop. This part explains which gates we have, why they are mechanical, and how to handle the exception: the false positive.

Why mechanical gates?

AI review scales, but it's not predictable. Two runs on exactly the same diff can produce different findings. And AI is especially weak at the boring checking that doesn't require judgement — for example: "does every new PHP file have an SPDX licence header at the top?". You can do that kind of check much faster and cheaper with a script.

The rule inside Hydra:

Whatever can be checked objectively, a mechanical gate handles — one script that passes or fails per check. Only where judgement is required do we bring in a reviewer.

Concretely: before Juan Claude or Clyde spend expensive model time on a PR, the build lane has already run the gates and put their output in the PR body. The reviewers then focus on what a tool can't catch.

Category 1: generic code-quality tools

Every Conduction app carries its own quality suite (composer check:strict on the PHP side, plus the frontend linters). For a local check, hydra's scripts/run-quality.sh runs these in a Docker php:X.Y-cli container:

ToolWhat it catches
lintSyntax errors in PHP.
phpcsCoding standard (PSR-12 + Nextcloud convention).
phpmdCode-mess detector (overlong methods, deep nesting, dead code).
psalmStatic type analysis, as configured in the app.
phpstanSecond static type analyser (catches things Psalm misses and vice versa).
phpmetricsComplexity metrics (cyclomatic, maintainability index).
composer auditCVE check on composer.lock dependencies.
eslintJS/TS lint.
stylelintCSS/SCSS lint.
npm auditCVE check on package-lock.json.
PHPUnitUnit + integration tests.
NewmanAPI tests.

Where they run matters. The build lane does not run them: it runs only the hydra-gates suite, and no flow calls scripts/run-quality.sh. The reviewer's brief asks Juan Claude to run the app's suite and count only findings on files the PR touches. So a build:pass means "the gates are green", not "PHPUnit is green".

Category 2: Hydra-specific gates

On top of the generic tools, Hydra has its own gate suite for things that are Conduction-specific. The gate logic lives in the public conduction/hydra-gates package, in the hydra-gates/ folder of ConductionNL/.github. Hydra's scripts/run-hydra-gates.sh is a delegator: it finds that package and hands over, with the same flags and exit codes. It holds no gate logic of its own, so a gate fix goes into the package, not into hydra.

The hydra-gate-* skills in hydra/.claude/skills/ (about 60 of them) describe individual gates for the agents; the hydra-gates skill runs the whole suite and summarises the failures.

How many gates are there? The runner declares about 96, numbered up to 116 with gaps, and the number changes as gates are added. Don't trust a total written down anywhere, including here: every run prints its own count.

[hydra-gates] COVERAGE: N of M declared gates reported a result (K not applicable to this repo/diff; N of L applicable gates ran).

A clean run ends with ALL L APPLICABLE GATES GREEN — and all L of them ran. Read the COVERAGE line: a gate that was skipped (for example because a tool was missing) is not a pass. The exit code is the number of failures; 0 is the only pass, and 99 means the gates could not run at all.

The first 22 gates give a good picture of the kind of thing the suite checks:

#GateWhat it checks
1spdx-headersEvery lib/**/*.php file carries an EUPL-1.2 SPDX licence header.
2forbidden-patternsNo var_dump / die / exit / error_log / print_r / dd / dump left behind in lib/.
3stub-scanNo "In a complete implementation" stubs, empty run() bodies, or auth methods that accept a caller-identity argument but never use it.
4composer-auditNo known CVEs or advisories in the composer.lock dependency tree.
5route-authEvery controller method in appinfo/routes.php declares its auth posture (#[PublicPage] / #[NoAdminRequired] / #[NoCSRFRequired] / #[AuthorizedAdminSetting]).
6orphan-authPublic is*/requires*/validate*/authorize*/check*/ensure*/verify*/assert* methods have at least one caller — no dead auth code.
7no-admin-idorEvery #[NoAdminRequired] method has a per-object or admin authorization guard in its body (blocks IDOR).
8unsafe-auth-resolverNo catch (\Throwable) { return null; } fail-open pattern in auth / permission / role resolvers.
9semantic-authThe auth annotation matches the method body's actual authorization requirement (not just syntactically present).
10initial-stateFrontend uses loadState() from @nextcloud/initial-state, never getElementById(...).dataset.*.
11admin-routerAdmin-settings Vue components are NOT registered as in-app vue-router routes (they render via AdminSettings.php).
12nc-input-labelsEvery <NcSelect> declares an inputLabel or ariaLabelCombobox (WCAG 2.1 AA).
13modal-isolation<NcModal> / <NcDialog> markup lives in its own src/modals/ or src/dialogs/ file, never inline in a parent.
14route-reachabilityEvery Response-returning controller method is registered in appinfo/routes.php, and every route resolves to an existing method.
15dashboard-antipatternNo dashboard-in-dashboard (<CnDashboardPage>) nesting.
16spec-coverageEvery changed public/protected + frontend method carries @spec openspec/... or @spec exclude <reason> (diff-scoped, ADR-020).
17redundant-controllerNo pass-through CRUD methods that just wrap OpenRegister's ObjectService (ADR-022 — apps consume abstractions).
18notification-dialectRegister files use the canonical x-openregister-notifications dialect (ADR-031); the legacy dialect FAILs, imperative dispatch WARNs.
19e2e-coverageEvery Scenario added/changed in an openspec spec is referenced by a Playwright test via @e2e, or carries @e2e exclude <reason> (diff-scoped, ADR-020).
20or-objectservice-apiNo calls to fabricated ObjectService methods — only find / findAll / saveObject / createObject / updateObject / deleteObject exist.
21conflict-markersNo unresolved git merge markers (<<<<<<<, =======, >>>>>>>) committed.
22manifest-validationsrc/manifest.json validates against the vendored app-manifest schema.

The later gates cover accessibility (WCAG 2.2 AA), manifest and register cross-references, supply-chain configuration, and more. Many gates were born out of incidents: a gate arises in response to a bug that slipped through all earlier checks. Take the caller-identity rule in stub-scan (gate-3). In the decidesk-44-45 retrospective, a builder satisfied gate-7 by calling an authorisation method and then creating that method as an empty stub. Gate-7 only checked that the call existed; the security reviewer caught the empty body. stub-scan now flags a method that accepts a caller identity and never uses it.

Gate-19: where specs meet the browser

Two gates in the list above don't check style or security — they check traceability, and they're the mechanical heart of spec-driven development:

  • spec-coverage (gate-16) enforces the backward link: every method you add or change must point at the requirement it implements with a @spec openspec/... annotation.
  • e2e-coverage (gate-19) enforces the forward link: every Scenario you add or change in an OpenSpec spec must be exercised by a Playwright test, tied back with an @e2e annotation — or explicitly excused with @e2e exclude <reason> (for pure-backend behaviour, which belongs in Newman or PHPUnit instead).

Together they make the loop auditable in both directions: scenario → @e2e → Playwright test, and code → @spec → requirement. A builder can't quietly ship a feature that isn't specified, and can't quietly leave a specified UI scenario untested. Both gates are diff-scoped (ADR-020), so legacy debt in untouched files never blocks a PR.

The exact annotation mechanics — how a #### Scenario: heading becomes a slug, the long vs. short @e2e forms, and the three exclusion scopes — live in the OpenSpec series: From scenario to Playwright test. That's the page to read if you want to understand why this is a safe sandbox for AI rather than free-for-all "vibe coding".

ADR-020: gate scope is the PR diff

An important rule you already brushed against in part 2: gates run on the PR diff, not on the whole repo. That's in ADR-020.

Why? Because many of our repos carry a backlog of technical debt. If you turn a gate loose on the whole repo you get hundreds of findings that have nothing to do with the current PR — and every PR ends up red. By limiting the scope to the PR diff, the gate only fires on what the current build just added or changed.

The flows hard-code this: the build lane calls the runner with --scope-to-diff --base origin/<base>. The override HYDRA_REVIEW_SCOPE=full only affects local and manual runs (scripts/dev-run.sh, scripts/manual-review.sh), where it flips the reviewer's scope to the whole repo. Use it when you onboard a new repo or do a dedicated tech-debt sweep. Expect a lot of red.

Recognising false positives

Mechanical gates are deterministic, but not always correct. A real example from gate-2 (forbidden-patterns): it used to be a set of text searches for \bdd\(, \bvar_dump\( and friends over the raw bytes of lib/. That flagged a comment warning against var_dump(, and a string literal such as "select dd(x)", as uses. It missed real ones too, such as var_dump ($value) with a space, or die; without brackets. One detection choice, a repeatable false positive in both directions.

How to recognise a false positive:

  1. The gate keeps firing on the same line on every build, while the line itself looks innocuous.
  2. Running the gate locally (scripts/run-hydra-gates.sh, and the gate's log) confirms it: yes, the pattern matches, but not for the intended reason.
  3. Nobody on the team can explain why this specific line should be complained about.

In that case the fix is not to re-queue the build and hope. The fix is to repair the detection, in the conduction/hydra-gates package (ConductionNL/.github, hydra-gates/). For gate-2 that fix was to judge code, not text: the file is now read with comment and string contents blanked out before the patterns are matched.

Where the gates run again

The gates run in the build lane, over the pushed branch. If they fail, there is one fix pass, and then the gates run once more over the fixed tree. That second run is the only automatic re-run of the gates.

There is no "quality recheck" after review in the flows. A reviewer's bounded fix is not re-gated by the pipeline; instead, each review stage has to run the gates itself and name hydra-gates in the checks_run of its verdict, or the stage ends in needs-input.

Test yourself

Four short questions to check whether you've understood this part. Stuck? Click Hint. Curious about the answer? Click Answer.

1. Why does Hydra have mechanical gates and AI reviewers, instead of just one of the two?

Hint

One kind of check is predictable and cheap, the other is expensive but can judge. What's the strength and weakness of each?

Answer

They complement each other exactly where the other is weak.

  • Mechanical gates are deterministic: same input → same outcome, pass or fail. Perfect for boring, objective checking — for example "does every new PHP file have an SPDX header?". Cheap and repeatable.
  • AI reviewers do judgement: "does this authorisation logic semantically match the route's purpose", "is this a security risk in this context". Not predictable and more expensive, so you deploy them where judgement is needed.

Mechanical-only misses context-dependent errors; AI-only is expensive, not predictable, and wastes model time on things a script can do.

2. What does ADR-020 say about gate scope, and where does HYDRA_REVIEW_SCOPE=full have an effect?

Hint

Think about what happens when you turn a gate loose on a repo with lots of old technical debt. And do the flows read that variable?

Answer

ADR-020 says: gates run on the PR diff, not on the whole repo.

Reason: many repos drag along technical debt. A gate across all of it produces hundreds of findings that have nothing to do with the current PR — every PR would be red. By checking only the diff, the gate fires on what the build just added or changed.

The flows hard-code --scope-to-diff. HYDRA_REVIEW_SCOPE=full only changes local and manual runs (dev-run.sh, manual-review.sh), where the review then spans the whole repo. Use it for:

  • Onboarding a new repo into Hydra.
  • A dedicated tech-debt sweep.

NOT for regular PRs — expect a lot of red.

3. How do you recognise a false-positive gate, and what's the right fix? What is NOT?

Hint

Three signals together point at a false positive. And the "tempting but wrong" reflex is doing the same thing again.

Answer

Recognition — three signals together:

  1. The gate fires on the same line on every build, while that line looks innocuous.
  2. Running the gate locally confirms: the pattern matches, but not for the intended reason (e.g. a text search that also matches a comment or a string literal).
  3. Nobody on the team can explain why this specific line should be complained about.

Right fix: repair the gate detection in the conduction/hydra-gates package (ConductionNL/.github, hydra-gates/). Hydra's scripts/run-hydra-gates.sh only delegates to it.

NOT: re-queueing the build. That reproduces the same false positive and wastes a run.

4. Where does Hydra re-run the gates automatically, and where does it NOT?

Hint

Think about the build lane's fix pass, and about what happens after a reviewer's bounded fix.

Answer
  • Yes: in the build lane. When the first gate run fails, there is one fix pass, and then the gates run once more over the fixed tree. That result decides build:pass or build:fail.
  • No: after review. There is no post-review "quality recheck" in the flows. Instead each review stage must run the gates itself and list hydra-gates in checks_run; a verdict that doesn't is not accepted and the issue lands on needs-input.

Next step

In part 4 we look at the skills and commands the personas use during their work — including how you can add a new skill yourself.