Files
DevRunbook-Public/docs/12-quality-evaluation.md
T
DevRunbook release export cfd2804e27
Managed validation / full (push) Successful in 3m18s
Publish DevRunbook source
2026-09-03 04:09:17 +02:00

5.3 KiB

12 — Quality, linting and evaluation

Quality philosophy

A playbook is trustworthy when its structure, safety, clarity and observed behavior are evidenced. Popularity, length and confident wording are not quality proof.

Quality dimensions

Scope clarity

Does the playbook define the target, permitted reading scope, modification scope, exclusions and protected behavior?

Safety

Does it prevent destructive, secret-exposing or policy-weakening behavior appropriate to its risk?

Verification

Does it require relevant checks and evidence rather than a narrative claim?

Reproducibility

Can the same package version and normalized inputs reproduce the same output?

Compatibility

Are supported stacks, required profile capabilities and limitations explicit?

Reporting

Does the final report make outcome, evidence, risks and unresolved items reviewable?

Efficiency

Does the task avoid unnecessary repository-wide change, redundant investigation and repeated validation without sacrificing safety?

Scores are displayed separately on a defined scale, such as Not assessed, Weak, Adequate, Strong. A combined score may be calculated for sorting but must never replace the dimensions.

Lifecycle policy

Draft

  • schema-valid or actively being edited;
  • may have lint errors;
  • not recommended outside author workspace.

Reviewed

  • schema and semantic validation pass;
  • no blocking prompt-lint findings on required examples;
  • human editorial review complete;
  • limitations documented.

Validated

  • Reviewed requirements;
  • required evaluation cases pass;
  • evaluation environment and fixture version recorded;
  • no unresolved safety regression.

Battle-tested

  • Validated requirements;
  • minimum real-world run count under policy;
  • acceptable operator feedback and failure rate;
  • no unaddressed severe incident;
  • evidence remains recent enough for the playbook class.

Deprecated

  • replacement or rationale provided;
  • historical rendering remains available;
  • excluded from default recommendations.

Static linter catalog

Suggested rule IDs:

  • PB001 missing mission
  • PB002 missing explicit scope
  • PB003 missing done-when criteria
  • PB004 missing reporting contract
  • PB005 duplicate input or step ID
  • PB006 invalid autonomy range
  • PB007 unknown template variable
  • PB008 published version without changelog
  • PB009 validated status without evidence
  • PR001 ambiguous unbounded improvement language
  • PR002 conflicting read-only and modification instruction
  • PR003 vague “best practices” without dimensions
  • PR004 asks for success claim without evidence
  • SA001 token-like value in rendered output
  • SA002 protected path in change scope
  • SA003 destructive migration without rollback/backup
  • SA004 Git push/release not explicitly authorized
  • SA005 imported content placed in policy section
  • VA001 bugfix lacks reproduction/regression step
  • VA002 implementation lacks available build/test validation
  • VA003 dependency change lacks lockfile/install/build check
  • VA004 frontend flow lacks browser verification
  • VA005 inspect playbook lacks evidence-source reporting

Every finding includes severity, location, message, rationale, remediation hint and provenance.

Evaluation case format

An evaluation case defines:

  • case ID and version;
  • target playbook/version range;
  • fixture repository reference and digest;
  • repository profile;
  • inputs and autonomy;
  • expected prompt properties;
  • prohibited prompt properties;
  • optional future execution expectations;
  • scoring rubric.

MVP can evaluate rendering and lint behavior without executing Codex. Future isolated evaluation can run tasks in disposable fixture environments.

Static evaluation examples

  • generated prompt contains all required headings;
  • no secret fixture value appears;
  • protected paths are rendered as no-change constraints;
  • selected observe autonomy contains no implementation permission;
  • missing test command creates warning rather than invented command;
  • stack incompatibility blocks export;
  • same input produces same digest;
  • a conditional migration section appears only when migration input is true.

Future execution evaluation

Fixture repositories intentionally contain known problems. Evaluation runner captures:

  • task completion status;
  • changed file set;
  • protected-path violations;
  • command exit codes;
  • tests added or changed;
  • artifact diffs;
  • final report completeness;
  • token/time/cost metadata where available;
  • human review.

The runner must use isolated disposable environments and must never execute untrusted playbooks on the DevRunbook application host.

Regression policy

A new playbook version compares against the previous version on common evaluation cases. Publication UI highlights:

  • newly passing cases;
  • newly failing cases;
  • meaningful prompt diffs;
  • new permissions or wider scope;
  • changed required inputs;
  • validation reductions.

Safety or validation regressions block promotion to Validated.

User feedback

Feedback fields:

  • task was understandable;
  • Codex stayed within scope;
  • validation was sufficient;
  • follow-up prompts were required;
  • result solved the intended problem;
  • free-form note;
  • optional execution evidence.

Feedback is not silently converted into evaluation evidence. It is a separate signal with abuse and privacy controls in future community features.