270 lines
9.5 KiB
Markdown
270 lines
9.5 KiB
Markdown
# 07 — Playbook Package specification
|
|
|
|
## Purpose
|
|
|
|
A Playbook Package is the portable, Git-reviewable unit of reusable development guidance. It combines machine-readable metadata with human-readable instruction content, examples, optional evaluation cases and a changelog.
|
|
|
|
The package is designed to support:
|
|
|
|
- deterministic rendering;
|
|
- schema validation;
|
|
- semantic versioning;
|
|
- stack and repository compatibility;
|
|
- explicit guardrails and completion criteria;
|
|
- future export as a Codex Skill without making Skills the internal storage model;
|
|
- human review through ordinary Git diffs.
|
|
|
|
## Directory layout
|
|
|
|
```text
|
|
my-playbook/
|
|
playbook.yaml required canonical manifest
|
|
prompt.md required detailed task instructions
|
|
README.md recommended author documentation
|
|
CHANGELOG.md required for published versions
|
|
examples/
|
|
minimal.yaml optional composition input examples
|
|
repository-aware.yaml
|
|
evaluations/
|
|
case-basic.yaml optional evaluation definitions
|
|
resources/ optional non-executable supporting files
|
|
scripts/ prohibited in MVP built-in runtime; reserved for future Skill export
|
|
```
|
|
|
|
Every file other than `playbook.yaml` is declared in `package.files` with a role, digest participation and default-export flag. Undeclared files, symlinks and non-regular files are rejected during import and packaging.
|
|
|
|
## Identity
|
|
|
|
- `metadata.id` is a stable reverse-domain-style logical ID, for example `audit.repository-health`.
|
|
- `metadata.slug` is a stable URL slug.
|
|
- `metadata.version` follows semantic versioning.
|
|
- Changing title text alone may be a patch version.
|
|
- Changing required inputs, safety behavior or output contract normally requires a minor version.
|
|
- Removing supported behavior or changing the meaning of existing inputs requires a major version.
|
|
|
|
Published versions are immutable. A corrected package receives a new version, even when the change appears editorial, because historical generated runs must remain reproducible.
|
|
|
|
## Lifecycle and quality status
|
|
|
|
Lifecycle values:
|
|
|
|
- `draft` — editable, not presented as generally ready;
|
|
- `reviewed` — structurally and editorially reviewed;
|
|
- `validated` — required evaluation evidence passes;
|
|
- `battle-tested` — validated plus sustained real-world evidence under the configured policy;
|
|
- `deprecated` — retained for historical runs but not recommended.
|
|
|
|
Lifecycle is not inferred from popularity.
|
|
|
|
## Playbook types
|
|
|
|
### `quick`
|
|
|
|
A compact task requiring few inputs and limited branching. Still contains scope and reporting behavior where relevant.
|
|
|
|
### `guided`
|
|
|
|
A form-driven playbook with repository context, autonomy and conditional sections.
|
|
|
|
### `run-pack`
|
|
|
|
A larger procedure intended to export several files, such as specification, plan, implementation rules and handoff template.
|
|
|
|
## Work modes
|
|
|
|
A package declares `modes` and one `defaultMode`, which must be present in `modes`.
|
|
|
|
- `inspect` — read-only evidence collection;
|
|
- `plan` — investigation and implementation plan, no code changes;
|
|
- `guided` — implementation with declared human checkpoints;
|
|
- `execute` — implementation and validation inside defined boundaries;
|
|
- `recovery` — diagnose and repair a failed or partial implementation.
|
|
|
|
## Autonomy levels
|
|
|
|
Ordered values:
|
|
|
|
1. `observe`
|
|
2. `diagnose`
|
|
3. `plan`
|
|
4. `implement`
|
|
5. `verify`
|
|
6. `repair`
|
|
|
|
A playbook declares minimum, maximum and default. The platform rejects a selected level outside this range.
|
|
|
|
### Behavioral contract
|
|
|
|
| Level | Changes | Validation | Failure behavior |
|
|
|---|---|---|---|
|
|
| Observe | None | Evidence checks | Report unknowns |
|
|
| Diagnose | None by default | Reproduction/analysis | Identify next evidence |
|
|
| Plan | Documentation/plan only if allowed | Plan consistency | Stop before implementation |
|
|
| Implement | Code/config in scope | Relevant targeted checks | Report failed checks |
|
|
| Verify | Code/config in scope | Full declared checks | Repair direct regressions when safe |
|
|
| Repair | Code/config in scope | Iterative full checks | Continue until done or genuine blocker |
|
|
|
|
The exact generated language comes from platform policy plus playbook constraints.
|
|
|
|
## Inputs
|
|
|
|
Each input declares:
|
|
|
|
- stable `key`;
|
|
- human label and description;
|
|
- type;
|
|
- required state;
|
|
- default where safe;
|
|
- validation constraints;
|
|
- optional declarative visibility condition;
|
|
- whether the value may be included in output;
|
|
- whether the field can contain sensitive data.
|
|
|
|
Supported MVP types:
|
|
|
|
- string;
|
|
- multiline;
|
|
- boolean;
|
|
- integer;
|
|
- enum;
|
|
- multiselect;
|
|
- path;
|
|
- command;
|
|
- string-list;
|
|
- key-value-list.
|
|
|
|
Secrets are not ordinary playbook inputs. A playbook may ask whether a credential-dependent validation is available, but it must not solicit or embed secret values in generated prompts.
|
|
|
|
## Compatibility
|
|
|
|
Compatibility is advisory and enforceable where declared. Capability identifiers come from the governed vocabulary in `schemas/playbook.schema.json`; `test-command` is satisfied by at least one confirmed test command:
|
|
|
|
- languages;
|
|
- frameworks;
|
|
- package managers;
|
|
- database technologies;
|
|
- deployment types;
|
|
- repository required/optional;
|
|
- required profile capabilities, such as a test command;
|
|
- known incompatible conditions.
|
|
|
|
A playbook can be stack-neutral. Empty lists mean no restriction, not unknown.
|
|
|
|
## Guardrails
|
|
|
|
Guardrails are structured and rendered in a platform-controlled section. Each package guardrail has:
|
|
|
|
- stable ID;
|
|
- severity: info, warning, blocking;
|
|
- instruction text;
|
|
- optional declarative condition;
|
|
- optional rationale.
|
|
|
|
The composition provenance records package guardrails as playbook-sourced; platform policy is stored and rendered separately.
|
|
|
|
Built-in platform guardrails always outrank package instructions. A package cannot opt out of secret redaction, archive safety or untrusted-context boundaries.
|
|
|
|
## Workflow
|
|
|
|
Workflow steps define intended order and can include conditions. Steps must be outcome-oriented and testable. Avoid micro-managing exact file edits when the repository may vary.
|
|
|
|
Good:
|
|
|
|
> Reproduce the reported failure using the narrowest existing test or a new focused regression test before changing production logic.
|
|
|
|
Weak:
|
|
|
|
> Open the code and fix the bug carefully.
|
|
|
|
## Validation
|
|
|
|
Validation consists of:
|
|
|
|
- command roles to resolve from the repository profile, such as `lint`, `typecheck`, `test`, `build`, `smoke`;
|
|
- explicit checks;
|
|
- conditional checks based on changed areas;
|
|
- blocking or advisory status;
|
|
- evidence requirements.
|
|
|
|
A command role does not hardcode a project-specific command in a generic playbook. The profile supplies the command.
|
|
|
|
## Completion contract
|
|
|
|
Every non-trivial playbook defines observable criteria. Examples:
|
|
|
|
- issue reproduced before modification;
|
|
- regression test demonstrates the fix;
|
|
- no protected path changed;
|
|
- declared commands pass;
|
|
- documentation matches implemented behavior;
|
|
- unresolved risks are explicitly reported.
|
|
|
|
## Failure policy
|
|
|
|
The package defines behavior for:
|
|
|
|
- validation failure;
|
|
- incomplete repository context;
|
|
- ambiguous requirements;
|
|
- unavailable external dependency;
|
|
- detected out-of-scope root cause;
|
|
- inability to reproduce.
|
|
|
|
The policy must not instruct the agent to hide failure, weaken checks or invent evidence.
|
|
|
|
## Reporting contract
|
|
|
|
The final report is structured, typically including:
|
|
|
|
- outcome;
|
|
- root cause or findings;
|
|
- changed files or inspected scope;
|
|
- validation evidence;
|
|
- risks and limitations;
|
|
- unresolved items;
|
|
- recommended follow-up.
|
|
|
|
Inspect-only playbooks replace changed-files reporting with evidence sources.
|
|
|
|
## Template rules
|
|
|
|
Conditions never use template expressions. They use the non-executable AST in `docs/28-conditions-and-policy-dsl.md`. `prompt.md` uses a restricted template syntax. It may reference normalized inputs and selected safe profile fields. It cannot read arbitrary filesystem paths, environment variables, integration secrets or database records.
|
|
|
|
Recommended syntax examples:
|
|
|
|
```text
|
|
{{ inputs.problemStatement }}
|
|
{{ repository.displayName }}
|
|
{{#if inputs.preserveCompatibility}}...{{/if}}
|
|
{{#each repository.validationCommands}}...{{/each}}
|
|
```
|
|
|
|
The implementation may use a maintained template engine but must expose only an allowlisted context and disable unsafe helpers or dynamic code execution.
|
|
|
|
## Canonicalization and digest
|
|
|
|
Canonical text normalization, RFC 8785 serialization, package-file inventory, package digest, render digest and Run Pack manifest digest are defined normatively in `docs/29-package-integrity-canonicalization.md`. Implementations must use that algorithm rather than an archive library's file order or YAML serialization.
|
|
|
|
## Semantic validation beyond JSON Schema
|
|
|
|
- semantic version parses correctly;
|
|
- min autonomy is not above max;
|
|
- default autonomy is within range;
|
|
- `defaultMode` is present in `modes`;
|
|
- input keys are unique and match allowed pattern;
|
|
- workflow, guardrail, check and report IDs are unique;
|
|
- template references only known variables;
|
|
- conditions use only declared inputs, safe roots and governed operators;
|
|
- sensitive inputs never set `includeInOutput: true`;
|
|
- enum and multiselect inputs declare options and defaults match their type;
|
|
- all declared files exist, every non-manifest file is declared and paths/roles are unique;
|
|
- published package includes a changelog;
|
|
- `validated` and `battle-tested` statuses meet evidence policy;
|
|
- deprecated package references a replacement when available;
|
|
- no path escapes the package root;
|
|
- no symlinks or executable package content in MVP imports.
|
|
|
|
## Example packages
|
|
|
|
See `examples/playbooks/`. They are normative examples for schema and rendering tests, not merely illustrative text.
|