Files
DevRunbook-Public/content/playbooks/health-readiness/playbook.yaml
T
DevRunbook release export cfd2804e27
Managed validation / full (push) Successful in 3m18s
Publish DevRunbook source
2026-09-03 04:09:17 +02:00

231 lines
8.8 KiB
YAML

apiVersion: devrunbook.io/v1alpha1
kind: Playbook
metadata:
id: release-operations.health-readiness
slug: health-readiness
version: 1.0.0
title: Implement Health and Readiness Checks
summary: Add accurate liveness, readiness and dependency health without hiding partial outages.
category: release-operations
tags:
- healthcheck
- operations
- reliability
lifecycle: reviewed
riskTier: moderate
authors:
- name: DevRunbook Core Team
license: MIT
package:
files:
- path: prompt.md
role: template
digest: true
exportByDefault: false
- path: README.md
role: documentation
digest: true
exportByDefault: false
- path: CHANGELOG.md
role: changelog
digest: true
exportByDefault: false
- path: examples/minimal.yaml
role: example
digest: true
exportByDefault: false
- path: evaluations/static-structure.yaml
role: evaluation
digest: true
exportByDefault: false
spec:
type: guided
intent:
problem: Development work around implement health and readiness checks is often underspecified, inconsistently executed
or reported without enough evidence.
outcome: Add accurate liveness, readiness and dependency health without hiding partial outages.
whenToUse:
- Use this playbook when the repository needs a bounded implement health and readiness checks task with explicit evidence
and completion criteria.
- Use it when Codex should follow a repeatable workflow rather than improvise from a one-line request.
whenNotToUse:
- Do not use it when the desired outcome or authority boundaries are still materially undecided.
- Do not use it to access unavailable production credentials, bypass safeguards or claim validation that cannot be performed.
modes:
- guided
- execute
- recovery
defaultMode: execute
autonomy:
min: implement
max: repair
default: verify
inputs:
- key: requiredDependencies
label: Required dependencies
description: List dependencies that determine readiness and their failure semantics.
type: string-list
required: true
sensitive: false
includeInOutput: true
- key: degradedComponents
label: Degraded components
description: List optional components that may fail without making the whole service unready.
type: string-list
required: false
sensitive: false
includeInOutput: true
default: []
compatibility:
repositoryRequired: true
languages: []
frameworks: []
packageManagers: []
databases: []
deploymentTypes: []
requiredProfileCapabilities: []
incompatibleConditions: []
guardrails:
- id: guardrail-1
severity: blocking
text: Keep liveness independent from optional downstream availability.
- id: guardrail-2
severity: blocking
text: Do not expose secrets, topology details or raw dependency errors in public health responses.
- id: guardrail-3
severity: blocking
text: Avoid health checks that create load or mutate external systems.
workflow:
- id: classify-dependencies
title: Classify dependencies
instruction: Separate process health, required readiness dependencies and optional degraded components.
required: true
- id: define-contract
title: Define endpoint contract
instruction: Specify status codes, response shape, timeouts, caching and authentication/exposure.
required: true
- id: implement-checks
title: Implement checks
instruction: Add bounded checks and aggregate them with clear required/degraded semantics.
required: true
- id: integrate-runtime
title: Integrate runtime
instruction: Configure container healthchecks and startup/shutdown behavior.
required: true
- id: add-observability
title: Add observability
instruction: Emit safe structured logs and metrics for state transitions.
required: true
- id: test-failures
title: Test failure matrix
instruction: Simulate required and optional dependency failures and recovery.
required: true
- id: document
title: Document operations
instruction: Explain how orchestrators and operators should use each endpoint.
required: true
validation:
commandRoles:
- lint
- typecheck
- unit-test
- integration-test
- build
- smoke-test
checks:
- id: check-1
type: assertion
description: Required dependency failure changes readiness without killing liveness.
blocking: true
evidence: Referenced files, command results or explicit review notes.
- id: check-2
type: assertion
description: Optional component failure is visible as degraded according to policy.
blocking: true
evidence: Referenced files, command results or explicit review notes.
- id: command-lint
type: command
description: Run the resolved lint command when the repository profile provides it and record the result.
blocking: true
evidence: Resolved command, exit status and concise result summary.
- id: command-typecheck
type: command
description: Run the resolved typecheck command when the repository profile provides it and record the result.
blocking: true
evidence: Resolved command, exit status and concise result summary.
- id: command-unit-test
type: command
description: Run the resolved unit-test command when the repository profile provides it and record the result.
blocking: true
evidence: Resolved command, exit status and concise result summary.
- id: command-integration-test
type: command
description: Run the resolved integration-test command when the repository profile provides it and record the result.
blocking: true
evidence: Resolved command, exit status and concise result summary.
- id: command-build
type: command
description: Run the resolved build command when the repository profile provides it and record the result.
blocking: true
evidence: Resolved command, exit status and concise result summary.
- id: command-smoke-test
type: command
description: Run the resolved smoke-test command when the repository profile provides it and record the result.
blocking: true
evidence: Resolved command, exit status and concise result summary.
completion:
criteria:
- Orchestrator behavior matches documented semantics.
- Optional integration outages do not misreport total failure.
- Validation evidence and unresolved limitations are reported honestly.
failurePolicy:
onValidationFailure: Investigate failures caused by the current work, repair them when they remain within scope, rerun
affected validation and report any genuine blocker without claiming success.
onAmbiguity: Use repository evidence and existing conventions for minor reversible choices. Preserve current behavior
and stop before any material irreversible decision that the specification does not resolve.
onMissingContext: Inspect the repository for missing non-sensitive context. Never invent commands, credentials, production
behavior or validation results; report what remains unavailable.
onOutOfScopeCause: Explain the evidenced out-of-scope cause, avoid unrelated changes and provide the smallest safe follow-up
recommendation.
onExternalDependencyUnavailable: Use an approved local substitute or fixture only when it preserves the behavior under
test. Otherwise record the blocked validation and do not claim the external path succeeded.
onUnableToReproduce: Record attempted reproduction, environment and observed evidence. Do not apply speculative production
changes; provide the narrowest next diagnostic action.
reporting:
sections:
- id: outcome
title: Outcome
required: true
description: State the delivered result or audit conclusion without overstating evidence.
- id: evidence
title: Evidence and scope
required: true
description: List inspected or changed areas and the evidence supporting the result.
- id: validation
title: Validation
required: true
description: Report commands, manual checks and their actual outcomes.
- id: risks
title: Risks and limitations
required: true
description: State residual risk, inaccessible evidence and untested conditions.
- id: follow-up
title: Recommended follow-up
required: true
description: List the smallest useful next actions or state None.
template:
main: prompt.md
partials: []
exports:
prompt: true
markdown: true
runPack: false
agentsSuggestion: false
quality:
reviewStatus: editorial-reviewed
testedStacks: []
knownLimitations:
- Repository-specific effectiveness depends on the accuracy of the selected profile and the evidence available to Codex.
evaluationCaseIds:
- health-readiness.static-structure