SprintFlow runbooks#
Each procedure: symptoms → check → act → verify. Commands assume kubectl -n sprintflow for Kubernetes
installs; use the equivalent docker compose / CLI commands on a single host.
1. Stop everything now (incident, release freeze, suspected misbehaviour)#
- Act: console → Health → Pause all automation (manager/admin), with a reason. Or from a shell:
python -c "from sprintflow.control import set_paused; set_paused('<runs_dir>', True, 'oncall', 'reason')". - Effect: no new runs, no sprint queueing, no PR follow-ups, webhooks ignored; runs in progress finish their current stage. A banner shows on every page. SprintFlow never merges, so nothing reaches main without a person.
- Resume: Health → Resume automation. Both actions are audited and alerted.
2. Jira, GitLab/GitHub or Claude is unreachable#
- Symptoms: supervisor alert "
<name>is unreachable", runs failing with network errors, Health shows a red check. - Check:
sprintflow pilot check(or Connections → Test connection) shows which permission or endpoint fails; status pages of Atlassian / your Git host / AWS Bedrock; egress rules (network policy allows 443). - Act: nothing, for short outages: transient failures auto-resume with backoff, and the supervisor resumes interrupted runs. For long ones, pause automation (1) so runs do not pile up failures.
- Verify: the "reachable again" alert arrives; failed runs → Resume on their run page.
3. Budget used up#
- Symptoms: alert "Daily/Monthly budget used up", runs stopped with
budget_exhausted, red banner. - Check: Health → Budget bars; Team page → spend by project/person; Runs sorted by cost for outliers.
- Act: if expected, raise
budgets(Settings → Budgets) or wait for the period to reset. If not, look for a runaway: a story looping on fixes (Timeline tab), a too-expensive model for simple steps (Settings → Model routing). - Verify: stopped runs keep their work: Resume continues from the last completed stage.
4. Stuck or failed runs#
- Symptoms: supervisor alert "run is stuck and needs a person", or a run in
failed. - Check: run page → Events tab (last event), Timeline (which step), Files → logs.
sprintflow status <run_id>. - Act: transient cause → Resume. Code/test problem → read the report, fix the ticket or
.sprintflow/config.yml, resume or start a new run. Worker died → the queue retries the job automatically (up tomax_job_attempts). - Verify: the run reaches a draft PR or a clear stop with a reason.
5. Alerts are not arriving#
- Check: Health → Recent alerts shows
sent to …ornot delivered; Settings → Alerts → Send test alert. - Act: Slack/Teams webhook revoked → create a new one and paste it in Settings → Alerts (stored encrypted).
Email → SMTP host/port/STARTTLS and the SMTP password. Check that the event is not filtered by the channel's
eventsormin_severity.
6. Scaling#
- Queue backing up (
sprintflow_sprint_queue, jobsqueuedgrowing): add workers — the HPA does this on CPU; raisemaxReplicasor--concurrency. Workers finish their jobs before stopping (grace period 1 h). - Console slow: add console replicas (safe: only the lease holder runs background work).
- Database: PostgreSQL connections ≈ (consoles + workers × concurrency) × 5; size
max_connectionsaccordingly.
7. Locked out of the console (SSO outage or misconfiguration)#
- Act: sign in with a break-glass account (
sso.break_glass_users, password login). If none exists:sprintflow console user add rescue --role adminon a host with the console key and database access, then fix Settings → Single sign-on (issuer, client id/secret, redirect URL, role map). - Verify: SSO sign-in works; remove or rotate the rescue account; audit log shows
login_sso.
8. Rotating keys and tokens#
- Console master key (encrypts stored secrets and signs sessions):
1.
sprintflow console init-key > new.key2.sprintflow console rotate-key --new-key-file new.key3. pointSPRINTFLOW_CONSOLE_KEY_FILE/_REF(and the Kubernetes Secret) at the new key; restart consoles and workers; everyone signs in again. Keep the old key until you have verified, then destroy it. - Jira / Git host / Anthropic / webhook secrets: create the new credential, paste it in Connections/Settings
(or update the secrets-manager entry — cached values refresh within
SPRINTFLOW_SECRET_TTL, default 5 min), test the connection, then revoke the old one.
9. Backup and restore#
- Backup (daily):
sprintflow backup create --out /backups/sprintflow-$(date +%F).tar.gz;sprintflow backup verify …. Store the console master key separately (it is not in the backup, by design). - Restore: stop consoles/workers, then
sprintflow backup restore <archive> --runs-dir … --console-dir …(withSPRINTFLOW_DATABASE_URLset to restore database rows). It verifies checksums first and refuses to overwrite existing data unless--force. Start consoles/workers; open a few runs to check.
10. Upgrading SprintFlow#
- Read the release notes; take a backup (9).
- Rolling upgrade: consoles (
maxUnavailable: 0) then workers (they drain). Run state is versioned: a new release migrates older runs as they load. - Rollback: allowed while no run has been written by the new version. If one has, the older release refuses to open it ("written by a newer SprintFlow") instead of corrupting it — roll forward, or restore the pre-upgrade backup.
AI SprintFlow 1.0.5 · Questions? Contact us