AI SprintFlow

SprintFlow runbooks#

Each procedure: symptoms → check → act → verify. Commands assume kubectl -n sprintflow for Kubernetes installs; use the equivalent docker compose / CLI commands on a single host.

1. Stop everything now (incident, release freeze, suspected misbehaviour)#

  • Act: console → Health → Pause all automation (manager/admin), with a reason. Or from a shell: python -c "from sprintflow.control import set_paused; set_paused('<runs_dir>', True, 'oncall', 'reason')".
  • Effect: no new runs, no sprint queueing, no PR follow-ups, webhooks ignored; runs in progress finish their current stage. A banner shows on every page. SprintFlow never merges, so nothing reaches main without a person.
  • Resume: Health → Resume automation. Both actions are audited and alerted.

2. Jira, GitLab/GitHub or Claude is unreachable#

  • Symptoms: supervisor alert "<name> is unreachable", runs failing with network errors, Health shows a red check.
  • Check: sprintflow pilot check (or Connections → Test connection) shows which permission or endpoint fails; status pages of Atlassian / your Git host / AWS Bedrock; egress rules (network policy allows 443).
  • Act: nothing, for short outages: transient failures auto-resume with backoff, and the supervisor resumes interrupted runs. For long ones, pause automation (1) so runs do not pile up failures.
  • Verify: the "reachable again" alert arrives; failed runs → Resume on their run page.

3. Budget used up#

  • Symptoms: alert "Daily/Monthly budget used up", runs stopped with budget_exhausted, red banner.
  • Check: Health → Budget bars; Team page → spend by project/person; Runs sorted by cost for outliers.
  • Act: if expected, raise budgets (Settings → Budgets) or wait for the period to reset. If not, look for a runaway: a story looping on fixes (Timeline tab), a too-expensive model for simple steps (Settings → Model routing).
  • Verify: stopped runs keep their work: Resume continues from the last completed stage.

4. Stuck or failed runs#

  • Symptoms: supervisor alert "run is stuck and needs a person", or a run in failed.
  • Check: run page → Events tab (last event), Timeline (which step), Files → logs. sprintflow status <run_id>.
  • Act: transient cause → Resume. Code/test problem → read the report, fix the ticket or .sprintflow/config.yml, resume or start a new run. Worker died → the queue retries the job automatically (up to max_job_attempts).
  • Verify: the run reaches a draft PR or a clear stop with a reason.

5. Alerts are not arriving#

  • Check: Health → Recent alerts shows sent to … or not delivered; Settings → Alerts → Send test alert.
  • Act: Slack/Teams webhook revoked → create a new one and paste it in Settings → Alerts (stored encrypted). Email → SMTP host/port/STARTTLS and the SMTP password. Check that the event is not filtered by the channel's events or min_severity.

6. Scaling#

  • Queue backing up (sprintflow_sprint_queue, jobs queued growing): add workers — the HPA does this on CPU; raise maxReplicas or --concurrency. Workers finish their jobs before stopping (grace period 1 h).
  • Console slow: add console replicas (safe: only the lease holder runs background work).
  • Database: PostgreSQL connections ≈ (consoles + workers × concurrency) × 5; size max_connections accordingly.

7. Locked out of the console (SSO outage or misconfiguration)#

  • Act: sign in with a break-glass account (sso.break_glass_users, password login). If none exists: sprintflow console user add rescue --role admin on a host with the console key and database access, then fix Settings → Single sign-on (issuer, client id/secret, redirect URL, role map).
  • Verify: SSO sign-in works; remove or rotate the rescue account; audit log shows login_sso.

8. Rotating keys and tokens#

  • Console master key (encrypts stored secrets and signs sessions): 1. sprintflow console init-key > new.key 2. sprintflow console rotate-key --new-key-file new.key 3. point SPRINTFLOW_CONSOLE_KEY_FILE / _REF (and the Kubernetes Secret) at the new key; restart consoles and workers; everyone signs in again. Keep the old key until you have verified, then destroy it.
  • Jira / Git host / Anthropic / webhook secrets: create the new credential, paste it in Connections/Settings (or update the secrets-manager entry — cached values refresh within SPRINTFLOW_SECRET_TTL, default 5 min), test the connection, then revoke the old one.

9. Backup and restore#

  • Backup (daily): sprintflow backup create --out /backups/sprintflow-$(date +%F).tar.gz; sprintflow backup verify …. Store the console master key separately (it is not in the backup, by design).
  • Restore: stop consoles/workers, then sprintflow backup restore <archive> --runs-dir … --console-dir … (with SPRINTFLOW_DATABASE_URL set to restore database rows). It verifies checksums first and refuses to overwrite existing data unless --force. Start consoles/workers; open a few runs to check.

10. Upgrading SprintFlow#

  • Read the release notes; take a backup (9).
  • Rolling upgrade: consoles (maxUnavailable: 0) then workers (they drain). Run state is versioned: a new release migrates older runs as they load.
  • Rollback: allowed while no run has been written by the new version. If one has, the older release refuses to open it ("written by a newer SprintFlow") instead of corrupting it — roll forward, or restore the pre-upgrade backup.
AI SprintFlow 1.0.5 · Questions? Contact us