AI SprintFlow

Deploying AI SprintFlow in an enterprise#

To try it on your own machine first, see getting-started.md.

Four stages, each building on the last:

Stage For Runs on Needs
1. Pilot one team, trying it one VM, Docker Compose nothing else
2. Kubernetes, single node one team, day to day one pod + a volume a cluster
3. Production one repository at scale consoles + autoscaled workers PostgreSQL, S3
4. Organisation-wide every team one installation per repository shared PostgreSQL server, bucket, SSO, monitoring

Before you start: what to request#

Most of the lead time is approvals, not installation. Request these in parallel; each needs least privilege as listed in security.md.

# Item Owner (typical) Details
1 Jira bot account + API token Jira admins Browse, Add comments, Create attachments, Transition issues; pilot projects only
2 Code host token Git platform team GitLab: project access token, role Developer, scopes api, read_repository, write_repository; protect the default branch
3 Claude access Cloud / AI platform team Bedrock: model access for the Claude models in ap-south-1 (data residency), a quota increase for tokens per minute, and an IAM role for the pods; or an Anthropic API key with a spend limit
4 SSO app registration Identity team Entra ID/Okta OIDC app, redirect URI https://<host>/api/auth/oidc/callback, groups claim in the ID token, groups such as sprintflow-admins, -managers, -operators, -viewers
5 DNS name + TLS certificate Network team e.g. sprintflow.company.com, internal CA is fine
6 PostgreSQL (production stage) DBA / cloud team managed PostgreSQL 14+, one database, a dedicated role, TLS
7 S3 bucket (production stage) Cloud team public access blocked, default encryption, IAM access to sprintflow/*
8 Secrets manager entries Security console key, database URL, tokens (AWS Secrets Manager, Azure Key Vault, Vault or Google)
9 Container registry + Kubernetes namespace Platform team or one Linux VM (4 vCPU, 16 GiB, 200 GiB disk) with Docker for the pilot
10 License key AI SprintFlow (ai-sprintflow.com) not needed for a pilot of up to 5 seats

Outbound network access (the pods need no inbound access except from your ingress and Prometheus): your Jira site (443), your code host (443, or 22 for SSH git), the Claude endpoint (bedrock-runtime.ap-south-1.amazonaws.com, api.anthropic.com or Vertex), your package registries for dependency installs (NuGet, npm, PyPI, Maven — or your internal mirror), S3, PostgreSQL (5432), your SSO issuer, and your Slack/Teams/SMTP endpoints. Build and test containers themselves get no network.

1. Pilot: single host, Docker Compose#

unzip ai-sprintflow.zip && cd ai-sprintflow/deploy
mkdir -p secrets tls
docker run --rm $(docker build -q -f Dockerfile.any ..) console init-key > secrets/console_key   # back it up!
# put fullchain.pem + privkey.pem for your hostname in tls/, set server_name in nginx-console.conf
docker compose up -d                      # builds the image (set BASE_IMAGE in docker-compose.yml to your toolchain)
docker compose exec -it console sprintflow console user add <you> --role admin    # prompts for a password

Open https://<host>, then: 1. Connections: Jira, your Git host, Claude → Test connection each. 2. Settings → General: turn shadow mode on (nothing is pushed or posted while you evaluate). 3. On the host: docker compose exec console sprintflow pilot check --from-console --issue <KEY> --baseline → fix everything marked ❌ (the report is written to pilot-report.md in the container's working folder). 4. License (admins): paste your key. Without one the installation runs on the 7-day free trial (up to 5 seats), enough to start a pilot. 5. Run a few real stories (New run). Compare the patches with what your developers would write.

The BASE_IMAGE build argument picks the toolchain image (e.g. mcr.microsoft.com/dotnet/sdk:8.0 for .NET). For container isolation of builds (sandbox.runtime: docker), the console needs access to a Docker daemon.

2. Kubernetes, single node#

# 1. image
docker build -f deploy/Dockerfile.any --build-arg BASE_IMAGE=mcr.microsoft.com/dotnet/sdk:8.0 \
  -t registry.company.com/sprintflow:1.0.0 . && docker push registry.company.com/sprintflow:1.0.0
# 2. secret (better: External Secrets / Vault / CSI driver from your secrets manager)
kubectl create namespace sprintflow
kubectl -n sprintflow create secret generic sprintflow-secrets \
  --from-literal=console-key="$(sprintflow console init-key)" --from-literal=metrics-token="$(openssl rand -hex 24)"
# 3. deploy (edit images: in deploy/kubernetes/base/kustomization.yaml first)
kubectl apply -k deploy/kubernetes/overlays/single-node
kubectl -n sprintflow exec -it deploy/sprintflow-console -- sprintflow console user add <you> --role admin

Expose the sprintflow-console service through your ingress with TLS. State lives on the sprintflow-data volume; back it up with sprintflow backup create (see §5).

3. Production: scale-out#

  1. PostgreSQL (managed: RDS / Azure Database / Cloud SQL) — one database, TLS, a dedicated role. Add database-url: postgresql+psycopg://sprintflow:<pw>@<host>:5432/sprintflow?sslmode=require to sprintflow-secrets.
  2. S3 bucket (or MinIO) with public access blocked and default encryption; give the pods IAM access to sprintflow/* (IRSA / workload identity) — see security.md.
  3. Deploy: kubectl apply -k deploy/kubernetes/overlays/production (edit the ingress host). You get 2 console replicas (one leads background work), 2–10 autoscaled workers, network policies and a disruption budget.
  4. Console settings: Settings → Storage & workers: execution: queue and the S3 bucket; Single sign-on; Webhooks; Alerts; Budgets; Data guard; Risk approvals; Sandbox.
  5. Point Jira and your Git host at the webhooks (Settings → Webhooks shows the URLs) for instant reactions.
  6. Monitoring: Prometheus scrapes /metrics with the metrics-token bearer token. Suggested alerts:
- alert: SprintFlowSupervisorStalled
  expr: sprintflow_supervisor_heartbeat_age_seconds > 300
- alert: SprintFlowBudgetNearlyUsed
  expr: sprintflow_budget_spent_usd / sprintflow_budget_limit_usd > 0.9 and sprintflow_budget_limit_usd > 0
- alert: SprintFlowQueueBackingUp
  expr: sprintflow_sprint_queue{state="queued"} > 20

Traces: set OTEL_EXPORTER_OTLP_ENDPOINT in the sprintflow-env ConfigMap (Jaeger, Tempo, Datadog, Langfuse).

4. Organisation-wide: every team, any number of people#

How it scales#

  • People: unlimited. Everyone signs in with company SSO; roles come from groups via Settings → Single sign-on → role map (e.g. sprintflow-payroll-devs → operator, sprintflow-payroll-leads → manager). No per-user setup: an account is created at first sign-in, and removing someone from the group removes their access.
  • Work: horizontal. Each run occupies one worker slot at a time; add worker replicas (HPA) and --concurrency for throughput. PostgreSQL and S3 are shared by all replicas.
  • Repositories: one installation each. An installation serves one Jira site and one repository. For many repositories, run one installation (a tenant) per repository or team. A single installation serving many repositories is not built yet.

What tenants share and what they don't#

Shared across the organisation Per tenant (repository/team)
Kubernetes cluster, image and registry Namespace, console URL
PostgreSQL server (one database per tenant) Database, console master key, secrets
S3 bucket (one prefix per tenant) Connections: Jira project/board, repository, tokens
SSO application (one redirect URL per tenant) and IdP groups Role map, users, audit log
Prometheus/Grafana, OpenTelemetry, Slack/Teams Budgets, alerts routing, settings, memory

Add a tenant (≈15 minutes)#

cp -r deploy/kubernetes/overlays/tenant-example deploy/kubernetes/overlays/payroll   # edit the 3 CHANGE lines
createdb -h <pg-host> sprintflow_payroll                                              # its own database
kubectl create namespace sprintflow-payroll
kubectl -n sprintflow-payroll create secret generic sprintflow-secrets \
  --from-literal=console-key="$(sprintflow console init-key)" \
  --from-literal=database-url="postgresql+psycopg://…/sprintflow_payroll?sslmode=require" \
  --from-literal=metrics-token="$(openssl rand -hex 24)"
kubectl apply -k deploy/kubernetes/overlays/payroll

Then in that tenant's console: Connections (its Jira board and repository), Single sign-on (add its redirect URL to the shared SSO app; map its groups), Storage & workers → S3 prefix payroll, budgets and alerts. Keep tenants in git (GitOps with Argo CD or Flux) so adding one is a reviewed pull request.

Capacity planning#

Measure in the pilot, then size: - Worker slots ≈ (stories per day × average story duration in hours) ÷ working hours, plus 30% headroom. A .NET monorepo build/test typically needs 2–4 CPU and 4–8 GiB per slot. - Claude throughput: request a Bedrock (or Anthropic) quota increase for tokens per minute ≈ concurrent runs × the peak tokens per minute of one run (visible in traces). - Spend: cost per merged story from the Team page × stories per month → set monthly budgets per tenant. - PostgreSQL: small (2 vCPU, 4–8 GiB) serves many tenants; connections ≈ (consoles + worker slots) × 5 per tenant.

Rollout#

  1. Pilot (4–6 weeks, one team): shadow mode → draft MRs with checkpoints. Agree success criteria first (merge rate, rework, cost per merged story, zero data-guard incidents).
  2. Waves of 2–3 teams: champion per team, 1-hour onboarding using user-guide.md, shadow mode for the first week of each team.
  3. Governance: an AI usage policy (people merge; approvals on risky paths), the evaluation gate (sprintflow eval compare … --max-regression 0.05) before any model or prompt change, monthly review of the Team page metrics.
  4. Support: platform team owns the cluster, database and upgrades (runbooks.md); each tenant's admins own its connections and settings.

Go-live checklist#

Before switching shadow mode off for a team:

  • [ ] sprintflow pilot check --from-console --baseline passes with no ❌
  • [ ] HTTPS only; SSO on; password sign-in off except named break-glass accounts
  • [ ] Default branch protected on the code host (people merge; AI SprintFlow has no merge permission)
  • [ ] Data guard on (default); risk approvals set for payroll maths, authentication and migrations
  • [ ] Sandbox runtime docker/podman, or workers isolated at pod level with network policies
  • [ ] Budgets (daily and monthly) and at least one alert channel; a test alert received
  • [ ] Backups scheduled (sprintflow backup create) and the console key stored separately; one restore rehearsed
  • [ ] Prometheus scraping /metrics; the three alerts above configured
  • [ ] Success criteria for the team agreed (merge rate, rework, cost per merged story)
  • [ ] License installed with enough seats (viewers are free)

5. Operating any installation#

  • Backups: daily sprintflow backup create --out … (a Kubernetes CronJob running the image with the same secrets and volumes) + your database's own backups; keep the console master key separately.
  • Upgrades: rolling (consoles then workers); run state is versioned — see runbooks.md.
  • Incidents: runbooks.md — pause all automation first, investigate second.
AI SprintFlow 1.0.5 · Questions? Contact us