Deploying AI SprintFlow in an enterprise#
To try it on your own machine first, see getting-started.md.
Four stages, each building on the last:
| Stage | For | Runs on | Needs |
|---|---|---|---|
| 1. Pilot | one team, trying it | one VM, Docker Compose | nothing else |
| 2. Kubernetes, single node | one team, day to day | one pod + a volume | a cluster |
| 3. Production | one repository at scale | consoles + autoscaled workers | PostgreSQL, S3 |
| 4. Organisation-wide | every team | one installation per repository | shared PostgreSQL server, bucket, SSO, monitoring |
Before you start: what to request#
Most of the lead time is approvals, not installation. Request these in parallel; each needs least privilege as listed in security.md.
| # | Item | Owner (typical) | Details |
|---|---|---|---|
| 1 | Jira bot account + API token | Jira admins | Browse, Add comments, Create attachments, Transition issues; pilot projects only |
| 2 | Code host token | Git platform team | GitLab: project access token, role Developer, scopes api, read_repository, write_repository; protect the default branch |
| 3 | Claude access | Cloud / AI platform team | Bedrock: model access for the Claude models in ap-south-1 (data residency), a quota increase for tokens per minute, and an IAM role for the pods; or an Anthropic API key with a spend limit |
| 4 | SSO app registration | Identity team | Entra ID/Okta OIDC app, redirect URI https://<host>/api/auth/oidc/callback, groups claim in the ID token, groups such as sprintflow-admins, -managers, -operators, -viewers |
| 5 | DNS name + TLS certificate | Network team | e.g. sprintflow.company.com, internal CA is fine |
| 6 | PostgreSQL (production stage) | DBA / cloud team | managed PostgreSQL 14+, one database, a dedicated role, TLS |
| 7 | S3 bucket (production stage) | Cloud team | public access blocked, default encryption, IAM access to sprintflow/* |
| 8 | Secrets manager entries | Security | console key, database URL, tokens (AWS Secrets Manager, Azure Key Vault, Vault or Google) |
| 9 | Container registry + Kubernetes namespace | Platform team | or one Linux VM (4 vCPU, 16 GiB, 200 GiB disk) with Docker for the pilot |
| 10 | License key | AI SprintFlow (ai-sprintflow.com) | not needed for a pilot of up to 5 seats |
Outbound network access (the pods need no inbound access except from your ingress and Prometheus):
your Jira site (443), your code host (443, or 22 for SSH git), the Claude endpoint
(bedrock-runtime.ap-south-1.amazonaws.com, api.anthropic.com or Vertex), your package registries for dependency
installs (NuGet, npm, PyPI, Maven — or your internal mirror), S3, PostgreSQL (5432), your SSO issuer, and your
Slack/Teams/SMTP endpoints. Build and test containers themselves get no network.
1. Pilot: single host, Docker Compose#
unzip ai-sprintflow.zip && cd ai-sprintflow/deploy
mkdir -p secrets tls
docker run --rm $(docker build -q -f Dockerfile.any ..) console init-key > secrets/console_key # back it up!
# put fullchain.pem + privkey.pem for your hostname in tls/, set server_name in nginx-console.conf
docker compose up -d # builds the image (set BASE_IMAGE in docker-compose.yml to your toolchain)
docker compose exec -it console sprintflow console user add <you> --role admin # prompts for a password
Open https://<host>, then:
1. Connections: Jira, your Git host, Claude → Test connection each.
2. Settings → General: turn shadow mode on (nothing is pushed or posted while you evaluate).
3. On the host: docker compose exec console sprintflow pilot check --from-console --issue <KEY> --baseline → fix
everything marked ❌ (the report is written to pilot-report.md in the container's working folder).
4. License (admins): paste your key. Without one the installation runs on the 7-day free trial (up to 5 seats), enough to start a pilot.
5. Run a few real stories (New run). Compare the patches with what your developers would write.
The BASE_IMAGE build argument picks the toolchain image (e.g. mcr.microsoft.com/dotnet/sdk:8.0 for .NET). For
container isolation of builds (sandbox.runtime: docker), the console needs access to a Docker daemon.
2. Kubernetes, single node#
# 1. image
docker build -f deploy/Dockerfile.any --build-arg BASE_IMAGE=mcr.microsoft.com/dotnet/sdk:8.0 \
-t registry.company.com/sprintflow:1.0.0 . && docker push registry.company.com/sprintflow:1.0.0
# 2. secret (better: External Secrets / Vault / CSI driver from your secrets manager)
kubectl create namespace sprintflow
kubectl -n sprintflow create secret generic sprintflow-secrets \
--from-literal=console-key="$(sprintflow console init-key)" --from-literal=metrics-token="$(openssl rand -hex 24)"
# 3. deploy (edit images: in deploy/kubernetes/base/kustomization.yaml first)
kubectl apply -k deploy/kubernetes/overlays/single-node
kubectl -n sprintflow exec -it deploy/sprintflow-console -- sprintflow console user add <you> --role admin
Expose the sprintflow-console service through your ingress with TLS. State lives on the sprintflow-data volume;
back it up with sprintflow backup create (see §5).
3. Production: scale-out#
- PostgreSQL (managed: RDS / Azure Database / Cloud SQL) — one database, TLS, a dedicated role.
Add
database-url: postgresql+psycopg://sprintflow:<pw>@<host>:5432/sprintflow?sslmode=requiretosprintflow-secrets. - S3 bucket (or MinIO) with public access blocked and default encryption; give the pods IAM access to
sprintflow/*(IRSA / workload identity) — see security.md. - Deploy:
kubectl apply -k deploy/kubernetes/overlays/production(edit the ingress host). You get 2 console replicas (one leads background work), 2–10 autoscaled workers, network policies and a disruption budget. - Console settings: Settings → Storage & workers:
execution: queueand the S3 bucket; Single sign-on; Webhooks; Alerts; Budgets; Data guard; Risk approvals; Sandbox. - Point Jira and your Git host at the webhooks (Settings → Webhooks shows the URLs) for instant reactions.
- Monitoring: Prometheus scrapes
/metricswith themetrics-tokenbearer token. Suggested alerts:
- alert: SprintFlowSupervisorStalled
expr: sprintflow_supervisor_heartbeat_age_seconds > 300
- alert: SprintFlowBudgetNearlyUsed
expr: sprintflow_budget_spent_usd / sprintflow_budget_limit_usd > 0.9 and sprintflow_budget_limit_usd > 0
- alert: SprintFlowQueueBackingUp
expr: sprintflow_sprint_queue{state="queued"} > 20
Traces: set OTEL_EXPORTER_OTLP_ENDPOINT in the sprintflow-env ConfigMap (Jaeger, Tempo, Datadog, Langfuse).
4. Organisation-wide: every team, any number of people#
How it scales#
- People: unlimited. Everyone signs in with company SSO; roles come from groups via Settings → Single sign-on →
role map (e.g.
sprintflow-payroll-devs→ operator,sprintflow-payroll-leads→ manager). No per-user setup: an account is created at first sign-in, and removing someone from the group removes their access. - Work: horizontal. Each run occupies one worker slot at a time; add worker replicas (HPA) and
--concurrencyfor throughput. PostgreSQL and S3 are shared by all replicas. - Repositories: one installation each. An installation serves one Jira site and one repository. For many repositories, run one installation (a tenant) per repository or team. A single installation serving many repositories is not built yet.
What tenants share and what they don't#
| Shared across the organisation | Per tenant (repository/team) |
|---|---|
| Kubernetes cluster, image and registry | Namespace, console URL |
| PostgreSQL server (one database per tenant) | Database, console master key, secrets |
| S3 bucket (one prefix per tenant) | Connections: Jira project/board, repository, tokens |
| SSO application (one redirect URL per tenant) and IdP groups | Role map, users, audit log |
| Prometheus/Grafana, OpenTelemetry, Slack/Teams | Budgets, alerts routing, settings, memory |
Add a tenant (≈15 minutes)#
cp -r deploy/kubernetes/overlays/tenant-example deploy/kubernetes/overlays/payroll # edit the 3 CHANGE lines
createdb -h <pg-host> sprintflow_payroll # its own database
kubectl create namespace sprintflow-payroll
kubectl -n sprintflow-payroll create secret generic sprintflow-secrets \
--from-literal=console-key="$(sprintflow console init-key)" \
--from-literal=database-url="postgresql+psycopg://…/sprintflow_payroll?sslmode=require" \
--from-literal=metrics-token="$(openssl rand -hex 24)"
kubectl apply -k deploy/kubernetes/overlays/payroll
Then in that tenant's console: Connections (its Jira board and repository), Single sign-on (add its redirect URL to
the shared SSO app; map its groups), Storage & workers → S3 prefix payroll, budgets and alerts. Keep tenants in git (GitOps with
Argo CD or Flux) so adding one is a reviewed pull request.
Capacity planning#
Measure in the pilot, then size: - Worker slots ≈ (stories per day × average story duration in hours) ÷ working hours, plus 30% headroom. A .NET monorepo build/test typically needs 2–4 CPU and 4–8 GiB per slot. - Claude throughput: request a Bedrock (or Anthropic) quota increase for tokens per minute ≈ concurrent runs × the peak tokens per minute of one run (visible in traces). - Spend: cost per merged story from the Team page × stories per month → set monthly budgets per tenant. - PostgreSQL: small (2 vCPU, 4–8 GiB) serves many tenants; connections ≈ (consoles + worker slots) × 5 per tenant.
Rollout#
- Pilot (4–6 weeks, one team): shadow mode → draft MRs with checkpoints. Agree success criteria first (merge rate, rework, cost per merged story, zero data-guard incidents).
- Waves of 2–3 teams: champion per team, 1-hour onboarding using user-guide.md, shadow mode for the first week of each team.
- Governance: an AI usage policy (people merge; approvals on risky paths), the evaluation gate
(
sprintflow eval compare … --max-regression 0.05) before any model or prompt change, monthly review of the Team page metrics. - Support: platform team owns the cluster, database and upgrades (runbooks.md); each tenant's admins own its connections and settings.
Go-live checklist#
Before switching shadow mode off for a team:
- [ ]
sprintflow pilot check --from-console --baselinepasses with no ❌ - [ ] HTTPS only; SSO on; password sign-in off except named break-glass accounts
- [ ] Default branch protected on the code host (people merge; AI SprintFlow has no merge permission)
- [ ] Data guard on (default); risk approvals set for payroll maths, authentication and migrations
- [ ] Sandbox runtime
docker/podman, or workers isolated at pod level with network policies - [ ] Budgets (daily and monthly) and at least one alert channel; a test alert received
- [ ] Backups scheduled (
sprintflow backup create) and the console key stored separately; one restore rehearsed - [ ] Prometheus scraping
/metrics; the three alerts above configured - [ ] Success criteria for the team agreed (merge rate, rework, cost per merged story)
- [ ] License installed with enough seats (viewers are free)
5. Operating any installation#
- Backups: daily
sprintflow backup create --out …(a Kubernetes CronJob running the image with the same secrets and volumes) + your database's own backups; keep the console master key separately. - Upgrades: rolling (consoles then workers); run state is versioned — see runbooks.md.
- Incidents: runbooks.md — pause all automation first, investigate second.