SprintFlow — full reference#
An AI coding agent that turns a Jira ticket into a reviewed draft pull request, under rules the team controls.
One ticket, one automated run:
Jira ticket ─┐
Code repo ─┼─► 1 Understand ─► 2 Locate & plan ─► 3 Implement ─► 4 Verify ─► 5 Independent review ─► 6 Browser proof ─┬─► Draft pull request
Team rules ─┘ │ ▲ ▲ │ │ │ ├─► Video proof on the Jira ticket
└─ asks a person only if it must │ └─ tests fail → fix (bounded) │ ├─► Run report
├────── review asks for changes → fix, re-verify, re-review └─► or a clear stop
└────── browser test finds an app defect → fix, re-verify, re-review, re-record
Steps 1, 2, 3, 5 and the test-writing part of 6 use Claude. Step 4, running the browser tests, recording and attaching the video are plain code. The pipeline, not the model, decides what happens next. Nothing merges by itself.
- Web console. A browser UI to connect Jira and your code host (with real "Test connection" checks), tune every setting, start runs, watch them live, approve checkpoints, play the 4K proof videos, manage memory, users and roles, and read the audit log. The CLI keeps working alongside it.
- Acts on PR reviews and CI. After the draft PR opens, SprintFlow answers reviewers' change requests and fixes failed CI on the same branch - verified and re-reviewed before every push, and replying in each thread.
- Sprint automation. People link their Jira account in the console; when a sprint starts, SprintFlow collects the stories assigned to them and implements them one by one, each ending in a draft PR for its owner to review.
- Any source host. GitHub (incl. Enterprise), GitLab (incl. self-managed), Bitbucket Cloud, Bitbucket Server/Data Center, Azure DevOps, Gitea/Forgejo, any other git remote (HTTPS or SSH, with an optional webhook to open the review request anywhere), or your own plugin.
- Browser proof on the ticket. For stories with user-visible behaviour, SprintFlow writes Playwright tests, runs them against the changed app with video, and attaches the recording (MP4, captioned per acceptance criterion) to the same Jira ticket.
- Any tech stack, zero config. Detects .NET, Node, Python, Java/Kotlin, Go, Rust, Ruby, PHP, Elixir, Swift,
Dart/Flutter, C/C++ and Makefile projects, including mixed-stack monorepos. A team can pin anything in
.sprintflow/config.yml. - Self-healing. Transient outages auto-resume; network blips retry; flaky tests are separated from real regressions; changed dependencies are re-installed; a moved base branch is rebased and re-verified; a lost workspace is rebuilt; a taken branch name is never overwritten. None of this spends model tokens.
- Token-frugal. Cheap model where it's enough, local repo map instead of exploration, transcript compaction, distilled failure logs, per-step knowledge slices and output caps, batched tools, prompt caching.
- Memory. Learns from every run (flaky tests, commands, where similar tickets were changed, recurring review findings) and periodically consolidates that into a few short lessons fed to later runs.
Contents#
- Quick start
- How each part of the architecture is implemented
- Claude credentials from Claude Code settings.json
- Web console
- PR follow-up: review comments and CI
- Sprint automation
- Manager dashboard
- Running unattended: alerts, budgets, supervisor
- Jira status updates
- Monorepos: build and test only what changed
- Evaluation harness
- Production readiness
- Browser proof (Playwright video on the ticket)
- Source providers
- Any tech stack
- Self-healing
- Token economy
- Memory
- Team knowledge (
.sprintflow/in the target repo) - Operator configuration (
sprintflow.yml) - Command line
- Outcomes
- Security model
- Deployment
- Operating it
- Development
- Before going live
Quick start#
pip install . # Python 3.11+; git required, ripgrep recommended
cp examples/sprintflow.yml sprintflow.yml # fill in Jira, source provider, model, pricing
sprintflow check-repo /path/to/your/repo # shows the detected stack(s) and the commands it will run
sprintflow validate-config
sprintflow run PROJ-123 --dry-run # everything except push + PR; keeps the verified patch
sprintflow run PROJ-123 # opens a draft PR
No .sprintflow/ folder is needed. Add one (see examples/team-knowledge/) to pin commands and give the
agent your team's rules, guides and skills.
How each part of the architecture is implemented#
| Architecture box | Implementation |
|---|---|
| What goes in: Jira ticket (Story/Bug, comments, screenshots) | integrations/jira.py - Cloud (v3, ADF) and Data Center (v2); image attachments are sent to Claude within count/size limits |
| Code repository (branch, build and test commands) | workspace.py clones the base branch into an isolated per-run workspace; commands are auto-detected (stack.py) and/or pinned in the repo's .sprintflow/config.yml |
| Team knowledge (guides, rules, reusable skills) | knowledge.py loads .sprintflow/{guides,rules,skills,prompts} plus CLAUDE.md/AGENTS.md, once, from the pristine base, and snapshots it |
| 1 Understand | steps/understand.py - writes testable acceptance criteria; ask posts questions to Jira and pauses; too_vague / already_implemented are clear stops |
| 2 Locate & plan | steps/plan.py - file-by-file plan; for bugs the prompt requires a traced root cause with evidence. The plan is posted to Jira |
| 3 Implement | steps/implement.py - code and tests via sandboxed file tools; no shell tool |
| 4 Verify | steps/verify.py - no AI. Runs setup/build/test/lint, compares per-test results (JUnit XML or .NET TRX) with a baseline taken before any change; optional security scan counts only new findings |
| Fix loop | pipeline._verify_loop - bounded by max_fix_attempts for the whole run; feedback names the exact regressions |
| 5 Independent review | steps/review.py - a fresh conversation with no memory of the work: ticket, criteria, rules, diff and the pipeline's verify result only; read-only tools; optionally a different model |
| Review loop | pipeline._review_loop - changes requested → fix → re-verify → re-review, bounded by max_review_rounds. An "approve" that lists blockers or unmet criteria is treated as "request changes" |
| Built-in safety rails | Every step recorded and resumable (state.py); per-step time/cost/turn caps and a run cost cap (budget.py); AI only touches the isolated copy (tools.py); human checkpoints; no merge capability exists |
| What comes out: draft PR | integrations/providers.py + one adapter per host - a draft PR/MR (native draft, or marked by title where the host can't), labels, reviewers; body has summary, root cause, criteria checklist, verification, review verdict |
| Run report | artifacts/report.md + report.json: what was done, what was skipped, cost, duration, full event log |
| Or a clear stop | Reason code + explanation, posted to Jira, patch kept in artifacts/<KEY>.patch |
| Systems it talks to: Claude | llm.py - Anthropic API or AWS Bedrock through the official SDK, prompt caching on system/tools/transcript |
| Jira; GitHub / GitLab / Bitbucket / Azure DevOps / Gitea / any git host | as above; retries with backoff and Retry-After |
| Code index (Kodify today, graphify as an option) | integrations/code_index.py - built-in (ripgrep / pure Python), or any external indexer by command; falls back to built-in and records it |
| Security scan (Aikido, optional) | integrations/security.py - any scanner that writes SARIF 2.1.0 |
Claude credentials: Automatic#
model.credentials: auto is the console's default. When the Claude card opens, it loads local credentials: Claude Code's
settings.json on the SprintFlow server (see below), or ANTHROPIC_API_KEY in its environment. If there are none, the
card offers Connect (browser sign-in, next section). In Docker, the container cannot see your own ~/.claude.
Mount it read-only (see the commented line in deploy/docker-compose.yml) or use Connect.
Claude credentials: browser sign-in (Connect)#
Set model.credentials: anthropic_login (or choose Connect with Anthropic on the console's Claude card) and click
Connect. It works like VS Code signing in on a remote machine. The SprintFlow server runs
ant auth login --no-browser, and the console opens the sign-in page in a new tab of your own browser. There you pick
the organisation and workspace to bill. You then paste the code Anthropic shows into the card and click
Finish sign-in. The Docker images include ant and keep the profile in
ANTHROPIC_CONFIG_DIR=/var/lib/sprintflow/anthropic, on the persistent volume.
The CLI stores an OAuth profile in ~/.config/anthropic/ (Windows: %APPDATA%\Anthropic\, or
$ANTHROPIC_CONFIG_DIR). The Anthropic SDK reads that profile and refreshes the token itself. SprintFlow only checks
that the profile exists and never reads the token. model.anthropic_profile selects a named profile. By default it
uses the CLI's active profile. A stale ANTHROPIC_API_KEY in the environment is ignored in this mode. The ant CLI
must be installed on the SprintFlow server
(releases, macOS: brew install anthropics/tap/ant).
Claude credentials from Claude Code settings.json#
Set model.credentials: claude_settings (or choose Use local Claude Code settings on the console's Claude card) and
SprintFlow authenticates exactly as the claude CLI on the same machine is configured - no separate key to manage.
It reads the same files with the same precedence as Claude Code: managed settings (managed-settings.json and
managed-settings.d/) > .claude/settings.local.json > .claude/settings.json > ~/.claude/settings.json
($CLAUDE_CONFIG_DIR respected), merging the env block per variable.
| In Claude settings | SprintFlow uses |
|---|---|
env.ANTHROPIC_API_KEY |
Anthropic API |
apiKeyHelper |
runs the script for the key; re-runs after CLAUDE_CODE_API_KEY_HELPER_TTL_MS (default 5 min) and immediately when a key is rejected |
env.ANTHROPIC_BASE_URL + ANTHROPIC_AUTH_TOKEN, ANTHROPIC_CUSTOM_HEADERS |
an LLM gateway |
env.CLAUDE_CODE_USE_BEDROCK=1 |
Amazon Bedrock with AWS_PROFILE / the default credential chain, region from AWS_REGION > AWS_DEFAULT_REGION > the profile > us-east-1, AWS_BEARER_TOKEN_BEDROCK (Bedrock API key), ANTHROPIC_BEDROCK_BASE_URL |
awsAuthRefresh (e.g. aws sso login --profile x) |
run when STS says the credentials are expired or Bedrock rejects them - this is the browser sign-in: it opens the browser on the SprintFlow host or prints a URL and code. The console's Sign in to AWS button runs it and shows the link live |
awsCredentialExport |
runs it and uses the returned credentials until 5 minutes before Expiration |
env.CLAUDE_CODE_USE_VERTEX=1, ANTHROPIC_VERTEX_PROJECT_ID, CLOUD_ML_REGION |
Google Vertex AI with Application Default Credentials (gcloud auth application-default login) |
env.ANTHROPIC_MODEL, ANTHROPIC_DEFAULT_HAIKU_MODEL |
the default model (turn off with claude_settings.use_models: false) |
sprintflow auth status shows what was detected (secrets masked); sprintflow auth aws-login runs the SSO sign-in.
Not supported: a claude.ai (Pro/Max) login. Anthropic does not allow third-party products to offer claude.ai login unless previously approved, and asks them to use API-key or cloud-provider authentication instead. If Claude Code on the machine is signed in only with a claude.ai account, SprintFlow says so and points to the options above - it checks only that such a login exists and never reads the stored token. Claude Code itself can keep using the subscription.
Security. Commands in settings (apiKeyHelper, awsAuthRefresh, awsCredentialExport) run only from managed and
user settings, or from a project directory you explicitly trust (claude_settings.trust_project). Settings inside the
repositories SprintFlow works on can never make it run a command.
Web console#
An enterprise web application served by sprintflow console serve: dark sidebar navigation (collapsible), breadcrumbs,
global search and command palette (Ctrl/⌘ K), a notifications bell for runs waiting on people, light and dark themes,
keyboard shortcuts (? lists them), and dialogs instead of browser pop-ups. It is a single build-free page with its
own SVG charts and icons (no third-party scripts), so it runs under a strict Content-Security-Policy. Install with pip install ".[console]" (the Docker
images include it).
sprintflow console init-key > /secure/console_key # master key: encrypts stored tokens. Back it up.
export SPRINTFLOW_CONSOLE_KEY_FILE=/secure/console_key
export SPRINTFLOW_CONSOLE_DIR=/var/lib/sprintflow/console # console config, encrypted secrets, users, audit
export SPRINTFLOW_RUNS_DIR=/var/lib/sprintflow/runs # run state, artifacts, memory (shared with the CLI)
sprintflow console user add mg --role admin # prompts for a password (12+ characters)
sprintflow console import sprintflow.yml # optional: bring over an existing CLI config
sprintflow console serve # 127.0.0.1:8700 - put HTTPS in front (deploy/)
"Just sign in" to the code host. On GitHub.com, GitLab.com and Azure DevOps people connect the repository by
pasting its URL and clicking Sign in (device sign-in: approve a code on the host). That uses AI SprintFlow's own
public OAuth apps; their client IDs are built in (BUILTIN_CLIENT_IDS in integrations/source_oauth.py) or set with
SPRINTFLOW_GITHUB_CLIENT_ID, SPRINTFLOW_GITLAB_CLIENT_ID and SPRINTFLOW_AZURE_CLIENT_ID. Each is registered once
by whoever ships SprintFlow — no secret, no callback URL:
| Host | Register | Settings |
|---|---|---|
| GitHub | github.com → Settings → Developer settings → OAuth Apps (owned by the SprintFlow organisation) | Enable Device Flow; any homepage/callback URL |
| GitLab.com | gitlab.com → group or user Settings → Applications | Confidential off; scopes api read_user write_repository |
| Azure DevOps | Microsoft Entra ID → App registrations, Accounts in any organizational directory | Allow public client flows on; API permission Azure DevOps → user_impersonation |
Without a client ID for a host, its card falls back to the access-token link (or an OAuth app the admin registers under Advanced settings).
Pages
| Page | What you do there |
|---|---|
| Dashboard | KPI cards with trend sparklines and period-over-period deltas (runs, success rate, draft PRs, median time to PR, spend), runs-over-time stacked by outcome, outcome donut, spend trend, cost by pipeline step, a "needs attention" queue and a live activity feed; 7/30/90-day range; readiness checklist |
| Runs | Data table with search, status filter chips with counts, sortable columns, pagination; New run dialog |
| Run | Header with status, branch, duration, cost and actions (open PR, patch, copy link); live stage stepper; Approve / Reject (with reason dialog), "check for answers", resume. Tabs: Overview (outcome, steps with tokens and cost, usage, self-heals, notes) · Timeline (Gantt of every step execution) · Changes (diff viewer: file list with +/− counts, line numbers, collapsible files) · Proof (4K video player with poster, scenario switcher, screenshot gallery with lightbox) · Report · Events (live, filterable) · Files |
| Connections | Jira, source host (all providers, fields adapt to the provider) and Claude, each with Test connection that really signs in and checks permissions (Jira project permissions, API + git access, every configured model), plus a runner toolchain check |
| Settings | Every setting, generated from the same models the pipeline validates with, so the UI never drifts from the engine: model routing, limits, self-healing, browser proof (incl. 4K and test credentials), memory, code index, security scan, sandbox, checkpoints. Import an existing sprintflow.yml |
| Memory | Lessons, flaky tests, commands, outcomes; forget individual items or everything; consolidate now |
| Users / Audit log | Admin only: add users, change roles, disable, reset passwords; every sign-in, change and action |
| Health | Readiness, toolchain, background jobs and the poller |
Roles: viewer reads everything, operator also starts, approves, rejects and resumes runs and manages memory, manager also sees the Team dashboard, admin also manages connections, settings and users. Enforced by the server on every request; the UI only hides what you can't use.
Security
- Passwords hashed with scrypt; 12-character minimum; five failed sign-ins lock that user/IP for 15 minutes;
unknown usernames take as long to reject as wrong passwords.
- Sessions are HMAC-signed cookies (HttpOnly, Secure, SameSite=Strict, 12 h), invalidated everywhere when a
password, role or status changes. Every change needs a CSRF token bound to the session.
- All tokens and test credentials are encrypted at rest (Fernet, key derived from the master key) in files readable
only by the service user. The API never returns a secret - only a mask like ••••1a2b. Blank keeps a secret,
typing replaces it.
- Strict Content-Security-Policy with no inline script, frame-deny, no-referrer, nosniff, HSTS. All untrusted text
(ticket content, model output, logs) is inserted as text, never HTML - a real-browser test injects
<img onerror> and javascript: links and checks nothing executes.
- Artifact downloads are confined to the run's artifacts folder (path traversal is rejected).
- Every change and action is written to the audit log with user and client IP (SPRINTFLOW_TRUSTED_PROXIES lets
the proxy pass the real IP).
Runs execute in a bounded background pool (--max-parallel-runs, default 2) inside the console process. A
built-in poller (--poll-minutes, default 5) resumes runs waiting on Jira answers or approvals and keeps memory
up to date - no separate cron needed. Runs share state with the CLI and are protected by the same per-run lock;
if the console restarts mid-run, the run page offers Resume from the last completed stage.
Deploy: deploy/docker-compose.yml runs the console behind deploy/nginx-console.conf (TLS). Plain HTTP is
refused except on localhost (--insecure-http) because session cookies are Secure.
PR follow-up: review comments and CI#
The draft PR is where people and SprintFlow meet. After it opens, the poller (sprintflow poll, or the console's
background poller) watches each open SprintFlow PR and acts on:
- line comments and changes-requested reviews from people;
- conversation comments that address SprintFlow (
/sprintflow …or@sprintflow …; configurable); - a failed CI pipeline / failed checks on the PR's latest commit, with the failing jobs' logs.
Each round: fetch the branch as it is now (commits people pushed are kept) → fix → verify with the same baseline
comparison → independent re-review (optional) → push one new commit on top (history is not rewritten, so reviewers
see exactly what changed) → reply in each thread: "Addressed in abc1234: …". A fix that does not verify is never
pushed - SprintFlow replies with the reason instead. The same failing CI commit is never retried; rounds per PR are
capped (followup.max_rounds, default 3), merged or closed PRs stop being watched, and SprintFlow ignores its own
comments and any authors in followup.ignore_authors (e.g. other bots). Every human review comment is also recorded
in memory, so repeated feedback becomes a lesson for future runs.
Supported on GitHub (review comments, reviews, issue comments, check runs and Actions job logs) and GitLab
(merge request discussions including diff notes, pipelines and failed job traces). Settings: Settings → PR
follow-up or followup: in sprintflow.yml. The run page shows each follow-up round with its commit and cost.
Sprint automation#
SprintFlow can work through a sprint on its own: when a sprint starts, it collects the stories assigned to people and implements them one by one.
- Link people to Jira. On Account, each user enters their Jira email or account id (SprintFlow resolves it through Jira's user search) and can opt in to "Implement my stories automatically when a sprint starts" (operator role or above).
- Point SprintFlow at the board. In Settings → Sprint automation an admin adds the Scrum board ids (or a project
key, to use all of its Scrum boards), the issue types to implement (default
Story), and the scope: -linked_users(default): stories assigned to console users who opted in; -everyone: every story in the sprint; -bot: stories assigned to a dedicated Jira account (e.g. "SprintFlow") - teams assign work to it explicitly. - Sprint start. The console poller reads each board's active sprint through the Jira Agile REST API
(
/rest/agile/1.0, identical on Cloud and Data Center). The first time it sees a sprint it queues the stories in scope that are not done, in board rank order. Withpick_up_added_stories, stories added to the sprint later are queued on the next poll. - One by one. A queue worker starts the next story only when the current one has finished - draft PR, clear stop
or failure. A story waiting on a person (approval checkpoint or clarification) does not block the queue unless
continue_when_waitingis off. Stories that already have a draft PR, a run in progress or are done in Jira are never queued again. - Watch and steer. After login, the dashboard's My sprint stories card lists your stories in the active sprint with their Jira status and SprintFlow status (not started, queued, implementing, draft PR, stopped) and an Implement all my stories button. The Sprint page shows the live queue - pause/resume, reorder, remove, Check Jira now, Queue sprint stories - and the whole sprint board with owners and statuses.
Everything is audited (sprint_started_detected, sprint_queued, pause/resume, links), the queue survives
restarts (<console_dir>/sprint.json), and Jira errors are shown in the UI instead of stopping the poller.
The bot needs Jira permission to browse the board's projects (the Agile API uses the same account as the rest of SprintFlow).
Manager dashboard#
The console's Team page (role manager or admin) tracks every sprint and every person in one place:
- Sprint completion per active sprint - a progress bar of all stories by state: draft PR, done in Jira, implementing, queued, waiting on people, stopped, failed, not started - and the completion percentage.
- Per-person progress in each sprint: stories, a progress bar, draft PRs, in progress, waiting, stopped/failed and cost; expand a person to see each story with its Jira status, SprintFlow state, run and PR. It shows who is not linked to the console and who has auto-run on, so gaps in adoption are visible.
- Project statistics by Jira project key over 7/30/90 days: runs, outcomes, draft PRs, success rate, median time to PR, PR follow-up rounds, reviewer comments handled, proof videos, cost and a daily trend.
- Team member statistics: sprint stories completed, runs, draft PRs, success rate, median time to PR, follow-ups, cost and last activity.
- A project filter, a date range, and Export CSV (sprint stories, projects and people in one file).
Runs are attributed to the person whose story it is: the story owner for sprint runs, the person who started it
otherwise (owner and requested_by on every run).
Roles: viewer reads everything; operator starts, approves, rejects and resumes runs and manages memory; manager is an operator who also sees the Team page; admin also manages connections, settings and users.
Running unattended: alerts, budgets, supervisor#
SprintFlow is built to run 24/7 without anyone watching - and to reach people the moment it needs them. It never merges: every change ends as a draft PR/MR for a person to review and merge (no provider adapter even has a merge, approve or mark-ready method, and a test enforces that).
Alerts (alerts: / Settings → Alerts) go to Slack (Block Kit), Microsoft Teams (Adaptive Card via an
incoming webhook / Workflows), email (SMTP + STARTTLS) and generic webhooks, each channel filtered by event and
minimum severity: approval_needed, clarification_needed, pr_opened, pr_updated, run_stopped, run_failed,
followup_failed, budget_warning, budget_exceeded, supervisor. Webhook URLs and the SMTP password are stored
encrypted. The same alert is not repeated within dedupe_minutes, a broken channel never blocks the others, and
every alert is listed on the Health page with where it was delivered. Send test alert checks the setup.
Budgets (budgets:): daily_usd and monthly_usd for all automation together, on top of the per-run cap. One
warning at warn_at (80%); when a limit is used up, runs stop at their next stage with their work saved
(budget_exhausted) and one alert goes out.
Stop switch: Pause all automation (manager or admin, with a reason) blocks new runs, sprint queueing and PR follow-ups - in the console and the CLI - while runs in progress finish their current work. A banner shows on every page until someone resumes. Both are audited and alerted.
Supervisor (supervisor:, on by default) runs every minute inside the console, or from cron with
sprintflow supervise --once:
- resumes runs left "running" with no live job and no progress for stuck_after_minutes (up to
max_auto_resumes), then alerts a person;
- reminds about runs waiting on approval or Jira answers longer than approval_reminder_hours (once a day);
- checks Jira and the source host every dependency_check_minutes, alerting after alert_after_failures
consecutive failures and once more on recovery;
- restarts the poller or sprint-queue thread if one has died;
- alerts when free disk under the runs directory is below min_free_disk_gb, and removes workspaces of runs
finished more than workspace_retention_days ago (state, reports, patches and videos are kept; open PRs keep
their workspace for follow-ups);
- writes a heartbeat; /healthz returns 503 when it goes stale, so Kubernetes (liveness probe) or an
autoheal container restarts a hung process. Plain Docker restart: unless-stopped restarts on exit only.
Jira status updates#
jira_workflow: / Settings → Jira status updates (off by default - workflows differ): move the story to
on_start (default In Progress) when work starts, to on_pr (In Review) when the draft PR opens - adding the
PR as a link on the issue - and optionally to on_stopped (e.g. Blocked) and on_merged (e.g. Done, only when
a person merges). Targets are status or transition names. SprintFlow never re-opens a done issue or moves one
backwards; a missing transition or a Jira error is noted on the run and never affects the code change.
Monorepos: build and test only what changed#
In a monorepo, SprintFlow builds and tests only the components a change affects (scope:; mode: auto scopes when
the repository has at least min_components, full never scopes, affected always does).
- Components: .NET project files with their
<ProjectReference>graph (test projects recognised byIsTestProject,Microsoft.NET.Test.Sdk, xUnit, NUnit or MSTest), Node workspaces with their dependency graph, other sub-projects - or declare them in.sprintflow/config.yml: ```yaml components:- {name: payroll, path: src/Payroll, test: "dotnet test tests/Payroll.Tests", depends_on: [shared]} ```
- Affected = components containing changed files + everything that depends on them, transitively. A changed
file outside every component (
Directory.Build.props,global.json, the solution file, root scripts) means test everything - scoping never under-tests. Documentation paths (neutral_paths) are ignored. - Commands: .NET builds only the affected roots (each builds its dependencies) and runs
dotnet testonly on affected test projects, with TRX reports; Node runs workspace scripts (npm, pnpm or yarn). - Baseline for the same components is taken in a separate git worktree at the base commit, cached per affected set - so "new failures vs pre-existing failures" stays exact. The run page shows the scope that was used.
Evaluation harness#
Prove a prompt, model or settings change makes SprintFlow better before rolling it out:
sprintflow eval harvest --out evals/suite.yml [--merged-only] # from past runs that reached a draft PR
sprintflow eval run evals/suite.yml --label baseline
sprintflow eval run evals/suite.yml --label opus --model claude-opus-5-5
sprintflow eval compare baseline opus --max-regression 0.05 # exit 1 on regression - use it in CI
Each case is a real ticket pinned to its base commit, with hidden tests - the tests from the final change,
which SprintFlow never sees. Cases replay in isolation (dry run: no push, no Jira writes, no alerts, no memory, no
browser proof) and a case is resolved only if the hidden tests pass (optionally within max_cost_usd). Results
report resolve rate, PR rate, cost, time, fix and review rounds and overlap with the files a good solution touched,
and are saved under <runs_dir>/evals/<label>/results.json. Suites are plain YAML, so you can also write cases by hand.
Production readiness#
Validate before you trust it. sprintflow pilot check [--issue PAY-123] [--baseline] [--send-test-alert] checks
your real Jira (permissions, workflow moves), source host (API, git, permission to open draft PRs/MRs, CI logs), every
configured Claude model, alert delivery and the repository (how it will be built and tested, toolchain, a timed
baseline) and writes pilot-report.md with a fix for every failure. Then run in shadow mode
(shadow_mode: true / Settings → General): the full pipeline on real stories, but nothing leaves SprintFlow — no
push, no PR, no Jira comment, attachment or status change. Compare the patches, measure with the evaluation harness,
then switch shadow mode off for one team.
Any stack, any project. Detection covers .NET, Node (npm/pnpm/yarn/bun), Python, Java/Kotlin (Maven, Gradle),
Scala (sbt), Go, Rust, Ruby, PHP, Elixir, Erlang, Swift, Dart/Flutter, C/C++ (CMake, Meson, Make), Bazel, Nx,
Turborepo, Haskell, Clojure, OCaml, Julia, R, Deno, Zig, Crystal, Nim, Perl, Terraform and PowerShell. Anything
else is built and tested with the commands from the project's own CI configuration (GitLab CI, GitHub Actions, Azure
Pipelines, Bitbucket, CircleCI, Jenkinsfile — deploy jobs and secret-handling lines ignored) or dev container, or
with a two-line .sprintflow/config.yml.
Payroll-grade data safety. The data guard (on by default) replaces personal data and secrets — email, phone,
card numbers (Luhn), IBAN (checksum), PAN, Aadhaar (Verhoeff), SSN, cloud keys, tokens, passwords, connection strings,
plus your own patterns — with placeholders before any text reaches the model, and restores them locally in the
model's edits. Credential and production files are withheld entirely. Risk approvals (risk:) pause changes to
paths you name, or files owned by chosen CODEOWNERS teams, for a person's approval before the PR opens.
Isolation. With sandbox.runtime: docker|podman, every setup/build/test command runs in a throwaway container
(the project's image; network only for dependency installs, none for build and test; non-root; all capabilities
dropped; CPU/memory limits; secrets passed by name; shared package caches). It fails closed.
Enterprise integrations. OIDC single sign-on (Entra ID/Azure AD, Okta, Google, Keycloak; roles from groups;
break-glass accounts), secrets managers (${aws-sm:…}, ${azure-kv:…}, ${vault:…}, ${gcp-sm:…} anywhere in
configuration, and the console master key via SPRINTFLOW_CONSOLE_KEY_REF), webhooks (Jira, GitLab, GitHub —
authenticated — for instant sprint checks and PR follow-ups), observability (Prometheus /metrics, JSON logs,
OpenTelemetry traces of every stage and Claude call), and scanner presets (Semgrep, Gitleaks, Trivy, Bandit,
DevSkim, OSV-Scanner; only new findings block).
Scale. Set SPRINTFLOW_DATABASE_URL (PostgreSQL) and run state, events and all console data move to the database
(advisory locks, so any number of machines). With storage.execution: queue the console enqueues and
sprintflow worker processes execute (SKIP LOCKED claims, heartbeats, retries, one job per run at a time); several
console replicas share the work through a leader lease. storage.s3 keeps artifacts in S3/MinIO, served with SigV4
signed links.
Operating it. Kubernetes manifests (deploy/kubernetes, Kustomize: hardened pods, probes, disruption budget,
worker autoscaling, network policies; production and single-node overlays, validated against the Kubernetes
schemas), versioned run state (older runs migrate; newer ones are refused rather than corrupted),
sprintflow backup create|verify|restore, sprintflow console rotate-key, CI for SprintFlow itself (GitLab and GitHub:
lint, types, tests with PostgreSQL, pip-audit, Bandit, Gitleaks, manifests, Trivy, SBOM, cosign), the
runbooks and the security guide with least-privilege permissions for every
integration.
Browser proof#
After the independent review approves, a Browser proof stage demonstrates the story in a real browser and attaches the recording to the same Jira ticket.
- Which tickets. Understand lists
ui_scenarios- concrete user-visible flows. Back-end-only tickets have none and skip this stage at zero cost (e2e.mode: alwaysuses the acceptance criteria instead). - Start the app. From
e2e.startin.sprintflow/config.yml, the repo's PlaywrightwebServer, or auto-detection (ASP.NET Core web projects, Next.js, Vite, Angular, Create React App, Django). SprintFlow waits untilbase_url + ready_pathanswers and stops the whole process group afterwards. - Write the tests. Claude writes Playwright specs (TypeScript), one test per scenario, each stage wrapped
in
proofStep(page, 'Criterion …', …)which shows an on-screen caption so the video explains itself. Withinspect_pagesit first reads the live page's accessibility tree, so selectors match the real UI instead of guesses. It can write only in the spec directory - never application code. - Run and record in 4K. Chromium with a 1920×1080 layout at device scale factor 2, recorded at
3840×2160 by default,
slowMoso people can follow it, trace kept on failure. The repo's own Playwright config is reused if it has one (via an overlay that sets the resolution, one browser project and a JSON report). - Fix loop (bounded). Failures go back with the error and the page's accessibility tree at the moment of
failure. If Claude concludes the application is wrong, it reports an
app_defect: the pipeline fixes the app through Implement, re-verifies and re-reviews, then runs the browser tests again. - Attach the proof. Each passing scenario's video is attached as
KEY-proof-N-<scenario>.mp4, with a full-resolution PNG for every proof step, followed by a comment listing each scenario → test with ✔/✘ and the video resolution. On final failure the failing video is attached as evidence and the run stops withe2e_failed(or opens the PR with a note whenrequired: false). - Keep the automation. Repos that already use Playwright get the specs committed in the PR under
<testDir>/sprintflow/. Otherwise SprintFlow uses a private, git-excluded harness with pinned@playwright/testand attaches the spec to Jira (commit_tests: alwaysalso commits it undere2e/sprintflow/).
Video quality. Playwright's built-in recorder encodes VP8 at a fixed 1 Mbit/s in real-time mode (verified in
Playwright 1.63's source) - text and fine UI detail smear at 1080p, and no setting changes it. So SprintFlow records
itself: the sprintflow-proof helper (which specs import instead of @playwright/test; SprintFlow rewrites the
import if a spec doesn't) captures Chromium's screencast frames at JPEG quality 95, and SprintFlow encodes them with
ffmpeg as H.264 High, yuv420p, CRF 18 (visually lossless) (preset slow up to 1440p, medium at 4K for
encode speed at the same quality), using the frames' real timestamps so
pauses and animations play at true speed. Resolution is viewport × device_scale_factor. The default is 4K
(3840×2160): the page keeps its normal 1920×1080 desktop layout, rendered at twice the pixel density, so text is
razor-sharp when zoomed or paused. Set 1.333 for 1440p or 1.0 for 1080p to trade sharpness for smaller files and
faster runs. To fit max_attachment_mb (default 50 MB) the quality factor is raised first
(CRF 22 → 26 → 30); only then does resolution step down one tier at a time (1440p → 1080p → 720p), and the
Jira comment says so. Playwright's own video is kept only as a fallback. browser_executable / browser_args
point Playwright at a specific Chromium or Chrome (corporate builds, air-gapped runners).
Measured in a real Chromium run (a payroll-style form; each video's final frame compared with the full-resolution screenshot of the same state): SprintFlow' 4K recording scored PSNR 47.6 dB / SSIM 0.9993; Playwright's built-in recording of the same run scored 36.5 dB / 0.9941 - it cannot exceed 1080p and leaves compression halos around text.
Test credentials come from e2e.env in sprintflow.yml (use ${file:...} secrets). The model is told the
variable names only; values never enter a prompt, and the app process doesn't receive them. Use a dedicated
test account and a seeded test database (e2e.env in the repo config, e.g. ASPNETCORE_ENVIRONMENT: E2E) -
never production data. Team-specific guidance (login flow, test tenant, data-testid conventions) goes in
.sprintflow/prompts/e2e.md.
The runner needs Node and Chromium with its system libraries: use the Playwright image
(mcr.microsoft.com/playwright:v1.63.0-noble) as BASE_IMAGE plus your toolchain, or let SprintFlow run
npx playwright install chromium. If Node, the browser or the app is unavailable, the stage is skipped with a
note - the change was still verified by the build and test suite.
Source providers#
Set source.provider in sprintflow.yml (the old github: block still works). Every provider opens (or
updates, never duplicates) a draft review request, and none of them has any way to merge or approve.
| Provider | repository |
Draft mechanism | Auth |
|---|---|---|---|
github (+ Enterprise: api_url, web_url) |
owner/name |
native draft PR; title prefix if the plan lacks drafts | fine-grained token |
gitlab (+ self-managed: web_url) |
group/subgroup/project |
Draft: title (GitLab's native mechanism; blocks merge) |
project access token, api scope |
bitbucket |
workspace/repo_slug |
native draft |
access token, or username + app password / API token |
bitbucket-server (web_url) |
PROJECT_KEY/repo_slug |
native draft on 8.18+, title prefix before |
HTTP access token |
azure-devops (+ Server: web_url) |
organization/project/repository |
native isDraft |
PAT, Code read/write |
gitea / Forgejo (web_url) |
owner/name |
WIP: title (native; blocks merge) |
access token + username |
generic |
clone_url (HTTPS or SSH) |
via your webhook_url, else the verified branch is pushed |
token or the runner's SSH key |
custom |
your choice | your plugin (plugin: "pkg.module:Class") |
your plugin |
Details that differ by host and are handled for you:
- Descriptions are cut to each host's limit (Azure DevOps: 4,000 characters), pointing to the run report.
- Labels: GitHub, GitLab, Azure DevOps and Gitea apply them; Bitbucket has no PR labels.
- Reviewers: GitHub/Gitea usernames; GitLab usernames resolved to IDs; Bitbucket Cloud account IDs or
{uuid}; Bitbucket Server user slugs; Azure DevOps identity GUIDs. An unknown reviewer never blocks the PR. draft_fallback: title(default) marks a request by title when a host can't create a native draft;failstops instead. In every case merging remains a human action.- SSH clone URLs use the runner's SSH agent/key, which is passed to git only - never to the team's build/test commands.
- Generic webhook contract. SprintFlow POSTs
{"event": "sprintflow.review_request", "repository", "branch", "base", "title", "body", "draft": true, "labels", "reviewers"}withAuthorization: Bearer <webhook_token>and expects{"url": "..."}back. That's enough to open a Gerrit change, a CodeCommit PR or anything else with a few lines of script. - Custom plugin. A class taking the
SourceConfig, withname,clone_url,git_auth_header(),default_branch()andopen_draft_pr(branch, base, title, body) -> PullRequest. SprintFlow refuses to load a plugin that exposes anything named merge or approve.
Any tech stack#
stack.py finds projects at the repository root, and in sub-directories (two levels) for stacks the root
doesn't already cover - so a .NET solution with a React ClientApp/, or a Go service beside a Node front end,
gets every stack built and tested. Same-stack sub-projects are left to the root build (solutions, Maven
reactors, npm/Cargo workspaces, go ./...). In a monorepo every project's tests run even if one fails.
| Stack | Detected from | Per-test report |
|---|---|---|
| .NET | *.sln, *.csproj/fsproj/vbproj |
TRX |
| Node (npm/yarn/pnpm/bun) | package.json + lockfile |
jest JSON, vitest JUnit, node:test JUnit, mocha xunit; other runners by exit code |
| Python (pip/uv/poetry) | pyproject.toml, requirements*.txt, setup.py |
pytest JUnit |
| Java/Kotlin | pom.xml (Maven, mvnw), build.gradle(.kts) (gradlew) |
Surefire / Gradle JUnit |
| Go | go.mod |
go test -json |
| Rust | Cargo.toml |
cargo test output |
| Ruby | Gemfile (+ spec/) |
RSpec JSON |
| PHP | composer.json + phpunit.xml |
PHPUnit JUnit |
| Swift / C, C++ | Package.swift / CMakeLists.txt |
xUnit / ctest JUnit |
| Elixir, Dart/Flutter, Makefile | mix.exs, pubspec.yaml, Makefile with test: |
exit code |
Reports are written under $SPRINTFLOW_REPORTS (a folder excluded from git through .git/info/exclude, never
the repo's .gitignore), so nothing SprintFlow produces can leak into a pull request. Build and test byproducts
that aren't git-ignored are also cleaned after every verification.
The toolchain has to be installed where SprintFlow runs. deploy/Dockerfile.any layers SprintFlow onto any
Debian/Ubuntu image - ideally the one your CI already uses - so whatever builds in CI builds here. If a tool is
missing, the run stops at the baseline with toolchain_missing, naming the tool, before any tokens are spent.
Self-healing#
Every remedy is decided in code and bounded; only auto-resume re-uses tokens already budgeted.
| Problem | What SprintFlow does |
|---|---|
| Claude / Jira / source-host outage, rate limit, network failure mid-run | Auto-resumes the run from its last completed stage with exponential backoff (auto_resume_attempts) |
| Routed model overloaded or unavailable | Retries (SDK), then switches to fallback_model for that call |
| Package registry / network blip in setup, build or tests | Retries the command with backoff (transient_command_retries) |
| The change edits a dependency manifest | Re-runs setup before building |
| Build fails on a missing dependency | Re-runs setup once and rebuilds |
| New test failures | Re-runs; tests that pass on re-run are reported as flaky instead of starting a paid fix round. Tests memory knows to be flaky get two extra re-runs |
| Transcript grows too large for the context window | Compacts aggressively and retries the call |
| Implement hits its turn/time cap after making edits | Hands the edits to Verify instead of stopping |
| Base branch moved during the run | Rebases and re-verifies; on conflict or failure keeps the original base and says so in the PR |
| A branch with our name exists with someone else's commits | Pushes under a new name; never overwrites |
| Workspace lost or corrupted between resumes | Re-clones at the base commit and re-applies the saved patch |
| Fix rounds cancel out to an empty change | Clear stop no_changes instead of an empty PR |
| Memory file corrupted | Quarantines it and starts fresh; memory errors never fail a run |
Everything healed is listed under Self-healing actions in the PR and the run report.
Token economy#
| Technique | Effect |
|---|---|
Model routing (model.routes) |
Understand runs on a small model by default in the example config; memory consolidation too |
Repo map (repomap.py) |
Ranked files with symbols, computed locally and boosted by memory, handed to Understand and Plan so the model doesn't spend turns exploring |
| Transcript compaction | Old bulky tool results and file contents become one-line stubs once the transcript passes ~22k tokens; measured 47% less re-sent context on a simulated 30-turn step. Done in one jump so the prompt cache is rarely invalidated |
Log distillation (healing.distill) |
Fix rounds get the lines that explain a failure, not the raw tail; measured 54x smaller on a noisy .NET build log |
| Lean fix rounds | No raw ticket, capped diff, distilled feedback |
| Per-step knowledge slices | Understand gets guides, Implement gets rules and skills, Review gets rules only; guides/skills ranked by relevance and capped |
| Output caps | 2k / 4k / 8k / 4k tokens for understand / plan / implement / review |
| Batched tools | read_files and apply_edits do in one turn what took several; reads capped at 400 lines unless a range is asked for |
| Flaky-test healing, toolchain checks | Avoid whole paid fix rounds and runs that could never succeed |
| Prompt caching | System prompt, tools and transcript prefix are cached |
| Memory | A fixed ≤1,500-character block replaces rediscovering the same facts every run |
Per-step token counts and cost are in the run report; totals are in every PR's run details.
Memory#
One memory file per repository in runs_dir/memory/, updated after every finished run and consolidated
from time to time.
- Learned deterministically, at zero token cost, after each run: command durations and failure rates, flaky tests, tests already failing on the base branch, self-heals needed, outcomes and cost, which files were changed for which ticket vocabulary (only from runs that produced a verified, reviewed PR), and raw observations (verify failures, blocker/major review findings, stops).
- Consolidated periodically: every
consolidate_every_runsruns, or when observations have waitedconsolidate_every_days(checked after runs and bysprintflow poll), one small capped model call merges observations into at mostmax_lessonsshort lessons. Lessons re-confirmed gain weight. It works on a snapshot and merges back, so parallel runs never block or lose updates. - Used: a block of at most
max_prompt_charsin Understand/Plan/Implement (lessons, flaky and pre-existing failures, files changed for similar tickets) and Review ("recurring issues to check"); a relevance boost in the repo map; extra re-runs for known-flaky tests. - Hygiene: facts that stop recurring expire after
ttl_days; sizes are capped; memory is presented to the model as data from past runs, never as instructions, and the team's written rules win on any conflict.
sprintflow memory show # lessons, flaky tests, commands, heals, outcomes, cost
sprintflow memory consolidate # consolidate now
sprintflow memory forget --lesson 3 # or --flaky <test>, --files, --all
Team knowledge#
Optional. Lives in the target repository, reviewed like code. The AI can read it but can never write to
.sprintflow/.
.sprintflow/
config.yml setup/build/test/lint commands, test report globs, protected paths, env (all optional)
rules/*.md coding rules - the reviewer checks every change against these (loaded first)
guides/*.md engineering guides
skills/*.md reusable how-tos ("adding an endpoint")
prompts/<step>.md extra instructions for understand | plan | implement | review
Pinning a command (other fields are still auto-detected):
test: "pnpm vitest run --reporter=junit --outputFile.junit=$SPRINTFLOW_REPORTS/vitest.xml"
All fields: setup, build, test, lint, test_reports, protected_paths, env, pr_title_template,
auto_detect (default true). See examples/team-knowledge/.sprintflow/.
Why test reports matter. With per-test reports, Verify knows exactly which tests regressed, which pre-existing failures to ignore, and whether tests that used to pass disappeared (a model "fixing" tests by deleting them fails verification). Without them, only exit codes can be compared, and if the suite already fails on the base branch the result is flagged as inconclusive in the PR.
Verify also refuses to start if the base branch does not set up or build (baseline_build_broken): a change
cannot be verified against a broken baseline.
Operator configuration#
examples/sprintflow.yml is fully commented. Highlights:
model.provider:anthropicorbedrock;model.review_modelfor a separate reviewer model.model.pricing: optional; it drives every cost cap. When it is left out, each model is priced at the Anthropic list price, and an unknown model at the highest list price. Set it (ormodel.prices) for Bedrock, Vertex or a negotiated rate.limits.steps.<step>:max_seconds,max_cost_usd,max_turnsfor understand, plan, implement, verify, review.limits:max_run_cost_usd,max_fix_attempts,max_review_rounds,max_clarification_rounds,command_timeout_seconds.model.routes,model.fallback_model,model.prices: per-step models, automatic fallback, per-model pricing.healing: bounds for every automatic recovery (see Self-healing).memory: consolidation cadence, lesson count, prompt ceiling, cost cap, expiry (see Memory).checkpoints: any ofafter_understand,after_plan,before_pr.sandbox.command_wrapper: how team commands are isolated (see Security model).- Secrets:
${file:/run/secrets/name}(preferred) or${ENV_VAR}/${ENV_VAR:-default}.
Command line#
sprintflow run PROJ-123 [--base develop] [--dry-run]
sprintflow resume <run-id>
sprintflow poll # resume runs waiting on a person + time-based memory upkeep
sprintflow approve <run-id> [--checkpoint after_plan]
sprintflow reject <run-id> --reason "..."
sprintflow status [<run-id>]
sprintflow validate-config
sprintflow check-repo <path> # detected stack(s), commands, reports, team knowledge
sprintflow memory show|consolidate|forget
sprintflow console serve | init-key | user add <name> --role admin | import sprintflow.yml
Global: --config (default sprintflow.yml or $SPRINTFLOW_CONFIG). Logging: SPRINTFLOW_LOG_LEVEL,
SPRINTFLOW_LOG_FORMAT=json.
Exit codes: 0 draft PR opened or waiting on a person, 2 clear stop, 1 failure.
People can also answer from Jira: reply to SprintFlow' questions in a comment, or reply
/sprintflow approve / /sprintflow reject <reason> at a checkpoint. sprintflow poll picks these up.
Outcomes#
| Outcome | Codes |
|---|---|
| Draft PR | - |
| Waiting on a person | clarification, checkpoint |
| Clear stop (patch kept if any) | too_vague, already_implemented, cannot_locate, no_changes, tests_still_failing, review_not_approved, budget_exceeded, rejected_at_checkpoint, baseline_build_broken, baseline_setup_failed, toolchain_missing, unknown_stack, invalid_team_config, e2e_failed, dry_run |
| Failed (unexpected error; resumable; auto-resumed if transient) | exception type + message |
Every run directory contains:
<runs_dir>/<run-id>/
state.json current state (atomic writes)
events.jsonl append-only audit log: every step, tool call, command, cost
workspace/ isolated clone the AI works in
artifacts/ report.md, report.json, pr_body.md, knowledge.json, baseline.json,
verify-N.json, logs/*.log, security/*.sarif, latest.patch, <KEY>.patch
<runs_dir>/memory/<owner__repo>.json per-repository memory
Security model#
What the AI can do. Read files in its workspace; in the implement step, create/edit/delete files there.
Paths are resolved (symlinks included) and must stay inside the workspace; .git, .sprintflow and
team-declared protected_paths are read-only. There is no shell tool. Refusals are returned to the model and
logged.
What it cannot do. Run commands, touch other directories, change its own rules or the commands used to
judge it, push anywhere but its own sprintflow/ branch, or merge. No provider has a merge, approve or
ready-for-review method (a test checks every adapter, and custom plugins that expose one are refused). The push uses an explicit lease, so commits a person pushed to the branch are never
overwritten.
Prompt injection. Ticket text, comments, file contents and diffs are wrapped as data and the prompts say to treat them as such. This lowers the risk; it does not remove it. The structural controls above (no shell, sandboxed paths, draft-only output, independent review, human merge) are what bound the impact.
The important residual risk: Verify runs AI-written code. Build and test commands execute whatever the model wrote. SprintFlow scrubs credentials from those commands' environment, but in production also:
- Keep secrets in files, not environment variables. A process can read the environment of other
processes running as the same user via
/proc/<pid>/environ. Use${file:/run/secrets/...}. - Run team commands as a different user with
sandbox.command_wrapper. The Docker images create abuilderuser that cannot read/run/secrets:command_wrapper: ["sudo", "-n", "-E", "-H", "-u", "builder", "--"]. - Run each ticket in an ephemeral container with restricted network egress (package feeds only).
- Use least-privilege tokens scoped to one repository that can push branches and open pull/merge
requests and nothing more (GitHub fine-grained
contents+pull_requestswrite; GitLab project access token; Bitbucket repository access token; Azure DevOps PAT with Code read/write); a Jira bot account that can only browse and comment on the relevant projects. - Protect
sprintflow/*branches from triggering deploys in your CI.
Deployment#
deploy/Dockerfile.any- recommended: SprintFlow layered onto any Debian/Ubuntu toolchain image via--build-arg BASE_IMAGE=...(use your CI image; a multi-toolchain image for polyglot monorepos).deploy/Dockerfile.dotnet- a ready-made .NET 8 SDK variant.deploy/Dockerfile- minimal image (Python, git, ripgrep) for repos whose toolchain is already there.deploy/docker-compose.yml- single host with Docker secrets and a persistent runs volume.- CI triggers for every platform:
deploy/github-actions-trigger.yml,gitlab-ci-trigger.yml,bitbucket-pipelines-trigger.yml,azure-pipelines-trigger.yml,Jenkinsfile.trigger- start runs (manual or a Jira automation webhook) and schedulesprintflow poll.
runs_dir must be persistent: it holds run state (for pause/resume) and the per-repository memory. One process drives a run at a time (file
lock); different runs can execute in parallel.
Operating it#
- Resume. A crashed or paused run continues from its last completed stage. Half-finished edits are rolled back to the last committed state first. Attempt counters persist, so resuming never resets the fix or review limits.
- Costs. Every LLM call is priced and recorded per step;
status <run-id>shows the breakdown. Prompt caching keeps multi-turn steps cheap. - Audit.
events.jsonlrecords every stage, step, tool call (arguments truncated), command (exit code, duration) and outcome. - Tuning. Start with
checkpoints: [after_plan]while the team builds trust, then remove it.
Development#
make install # pip install -e ".[dev]"
make check # ruff + mypy + pytest
The test suite (403 tests) runs the whole pipeline end to end against real git repositories and real test
runners (pytest, and Node's built-in runner for a zero-config non-Python repo), with a scripted Claude and
in-memory Jira and source host. It covers the happy path, fix and review loops, clarification and checkpoints,
crash/resume, sandbox escapes and budget stops. It also covers every self-healing path in the table above,
detection for every listed stack, parsers for every report format, token-economy behaviour (routing, caps,
compaction, distillation, lean fix rounds, repo map) and memory (learning across runs, consolidation,
concurrency, decay, corruption, CLI). Every source provider (GitHub, GitLab, Bitbucket Cloud and Server, Azure
DevOps, Gitea, generic webhook, custom plugin) is tested at the HTTP level, and a full run goes through the
generic provider to a real git remote. Browser proof is tested end to end with a real app server, real ffmpeg
video conversion, Playwright-shaped reports (from a real Playwright 1.63 run) and multipart Jira upload, and the
generated Playwright configs are run by genuine Playwright 1.63. With a Chromium available (SPRINTFLOW_TEST_CHROMIUM),
a real-browser test runs the whole pipeline through real Chromium and asserts 4K H.264 output and measured
quality against the full-resolution screenshot. The web console has API tests (auth, rate limiting, CSRF, roles,
session invalidation, encrypted secrets, validation, connection tests, runs, approvals, artifact ranges and path
traversal, audit, analytics, diff) and a 16-step real-browser walkthrough in Chromium: sign in, configure and test
connections, start a run and follow it to a draft PR, timeline and diff tabs, play the 4K proof video and open the
screenshot lightbox, approve a checkpoint, command palette, dashboard charts, table search/filter/sort, users and audit,
dark mode, and confirm a viewer is read-only and injected markup never executes. Sprint automation is tested
against the Jira Agile API at the HTTP level with real pipeline runs (sprint detection, scope, rank order, strictly
sequential execution, mid-sprint additions, pause, reorder, no re-runs, error reporting), plus a real-browser
walkthrough: link Jira, see my stories, check Jira, and watch the queue implement them one by one. Claude Code settings
authentication is tested with real settings files and real helper scripts (precedence, apiKeyHelper rotation on a
rejected key, Bedrock by profile/API key/credential export, SSO refresh only when expired, Vertex, gateway, untrusted
repository settings, the claude.ai-login case without reading the token), including the exact Bedrock request sent,
plus a real-browser walkthrough of detection and AWS sign-in in the console. PR follow-up is tested against real git
branches (fix pushed on top, human commits kept, unverified fixes never pushed, CI failures, mentions, other bots,
round limits, merged PRs) and at the HTTP level against the GitHub and GitLab APIs. The manager dashboard is tested
for sprint progress, per-person grouping, project and people statistics, attribution, the project filter and roles,
plus a real-browser walkthrough (progress, expanding a person, CSV export, filter, no page overflow). The Jira client is tested at the HTTP level; the Claude client
(including fallback) is tested through the real Anthropic SDK with only the transport mocked.
Before going live#
Verified in development: everything in the test suite above.
Not yet verified against live systems - do these once before production:
- A dry run and then a real run against a sandbox Jira project and a throwaway repository on each host you
use. Provider adapters were built from each host's API documentation and tested against mocked HTTP,
not against live GitLab/Bitbucket/Azure DevOps/Gitea instances. Check draft behaviour, reviewers and
git-over-HTTPS auth (especially Gitea's
usernameand Bitbucket Server's Bearer header for git). - The model IDs for your provider (Bedrock uses its own model / inference-profile IDs), including the routed and fallback models, and your pricing for each.
sprintflow check-repoon each target repository: confirm the detected commands match what CI runs (only Python and Node repositories were executed for real during development; other stacks' commands are tested as generated strings, not executed).- Your external code indexer's output format (JSON
[{path,line,text}]orpath:line:text). - That your security scanner (e.g. Aikido) can write SARIF to a file path for
security_scan.command. - Building the Docker images and the
sudosandbox wrapper on your runner (not exercised in the test suite). - Browser proof on your runner. Verified in development with a real Chromium 153 (full pipeline: app server,
Playwright 1.63, recorder, captions, step screenshots, 4K encode,
inspect_page). Still run one UI ticket on your runner image to confirm login via.sprintflow/prompts/e2e.md, your app's start command, and the MP4 preview and size limit in your Jira. Confirm thee2e.startcommand and thatmax_attachment_mbis under your Jira attachment limit. - After ~10 real runs,
sprintflow memory show: check the lessons are specific and correct;forgetany that aren't.