A team of agents working one plan.

Hand Sprint Engine a backlog item or an epic. An architect plans it against your codebase, specialists build the tasks in parallel, and reviewers gate every one — you watch the board move and stay the final gate.

How Sprint Engine works.

  1. A plan becomes a graph

    The architect takes an approved plan — or writes one against your codebase — and splits it into tasks with dependencies, owners, and acceptance criteria.

  2. The engine takes the wheel

    Sprint Engine is an MCP server. It decides which tasks unlock and when, owns all the concurrency, and drives the sprint from the first task to the last.

  3. Specialists run in parallel

    Several agents work different tasks at the same time. The engine schedules the right specialist for each one — from the built-in roster, or a custom roster you define yourself.

  4. Epic in, shipped work out

    Hand it a two-week, whole-team epic. It runs autonomously to done — every task gated by reviewers, every step on the record.

plansbuilds in parallelgatesdoneAArchitectwrites the planplan.mdFFrontendown tasksDDeveloperown tasksDDeveloper 2own tasksRReviewevery taskTTestingacceptanceDonewith evidence

Tasks flow from the plan through parallel specialists, review, and testing — all at the same time.

Sprints: a team of agents working one plan.

A sprint takes a backlog item from plan to reviewed code. You can inspect all of it while it runs — the board, each agent's terminal, and an inbox where every decision waits for you.

The architect plans first

Hand Sprint Engine a task, a backlog item, or an epic. An architect agent reads your codebase and knowledge graph, writes a plan, and splits it into a dependency graph of tasks.

Specialists, not one agent

Developers, frontend engineers, testers, and reviewers each hold their own terminal and their own slice of the work. The roster shows who is running, what they own, and when they need you.

Reviewers gate every task

Reviewer agents check each task and leave the evidence behind — commands run, tests passed, files touched. Approvals land in one inbox; you stay the final gate.

Sprints6 runs

The Sprints panel, straight out of the studio — filter the lenses, search, pick a run.

Configure the roster: any CLI, any model, per role.

Every role on the roster picks its own runtime and model — Claude Code, Codex, Cursor, Kimi, Grok, GLM, Qwen, or any Claude Code-compatible CLI — billed through the subscriptions you already pay for, not metered API keys.

Sharp models hold the pen

Put your most capable model — Fable-class — in the architect seat. The plan it writes is the plan every task inherits.

Cheaper models carry the build

Implementation burns most of the tokens. Hand it to faster, cheaper models working from a plan sharp enough to follow.

Sharp models gate the result

The same top-tier models review every task and send focused fixes back — steering the builders instead of babysitting them.

The result: the expensive model's judgment applied to every task, at a fraction of the tokens.

Roster5 members · 3 working
  • Architect1 seat
    • ArchitectReviewing plan sign-off · MC-1502
  • Developer2 seats
    • Developer 2Wizard scroll fix · MC-1502
    • Developer 3Creation hub modal · MC-1506
  • Tester1 seat
    • Tester 1Waiting for testable work
  • Code reviewer1 seat
    • ReviewerWaiting on you · gate on MC-1499

The roster from a live run — click a role's model chip to change what sits in the seat.

What the reviewer gates actually catch.

Every number below comes from our own run records — the sprints that built this product. The single biggest thing reviewers catch is not a typo or a crash: it's work that doesn't match the requirements.

17%
of review attempts sent work back
roughly 1 in 6 — changes requested by a reviewer agent
40%
of gated tasks were reworked
at least one round of changes before merging
30%
of gates took 2+ review rounds
reviewers keep pushing until it actually passes
23%
of findings were requirements violations
the #1 category — built-the-wrong-thing, not typos

Findings by kind

Reviewer agents record what they catch at verdict time. Nearly a quarter of everything caught is a requirements violation — the work ran, the tests passed, and it still wasn't what was asked for. That's the class of defect a single-agent workflow ships.

  • Requirements violations23%
  • Uncategorised (other)19%
  • Test gaps17%
  • Code bugs13%
  • Reliability issues12%
  • Documentation gaps6%
  • Security issues5%
  • Performance issues3%
  • Accessibility issues3%

Measured across a sample of 136 sprint runs — 1,401 tasks, 2,981 review attempts, and 876 reviewer findings from the projects that build SprintEngine itself. Aggregated 2026-07-08 from each run's on-disk review records. Per-model comparisons will follow once runtime stamping covers the full history.

The run summary.

When the sprint finishes you get the whole story: who worked on what and for how long, which gates passed, what the reviewers caught, and a delivery score per role — then one button to open the pull request.

cairn
41
InboxRosterTasksSummary
14/14Ready for review
Run complete
Review uncommitted workspace changes and manually test the feature.
52 files touched · 41 commands recorded · 60 validation results
Open pull request
0%0 / 14
Tasks done
0%0 / 21
Gates passed
Tasks completed over the run0 → 14
start2h 10m
Agent activity 9click an agent to solo
Files touched
0
Commands
0
Validations
0
Gates passed
0 / 21
Issues caught in review 13flagged by reviewers during the run
Bugs0
Factual errors0
Missed requirements0
Implementation mistakes0
Unsafe changes0
Regressions0
Delivery score

Weighted review issues per completed task — lower is better.

Developer
12
2 issues · 6 tasks
Frontend
36
8 issues · 5 tasks
Task clarity92%
Acceptance clarity90%
Context fit85%
Role fit100%

Download SprintEngine Studio.

Download the studio for free and bring the subscriptions you already have.