GitHub Spec Kit vs. phax
A side-by-side comparison. Spec Kit details were fetched from the current upstream docs (
/github/spec-kit, Context7) rather than from memory.
The one-line difference
Spec Kit is a methodology + prompt/template layer that runs inside your AI agent. phax is a standalone orchestration engine that runs your AI agent as a subprocess. Same spec → plan → implement philosophy; very different altitude and enforcement.
What Spec Kit is, concretely
You run specify init, which scaffolds your repo with slash-command templates
for whichever agent you use. Then, inside Copilot / Claude / Gemini / Cursor /
etc., you drive a sequence of slash commands:
Each produces a markdown artifact in your repo (constitution, spec, plan, tasks). It is agent-agnostic by design — essentially a shared vocabulary of prompts + templates plus a workflow definition with human approve/reject gates between stages. The runtime is your agent, working in your normal tree with its native permissions.
What phax is, concretely
A compiled CLI (Node / Effect / TypeScript) that drives an AI coding agent
(Claude Code by default, Codex, or Mistral Vibe) through isolated, gated phases.
You author a plan.md, phax deterministically extracts it to phax-plan.json,
then phax run executes each phase in its own git worktree, runs a
mechanical gate profile after each phase with an automatic same-session fix
loop, reconciles the planned files against the actual diff, and (per phax.json)
produces a compliance review and a GitHub PR. phax is the runtime; it spawns
the agent headlessly.
Side-by-side
The key distinction
The workflows rhyme — both are spec → plan → tasks/phases → implement with review gates. The difference is what a gate means:
- Spec Kit gate = "a human reads the artifact and clicks approve."
Enforcement is social/manual;
/analyzeand/checklistare AI consistency prompts, not executable checks. - phax gate = "the phase's code typechecks and tests pass, or the phase fails and the agent loops to fix it." Enforcement is a real process exit code, in an isolated worktree, with deterministic reconciliation of plan-vs-actual.
So Spec Kit is closer to a shared SDD discipline + prompt library you layer onto any agent; phax is closer to a build system / CI harness for agent work. Spec Kit standardizes the conversation and artifacts; phax standardizes the execution and verification.
When each fits
- Spec Kit if you want a lightweight, agent-agnostic methodology that meets you where you already work, across many assistants, with humans in the review loop. Low ceremony, broad reach.
- phax if you want mechanical gates, worktree isolation, security sandboxing,
provider routing, and a deterministic, auditable trajectory (compliance report
- PR) — i.e., you trade agent-breadth and lightness for enforcement and safety.
Two things worth noting
- They are composable, and phax already mirrors Spec Kit's front half. phax's
/phax-specand/phax-planningskills are direct analogs of/speckit.specifyand/speckit.plan. You could author Spec-Kit-style specs and execute them through phax's gated harness — spec discipline on the front, mechanical execution on the back. - Relative to the "review-by-trajectory" desktop idea (see
docs/ideas/desktop-app.md): Spec Kit lives at "structure the prompts and artifacts"; phax lives at "make execution isolated and gates mechanical." The idea's "approve the trajectory, not the diff" is a step beyond both — it presumes exactly the mechanical gates + reconciliation phax has (and Spec Kit does not), then makes them the primary review surface. You cannot review-by-trajectory credibly when your gates are human approve/reject prompts; you need real green-or-red evidence, which is precisely what phax produces.