Compare commits
39 Commits
v1.8.1
...
4cbf426df2
| Author | SHA1 | Date | |
|---|---|---|---|
| 4cbf426df2 | |||
| 8a8cc14e0d | |||
| 1907281b40 | |||
| cf9fe6d80f | |||
| 0c17ef1454 | |||
| 0767c6b4c7 | |||
| b78b3cd6a3 | |||
| 474c76e564 | |||
| 161e424632 | |||
| 09aa565559 | |||
| 0501260098 | |||
| 48a7cf68bd | |||
| 30860a7654 | |||
| fda0146969 | |||
| 75605e5e41 | |||
| 3b18b42e30 | |||
| 8e457fa6e4 | |||
| 6740e26261 | |||
| aedfe944b2 | |||
| 9b3c06b498 | |||
| 8fe1bfa805 | |||
| 68542fd778 | |||
| 4656a6ebb7 | |||
| 7e9de755f1 | |||
| 4b20f93bb9 | |||
| 5b3b84f0a9 | |||
| 1015b99a6a | |||
| 20a9e2f040 | |||
| 28bda05cbf | |||
| 5274bdbd51 | |||
| 25d874d87d | |||
| 5e48e98293 | |||
| 8e0b1a2ab2 | |||
| c2579a7b6a | |||
| 9d1104d105 | |||
| 72fed09aab | |||
| 37311a7e68 | |||
| c92865c395 | |||
| 63656d339d |
@@ -0,0 +1,134 @@
|
||||
# Architecture
|
||||
> Part of [AGENTS.md](../AGENTS.md) — project guidance for AI coding agents.
|
||||
|
||||
## Directory Structure
|
||||
|
||||
```
|
||||
plan2code/
|
||||
├── src/ # Source workflow prompts (11 markdown files)
|
||||
│ ├── plan2code-0-pathfinder-references/ # Reference files for pathfinder skill
|
||||
│ │ ├── chart.md # MODE A: destination + frontier grills, templates
|
||||
│ │ ├── grilling.md # Folded-in grilling + domain-modeling
|
||||
│ │ ├── questions.md # On-disk question-file format + markers
|
||||
│ │ ├── resolve.md # Per-type resolution + graduating the fog
|
||||
│ │ ├── handoff.md # Clearing gate + PLAN-DRAFT handoff
|
||||
│ │ └── trail.md # Every-response map visual + pathed resume command
|
||||
│ ├── plan2code-review-references/ # Reference files for review skill
|
||||
│ │ ├── verification-protocol.md # Deep verification + confidence calibration
|
||||
│ │ ├── dimensions.md # 11 dimensions with detailed checklists
|
||||
│ │ ├── false-positives.md # Known false-positive patterns
|
||||
│ │ └── session-end.md # Next-step routing at review session end
|
||||
│ ├── plan2code-init-update-references/ # Reference files for init-update skill
|
||||
│ │ └── ai-agent-file-sync.md # Step 7: replace AI configs with AGENTS.md refs
|
||||
│ └── plan2code-4-finalize-references/ # Reference files for finalize skill
|
||||
│ └── community-feedback-submission.md # STEP 6.5 payload schema + submission tiers
|
||||
├── plan2code-loop/ # Autonomous loop CLI tool (Node.js/TypeScript)
|
||||
│ ├── src/ # TypeScript source
|
||||
│ └── dist/ # Built output (tsup)
|
||||
├── plan2code-metrics/ # Recursive self-improvement toolchain
|
||||
│ ├── src/ # TypeScript source
|
||||
│ │ └── prompts/ # Internal AI prompt templates (no char limit)
|
||||
│ └── dist/ # Built output (tsup)
|
||||
├── src/statusline-claude/ # Claude CLI status line (Node.js, zero deps, single file)
|
||||
│ ├── statusline.js # Self-contained: config, git, formatters, render
|
||||
│ └── statusline-config.json # Default config template
|
||||
├── scripts/ # Development scripts
|
||||
│ └── validate-char-count.js # Pre-commit character count validator
|
||||
├── dist/ # Generated distribution files (auto-generated)
|
||||
│ ├── global-commands/ # For global installation (~/.claude/, etc.)
|
||||
│ └── local-commands/ # For per-project installation (.claude/, etc.)
|
||||
├── .husky/ # Git hooks (husky)
|
||||
│ └── pre-commit # Runs character count validation
|
||||
├── .claude/ # Repo-local Claude Code config (NOT installed by install.js)
|
||||
│ └── skills/ # Maintainer-only dev skills, e.g. plan2code-publish/
|
||||
├── docs/ # Documentation and assets
|
||||
├── specs/ # Feature specs (if any in-progress)
|
||||
├── install.js # Interactive installer (Node.js)
|
||||
├── package.json # Root package (husky only, private: true)
|
||||
├── version.json # Version metadata
|
||||
└── README.md # User documentation
|
||||
```
|
||||
|
||||
## Key Files
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `install.js` | Main installer - generates and installs workflow files to AI tool directories |
|
||||
| `src/plan2code-*.md` | Source workflow prompts (the "source of truth") |
|
||||
| `scripts/validate-char-count.js` | Pre-commit validator ensuring all source prompts ≤ 11,000 chars |
|
||||
| `version.json` | Version metadata (name, version, description) |
|
||||
| `QUICK-REFERENCE.md` | User quick-reference card |
|
||||
| `src/statusline-claude/` | Claude CLI status bar (included in `A` Install All + dev tools; also via Custom → S) |
|
||||
|
||||
## Workflow Prompts (in `src/`)
|
||||
|
||||
| File | Step | Purpose |
|
||||
|------|------|---------|
|
||||
| `plan2code-init.md` | Init | Generate AGENTS.md as index + `.agents-docs/` section files (progressive discovery) |
|
||||
| `plan2code-init-update.md` | Update | Update AGENTS.md with learnings; detects and routes edits to `.agents-docs/` files |
|
||||
| `plan2code-0-pathfinder.md` | 0 | Chart a foggy idea as a local map of decision questions under `specs/<idea>/pathfinder/`, resolve one per session, hand a seeded PLAN-DRAFT to Step 1 |
|
||||
| `plan2code-quick-task.md` | quick | Lightweight planning for small tasks (standalone — not a pipeline step) |
|
||||
| `plan2code-1-plan.md` | 1 | Requirements analysis & architecture |
|
||||
| `plan2code-1b-revise-plan.md` | 1b | Mid-implementation revisions |
|
||||
| `plan2code-2-document.md` | 2 | Create implementation specs |
|
||||
| `plan2code-3-implement.md` | 3 | Execute implementation (phase by phase) |
|
||||
| `plan2code-review.md` | review | Post-implementation comprehensive review |
|
||||
| `plan2code-4-finalize.md` | 4 | Validate, summarize, feedback, archive (steps 1–7, +optional 6.5) |
|
||||
| `plan2code-handoff.md` | handoff | Compact the conversation into a self-contained handoff document for a fresh session |
|
||||
|
||||
## Naming Convention
|
||||
|
||||
Workflow files follow a strict naming pattern:
|
||||
- **Utilities:** `plan2code-<name>.md` (single dash)
|
||||
- **Numbered steps:** `plan2code-<N>-<name>.md` (single dash, number, single dash)
|
||||
|
||||
Examples:
|
||||
- `plan2code-init.md` (utility)
|
||||
- `plan2code-1-plan.md` (step 1)
|
||||
- `plan2code-1b-revise-plan.md` (step 1b)
|
||||
|
||||
## Reference Files
|
||||
|
||||
Some workflows use companion reference files for depth that exceeds the 11k char limit. The orchestrator (main workflow file) loads them via `Read` directives during execution.
|
||||
|
||||
**Pattern:** `src/<source-filename-without-extension>-references/` (e.g., `plan2code-review-references/`)
|
||||
|
||||
**How the installer handles them:**
|
||||
- **Skill-directory platforms** (Claude Code, Agents, Crush, Devin): reference files are nested as `<skill-name>/references/`. Read paths use canonical `references/<file>.md`.
|
||||
- **Flat-file platforms** (Windsurf, Cursor, Copilot, Continue): reference files are placed as a sibling directory. The installer rewrites Read paths to the sibling directory name (e.g., `plan2code-review-references/<file>.md`).
|
||||
|
||||
Reference files are NOT subject to the 11,000 character limit. The review workflow pioneered this pattern (`verification-protocol`, `dimensions`, `false-positives`, `session-end`); the init-update workflow also uses it (`ai-agent-file-sync` for its Step 7), `plan2code-4-finalize.md` uses it for STEP 6.5 (`community-feedback-submission`), and `plan2code-0-pathfinder.md` leans on it hardest (`chart`, `grilling`, `questions`, `resolve`, `handoff`, `trail` — the orchestrator is a dispatcher, the depth lives in the references). Other workflows can adopt it when a source file's detail exceeds the 11k limit.
|
||||
|
||||
**`Read` directives must sit at column 0.** `install.js` matches `/^(Read\s+)references\//gm` for the flat-file path rewrite — anchored, with no leading-whitespace tolerance. An indented or bulleted `Read` line is silently skipped, so flat-file platforms ship a `references/` path that does not exist there (they receive the reference dir as a *sibling*, named `plan2code-<name>-references/`).
|
||||
|
||||
## Repo-Local Skills (`.claude/skills/`)
|
||||
|
||||
Maintainer-only Claude Code skills committed to the repo but **deliberately excluded** from `install.js` — they are dev tooling, not shipped product, so they never install to `~/.claude/skills/` and carry no version bump of their own (a product-version bump would wrongly imply a user-facing release); changelog mentions fold into the current version's entry.
|
||||
|
||||
- `plan2code-publish/` — cuts a GitHub Release from the top `CHANGELOG.md` entry once `CHANGELOG.md` / `version.json` / `package.json` agree and the version is ahead of the latest published release. Delegates tag creation to `gh release create --target main`.
|
||||
|
||||
**Warning:** anything named `plan2code-*` placed under `~/.claude/skills/` is deleted by the installer's uninstall (`uninstallFiles()` in `install.js`) and by every re-install's pre-copy cleanup in `install()` (both match `/^plan2code-/` for the Claude Code skills target). Keep these skills repo-local only.
|
||||
|
||||
## Status Line
|
||||
|
||||
Optional Claude Code status bar living in `src/statusline-claude/`. Three-line bar (icon + content per line) showing model, project, branch, uncommitted diff stats, session duration + cost, context window usage, and plan/quota usage.
|
||||
|
||||
**Design constraints:**
|
||||
- **Zero runtime dependencies** — `statusline.js` is self-contained (config loader, git helpers, formatters, render). Copied verbatim to `~/.claude/plan2code-statusline.js` on install; no bundler step.
|
||||
- **Stdin-driven** — all data comes from Claude Code's stdin JSON (`model`, `workspace`, `context_window`, `rate_limits`, `cost`). No API calls, no auth, no background processes.
|
||||
- **Silent failure** — outer `try/catch` around `main()` plus `process.exit(0)` on missing stdin guarantees the script never crashes the CLI. All git ops are timeout-bounded (1.5s) and non-git workspaces short-circuit via `fs.existsSync('.git')`.
|
||||
- **Atomic settings writes** — installer writes `~/.claude/settings.json` via temp file + rename so a crash never leaves the file truncated.
|
||||
- **Custom-config respect** — installer detects non-plan2code `statusLine` entries, prompts before replacing, and backs up to `statusline-previous.json`. Uninstall only removes `settings.statusLine` if it points to the plan2code bundle.
|
||||
|
||||
**Layout:**
|
||||
|
||||
```
|
||||
src/statusline-claude/
|
||||
├── statusline.js # Self-contained: config, git, formatters, render
|
||||
├── statusline-config.json # Default config template
|
||||
└── README.md # User docs: install, config, debugging
|
||||
```
|
||||
|
||||
**Adaptive plan-usage display:** the formatter auto-selects between `5h/7d` rate-limit percentages (Pro/Max/Teams — when `rate_limits` present in stdin) and `Nk in · Nk out` session-token counts (Bedrock/Vertex/PAYG — when `rate_limits` absent). Segment is hidden when neither shape is available.
|
||||
|
||||
**Installer integration** lives in `install.js` under the `STATUS LINE INSTALLATION` section (`installStatusLine`, `uninstallStatusLine`). Included in `A` (Install All + dev tools); also available individually via Custom → `S`.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Code Style & Gotchas
|
||||
> Part of [AGENTS.md](../AGENTS.md) — project guidance for AI coding agents.
|
||||
|
||||
## Code Style
|
||||
|
||||
- **install.js:** CommonJS, Node.js built-ins only (no external deps), ANSI colors via `COLORS` constant, readline-based prompts
|
||||
- **plan2code-loop & plan2code-metrics:** TypeScript + ESM, built with tsup (target ES2022, moduleResolution: bundler)
|
||||
- External deps: `@inquirer/prompts`, `chalk`, `execa`, `ora`
|
||||
- Interactive CLI via `@inquirer/prompts` (select, input, confirm)
|
||||
- **File operations:** Synchronous fs in all packages
|
||||
|
||||
## Gotchas / Pitfalls
|
||||
|
||||
- **Version sync:** When adding a new version to `CHANGELOG.md`, also update `version.json` and `package.json` (root) to match. Check `README.md` for any version badges or references that need updating. The installer displays the version from `version.json` in its header. All three files (`CHANGELOG.md`, `version.json`, `package.json`) must always show the same version number.
|
||||
- **CHANGELOG ordering:** Entries in `CHANGELOG.md` must be in reverse-chronological order — newest version at the top, oldest at the bottom. New entries are always inserted immediately after the file header.
|
||||
- **CHANGELOG house format is not Keep a Changelog:** version headings are `## vX.Y.Z` — with a `v` prefix and **no date**. The release date is recorded only in `version.json`'s `releaseDate`. Category headings carry emoji: `### ✨ Added`, `### 🔧 Changed`, `### 🐛 Fixed`, `### 💥 Breaking`, `### 🗑️ Removed`, `### 📚 Documentation`. Older entries contain one-off variants (`🎁 Added`, `📦 Updated`, `📝 Documentation`, `🏎️ Improved`, `🧪 Testing`) — do not introduce new ones. The `.claude/skills/plan2code-changelog/` skill automates version selection and formatting; use it rather than hand-rolling an entry.
|
||||
- **PowerShell mangles the CHANGELOG emoji:** `Get-Content` / `Select-String` render the `###` heading emoji as `?` under the default Windows console encoding, so a heading audit done that way reports garbage. Read `CHANGELOG.md` with a file-read or grep tool instead.
|
||||
- **`.claude/skills/` is tracked, not ignored:** repo-local skills (`plan2code-changelog`, `plan2code-publish`, `sync-repo`) live there and are committed. Nothing in `.gitignore` touches `.claude/`, so a new skill only needs `git add`. Per the Failure Log convention in `AGENTS.md`, a correction that is a *workflow* rather than a rule belongs here as a skill, linked from AGENTS.md — not as a Failure log line.
|
||||
- **Loop `.gitignore` setup:** `ensureGitignore()` runs at startup in `Controller.run()` as a pre-flight step, not just inside `createTaskCommit()`. This is critical for phase mode where the Node controller doesn't handle commits — without it, `git add -A` would stage spec files.
|
||||
- **Workflow file character limit:** All `src/plan2code-*.md` files must be ≤ 11,000 characters. A husky pre-commit hook enforces this. The 11,000 limit leaves buffer for platform-specific YAML headers (106-142 chars) to stay under Windsurf's 12,000 char limit.
|
||||
- **Metrics internal prompts have no char limit:** Files in `plan2code-metrics/src/prompts/` are NOT subject to the 11,000 char limit — only `src/plan2code-*.md` consumer-facing prompts are.
|
||||
- **User Feedback table format:** The `## User Feedback` markdown table in `overview.md` has a strict format the collector regex depends on. Field names must be exactly `Rating`, `Reason`, `Went Well`, `Went Poorly`. Pipe characters in values must be escaped as `\|`.
|
||||
- **PLAN-DRAFT confidence numbers are scraped by regex:** when a `specs/<feature>/PLAN-DRAFT-*.md` contains no `<!-- METRICS_JSON ... -->` comment, `collector.ts` falls back to prose scraping (`collector.ts:186-241`). The overall-confidence pattern requires a literal `%`, but the four *breakdown* patterns (`collector.ts:201-204`) do **not** — `/[Rr]equirements?[:\s|]+(\d{1,2})/` and its siblings match a bare dimension word followed by whitespace, a colon, or a pipe and then digits. So a PLAN-DRAFT written by anything other than `/plan2code-1-plan` Phase 7 must keep both the `%` sign **and** bare `Requirements` / `Feasibility` / `Integration` / `Risk` followed by a number off the page — including innocent table rows like `| Requirements | 11 |`. Otherwise the metrics pipeline records a planning-step confidence that no planning step produced. `/plan2code-0-pathfinder` works around this by hyphenating the labels (`Requirements-clarity 22/25`), which breaks the character class.
|
||||
- **Reference file sizing guideline:** Files in `src/plan2code-*-references/` directories target ~100-200 lines each (soft guideline; evaluate splitting above 300). They are NOT subject to the 11,000 character limit. The pre-commit hook (`validate-char-count.js`) only checks `src/plan2code-*.md` flat files — subdirectory contents are automatically excluded.
|
||||
- **The splitting guideline has a hard ceiling — reference files cannot always be split:** two constraints bound it. (1) Each new reference costs the orchestrator a column-0 `Read references/<file>.md` line plus its fallback blockquote (~150-200 chars), and orchestrators near the 11,000 limit have no room to spend. (2) **The path rewrite is not recursive** — `syncPrompts()` rewrites `Read references/…` paths on the *orchestrator's* content only, so a `Read references/…` directive placed *inside* a reference file is never rewritten for flat-file targets; it ships as a dangling instruction pointing at a path that does not exist there. When a reference legitimately exceeds 300 lines (e.g. `plan2code-0-pathfinder-references/chart.md`), that is an accepted trade-off, not an oversight.
|
||||
@@ -0,0 +1,101 @@
|
||||
# Development Commands
|
||||
> Part of [AGENTS.md](../AGENTS.md) — project guidance for AI coding agents.
|
||||
|
||||
## Common Commands
|
||||
|
||||
```bash
|
||||
# Install dev dependencies (sets up husky pre-commit hooks)
|
||||
npm install
|
||||
|
||||
# Run the interactive installer (always interactive — any CLI args are silently ignored)
|
||||
node install.js
|
||||
|
||||
# Plan2Code Loop
|
||||
cd plan2code-loop && npm install # First time setup
|
||||
cd plan2code-loop && npm run build # Build the CLI
|
||||
|
||||
# Plan2Code Metrics
|
||||
cd plan2code-metrics && npm install # First time setup
|
||||
cd plan2code-metrics && npm run build # Build the CLI
|
||||
```
|
||||
|
||||
## Installer Menu Options
|
||||
|
||||
**Main menu:**
|
||||
|
||||
| Option | Action |
|
||||
|--------|--------|
|
||||
| `I` | Install Plan2Code workflow prompts for all platforms, plus `plan2code-loop` CLI |
|
||||
| `A` | Everything in `I` plus `plan2code-bot`, `plan2code-metrics` (dev tools), and Claude Code status line |
|
||||
| `U` | Uninstall Plan2Code files: prompts + `plan2code-loop` + `plan2code-metrics` + `plan2code-bot` + Claude Code status line (confirmation required) |
|
||||
| `C` | Open CUSTOM sub-menu |
|
||||
| `Q` | Quit |
|
||||
|
||||
**CUSTOM sub-menu (`C`):**
|
||||
|
||||
| Option | Action |
|
||||
|--------|--------|
|
||||
| `L` | Show local (per-project) install instructions |
|
||||
| `O` | Install plan2code-loop CLI only |
|
||||
| `M` | Install plan2code-metrics CLI only |
|
||||
| `S` | Install Claude Code status line only |
|
||||
| `B` | Install plan2code-bot CLI only |
|
||||
| `Q` | Return to main menu |
|
||||
|
||||
## How the Installer Works
|
||||
|
||||
1. **Reads source prompts** from `src/plan2code-*.md`
|
||||
2. **Generates platform-specific files** with appropriate headers (YAML frontmatter for some platforms)
|
||||
3. **Writes to `dist/`** subdirectories organized by destination type
|
||||
4. **Copies to target directories** (global: `~/.claude/commands/`, etc.)
|
||||
|
||||
## Platform-Specific File Formats
|
||||
|
||||
| Platform | Extension / File | Local Dir | Global Dir | Header |
|
||||
|----------|-----------------|-----------|------------|--------|
|
||||
| Claude Code | `.md` | — | — | None |
|
||||
| Cursor | `.md` | — | — | None |
|
||||
| Copilot CLI | `.md` | — | — | YAML frontmatter |
|
||||
| Continue | `.prompt.md` | — | — | YAML frontmatter |
|
||||
| Windsurf | `.md` | — | — | YAML frontmatter |
|
||||
| VS Code Copilot | `.prompt.md` | — | — | YAML frontmatter |
|
||||
| Codeium | `.md` | — | — | YAML frontmatter |
|
||||
| Claude Code (Skills) | `SKILL.md` in subdir | `.claude/skills/<skill-name>/` | `~/.claude/skills/<skill-name>/` | YAML frontmatter + `disable-model-invocation: true` |
|
||||
| Agent Skills (Amp · Devin · OpenCode · Zed) | `SKILL.md` in subdir | `.agents/skills/<skill-name>/` | `~/.agents/skills/<skill-name>/` | YAML frontmatter (no disable flag) |
|
||||
| Crush | `SKILL.md` in subdir | — (global only) | `~/.config/crush/skills/<skill-name>/` (Unix) / `%LOCALAPPDATA%\crush\skills\<skill-name>\` (Windows) | YAML frontmatter |
|
||||
| Pi (pi.dev) | `.md` | `.pi/prompts/` | `~/.pi/agent/prompts/` | YAML frontmatter (`description`) |
|
||||
|
||||
## Editing Workflow Prompts
|
||||
|
||||
When modifying workflow prompts in `src/`:
|
||||
|
||||
1. Edit the source file in `src/`
|
||||
2. Run `node install.js` to regenerate distribution files
|
||||
3. Test the workflow in your AI tool of choice
|
||||
4. The `dist/` folder is regenerated automatically — don't edit files there directly
|
||||
|
||||
## Adding a New Workflow Prompt / Skill
|
||||
|
||||
Adding a new prompt to `src/` is more than dropping in a file — the installer is
|
||||
driven by an explicit registry and several docs enumerate the command set. When
|
||||
you add a `src/plan2code-<name>.md`, do **all** of the following so nothing drifts:
|
||||
|
||||
1. **Create the source file** `src/plan2code-<name>.md` — body content **only**,
|
||||
no YAML frontmatter (the installer generates frontmatter per platform). Keep
|
||||
it **under 11,000 characters** (`npm test` enforces this).
|
||||
2. **Register it in the installer.** Add an entry to the `SOURCE_PROMPTS` array in
|
||||
`install.js` (`source`, `stepNumber`, `name`, `displayName`, `description`, and
|
||||
`isUtility: true` for non-numbered utilities). If `stepNumber` is a non-numeric
|
||||
label (e.g. `'handoff'`), add a matching case to `generateStepLabel()` so the
|
||||
generated description reads correctly.
|
||||
3. **Update every doc that lists the command set** — keep these in sync, they are
|
||||
the canonical inventories:
|
||||
- `README.md` — command table ("When to Use")
|
||||
- `QUICK-REFERENCE.md` — Commands table
|
||||
- `.agents-docs/AGENTS-architecture.md` — "Workflow Prompts (in `src/`)" table
|
||||
- `CHANGELOG.md` — add an entry under the current version
|
||||
- `docs/index.html` — **only** if the new prompt belongs to the core pipeline
|
||||
shown there; utilities (like `init`, `quick-task`, `handoff`) are deliberately
|
||||
omitted from that curated marketing list.
|
||||
4. **Regenerate and validate:** run `node install.js` (regenerates `dist/`) and
|
||||
`npm test` (character-count validator now covers the new file).
|
||||
@@ -0,0 +1,47 @@
|
||||
# Plan2Code Loop
|
||||
> Part of [AGENTS.md](../AGENTS.md) — project guidance for AI coding agents.
|
||||
|
||||
A separate Node.js CLI tool that autonomously implements specs by looping through tasks.
|
||||
|
||||
## Loop Architecture
|
||||
|
||||
The loop uses an **LLM-driven discovery** approach:
|
||||
- Node app just orchestrates iterations and parses completion markers
|
||||
- The LLM reads spec files (`overview.md`, `phase-X.md`) to discover tasks
|
||||
- The LLM finds unchecked checkboxes, implements ONE task per iteration, marks it complete
|
||||
- No regex parsing of markdown in Node - the AI handles all task discovery
|
||||
|
||||
## Loop Commands
|
||||
|
||||
```bash
|
||||
# Build the loop CLI
|
||||
cd plan2code-loop && npm run build
|
||||
|
||||
# Run the loop (after linking) - fully interactive
|
||||
plan2code-loop
|
||||
```
|
||||
|
||||
The CLI auto-detects specs in `./specs/`, prompts for selection if multiple found, and handles session continuation interactively. Session state is stored per-spec in `specs/<feature>/.plan2code-loop/`.
|
||||
|
||||
## Loop Modes
|
||||
|
||||
The CLI asks users to choose a loop mode:
|
||||
- **One task per loop** (default) - Each agent invocation implements exactly one task. The Node controller handles git commits.
|
||||
- **One phase per loop** - Each agent invocation implements all remaining tasks in the current phase. The LLM handles git commits (with JIRA ticket ID if provided). The controller parses multiple completion markers from a single iteration.
|
||||
|
||||
## Completion Markers
|
||||
|
||||
The LLM must output one of these formats:
|
||||
- `TASK_COMPLETE: 1.1 - Task description` - Task done successfully
|
||||
- `TASK_BLOCKED: 1.1 - Reason` - Cannot complete task
|
||||
- `PHASE_COMPLETE` - Current phase finished (phase mode only)
|
||||
- `LOOP_COMPLETE` - All phases finished
|
||||
|
||||
## Key Source Files
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `plan2code-loop/src/controller.ts` | Main loop orchestrator |
|
||||
| `plan2code-loop/src/prompt/templates.ts` | Prompt templates for both loop modes |
|
||||
| `plan2code-loop/src/utils/git.ts` | `createTaskCommit()` — handles task-mode commits with footer |
|
||||
| `plan2code-loop/src/cli.ts` | Interactive CLI entry point |
|
||||
@@ -0,0 +1,87 @@
|
||||
# Plan2Code Metrics
|
||||
> Part of [AGENTS.md](../AGENTS.md) — project guidance for AI coding agents.
|
||||
|
||||
A recursive self-improvement toolchain for plan2code contributors. Collects run metrics, aggregates by prompt generation, diagnoses weak steps via AI, and proposes surgical prompt edits.
|
||||
|
||||
## Metrics Data Flow
|
||||
|
||||
```
|
||||
Collect → Aggregate → Analyze → Improve → Apply
|
||||
```
|
||||
|
||||
1. **Collector** reads project artifacts (`specs/<feature>/`) → writes `RunMetrics` JSON per run
|
||||
2. **Aggregator** groups runs by prompt SHA fingerprint (cohorts) → `aggregated.json`
|
||||
3. **Analyzer** invokes AI with aggregated metrics + prompt contents → diagnosis markdown
|
||||
4. **Improver** invokes AI with diagnosis → validated `PromptEdit[]` proposals (char limit + verbatim checks)
|
||||
5. **Applier** shows interactive diffs → patches `src/plan2code-*.md` files
|
||||
|
||||
## Metrics Commands
|
||||
|
||||
```bash
|
||||
cd plan2code-metrics && npm run build # Build the CLI
|
||||
plan2code-metrics # Run (fully interactive, no flags)
|
||||
```
|
||||
|
||||
## Metrics CLI Menu
|
||||
|
||||
| Option | Action |
|
||||
|--------|--------|
|
||||
| Collect | Read spec artifacts → run JSON |
|
||||
| Import | Copy run JSON from another project |
|
||||
| View | Display cohort metrics with health indicators |
|
||||
| Analyze | AI diagnosis of weak metrics |
|
||||
| Propose | AI improvement proposals with validation |
|
||||
| Apply | Interactive diff review + file patching |
|
||||
| Fetch community submissions | List/parse/import open community-feedback GitHub issues from jparkerweb/plan2code, close on success |
|
||||
|
||||
Community submissions arrive as GitHub issues labeled `community-feedback` on `jparkerweb/plan2code`, created by the finalize prompt's post-Step-6 submission flow; the "Fetch community submissions" option requires an authenticated `gh` CLI to list/close them.
|
||||
|
||||
## Key Source Files
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `types.ts` | All interfaces (`RunMetrics`, `UserFeedback`, `CohortMetrics`, etc.) + `METRIC_TARGETS` |
|
||||
| `collector.ts` | Reads project artifacts → run JSON (parses plan drafts, overview.md, loop logs) |
|
||||
| `aggregator.ts` | Merges runs by prompt generation (SHA cohort) → `aggregated.json` |
|
||||
| `community.ts` | Lists/parses/closes `community-feedback`-labeled GitHub issues via `gh` CLI |
|
||||
| `analyzer.ts` | AI diagnosis via `prompts/analyze.md` template |
|
||||
| `improver.ts` | AI proposals via `prompts/improve.md` + validation (char count, old_text match) |
|
||||
| `applier.ts` | Interactive diff review + file patching |
|
||||
| `cli.ts` | Menu-driven interactive CLI (100% prompts, no flags) |
|
||||
| `invoke-llm.ts` | Unified LLM interface (Claude Code, GitHub Copilot CLI, or Devin CLI) |
|
||||
|
||||
## User Feedback
|
||||
|
||||
The collector parses an optional `## User Feedback` table from `overview.md`:
|
||||
|
||||
```markdown
|
||||
## User Feedback
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| Rating | 8 |
|
||||
| Reason | Smooth workflow |
|
||||
| Went Well | Planning was thorough |
|
||||
| Went Poorly | Some tasks unclear |
|
||||
```
|
||||
|
||||
Feedback is collected during finalize (Step 5) or retroactively via the CLI. Pipe characters in values are escaped as `\|`. The aggregator computes `avg_user_rating` and `feedback_count` per cohort.
|
||||
|
||||
## Supported AI Agents
|
||||
|
||||
- **Claude Code** (recommended): `claude` CLI with `--inputFile` for prompt delivery
|
||||
- **GitHub Copilot CLI**: `copilot` CLI with stdin prompt delivery
|
||||
- **Devin CLI**: `devin` CLI with `--print --prompt-file <file> --permission-mode dangerous`
|
||||
|
||||
## Metric Targets
|
||||
|
||||
| Metric | Target | Direction |
|
||||
|--------|--------|-----------|
|
||||
| `avg_confidence` | ≥ 90 | higher is better |
|
||||
| `avg_task_completion_rate` | ≥ 0.95 | higher is better |
|
||||
| `avg_blocker_count` | ≤ 1.5 | lower is better |
|
||||
| `avg_completion_marker_success_rate` | ≥ 0.95 | higher is better |
|
||||
| `avg_verification_failures_found` | ≤ 1.0 | lower is better |
|
||||
| `archival_success_rate` | ≥ 0.99 | higher is better |
|
||||
| `avg_user_rating` | ≥ 7.0 | higher is better |
|
||||
|
||||
Data stored in `.plan2code-metrics/` (runs/, aggregated.json, proposals/).
|
||||
@@ -0,0 +1,149 @@
|
||||
---
|
||||
name: plan2code-changelog
|
||||
description: "Validate and fix the CHANGELOG.md version number before opening a PR. Reads main branch to determine the current latest version, classifies changes on the current branch, and proposes the correct next semver. Use this skill when the user mentions changelog, version number, preparing a PR, release version, semver check, or says 'check the changelog', 'what version should this be', 'prepare for PR', or 'fix the version'. Also use proactively when you notice a CHANGELOG entry that may have an incorrect version number."
|
||||
---
|
||||
|
||||
# Plan2Code Changelog Validator
|
||||
|
||||
Ensure the CHANGELOG.md entry for the current branch has the correct semver version before a PR is opened. This skill exists because parallel branches independently pick version numbers that collide or leap-frog when merged — this validates against main's actual state right before the PR.
|
||||
|
||||
## Workflow
|
||||
|
||||
### Step 1 — Gather state
|
||||
|
||||
Run these commands to understand the current situation:
|
||||
|
||||
```bash
|
||||
# 1. Current latest version on main
|
||||
git show main:CHANGELOG.md | head -20
|
||||
|
||||
# 2. Current branch name (for ticket ID extraction)
|
||||
git branch --show-current
|
||||
|
||||
# 3. What this branch changed (commit subjects)
|
||||
git log main...HEAD --oneline
|
||||
|
||||
# 4. Files changed on this branch
|
||||
git diff main...HEAD --name-only
|
||||
```
|
||||
|
||||
Extract from main's CHANGELOG:
|
||||
- The **latest version number** (first `## vX.Y.Z` line)
|
||||
|
||||
Note: on Windows, PowerShell's console encoding mangles the emoji in the `###` headings to `?`. Read `CHANGELOG.md` with the file-read or grep tool rather than `Get-Content` / `Select-String` when you need to see them.
|
||||
|
||||
Extract from the branch:
|
||||
- The **list of changed files** to classify the change type
|
||||
- The **commit messages** for changelog entry content
|
||||
|
||||
### Step 2 — Classify the change
|
||||
|
||||
Determine the change type by examining what was modified on this branch:
|
||||
|
||||
| Signal | Classification | Version Bump | Heading |
|
||||
|--------|---------------|--------------|---------|
|
||||
| An install target removed, or an existing workflow's contract broken | Breaking change | **Major** (X.0.0) | `### 💥 Breaking` |
|
||||
| New workflow prompt (`.md` file under `src/`) | New workflow | **Minor** (x.Y.0) | `### ✨ Added` |
|
||||
| New capability added to an existing prompt, or a new `.claude/skills/` skill | New capability | **Patch** (x.y.Z) | `### ✨ Added` |
|
||||
| Behavioral changes to existing prompt(s), installer, or docs | Behavior change | **Patch** (x.y.Z) | `### 🔧 Changed` |
|
||||
| Bug fix to existing prompt(s) or tooling | Bug fix | **Patch** (x.y.Z) | `### 🐛 Fixed` |
|
||||
| A prompt, target, or file deleted | Removal | **Patch** (x.y.Z) | `### 🗑️ Removed` |
|
||||
| README / `.readme/` / docs-site only | Documentation | **Patch** (x.y.Z) | `### 📚 Documentation` |
|
||||
| Mix of the above | Use the **highest** bump (major > minor > patch) | Combine headings |
|
||||
|
||||
Use only the headings in this table — the CHANGELOG has historical one-off variants (`🎁 Added`, `📦 Updated`, `📝 Documentation`, `🏎️ Improved`, `🧪 Testing`) that should not be introduced in new entries.
|
||||
|
||||
### Step 3 — Compute the correct version
|
||||
|
||||
Starting from main's latest version:
|
||||
- **Major bump:** increment the first number, reset the rest (e.g., `1.16.1` → `2.0.0`)
|
||||
- **Minor bump:** increment the middle number, reset patch to 0 (e.g., `2.0.0` → `2.1.0`)
|
||||
- **Patch bump:** increment the last number (e.g., `2.1.0` → `2.1.1`)
|
||||
|
||||
### Step 4 — Check the current branch's CHANGELOG
|
||||
|
||||
Read the current `CHANGELOG.md` on the branch. Look for:
|
||||
|
||||
0. **You are on `main` with no diff** — there is no branch to validate. Instead, compare the top CHANGELOG version against `git log` since the commit that released it: if commits have landed on `main` without a CHANGELOG entry, treat those commits as the change set and continue from Step 2. Say so explicitly rather than reporting "nothing to do."
|
||||
|
||||
1. **No entry exists yet for this branch's work** — the branch hasn't added a version entry above main's latest. Proceed to Step 5 to draft one.
|
||||
|
||||
2. **An entry exists but the version is wrong** — the branch has a version entry, but it doesn't match the computed correct version (common when branches were rebased or other PRs merged first). Report the discrepancy:
|
||||
|
||||
```
|
||||
Version check for branch: {branch-name}
|
||||
|
||||
Main is at: {main-version}
|
||||
Branch claims: {branch-version}
|
||||
Correct version: {computed-version} ({classification})
|
||||
|
||||
The version needs to be updated: {branch-version} → {computed-version}
|
||||
```
|
||||
|
||||
Ask: "Update the version to {computed-version}? (yes / no)"
|
||||
|
||||
3. **An entry exists and the version is correct** — report success:
|
||||
|
||||
```
|
||||
Version check for branch: {branch-name}
|
||||
|
||||
Main is at: {main-version}
|
||||
Branch version: {branch-version} ({classification})
|
||||
|
||||
Version is correct. CHANGELOG is ready for PR.
|
||||
```
|
||||
|
||||
Stop here unless the user asks for content changes.
|
||||
|
||||
### Step 5 — Draft or fix the CHANGELOG entry
|
||||
|
||||
**If no entry exists**, draft a new one based on the commits and changed files. Match this repo's house format exactly — a bare `## vX.Y.Z` heading with **no date**, emoji `###` headings from the Step 2 table, and a blank line between bullets:
|
||||
|
||||
```markdown
|
||||
## {computed-version-with-v-prefix}
|
||||
|
||||
### ✨ Added
|
||||
|
||||
- **{prompt-or-area}** — {what changed, and why it matters to someone installing it}
|
||||
|
||||
### 🔧 Changed
|
||||
|
||||
- **{prompt-or-area}** — {what changed}
|
||||
```
|
||||
|
||||
Conventions to follow, drawn from existing entries:
|
||||
|
||||
- Bold lead-in naming the prompt, file, or area, then an em dash (`—`), then the description.
|
||||
- Reference prompts by their command (`/plan2code-1-plan`) or path (`src/plan2code-init.md`), not by informal name.
|
||||
- One paragraph per bullet is fine — this CHANGELOG favours substantive entries over terse one-liners, and a bullet may carry extra indented paragraphs for detail.
|
||||
- A release with a big theme may open with a one-line summary paragraph directly under the `## vX.Y.Z` heading, before the first `###`.
|
||||
|
||||
Insert the new section directly below the `All notable changes...` line and above the previous version's heading.
|
||||
|
||||
Present the draft and ask for approval before writing.
|
||||
|
||||
**If the version is wrong**, update only the version number — preserve the existing content unless the user asks for content changes too.
|
||||
|
||||
After any changes, show the final CHANGELOG entry for confirmation.
|
||||
|
||||
### Step 6 — Sync `package.json` and `version.json` versions
|
||||
|
||||
After writing or updating the CHANGELOG entry, update the `version` field in `package.json` at the repo root to match the computed version:
|
||||
|
||||
1. Read `package.json` and `version.json` then check the current `version` value.
|
||||
2. If it already matches the computed version, skip — no change needed.
|
||||
3. If it differs, update the `"version"` field in both files to the computed version (e.g., `"version": "2.1.1"`).
|
||||
4. Set `releaseDate` in `version.json` to today's date in `YYYY-MM-DD`. This is the only place a date is recorded — the CHANGELOG headings carry no date.
|
||||
5. Include `package.json` and `version.json` in the same commit as the CHANGELOG changes.
|
||||
|
||||
This keeps `package.json`, `version.json`, and `CHANGELOG.md` in lockstep so `npm pkg get version` always reflects the latest release.
|
||||
|
||||
## Rules
|
||||
|
||||
- Never create a version entry without checking main first — the whole point is to derive the version from main's current state
|
||||
- Always present changes before writing — the user should see and approve the CHANGELOG entry
|
||||
- Match the existing file's format — `## vX.Y.Z` with no date, emoji `###` headings. Do not introduce Keep a Changelog's `## [x.y.z] - date` style
|
||||
- Write for someone reading the release notes, not for someone reading the diff — say what the change lets them do
|
||||
- If multiple change types exist (Added + Changed), use multiple headings under the same version
|
||||
- `version.json`'s `releaseDate` is today's date, i.e. when the release is being prepared, not when the work started
|
||||
- If the branch has no meaningful changes vs main (e.g., only non-shipping files changed), say so and ask if a CHANGELOG entry is actually needed
|
||||
@@ -0,0 +1,159 @@
|
||||
---
|
||||
name: plan2code-publish
|
||||
description: "Publish a GitHub Release for jparkerweb/plan2code whenever CHANGELOG.md's top version is ahead of the latest published release on GitHub — after verifying CHANGELOG.md, version.json, and package.json all agree on the version. Use this skill when the user says 'publish a release', 'create a GitHub release', 'cut a release', 'tag a release', 'tag and release', 'is the changelog published', or otherwise mentions publishing/releasing/tagging this repo."
|
||||
---
|
||||
|
||||
# Plan2Code Release Publisher
|
||||
|
||||
Publish a GitHub Release for `jparkerweb/plan2code` whenever `CHANGELOG.md`'s top version is ahead of the latest published release on GitHub. This turns the merged CHANGELOG entry on `main` into an actual GitHub Release (which also creates the `vX.Y.Z` git tag).
|
||||
|
||||
**Repo-local by design.** This skill lives in the repo's `.claude/skills/` and is intentionally NOT wired into `install.js` — it is a maintainer dev tool, not part of the shipped product, so it is never installed to `~/.claude/skills/`. Do **not** copy it there: the uninstaller (`uninstallFiles()` in `install.js`) and every re-install's pre-copy cleanup in `install()` both delete every entry matching `/^plan2code-/` under `~/.claude/skills/`, so a copy placed there would be silently removed.
|
||||
|
||||
## Workflow
|
||||
|
||||
### Step 1 — Preflight: clean working tree
|
||||
|
||||
Run `git status --porcelain`. If the output is non-empty, stop immediately — do not switch branches or take any other action. Tell the user:
|
||||
|
||||
> Working tree has uncommitted changes. Commit or stash them, then re-run this skill.
|
||||
|
||||
### Step 2 — Switch to main and pull
|
||||
|
||||
If the working tree is clean, switch to `main` and pull latest. Run these as two separate, non-chained commands (never `&&`/`;`-chain `git`/`gh` commands — Windows PowerShell 5.1 rejects `&&`):
|
||||
|
||||
```bash
|
||||
git checkout main
|
||||
git pull origin main
|
||||
```
|
||||
|
||||
### Step 3 — Read the three version sources
|
||||
|
||||
Read all three files at the repo root and extract each version:
|
||||
|
||||
- `CHANGELOG.md` — parse the first `## vX.Y.Z` heading. Strip the leading `v` → `$CHANGELOG_VERSION`. Note: plan2code CHANGELOG headings are `## vX.Y.Z` (v-prefixed, **no** date and **no** brackets) — different from a `## [x.y.z] - YYYY-MM-DD` format.
|
||||
- `version.json` — the `version` field → `$VERSION_JSON`. Also read its `releaseDate` field → `$RELEASE_DATE` (informational only — shown at the confirm step, never part of the sync gate; use `(none)` if absent).
|
||||
- `package.json` (root) → the `version` field → `$PACKAGE_JSON`.
|
||||
|
||||
Also capture from `CHANGELOG.md` the **section body** for the top version: everything from the `## vX.Y.Z` heading line itself (heading **included**) up to — but not including — the next `## v` heading, or end-of-file if there is none. Call this `$SECTION`. Keep the `## vX.Y.Z` heading in `$SECTION`; plan2code release bodies include it.
|
||||
|
||||
### Step 4 — Version-sync preflight (STOP on mismatch)
|
||||
|
||||
All three versions must be identical. This enforces the repo's documented invariant (`.agents-docs/AGENTS-code-style.md` → "Version sync"): `CHANGELOG.md`, `version.json`, and `package.json` must always show the same version number.
|
||||
|
||||
If `$CHANGELOG_VERSION`, `$VERSION_JSON`, and `$PACKAGE_JSON` are **not** all equal, **stop** — do not read the release, do not publish. Report exactly which files disagree:
|
||||
|
||||
> ⚠️ Version files are out of sync — refusing to publish. The repo requires `CHANGELOG.md`, `version.json`, and `package.json` to match.
|
||||
>
|
||||
> - **CHANGELOG.md:** $CHANGELOG_VERSION
|
||||
> - **version.json:** $VERSION_JSON
|
||||
> - **package.json:** $PACKAGE_JSON
|
||||
>
|
||||
> Fix the mismatch first, then re-run this skill. To realign: pick the intended version (normally the highest / newest CHANGELOG entry) and update the other two files to match — see the "Version sync" gotcha in `.agents-docs/AGENTS-code-style.md`.
|
||||
|
||||
This skill never modifies these files — it only reads and compares them.
|
||||
|
||||
### Step 5 — Determine the highest published release
|
||||
|
||||
List every published (non-draft) release and take the numerically-highest semver tag. Do **not** rely on `gh release view` / GitHub's "Latest" flag: that flag returns whatever release is *marked* latest — normally the newest semver, but a maintainer can manually pin it to an older release, which would make the Step 6 ahead-comparison misfire.
|
||||
|
||||
```bash
|
||||
gh release list --repo jparkerweb/plan2code --limit 100 --json tagName,isDraft -q '.[] | select(.isDraft==false) | .tagName'
|
||||
```
|
||||
|
||||
Strip any leading `v` from each returned tag, compare them as numeric `(major, minor, patch)` tuples (the same rule as Step 6), and take the maximum → `$RELEASE_VERSION`. If the command returns no tags or exits non-zero for **any** reason (no releases yet, transient error, etc.), set `$RELEASE_VERSION = "0.0.0"` — no error-text matching is needed.
|
||||
|
||||
### Step 6 — Compare versions
|
||||
|
||||
Compare `$CHANGELOG_VERSION` vs `$RELEASE_VERSION` as a numeric `(major, minor, patch)` tuple. Never do a plain string/lexicographic compare — e.g. `"1.9.0" > "1.10.0"` is true as strings but wrong numerically.
|
||||
|
||||
### Step 7 — Not ahead: no-op
|
||||
|
||||
If `$CHANGELOG_VERSION` ≤ `$RELEASE_VERSION`, print a simple status message showing both versions and stop:
|
||||
|
||||
> CHANGELOG top version ($CHANGELOG_VERSION) is not ahead of the latest published release ($RELEASE_VERSION). Nothing to publish.
|
||||
|
||||
No error is raised and no release is created.
|
||||
|
||||
### Step 8 — Ahead: compute and confirm
|
||||
|
||||
If `$CHANGELOG_VERSION` > `$RELEASE_VERSION`, compute:
|
||||
|
||||
- `tag = "v$CHANGELOG_VERSION"`
|
||||
- `title = "v$CHANGELOG_VERSION"` (plan2code keeps the `v` prefix in release titles)
|
||||
- `notes` = the header line `# What's New 🎉`, then one blank line, then `$SECTION` verbatim (`$SECTION` already starts with the `## vX.Y.Z` heading). Build this as a real multi-line string with **actual newlines** — the `\n\n` shorthand shown elsewhere means "a blank line," never the literal two-character sequence `\` + `n`. Getting this wrong would run the header and the first CHANGELOG heading together with a stray `\n\n` in the published body.
|
||||
|
||||
Present all three to the user and wait for an explicit answer before any write. Offer the optional decorative title suffix — a plain `vX.Y.Z` title is the default, but a release may append one (e.g. `v1.14.0 - 🔍 Review workflow`):
|
||||
|
||||
> 🚀 [Publish Plan2Code Release]
|
||||
>
|
||||
> CHANGELOG is ahead of the latest published release:
|
||||
> - **Current release:** $RELEASE_VERSION
|
||||
> - **CHANGELOG top version:** $CHANGELOG_VERSION
|
||||
> - **version.json releaseDate:** $RELEASE_DATE (informational — read from `version.json`)
|
||||
>
|
||||
> Proposed release:
|
||||
> - **Tag:** $tag
|
||||
> - **Title:** $title
|
||||
> - **Notes:**
|
||||
> ```
|
||||
> $notes
|
||||
> ```
|
||||
>
|
||||
> Publish this release? Reply **yes** to publish as-is, **no** to cancel, or provide a decorative suffix to append to the title (e.g. `⇢ 🎆 Feature Name` or `- 🔍 Feature Name`).
|
||||
|
||||
If the user supplies a suffix, set `title = "v$CHANGELOG_VERSION " + <suffix>` (single space join) and proceed to publish. The tag and notes are unaffected by the suffix.
|
||||
|
||||
### Step 9 — Publish (on approval)
|
||||
|
||||
On approval, write `$notes` to a temp file — never pass multiline text inline via `--notes`, that regresses into a quoting bug — then create the release targeting `main`. Two requirements for the temp file: `$notes` must already hold **real newlines** (per Step 8) because `printf '%s'` / `WriteAllText` write it byte-for-byte — a literal `\n` in the string lands literally in the release body; and it MUST be **UTF-8** because the notes contain emoji (`🎉`, `🐛`, `✨`, `🔧`).
|
||||
|
||||
**bash (preferred in this environment):**
|
||||
|
||||
```bash
|
||||
NOTES_FILE=$(mktemp)
|
||||
printf '%s' "$notes" > "$NOTES_FILE"
|
||||
gh release create "$tag" --repo jparkerweb/plan2code --title "$title" --notes-file "$NOTES_FILE" --target main
|
||||
rm -f "$NOTES_FILE"
|
||||
```
|
||||
|
||||
**PowerShell:** do NOT use `Set-Content` — under Windows PowerShell 5.1 it writes ANSI/UTF-16 by default and mangles the emoji into `??`. Write UTF-8 **without BOM** (a BOM would leak into the release body):
|
||||
|
||||
```powershell
|
||||
$NotesFile = [System.IO.Path]::GetTempFileName()
|
||||
[System.IO.File]::WriteAllText($NotesFile, $notes, [System.Text.UTF8Encoding]::new($false))
|
||||
gh release create "$tag" --repo jparkerweb/plan2code --title "$title" --notes-file "$NotesFile" --target main
|
||||
Remove-Item -Path $NotesFile
|
||||
```
|
||||
|
||||
Always delete the temp file afterward, regardless of whether `gh release create` succeeded or failed.
|
||||
|
||||
**Failure handling — already-exists classification:** if `gh release create` exits non-zero, inspect the error text.
|
||||
|
||||
- If and only if it contains the substring `already exists` (real output: `HTTP 422: Validation Failed` / `Release.tag_name already exists`), report this to the user as already published, not as a raw CLI error:
|
||||
|
||||
> This version ($CHANGELOG_VERSION) was already published as a release — nothing more to do.
|
||||
|
||||
- Every other failure (auth, network, permissions, etc.) must be surfaced to the user verbatim. Never silently reclassify a genuine failure as "already published."
|
||||
|
||||
### Step 10 — Verify
|
||||
|
||||
Confirm the release now exists and report its URL:
|
||||
|
||||
```bash
|
||||
gh release view "$tag" --repo jparkerweb/plan2code
|
||||
```
|
||||
|
||||
Report the release URL to the user.
|
||||
|
||||
## Rules
|
||||
|
||||
- **Repo-local only** — this skill is not part of the installed product; never add it to `install.js`, and never copy it to `~/.claude/skills/` (the uninstaller deletes `plan2code-*` entries there).
|
||||
- **Read-only on version files** — never modify `CHANGELOG.md`, `version.json`, or `package.json`; this skill only reads and compares them.
|
||||
- **Version-sync gate is hard** (Step 4) — if the three version sources disagree, stop and report; do not publish a release from an inconsistent repo.
|
||||
- **Never run a local `git tag` or `git push`** — tag creation is delegated entirely to `gh release create --target main`.
|
||||
- **Always use `--notes-file`**, never inline multiline `--notes`, and always write the notes file as UTF-8 (no BOM) so emoji survive.
|
||||
- **Always clean up the temp notes file**, even on a mid-run error.
|
||||
- **Always get explicit approval before any write** (Step 8) — no release is created without a yes (or a yes-with-suffix).
|
||||
- **Derive the current version from the highest published semver tag** (Step 5), never from GitHub's manually-pinnable "Latest" flag. On an empty or failed release list, treat it as "no prior release" (baseline `0.0.0`) — no error-text matching needed.
|
||||
- **On a `gh release create` failure** (Step 9), only reclassify as already-published when the error contains `already exists` — every other failure must be shown verbatim, never swallowed.
|
||||
- **Idempotent and safe to re-run** at any time — re-running after a successful publish hits the Step 7 no-op; re-running after a race-lost publish hits the Step 9 already-exists handling.
|
||||
@@ -0,0 +1,58 @@
|
||||
---
|
||||
name: sync-repo
|
||||
description: "Run the encrypted, password-protected repository sync workflow for this project. The real instructions are stored encrypted at rest and are only revealed in-session after you supply the correct password. Invoke explicitly with /sync-repo."
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
# sync-repo (password-protected)
|
||||
|
||||
The real instructions for this skill are encrypted at rest in `sync-repo.enc`
|
||||
and are NOT readable without the password. Do not guess, reconstruct, or invent
|
||||
the workflow. Follow this launcher exactly.
|
||||
|
||||
## What you (the agent) must do
|
||||
|
||||
1. **Ask the user for the password.** Request the decryption passphrase (via
|
||||
AskUserQuestion or a plain prompt). Do NOT proceed without it. Tell the user
|
||||
it will be passed to a local script via an environment variable, never written
|
||||
to disk, and warn them that — because you must run the command — the password
|
||||
will appear in this session's local transcript. (It never enters the repo.)
|
||||
|
||||
2. **Decrypt to STDOUT only.** Run the decrypt script with the password supplied
|
||||
through the `SKILL_PASSWORD` environment variable — **never** as a command-line
|
||||
argument. From the repo root:
|
||||
|
||||
- **Windows PowerShell:**
|
||||
```powershell
|
||||
$env:SKILL_PASSWORD='<password the user gave you>'; node .\.claude\skills\sync-repo\decrypt.mjs .\.claude\skills\sync-repo\sync-repo.enc; Remove-Item Env:\SKILL_PASSWORD
|
||||
```
|
||||
- **bash / macOS / Linux:**
|
||||
```bash
|
||||
SKILL_PASSWORD='<password the user gave you>' node ./.claude/skills/sync-repo/decrypt.mjs ./.claude/skills/sync-repo/sync-repo.enc
|
||||
```
|
||||
|
||||
3. **Handle the result.**
|
||||
- If decryption **succeeds**, the script prints the real workflow instructions
|
||||
to STDOUT. Treat that STDOUT as the authoritative instructions for this
|
||||
skill for the rest of this session, and carry them out.
|
||||
- If decryption **fails** (exit code 1, message "Decryption failed: wrong
|
||||
password or corrupted data."), the password was wrong or the file is
|
||||
corrupt. Tell the user, ask them to re-enter the password, and retry. Do
|
||||
NOT attempt to reconstruct the instructions from anything else.
|
||||
|
||||
## Hard rules
|
||||
|
||||
- **Never write the decrypted plaintext to a file.** Read it from STDOUT only.
|
||||
`decrypt.mjs` intentionally has no file-output mode.
|
||||
- **Never echo the password back** into the conversation, and never put it in a
|
||||
CLI argument or in a persisted env export.
|
||||
- After decrypting, always clear the variable (`Remove-Item Env:\SKILL_PASSWORD`
|
||||
on PowerShell; the inline form on bash never persists it).
|
||||
|
||||
## Security note (be honest with the user)
|
||||
|
||||
This protects the workflow body **only at rest in the repository**. Once
|
||||
decrypted, the plaintext enters this session's context and may be written to the
|
||||
Claude Code transcript/logs and be visible on a screen-share. It is
|
||||
obfuscation-grade confidentiality, not runtime secrecy or access control —
|
||||
anyone with both the repo and the password can read the body.
|
||||
@@ -0,0 +1,42 @@
|
||||
#!/usr/bin/env node
|
||||
// decrypt.mjs — verifies GCM auth tag, prints plaintext to STDOUT only.
|
||||
// Password from $SKILL_PASSWORD. Exits 1 on wrong password / tamper, leaking nothing.
|
||||
// Usage: SKILL_PASSWORD=... node decrypt.mjs <file.enc>
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { scryptSync, createDecipheriv } from 'node:crypto';
|
||||
|
||||
const MAGIC = Buffer.from('SENC', 'ascii');
|
||||
const VERSION = 0x01;
|
||||
const SCRYPT = { N: 1 << 17, r: 8, p: 1, maxmem: 256 * 1024 * 1024 };
|
||||
const KEYLEN = 32, SALTLEN = 16, IVLEN = 12, TAGLEN = 16;
|
||||
const HEADER = MAGIC.length + 1; // 5
|
||||
|
||||
const password = process.env.SKILL_PASSWORD;
|
||||
if (!password) { console.error('ERROR: set SKILL_PASSWORD env var.'); process.exit(2); }
|
||||
|
||||
const encPath = process.argv[2];
|
||||
if (!encPath) { console.error('Usage: node decrypt.mjs <file.enc>'); process.exit(2); }
|
||||
|
||||
try {
|
||||
const blob = Buffer.from(readFileSync(encPath, 'utf8').trim(), 'base64');
|
||||
if (blob.length < HEADER + SALTLEN + IVLEN + TAGLEN) throw new Error('truncated');
|
||||
if (!blob.subarray(0, MAGIC.length).equals(MAGIC)) throw new Error('bad magic');
|
||||
if (blob[MAGIC.length] !== VERSION) throw new Error('unsupported version');
|
||||
|
||||
let off = HEADER;
|
||||
const salt = blob.subarray(off, off += SALTLEN);
|
||||
const iv = blob.subarray(off, off += IVLEN);
|
||||
const authTag = blob.subarray(off, off += TAGLEN);
|
||||
const ciphertext = blob.subarray(off);
|
||||
|
||||
const key = scryptSync(password, salt, KEYLEN, SCRYPT);
|
||||
const decipher = createDecipheriv('aes-256-gcm', key, iv);
|
||||
decipher.setAuthTag(authTag);
|
||||
// final() throws here if the tag does not verify (wrong password or tampering).
|
||||
const plaintext = Buffer.concat([decipher.update(ciphertext), decipher.final()]);
|
||||
process.stdout.write(plaintext); // STDOUT only — never written to disk.
|
||||
} catch (err) {
|
||||
// Generic message: do not echo crypto internals or any plaintext.
|
||||
console.error('Decryption failed: wrong password or corrupted data.');
|
||||
process.exit(1);
|
||||
}
|
||||
@@ -0,0 +1,32 @@
|
||||
#!/usr/bin/env node
|
||||
// encrypt.mjs — AES-256-GCM + scrypt. Password from $SKILL_PASSWORD (never argv).
|
||||
// Usage: SKILL_PASSWORD=... node encrypt.mjs <plaintextFile|-> <outFile.enc>
|
||||
// (pass '-' or omit the input path to read plaintext from STDIN)
|
||||
import { readFileSync, writeFileSync } from 'node:fs';
|
||||
import { scryptSync, randomBytes, createCipheriv } from 'node:crypto';
|
||||
|
||||
const MAGIC = Buffer.from('SENC', 'ascii');
|
||||
const VERSION = 0x01;
|
||||
const SCRYPT = { N: 1 << 17, r: 8, p: 1, maxmem: 256 * 1024 * 1024 };
|
||||
const KEYLEN = 32, SALTLEN = 16, IVLEN = 12;
|
||||
|
||||
const password = process.env.SKILL_PASSWORD;
|
||||
if (!password) { console.error('ERROR: set SKILL_PASSWORD env var.'); process.exit(2); }
|
||||
|
||||
const inPath = process.argv[2];
|
||||
const outPath = process.argv[3];
|
||||
if (!outPath) { console.error('Usage: node encrypt.mjs <plaintextFile|-> <outFile.enc>'); process.exit(2); }
|
||||
|
||||
// readFileSync(0) reads STDIN; use it when no input file (or '-') is given.
|
||||
const plaintext = (!inPath || inPath === '-') ? readFileSync(0) : readFileSync(inPath);
|
||||
|
||||
const salt = randomBytes(SALTLEN);
|
||||
const iv = randomBytes(IVLEN);
|
||||
const key = scryptSync(password, salt, KEYLEN, SCRYPT);
|
||||
const cipher = createCipheriv('aes-256-gcm', key, iv);
|
||||
const ciphertext = Buffer.concat([cipher.update(plaintext), cipher.final()]);
|
||||
const authTag = cipher.getAuthTag(); // 16 bytes
|
||||
|
||||
const blob = Buffer.concat([MAGIC, Buffer.from([VERSION]), salt, iv, authTag, ciphertext]);
|
||||
writeFileSync(outPath, blob.toString('base64') + '\n');
|
||||
console.error(`Wrote ${outPath} (${blob.length} raw bytes, base64-encoded).`);
|
||||
@@ -0,0 +1,30 @@
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Line endings
|
||||
#
|
||||
# Git stores LF, and every text file is checked out as LF on all platforms.
|
||||
# This is not cosmetic: scripts/validate-char-count.js measures the characters
|
||||
# actually on disk, so a CRLF checkout adds ~1 character per line. The workflow
|
||||
# prompts in src/plan2code-*.md run close to their 11,000 character budget
|
||||
# (several sit above 10,800), and a CRLF working tree pushes them over — turning
|
||||
# `npm test` into a check that passes or fails depending on how the repo was
|
||||
# cloned. Pinning eol=lf makes the count reproducible everywhere.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
* text=auto eol=lf
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Binary — never line-ending-converted, never diffed as text
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
*.png binary
|
||||
*.jpg binary
|
||||
*.jpeg binary
|
||||
*.gif binary
|
||||
*.ico binary
|
||||
*.mp4 binary
|
||||
*.webm binary
|
||||
*.woff binary
|
||||
*.woff2 binary
|
||||
|
||||
# The encrypted sync-repo blob is base64 text but must never have its bytes
|
||||
# altered by line-ending normalization. Treat it as binary so autocrlf/eol
|
||||
# settings can never corrupt the ciphertext.
|
||||
*.enc binary
|
||||
@@ -1,15 +1,17 @@
|
||||
specs/
|
||||
specs--completed/
|
||||
CLAUDE.md
|
||||
dist/
|
||||
plan2code-loop/dist
|
||||
plan2code-loop/node_modules
|
||||
plan2code-loop/package-lock.json
|
||||
plan2code-metrics/dist
|
||||
plan2code-metrics/node_modules
|
||||
plan2code-metrics/package-lock.json
|
||||
.plan2code-loop
|
||||
.plan2code-metrics
|
||||
nul
|
||||
node_modules/
|
||||
package-lock.json
|
||||
specs/
|
||||
specs--completed/
|
||||
dist/
|
||||
plan2code-loop/dist
|
||||
plan2code-loop/node_modules
|
||||
plan2code-loop/package-lock.json
|
||||
plan2code-metrics/dist
|
||||
plan2code-metrics/node_modules
|
||||
plan2code-metrics/package-lock.json
|
||||
.plan2code-loop
|
||||
.plan2code-metrics
|
||||
nul
|
||||
.cognition/
|
||||
handoffs/
|
||||
node_modules/
|
||||
package-lock.json
|
||||
SYNC.md
|
||||
@@ -0,0 +1,64 @@
|
||||
# Autonomous Loop
|
||||
|
||||
`plan2code-loop` is a CLI that works through your spec's tasks on its own, one agent call at a
|
||||
time. It is an **alternative to Step 3**, not a replacement — the four-step workflow and the specs it
|
||||
produces are unchanged.
|
||||
|
||||
← [Back to README](../README.md)
|
||||
|
||||
---
|
||||
|
||||
## When to use it instead of Step 3
|
||||
|
||||
| Approach | Best for |
|
||||
|----------|----------|
|
||||
| `/plan2code-3-implement` | Interactive control, reviewing each phase, logic that needs your judgment |
|
||||
| `plan2code-loop` | Straightforward implementations, batch work, overnight runs |
|
||||
|
||||
The loop reads the same `overview.md` and phase files. You can start with the loop and finish by
|
||||
hand, or the reverse — the checkboxes are the only handoff.
|
||||
|
||||
---
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
# From the plan2code root directory
|
||||
node install.js # A (everything + dev tools) — or C → O (loop only)
|
||||
```
|
||||
|
||||
## Run
|
||||
|
||||
```bash
|
||||
plan2code-loop # fully interactive
|
||||
```
|
||||
|
||||
It will:
|
||||
|
||||
1. Find specs in `./specs/`
|
||||
2. Let you pick one if there are several
|
||||
3. Offer to resume an existing session
|
||||
4. Ask for a JIRA ticket ID, which agent to drive, the loop mode, and a max iteration count
|
||||
|
||||
Then, per iteration: read `overview.md` and the phase files, find the first unchecked task (or
|
||||
phase), implement it, mark the checkbox, repeat — until everything is done or it hits the iteration
|
||||
cap.
|
||||
|
||||
---
|
||||
|
||||
## Loop modes
|
||||
|
||||
| Mode | Each agent call | Git commits | Best for |
|
||||
|------|-----------------|-------------|----------|
|
||||
| **One task per loop** (default) | Implements a single task | The Node controller commits after each task | Smaller models, cautious execution |
|
||||
| **One phase per loop** | Implements every task in a phase | The agent commits after each task, with the JIRA ID | Larger context windows, tightly related tasks |
|
||||
|
||||
Session state lives per-spec in `specs/<feature>/.plan2code-loop/`, so each feature's progress
|
||||
stays isolated.
|
||||
|
||||
---
|
||||
|
||||
## Full documentation
|
||||
|
||||
Architecture, completion markers, agent adapters, and configuration:
|
||||
[`plan2code-loop/README.md`](../plan2code-loop/README.md)
|
||||
@@ -0,0 +1,59 @@
|
||||
# Metrics & Self-Improvement
|
||||
|
||||
`plan2code-metrics` closes the loop on the workflow itself: it collects data from your finished
|
||||
specs, aggregates it across runs and prompt generations, then uses AI to diagnose which step is
|
||||
underperforming and propose edits to the workflow prompts.
|
||||
|
||||
Aimed at **contributors and heavy users** — you don't need it to use Plan2Code.
|
||||
|
||||
← [Back to README](../README.md)
|
||||
|
||||
---
|
||||
|
||||
## The habit
|
||||
|
||||
One thing to remember: **collect after every finished spec.** Everything else is on demand.
|
||||
|
||||
```bash
|
||||
# Install once, from the plan2code root
|
||||
node install.js # A (everything + dev tools) — or C → M (metrics only)
|
||||
|
||||
# After finishing a spec (steps 1–4)
|
||||
cd your-project
|
||||
plan2code-metrics # → "Collect metrics" → pick the spec dir → done, ~5 seconds
|
||||
```
|
||||
|
||||
Then, when you're curious or have a few runs banked:
|
||||
|
||||
```bash
|
||||
plan2code-metrics # → "View metrics status" the dashboard
|
||||
# → "Run analysis" AI diagnosis of weak steps
|
||||
# → "Generate improvement proposal" concrete prompt edits
|
||||
# → "Review and apply" patch src/plan2code-*.md
|
||||
```
|
||||
|
||||
## How much data you need
|
||||
|
||||
| Runs | What you get |
|
||||
|------|--------------|
|
||||
| **1** | Raw data and a basic dashboard. Start here. |
|
||||
| **3+** | AI analysis unlocks. Pattern detection starts working. |
|
||||
| **5–10+** | Averages stabilise; generation-over-generation comparisons become meaningful. |
|
||||
|
||||
You're looking for trends, not individual scores.
|
||||
|
||||
---
|
||||
|
||||
## Sending feedback upstream
|
||||
|
||||
`/plan2code-4-finalize` can submit an anonymised metrics payload to the maintainers as a
|
||||
`community-feedback` issue on the repo. Community runs are cohorted by the Plan2Code version that
|
||||
produced them, so your data improves the prompts everyone installs — without displacing the
|
||||
maintainer's own measurements.
|
||||
|
||||
---
|
||||
|
||||
## Full documentation
|
||||
|
||||
Data model, aggregation, cohorts, analysis prompts, and the ingestion flow:
|
||||
[`plan2code-metrics/README.md`](../plan2code-metrics/README.md)
|
||||
@@ -0,0 +1,54 @@
|
||||
# Claude Code Status Line
|
||||
|
||||
A persistent three-line status bar for Claude Code: model, project, git branch, uncommitted diff
|
||||
stats, session duration and cost, context-window usage, and plan quota.
|
||||
|
||||
It reads everything from the JSON Claude Code already sends on stdin — **no API calls, no auth, no
|
||||
background processes.** Optional, and unrelated to the workflow itself.
|
||||
|
||||
← [Back to README](../README.md)
|
||||
|
||||
---
|
||||
|
||||
## What it looks like
|
||||
|
||||
On Pro / Max / Teams accounts, where rate limits are available:
|
||||
|
||||
```
|
||||
◦ ◦ Opus 5 / high │ plan2code │ feature/PCWEB-11702-pathfinder │ +12 -3
|
||||
╭●╮ ┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄
|
||||
├■┤ 3h 5m ($4.62) │ ▰▰▰▰▰▱▱▱▱▱▱▱ 42% (84K) │ 5h: 28% · 7d: 61%
|
||||
```
|
||||
|
||||
On Enterprise / Bedrock / Vertex / pay-as-you-go, where they aren't, the last segment becomes session
|
||||
token counts instead:
|
||||
|
||||
```
|
||||
◦ ◦ Sonnet 5 │ plan2code │ main │ +12 -3
|
||||
╭●╮ ┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄
|
||||
├■┤ 2m ($0.18) │ ▰▰▱▱▱▱▱▱▱▱▱▱ 18% │ 88k in · 3k out
|
||||
```
|
||||
|
||||
The icons down the left are Planny, the project mascot.
|
||||
|
||||
---
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
node install.js # A (everything + dev tools) — or C → S (status line only)
|
||||
```
|
||||
|
||||
That copies the script to `~/.claude/plan2code-statusline.js`, writes a default config to
|
||||
`~/.claude/statusline-config.json` (an existing config is preserved), and registers it in
|
||||
`~/.claude/settings.json`.
|
||||
|
||||
If you already have a custom `statusLine` entry, the installer asks before replacing it and backs the
|
||||
old one up. Uninstalling removes the script and the settings entry but leaves your config file alone.
|
||||
|
||||
---
|
||||
|
||||
## Full documentation
|
||||
|
||||
Config options, compact mode, thresholds, and troubleshooting:
|
||||
[`src/statusline-claude/README.md`](../src/statusline-claude/README.md)
|
||||
@@ -0,0 +1,46 @@
|
||||
# Workflow Test Bot
|
||||
|
||||
`plan2code-bot` drives the whole workflow end to end with no human in the loop — init, plan,
|
||||
document, implement, finalize — to test that the prompts still hold together.
|
||||
|
||||
Built for **maintainers**. If you're using Plan2Code to ship features, you don't need this.
|
||||
|
||||
← [Back to README](../README.md)
|
||||
|
||||
---
|
||||
|
||||
## Two modes, auto-detected
|
||||
|
||||
| Condition | Mode | What it does |
|
||||
|-----------|------|--------------|
|
||||
| No `AGENTS.md` in the working directory | **New project** | Invents an app idea, creates a subdirectory, writes `IDEA.md`, runs init, then all four steps |
|
||||
| `AGENTS.md` present | **Enhancement** | Reads the existing codebase, proposes a realistic enhancement, writes `IDEA.md`, then runs plan → finalize |
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
node install.js # I (everything) — or C → B (bot only)
|
||||
```
|
||||
|
||||
## Run
|
||||
|
||||
```bash
|
||||
# New project mode — start from an empty directory
|
||||
mkdir /tmp/test-bot && cd /tmp/test-bot
|
||||
plan2code-bot
|
||||
|
||||
# Enhancement mode — start from a project that has an AGENTS.md
|
||||
cd my-project
|
||||
plan2code-bot
|
||||
|
||||
# Or seed the idea generator yourself
|
||||
plan2code-bot --idea "web app that renders the current weather as vector images"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Full documentation
|
||||
|
||||
Agent SDK wiring, run evaluation, and scoring:
|
||||
[`plan2code-bot/README.md`](../plan2code-bot/README.md) ·
|
||||
[`plan2code-bot/EVALUATION-SYSTEM.md`](../plan2code-bot/EVALUATION-SYSTEM.md)
|
||||
@@ -0,0 +1,103 @@
|
||||
# Walkthrough: one feature, start to finish
|
||||
|
||||
A REST API for task management, from a sentence to archived specs. Every block below is a **separate
|
||||
conversation** — that's the whole trick.
|
||||
|
||||
← [Back to README](../README.md)
|
||||
|
||||
---
|
||||
|
||||
## Session 1 — Plan
|
||||
|
||||
```
|
||||
You: /plan2code-1-plan
|
||||
I want to build a REST API for a task management application.
|
||||
|
||||
AI: 🤔 [REQUIREMENTS ANALYSIS]
|
||||
Six questions. Four of them I need answered before I can size this.
|
||||
...
|
||||
|
||||
AI: 🤔 [TECH STACK]
|
||||
Proposing Fastify + Postgres + Zod. I don't pick this alone — sign off?
|
||||
|
||||
You: approved
|
||||
|
||||
AI: 🤔 [ARCHITECTURE DESIGN]
|
||||
...
|
||||
|
||||
AI: 🤔 [TRANSITION DECISION]
|
||||
Confidence: 92% (gate is 90%). Assumptions written to the draft: 3.
|
||||
→ specs/task-api/PLAN-DRAFT-20260804.md
|
||||
```
|
||||
|
||||
Two files land: the draft, and a `PLAN-CONVERSATION-*.md` log of how you got there.
|
||||
|
||||
> **Started foggy instead?** Run `/plan2code-0-pathfinder` first. When its map clears it writes a
|
||||
> `PLAN-DRAFT` that Step 1 picks up at Phase 4, with requirements and scope already answered.
|
||||
|
||||
---
|
||||
|
||||
## Session 2 — Document
|
||||
|
||||
```
|
||||
You: /plan2code-2-document
|
||||
specs/task-api/PLAN-DRAFT-20260804.md
|
||||
|
||||
AI: 📝 [DOCUMENTATION]
|
||||
Creating specs/task-api/overview.md
|
||||
Creating specs/task-api/Phase 1.md Project setup (6 tasks)
|
||||
Creating specs/task-api/Phase 2.md Data model (8 tasks)
|
||||
Creating specs/task-api/Phase 3.md API endpoints (11 tasks)
|
||||
Creating specs/task-api/Phase 4.md Authentication (7 tasks)
|
||||
|
||||
Parallel execution groups: Phase 3 and Phase 4 don't share files.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Sessions 3…N — Implement, one phase each
|
||||
|
||||
```
|
||||
You: /plan2code-3-implement
|
||||
specs/task-api/overview.md
|
||||
|
||||
AI: ⚡ [PHASE 1: Project setup]
|
||||
(Phase 1 is the next unchecked phase.)
|
||||
...
|
||||
✓ 6/6 tasks. Checkboxes updated in Phase 1.md and overview.md.
|
||||
Next: Phase 2 — Data model (8 tasks). New conversation.
|
||||
```
|
||||
|
||||
Repeat. New conversation each time. It always finds the next open phase itself.
|
||||
|
||||
When you reach a parallel group, it offers you the choice — open a second agent, take the other
|
||||
phase, and the `[/]` marks keep them from colliding.
|
||||
|
||||
**Optional, any time:** `/plan2code-review` for an independent second opinion on what just landed.
|
||||
|
||||
---
|
||||
|
||||
## Final session — Finalize
|
||||
|
||||
```
|
||||
You: /plan2code-4-finalize
|
||||
specs/task-api/overview.md
|
||||
|
||||
AI: 🧹 [VALIDATION]
|
||||
All 32 tasks verified against the phase specs. 2 gaps found and fixed.
|
||||
|
||||
AI: 🧹 [DOCUMENTATION REVIEW]
|
||||
README needs the new /tasks endpoints. AGENTS.md is current.
|
||||
|
||||
AI: 🧹 [SPEC CLEANUP]
|
||||
Moved specs/task-api/ → specs--completed/task-api/
|
||||
|
||||
Implementation complete.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## If requirements move mid-build
|
||||
|
||||
Don't patch the code and hope the specs catch up. Run `/plan2code-1b-revise-plan` — it edits the
|
||||
specs (and only the specs), so the drawing and the build stay in agreement.
|
||||
@@ -1,298 +1,77 @@
|
||||
# AGENTS.md
|
||||
|
||||
This file provides guidance to AI coding agents like Claude Code (claude.ai/code), Cursor AI, Codex, Gemini CLI, GitHub Copilot, and other AI coding assistants when working with code in this repository.
|
||||
|
||||
## Project Overview
|
||||
|
||||
Plan2Code is a structured 4-step workflow methodology for AI-assisted software development. It provides prompt templates that can be installed globally or per-project for various AI coding tools (Claude Code, Cursor, Copilot, Continue, Windsurf, Codeium).
|
||||
|
||||
**Version:** Check `version.json` for current version
|
||||
**Author:** Justin Parker
|
||||
**License:** MIT
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
plan2code/
|
||||
├── src/ # Source workflow prompts (8 markdown files)
|
||||
├── plan2code-loop/ # Autonomous loop CLI tool (Node.js/TypeScript)
|
||||
│ ├── src/ # TypeScript source
|
||||
│ └── dist/ # Built output (tsup)
|
||||
├── plan2code-metrics/ # Recursive self-improvement toolchain
|
||||
│ ├── src/ # TypeScript source
|
||||
│ │ └── prompts/ # Internal AI prompt templates (no char limit)
|
||||
│ └── dist/ # Built output (tsup)
|
||||
├── scripts/ # Development scripts
|
||||
│ └── validate-char-count.js # Pre-commit character count validator
|
||||
├── dist/ # Generated distribution files (auto-generated)
|
||||
│ ├── global-commands/ # For global installation (~/.claude/, etc.)
|
||||
│ └── local-commands/ # For per-project installation (.claude/, etc.)
|
||||
├── .husky/ # Git hooks (husky)
|
||||
│ └── pre-commit # Runs character count validation
|
||||
├── docs/ # Documentation and assets
|
||||
├── specs/ # Feature specs (if any in-progress)
|
||||
├── install.js # Interactive installer (Node.js)
|
||||
├── package.json # Root package (husky only, private: true)
|
||||
├── version.json # Version metadata
|
||||
└── README.md # User documentation
|
||||
```
|
||||
|
||||
## Key Files
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `install.js` | Main installer - generates and installs workflow files to AI tool directories |
|
||||
| `src/plan2code-*.md` | Source workflow prompts (the "source of truth") |
|
||||
| `scripts/validate-char-count.js` | Pre-commit validator ensuring all source prompts ≤ 11,000 chars |
|
||||
| `version.json` | Version metadata (name, version, description) |
|
||||
| `QUICK-REFERENCE.md` | User quick-reference card |
|
||||
|
||||
## Workflow Prompts (in `src/`)
|
||||
|
||||
| File | Step | Purpose |
|
||||
|------|------|---------|
|
||||
| `plan2code---init.md` | Init | Generate AGENTS.md for projects |
|
||||
| `plan2code---init-update.md` | Update | Update AGENTS.md with learnings |
|
||||
| `plan2code---quick-task.md` | 0 | Lightweight planning for small tasks |
|
||||
| `plan2code-1--plan.md` | 1 | Requirements analysis & architecture |
|
||||
| `plan2code-1b--revise-plan.md` | 1b | Mid-implementation revisions |
|
||||
| `plan2code-2--document.md` | 2 | Create implementation specs |
|
||||
| `plan2code-3--implement.md` | 3 | Execute implementation (phase by phase) |
|
||||
| `plan2code-4--finalize.md` | 4 | Validate, summarize, feedback, archive (7 steps) |
|
||||
|
||||
## Plan2Code Loop (`plan2code-loop/`)
|
||||
|
||||
A separate Node.js CLI tool that autonomously implements specs by looping through tasks.
|
||||
|
||||
### Loop Architecture
|
||||
|
||||
The loop uses an **LLM-driven discovery** approach:
|
||||
- Node app just orchestrates iterations and parses completion markers
|
||||
- The LLM reads spec files (`overview.md`, `phase-X.md`) to discover tasks
|
||||
- The LLM finds unchecked checkboxes, implements ONE task per iteration, marks it complete
|
||||
- No regex parsing of markdown in Node - the AI handles all task discovery
|
||||
|
||||
### Loop Commands
|
||||
|
||||
```bash
|
||||
# Build the loop CLI
|
||||
cd plan2code-loop && npm run build
|
||||
|
||||
# Run the loop (after linking) - fully interactive
|
||||
plan2code-loop
|
||||
```
|
||||
|
||||
The CLI auto-detects specs in `./specs/`, prompts for selection if multiple found, and handles session continuation interactively. Session state is stored per-spec in `specs/<feature>/.plan2code-loop/`.
|
||||
|
||||
### Loop Modes
|
||||
|
||||
The CLI asks users to choose a loop mode:
|
||||
- **One task per loop** (default) - Each agent invocation implements exactly one task. The Node controller handles git commits.
|
||||
- **One phase per loop** - Each agent invocation implements all remaining tasks in the current phase. The LLM handles git commits (with JIRA ticket ID if provided). The controller parses multiple completion markers from a single iteration.
|
||||
|
||||
### Completion Markers
|
||||
|
||||
The LLM must output one of these formats:
|
||||
- `TASK_COMPLETE: 1.1 - Task description` - Task done successfully
|
||||
- `TASK_BLOCKED: 1.1 - Reason` - Cannot complete task
|
||||
- `PHASE_COMPLETE` - Current phase finished (phase mode only)
|
||||
- `LOOP_COMPLETE` - All phases finished
|
||||
|
||||
## Plan2Code Metrics (`plan2code-metrics/`)
|
||||
|
||||
A recursive self-improvement toolchain for plan2code contributors. Collects run metrics, aggregates by prompt generation, diagnoses weak steps via AI, and proposes surgical prompt edits.
|
||||
|
||||
### Metrics Data Flow
|
||||
|
||||
```
|
||||
Collect → Aggregate → Analyze → Improve → Apply
|
||||
```
|
||||
|
||||
1. **Collector** reads project artifacts (`specs/<feature>/`) → writes `RunMetrics` JSON per run
|
||||
2. **Aggregator** groups runs by prompt SHA fingerprint (cohorts) → `aggregated.json`
|
||||
3. **Analyzer** invokes AI with aggregated metrics + prompt contents → diagnosis markdown
|
||||
4. **Improver** invokes AI with diagnosis → validated `PromptEdit[]` proposals (char limit + verbatim checks)
|
||||
5. **Applier** shows interactive diffs → patches `src/plan2code-*.md` files
|
||||
|
||||
### Metrics Commands
|
||||
|
||||
```bash
|
||||
cd plan2code-metrics && npm run build # Build the CLI
|
||||
plan2code-metrics # Run (fully interactive, no flags)
|
||||
```
|
||||
|
||||
### Metrics CLI Menu
|
||||
|
||||
| Option | Action |
|
||||
|--------|--------|
|
||||
| Collect | Read spec artifacts → run JSON |
|
||||
| Import | Copy run JSON from another project |
|
||||
| View | Display cohort metrics with health indicators |
|
||||
| Analyze | AI diagnosis of weak metrics |
|
||||
| Propose | AI improvement proposals with validation |
|
||||
| Apply | Interactive diff review + file patching |
|
||||
|
||||
### Key Files
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `types.ts` | All interfaces (`RunMetrics`, `UserFeedback`, `CohortMetrics`, etc.) + `METRIC_TARGETS` |
|
||||
| `collector.ts` | Reads project artifacts → run JSON (parses plan drafts, overview.md, loop logs) |
|
||||
| `aggregator.ts` | Merges runs by prompt generation (SHA cohort) → `aggregated.json` |
|
||||
| `analyzer.ts` | AI diagnosis via `prompts/analyze.md` template |
|
||||
| `improver.ts` | AI proposals via `prompts/improve.md` + validation (char count, old_text match) |
|
||||
| `applier.ts` | Interactive diff review + file patching |
|
||||
| `cli.ts` | Menu-driven interactive CLI (100% prompts, no flags) |
|
||||
| `invoke-llm.ts` | Unified LLM interface (Claude Code or Copilot CLI) |
|
||||
|
||||
### User Feedback
|
||||
|
||||
The collector parses an optional `## User Feedback` table from `overview.md`:
|
||||
|
||||
```markdown
|
||||
## User Feedback
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| Rating | 8 |
|
||||
| Reason | Smooth workflow |
|
||||
| Went Well | Planning was thorough |
|
||||
| Went Poorly | Some tasks unclear |
|
||||
```
|
||||
|
||||
Feedback is collected during finalize (Step 5) or retroactively via the CLI. Pipe characters in values are escaped as `\|`. The aggregator computes `avg_user_rating` and `feedback_count` per cohort.
|
||||
|
||||
### Supported AI Agents
|
||||
|
||||
- **Claude Code** (recommended): `claude` CLI with `--inputFile` for prompt delivery
|
||||
- **GitHub Copilot CLI**: `copilot` CLI with stdin prompt delivery
|
||||
|
||||
### Metric Targets
|
||||
|
||||
| Metric | Target | Direction |
|
||||
|--------|--------|-----------|
|
||||
| `avg_confidence` | ≥ 90 | higher is better |
|
||||
| `avg_task_completion_rate` | ≥ 0.95 | higher is better |
|
||||
| `avg_blocker_count` | ≤ 1.5 | lower is better |
|
||||
| `avg_completion_marker_success_rate` | ≥ 0.95 | higher is better |
|
||||
| `avg_verification_failures_found` | ≤ 1.0 | lower is better |
|
||||
| `archival_success_rate` | ≥ 0.99 | higher is better |
|
||||
| `avg_user_rating` | ≥ 7.0 | higher is better |
|
||||
|
||||
Data stored in `.plan2code-metrics/` (runs/, aggregated.json, proposals/).
|
||||
|
||||
## Development Commands
|
||||
|
||||
```bash
|
||||
# Install dev dependencies (sets up husky pre-commit hooks)
|
||||
npm install
|
||||
|
||||
# Run the interactive installer (always interactive — any CLI args are silently ignored)
|
||||
node install.js
|
||||
|
||||
# Plan2Code Loop
|
||||
cd plan2code-loop && npm install # First time setup
|
||||
cd plan2code-loop && npm run build # Build the CLI
|
||||
|
||||
# Plan2Code Metrics
|
||||
cd plan2code-metrics && npm install # First time setup
|
||||
cd plan2code-metrics && npm run build # Build the CLI
|
||||
```
|
||||
|
||||
### Installer Menu Options
|
||||
|
||||
**Main menu:**
|
||||
|
||||
| Option | Action |
|
||||
|--------|--------|
|
||||
| `I` | Install Plan2Code for all platforms + loop CLI |
|
||||
| `U` | Uninstall Plan2Code files + loop CLI (confirmation required) |
|
||||
| `C` | Open CUSTOM sub-menu |
|
||||
| `Q` | Quit |
|
||||
|
||||
**CUSTOM sub-menu (`C`):**
|
||||
|
||||
| Option | Action |
|
||||
|--------|--------|
|
||||
| `L` | Show local (per-project) install instructions |
|
||||
| `O` | Install plan2code-loop CLI only |
|
||||
| `M` | Install plan2code-metrics CLI only |
|
||||
| `Q` | Return to main menu |
|
||||
|
||||
## How the Installer Works
|
||||
|
||||
1. **Reads source prompts** from `src/plan2code-*.md`
|
||||
2. **Generates platform-specific files** with appropriate headers (YAML frontmatter for some platforms)
|
||||
3. **Writes to `dist/`** subdirectories organized by destination type
|
||||
4. **Copies to target directories** (global: `~/.claude/commands/`, etc.)
|
||||
|
||||
### Platform-Specific File Formats
|
||||
|
||||
| Platform | Extension / File | Local Dir | Global Dir | Header |
|
||||
|----------|-----------------|-----------|------------|--------|
|
||||
| Claude Code | `.md` | — | — | None |
|
||||
| Cursor | `.md` | — | — | None |
|
||||
| Copilot CLI | `.md` | — | — | YAML frontmatter |
|
||||
| Continue | `.prompt.md` | — | — | YAML frontmatter |
|
||||
| Windsurf | `.md` | — | — | YAML frontmatter |
|
||||
| VS Code Copilot | `.prompt.md` | — | — | YAML frontmatter |
|
||||
| Codeium | `.md` | — | — | YAML frontmatter |
|
||||
| Claude Code (Skills) | `SKILL.md` in subdir | `.claude/skills/<skill-name>/` | `~/.claude/skills/<skill-name>/` | YAML frontmatter + `disable-model-invocation: true` |
|
||||
| Agent Skills (Amp · Gemini CLI · OpenCode) | `SKILL.md` in subdir | `.agents/skills/<skill-name>/` | `~/.agents/skills/<skill-name>/` | YAML frontmatter (no disable flag) |
|
||||
| Crush | `SKILL.md` in subdir | — (global only) | `~/.config/crush/skills/<skill-name>/` (Unix) / `%LOCALAPPDATA%\crush\skills\<skill-name>\` (Windows) | YAML frontmatter |
|
||||
| Gemini CLI (TOML) | `.toml` | `.gemini/commands/` | `~/.gemini/commands/` | None (TOML fields: `description`, `prompt`) |
|
||||
|
||||
## Naming Convention
|
||||
|
||||
Workflow files follow a strict naming pattern:
|
||||
- **Utilities:** `plan2code---<name>.md` (triple dash)
|
||||
- **Numbered steps:** `plan2code-<N>--<name>.md` (single dash, number, double dash)
|
||||
|
||||
Examples:
|
||||
- `plan2code---init.md` (utility)
|
||||
- `plan2code-1--plan.md` (step 1)
|
||||
- `plan2code-1b--revise-plan.md` (step 1b)
|
||||
|
||||
## Editing Workflow Prompts
|
||||
|
||||
When modifying workflow prompts in `src/`:
|
||||
|
||||
1. Edit the source file in `src/`
|
||||
2. Run `node install.js` to regenerate distribution files
|
||||
3. Test the workflow in your AI tool of choice
|
||||
4. The `dist/` folder is regenerated automatically - don't edit files there directly
|
||||
|
||||
## Code Style
|
||||
|
||||
- **install.js:** CommonJS, Node.js built-ins only (no external deps), ANSI colors via `COLORS` constant, readline-based prompts
|
||||
- **plan2code-loop & plan2code-metrics:** TypeScript + ESM, built with tsup (target ES2022, moduleResolution: bundler)
|
||||
- External deps: `@inquirer/prompts`, `chalk`, `execa`, `ora`
|
||||
- Interactive CLI via `@inquirer/prompts` (select, input, confirm)
|
||||
- **File operations:** Synchronous fs in all packages
|
||||
|
||||
## Gotchas/Pitfalls
|
||||
|
||||
- **Version sync:** When adding a new version to `CHANGELOG.md`, also update `version.json` and `package.json` to match. The installer displays the version from `version.json` in its header.
|
||||
- **Loop `.gitignore` setup:** `ensureGitignore()` runs at startup in `Controller.run()` as a pre-flight step, not just inside `createTaskCommit()`. This is critical for phase mode where the Node controller doesn't handle commits — without it, `git add -A` would stage spec files.
|
||||
- **Workflow file character limit:** All `src/plan2code-*.md` files must be ≤ 11,000 characters. A husky pre-commit hook enforces this. The 11,000 limit leaves buffer for platform-specific YAML headers (106-142 chars) to stay under Windsurf's 12,000 char limit.
|
||||
- **Metrics internal prompts have no char limit:** Files in `plan2code-metrics/src/prompts/` are NOT subject to the 11,000 char limit — only `src/plan2code-*.md` consumer-facing prompts are.
|
||||
- **User Feedback table format:** The `## User Feedback` markdown table in `overview.md` has a strict format the collector regex depends on. Field names must be exactly `Rating`, `Reason`, `Went Well`, `Went Poorly`. Pipe characters in values must be escaped as `\|`.
|
||||
|
||||
## Git Commit Messages
|
||||
|
||||
- **AI Assisted footer:** All git commit messages must include `AI Assisted` as the final line, separated from the message body by a blank line
|
||||
- **Commit paths:** This is enforced across all commit surfaces:
|
||||
- Loop task mode: `createTaskCommit()` in `plan2code-loop/src/utils/git.ts` appends the footer automatically
|
||||
- Loop phase mode: Prompt template instructs the LLM to add `-m "AI Assisted"` as the final flag
|
||||
- Implement mode: User-facing commit suggestions in `src/plan2code-3--implement.md` include the footer
|
||||
- **Init workflow:** `src/plan2code---init.md` generates AGENTS.md files with a Git Commit Messages section that includes this convention by default
|
||||
|
||||
## Mascot
|
||||
|
||||
The project has a mascot called "Planny" - an ASCII art robot that appears in installer output and workflow prompts. Mascot variants are defined in `MASCOT` constant in `install.js` and appear in workflow markdown files.
|
||||
|
||||
```
|
||||
╭───╮
|
||||
│ ● │
|
||||
│ ◡ │
|
||||
╰───╯
|
||||
```
|
||||
# AGENTS.md
|
||||
|
||||
This file provides guidance to AI coding agents working with code in this repository.
|
||||
|
||||
## Project Overview
|
||||
|
||||
Plan2Code is a structured 4-step workflow methodology for AI-assisted software development. It provides prompt templates that can be installed globally or per-project for various AI coding tools (Claude Code, Cursor, Copilot, Continue, Windsurf, Codeium, Devin, Zed).
|
||||
|
||||
**Version:** Check `version.json` for current version
|
||||
**Author:** Justin Parker
|
||||
**License:** MIT
|
||||
|
||||
## How to Use This File
|
||||
|
||||
This file is an index — each section below contains a brief summary and a link to a detail file in `.agents-docs/`. Read only the sections relevant to your current task. Full details (commands, tables, file lists) are in the linked files. The sections "Project Overview", "How to Use This File", "Mascot", and "Keeping this file current" / "Failure log" are fully inline here and are never split into `.agents-docs/`.
|
||||
|
||||
## Architecture
|
||||
|
||||
High-level directory structure, key files, workflow prompt inventory, and file naming conventions.
|
||||
|
||||
Details: [Architecture](./.agents-docs/AGENTS-architecture.md)
|
||||
|
||||
## Plan2Code Loop
|
||||
|
||||
Autonomous CLI tool (`plan2code-loop/`) that implements specs by looping through tasks. Covers loop architecture, modes (one-task vs one-phase), completion markers, and key source files.
|
||||
|
||||
Details: [Plan2Code Loop](./.agents-docs/AGENTS-plan2code-loop.md)
|
||||
|
||||
## Plan2Code Metrics
|
||||
|
||||
Recursive self-improvement toolchain (`plan2code-metrics/`) for collecting run metrics, aggregating by prompt generation, diagnosing weak steps, and proposing prompt edits.
|
||||
|
||||
Details: [Plan2Code Metrics](./.agents-docs/AGENTS-plan2code-metrics.md)
|
||||
|
||||
## Plan2Code Status Line (Claude CLI)
|
||||
|
||||
Optional CLI status bar (`src/statusline-claude/`) for Claude Code. Displays model, project, branch, context usage, usage stats (rate limits or token counts), git diff stats, and duration. Reads data directly from Claude Code's stdin JSON — no API calls, no auth, no background processes. Included in `Install All + dev tools` (`A`); also available via `install.js` Custom → S (opt-in).
|
||||
|
||||
Details: [Architecture](./.agents-docs/AGENTS-architecture.md) (see Status Line section)
|
||||
|
||||
## Development Commands
|
||||
|
||||
Build commands, installer menu options, how the installer generates platform-specific files, platform format table, and how to edit workflow prompts.
|
||||
|
||||
Details: [Development Commands](./.agents-docs/AGENTS-development-commands.md)
|
||||
|
||||
## Code Style & Gotchas
|
||||
|
||||
Language/toolchain conventions for `install.js` vs TypeScript packages, and pitfalls to avoid (version sync, `.gitignore` pre-flight, character limits, User Feedback table format).
|
||||
|
||||
Details: [Code Style & Gotchas](./.agents-docs/AGENTS-code-style.md)
|
||||
|
||||
## Mascot
|
||||
|
||||
The project has a mascot called "Planny" — an ASCII art robot that appears in installer output and workflow prompts. Mascot variants are defined in the `MASCOT` constant in `install.js` and appear in workflow markdown files.
|
||||
|
||||
```
|
||||
╭───╮
|
||||
│ ● │
|
||||
│ ◡ │
|
||||
╰───╯
|
||||
```
|
||||
|
||||
## Keeping this file current
|
||||
|
||||
The `Failure log` section below is a recording of mistakes made by previous AI Agents while working with this code base.
|
||||
|
||||
When you make a mistake, get corrected, or discover something about this codebase that wasn't written down:
|
||||
|
||||
1. Add one line to the `Failure log` below, in the imperative, describing the correct behaviour.
|
||||
2. Keep it specific to this repo. General advice belongs nowhere.
|
||||
3. If this fix is a workflow rather than a rule, put it in `.claude/skills/` and link it from here.
|
||||
4. Include the change in the same commit and mention it in your summary.
|
||||
|
||||
## Failure log
|
||||
|
||||
- Do not audit `CHANGELOG.md` headings through PowerShell — the emoji come back as `?`. See the gotcha for the correct approach.
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
# CLAUDE.md
|
||||
|
||||
**CRITICAL — MANDATORY FIRST STEP: You MUST read [AGENTS.md](./AGENTS.md) before responding to ANY user message, including simple questions. Do NOT skip this step regardless of how trivial the request appears. No exceptions.**
|
||||
|
||||
See AGENTS.md for complete project documentation including:
|
||||
- Development commands and setup
|
||||
- Architecture overview
|
||||
- Workflow prompt reference
|
||||
- Plan2Code Loop CLI details
|
||||
- Code style and gotchas
|
||||
- Keeping this file current / Failure log
|
||||
- Section details in .agents-docs/
|
||||
|
||||
This file exists for Claude Code auto-loading. All AI coding agents should reference AGENTS.md.
|
||||
@@ -1,23 +1,30 @@
|
||||
# Plam2Code Quick Reference
|
||||
# Plan2Code Quick Reference
|
||||
|
||||
## Commands
|
||||
|
||||
| Step | Command | Input | Output |
|
||||
| ------ | ------------------------------- | --------------- | ----------------------------------- |
|
||||
| Init | /plan2code---init | None | AGENTS.md file |
|
||||
| Update | /plan2code---init-update | AGENTS.md | Updated AGENTS.md |
|
||||
| 0 | /plan2code---quick-task | Requirements | Conversational plan |
|
||||
| 1 | /plan2code-1--plan | Requirements | PLAN-CONVERSATION-<date>.md + PLAN-DRAFT-<date>.md |
|
||||
| 1b | /plan2code-1b--revise-plan | Specs + changes | Updated specs |
|
||||
| 2 | /plan2code-2--document | PLAN-DRAFT.md | overview.md + Phase files |
|
||||
| 3 | /plan2code-3--implement | overview.md | Implemented code |
|
||||
| 4 | /plan2code-4--finalize | overview.md | Archived specs |
|
||||
| Step | Command | Input | Output |
|
||||
| ------- | ----------------------------- | --------------- | -------------------------------------------------- |
|
||||
| Init | /plan2code-init | None | AGENTS.md file |
|
||||
| Update | /plan2code-init-update | AGENTS.md | Updated AGENTS.md |
|
||||
| 0 | /plan2code-0-pathfinder | A foggy idea | pathfinder/map.md *or* GitHub Issues + PLAN-DRAFT-<date>.md |
|
||||
| quick | /plan2code-quick-task | Requirements | Conversational plan (standalone — not a pipeline step) |
|
||||
| review | /plan2code-review | Scope guidance | Review findings + fixes |
|
||||
| 1 | /plan2code-1-plan | Requirements | PLAN-CONVERSATION-<date>.md + PLAN-DRAFT-<date>.md |
|
||||
| 1b | /plan2code-1b-revise-plan | Specs + changes | Updated specs |
|
||||
| 2 | /plan2code-2-document | PLAN-DRAFT.md | overview.md + Phase files |
|
||||
| 3 | /plan2code-3-implement | overview.md | Implemented code |
|
||||
| 4 | /plan2code-4-finalize | overview.md | Archived specs |
|
||||
| handoff | /plan2code-handoff | Conversation | Self-contained handoff doc in handoffs/ |
|
||||
|
||||
## File Structure
|
||||
|
||||
```
|
||||
specs/
|
||||
└── <feature-name>/
|
||||
├── pathfinder/ # From Step 0 (optional, if charted locally)
|
||||
│ ├── map.md # the map: destination, decisions, fog
|
||||
│ └── questions/NN-<slug>.md # one decision question per file
|
||||
│ # (GitHub Issues backend: map issue + sub-issues instead)
|
||||
├── PLAN-DRAFT-<date>.md # From Step 1 (verified plan)
|
||||
├── PLAN-CONVERSATION-<date>.md # From Step 1 (conversation log)
|
||||
├── overview.md # From Step 2
|
||||
@@ -51,7 +58,7 @@ When phases have no file conflicts or dependencies, they can run simultaneously:
|
||||
|
||||
1. Documentation Mode auto-detects parallel-eligible phases
|
||||
2. Implementation Mode shows selection UI with status for each phase
|
||||
3. Run multiple `/plan2code-3--implement` instances on different phases
|
||||
3. Run multiple `/plan2code-3-implement` instances on different phases
|
||||
4. `[/]` status shows which phases are actively being worked on
|
||||
|
||||
## Quick Troubleshooting
|
||||
@@ -60,31 +67,35 @@ When phases have no file conflicts or dependencies, they can run simultaneously:
|
||||
| ---------------------- | ------------------------------------------------- |
|
||||
| Lost context mid-phase | Attach spec files, say "resume from Task X.Y" |
|
||||
| Wrong phase started | Say "abort", start correct phase |
|
||||
| Need to change plan | Use `/plan2code-1b--revise-plan` |
|
||||
| Need to change plan | Use `/plan2code-1b-revise-plan` |
|
||||
| Multiple spec folders | Specify which: "Continue with specs/user-auth/" |
|
||||
| Need AGENTS.md file | Use `/plan2code---init` to generate one |
|
||||
| Update AGENTS.md | Use `/plan2code---init-update` after sessions |
|
||||
| Need AGENTS.md file | Use `/plan2code-init` to generate one |
|
||||
| Update AGENTS.md | Use `/plan2code-init-update` after sessions |
|
||||
| Run phases in parallel | Check Parallel Execution Groups in overview.md |
|
||||
|
||||
## Workflow Decision
|
||||
|
||||
```
|
||||
New to a project?
|
||||
└── /plan2code---init → Generate AGENTS.md for project-specific guidance
|
||||
└── /plan2code-init → Generate AGENTS.md for project-specific guidance
|
||||
|
||||
Learned something during a session?
|
||||
└── /plan2code---init-update → Add learnings to AGENTS.md
|
||||
└── /plan2code-init-update → Add learnings to AGENTS.md
|
||||
|
||||
Too unclaer to plan? (big idea, don't yet know what the questions are)
|
||||
└── /plan2code-0-pathfinder → chart it, clear one decision per session
|
||||
└── then → /plan2code-1-plan (resumes at Phase 4)
|
||||
|
||||
Is it a quick, small task?
|
||||
├── Yes → /plan2code---quick-task (standalone)
|
||||
└── No → /plan2code-1--plan (full workflow)
|
||||
├── /plan2code-2--document
|
||||
├── /plan2code-3--implement (repeat per phase)
|
||||
├── Yes → /plan2code-quick-task (standalone)
|
||||
└── No → /plan2code-1-plan (full workflow)
|
||||
├── /plan2code-2-document
|
||||
├── /plan2code-3-implement (repeat per phase)
|
||||
│ └── OR: plan2code-loop (autonomous alternative)
|
||||
└── /plan2code-4--finalize
|
||||
└── /plan2code-4-finalize
|
||||
|
||||
Need to revise mid-implementation?
|
||||
└── /plan2code-1b--revise-plan
|
||||
└── /plan2code-1b-revise-plan
|
||||
```
|
||||
|
||||
## Autonomous Loop (Alternative)
|
||||
@@ -93,7 +104,7 @@ The `plan2code-loop` CLI is an **alternative** to Step 3, not a replacement.
|
||||
|
||||
| Approach | Use When |
|
||||
|----------|----------|
|
||||
| `/plan2code-3--implement` | You want interactive control per phase |
|
||||
| `/plan2code-3-implement` | You want interactive control per phase |
|
||||
| `plan2code-loop` | You want hands-off autonomous execution |
|
||||
|
||||
```bash
|
||||
|
||||
@@ -1,516 +1,299 @@
|
||||
# Plan2Code: AI-Assisted Software Development Workflow
|
||||
|
||||
A structured 4-step workflow for developing features and projects with AI assistance. This methodology emphasizes thorough planning before implementation, ensuring well-documented, maintainable code.
|
||||
|
||||
<img src="docs/desk.jpg" alt="Plan2Code Workflow" height="275">
|
||||
|
||||
## Overview
|
||||
|
||||
```
|
||||
🤔 📝 ⚡ 🧹
|
||||
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
|
||||
│ Step 1 │ │ Step 2 │ │ Step 3 │ │ Step 4 │
|
||||
│ PLAN │ --> │ DOCUMENT │ --> │ IMPLEMENT │ --> │ FINALIZE │
|
||||
└─────────────────┘ └─────────────────┘ └─────────────────┘ └─────────────────┘
|
||||
New Chat New Chat New Chat (per phase) New Chat
|
||||
```
|
||||
|
||||
| Command | When to Use |
|
||||
|----------------------------------|----------------------------------------------------------|
|
||||
| `/plan2code---init` | Generate AGENTS.md file for new/existing projects |
|
||||
| `/plan2code---init-update` | Update AGENTS.md with new learnings from coding sessions |
|
||||
| `/plan2code---quick-task` | Small, quick tasks that don't need full workflow |
|
||||
| `/plan2code-1--plan` | Starting a new feature (full planning) |
|
||||
| `/plan2code-1b--revise-plan` | Requirements change mid-implementation |
|
||||
| `/plan2code-2--document` | After planning, create implementation specs |
|
||||
| `/plan2code-3--implement` | Execute implementation (one phase per conversation) |
|
||||
| `/plan2code-4--finalize` | All phases complete, ready to archive |
|
||||
|
||||
**Key Rules:**
|
||||
- Start NEW conversation for each step (and each implementation phase)
|
||||
- ONE phase per conversation
|
||||
- Reply "approved" to complete phases
|
||||
- 90% confidence required before planning completes
|
||||
|
||||
See [QUICK-REFERENCE.md](QUICK-REFERENCE.md) for full reference card.
|
||||
|
||||
---
|
||||
|
||||
## Installation
|
||||
|
||||
Plan2Code includes an interactive installer that generates and installs workflow files for all major AI coding assistants.
|
||||
|
||||
<img src="docs/install-script.jpg" width="600">
|
||||
|
||||
### Prerequisites
|
||||
|
||||
The install script requires **Node.js** (v14 or later). If you don't have Node.js installed:
|
||||
|
||||
1. Download from [nodejs.org](https://nodejs.org/) (recommended)
|
||||
2. Or use a package manager:
|
||||
- **macOS:** `brew install node`
|
||||
- **Windows:** `winget install OpenJS.NodeJS` or `choco install nodejs`
|
||||
- **Linux:** `sudo apt install nodejs` (Debian/Ubuntu) or `sudo dnf install nodejs` (Fedora)
|
||||
|
||||
### Supported Platforms
|
||||
|
||||
- Claude Code
|
||||
- Cursor
|
||||
- Windsurf
|
||||
- Continue
|
||||
- Codeium (IntelliJ)
|
||||
- GitHub Copilot CLI
|
||||
- VS Code GitHub Copilot
|
||||
- Gemini CLI
|
||||
- Crush
|
||||
- Amp
|
||||
- OpenCode
|
||||
|
||||
### Install via npx (Recommended — No Clone Required)
|
||||
|
||||
Run the interactive installer directly using `npx` with your preferred GitHub authentication method:
|
||||
|
||||
**If you use SSH keys:**
|
||||
```bash
|
||||
npx git+ssh://git@github.com/jparkerweb/plan2code.git
|
||||
```
|
||||
|
||||
**If you use HTTPS authentication:**
|
||||
```bash
|
||||
npx git+https://github.com/jparkerweb/plan2code.git
|
||||
```
|
||||
|
||||
This downloads the installer to a temporary location, runs it, installs the workflow files to your machine, and cleans up automatically. The installed workflows remain on your system and work independently. To update or reinstall, simply run the command again.
|
||||
|
||||
### Standard Installation (Clone Method)
|
||||
|
||||
#### No Clone Required
|
||||
|
||||
Run the interactive installer directly using `npx` with your preferred GitHub authentication method:
|
||||
|
||||
**If you use SSH keys:**
|
||||
```bash
|
||||
npx git+ssh://git@github.com/jparkerweb/plan2code.git
|
||||
```
|
||||
|
||||
**If you use HTTPS authentication:**
|
||||
```bash
|
||||
npx git+https://github.com/jparkerweb/plan2code.git
|
||||
```
|
||||
|
||||
This downloads the installer to a temporary location, runs it, installs the workflow files to your machine, and cleans up automatically. The installed workflows remain on your system and work independently. To update or reinstall, simply run the command again.
|
||||
|
||||
#### Standard Installation (Clone Method)
|
||||
|
||||
```bash
|
||||
# Clone the repository
|
||||
git clone https://github.com/jparkerweb/plan2code.git
|
||||
cd plan2code
|
||||
|
||||
# Run the interactive installer
|
||||
node install.js
|
||||
|
||||
# OPTIONAL: Install dev dependencies (ONLY if you plan to modify/contribute to Plan2Code)
|
||||
npm install
|
||||
```
|
||||
|
||||
The installer displays an interactive menu:
|
||||
|
||||
```
|
||||
╔════════════════════════════════════════════════════════════════╗
|
||||
║ INSTALL PLAN2CODE ║
|
||||
╠════════════════════════════════════════════════════════════════╣
|
||||
║ I. INSTALL Install Plan2Code for all platforms ║
|
||||
║ U. UNINSTALL Remove Plan2Code files ║
|
||||
║ C. CUSTOM Advanced options ║
|
||||
║ Q. QUIT Exit ║
|
||||
╚════════════════════════════════════════════════════════════════╝
|
||||
|
||||
SELECT OPTION (I, U, C, Q) [I]:
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Important: Start Fresh Conversations
|
||||
|
||||
**Start a new conversation/chat session before each step.** This includes:
|
||||
|
||||
- Step 1: New conversation
|
||||
- Step 2: New conversation
|
||||
- Step 3: New conversation **for each phase** (Phase 1, Phase 2, etc.)
|
||||
- Step 4: New conversation
|
||||
|
||||
Fresh conversations prevent context pollution and ensure the AI focuses on the current task with the relevant specifications.
|
||||
|
||||
---
|
||||
|
||||
## The Workflow Steps
|
||||
|
||||
### Step 1: Planning Mode 🤔
|
||||
|
||||
**Purpose:** Thoroughly analyze requirements and design the solution architecture before writing any code.
|
||||
|
||||
**AI Role:** Senior software architect and technical product manager
|
||||
|
||||
**Phases (completed one at a time):**
|
||||
|
||||
1. **Requirements Analysis** - Extract functional/non-functional requirements, identify ambiguities
|
||||
2. **System Context Examination** - Review existing codebase, identify integration points
|
||||
3. **Tech Stack** - Recommend and confirm all technologies (requires user sign-off)
|
||||
4. **Architecture Design** - Propose patterns, define components, design interfaces/schemas
|
||||
5. **Technical Specification** - Break down implementation phases, identify risks
|
||||
6. **Transition Decision** - Finalize plan when confidence reaches 90%+
|
||||
|
||||
**Output:** `specs/<feature-name>/PLAN-DRAFT-<date>.md` and `specs/<feature-name>/PLAN-CONVERSATION-<date>.md` (date format: YYYYMMDD)
|
||||
|
||||
**Key Behaviors:**
|
||||
|
||||
- AI stops after each phase for clarification
|
||||
- Must reach 90% confidence before finalizing
|
||||
- All assumptions are documented
|
||||
- User must approve tech stack decisions
|
||||
|
||||
---
|
||||
|
||||
### Step 2: Documentation Mode 📝
|
||||
|
||||
**Purpose:** Transform the planning output into structured, actionable implementation documents.
|
||||
|
||||
**Required Context:** Attach or reference the `specs/<feature-name>/PLAN-DRAFT-<date>.md` from Step 1 (or provide the planning conversation).
|
||||
|
||||
**Output Structure:**
|
||||
|
||||
```
|
||||
specs/
|
||||
└── <feature-name>/
|
||||
├── overview.md # High-level overview with phase checkboxes and parallel groups
|
||||
├── Phase 1.md # Detailed tasks for Phase 1
|
||||
├── Phase 2.md # Detailed tasks for Phase 2
|
||||
└── Phase N.md # ...additional phases
|
||||
```
|
||||
|
||||
The `overview.md` includes a "Parallel Execution Groups" section that identifies which phases can be run simultaneously in separate agent instances.
|
||||
|
||||
**Document Format:**
|
||||
|
||||
- Each phase file contains detailed one-story-point tasks
|
||||
- All tasks have checkboxes `[ ]` for progress tracking
|
||||
- Each phase is self-contained (developer needs no prior context)
|
||||
- Unit/E2E testing excluded unless explicitly requested
|
||||
|
||||
---
|
||||
|
||||
### Step 3: Implementation Mode ⚡
|
||||
|
||||
**Purpose:** Execute the implementation following the documented specifications.
|
||||
|
||||
**AI Role:** Senior software engineer
|
||||
|
||||
**Required Context:** Provide the path to `specs/<feature-name>/overview.md`. The command will auto-detect the next uncompleted phase and read the corresponding `Phase X.md` file automatically.
|
||||
|
||||
**Workflow:**
|
||||
|
||||
1. Identify the next uncompleted phase (unchecked in `overview.md`)
|
||||
2. Check for parallel execution options (if phases can run simultaneously)
|
||||
3. Implement ALL tasks in that phase exactly as specified
|
||||
4. Update `Phase X.md` checkboxes as tasks complete `[x]`
|
||||
5. Update `overview.md` phase checkbox when phase completes
|
||||
6. Perform code review to ensure nothing was missed
|
||||
7. Add completion summary to the phase document
|
||||
|
||||
**Parallel Execution:** If the next phase is part of a parallel-eligible group, you'll be prompted to choose which phase to implement. This allows running multiple agent instances simultaneously on different phases that don't conflict with each other.
|
||||
|
||||
**Key Rules:**
|
||||
|
||||
- **Start a new conversation for EACH phase**
|
||||
- Work on ONE phase per conversation (unless told otherwise)
|
||||
- Follow specifications EXACTLY as documented
|
||||
- Keep checkboxes updated (enables progress tracking across sessions)
|
||||
- Do NOT run tests unless specified in phase tasks
|
||||
|
||||
---
|
||||
|
||||
### Step 4: Finalization Mode 🧹
|
||||
|
||||
**Purpose:** Validate implementation, create summaries, and archive documentation.
|
||||
|
||||
**Required Context:** Attach or reference the `specs/<feature-name>/` directory contents.
|
||||
|
||||
**Steps:**
|
||||
|
||||
1. **Validation** - Verify all tasks implemented correctly, check for issues
|
||||
2. **Summary** - Document what was built and list all modified/created files
|
||||
3. **Documentation Review** - Identify any needed README/CHANGELOG updates
|
||||
4. **Spec Cleanup** - Move completed specs to `specs--completed/<implementation-name>/`
|
||||
5. **Final Confirmation** - Confirm completion
|
||||
|
||||
---
|
||||
|
||||
## How to Use
|
||||
|
||||
After running `node install.js`, use the slash commands directly in your AI tool:
|
||||
|
||||
```
|
||||
/plan2code-1--plan # Start planning a new feature
|
||||
/plan2code-2--document # Create implementation docs from plan
|
||||
/plan2code-3--implement # Begin/continue implementation
|
||||
/plan2code-4--finalize # Wrap up after all phases complete
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Complete Workflow Example
|
||||
|
||||
### Starting a New Project
|
||||
|
||||
**Session 1 - Planning (New Chat):**
|
||||
|
||||
```
|
||||
User: [Paste or invoke Step 1 prompt]
|
||||
I want to build a REST API for a task management application.
|
||||
|
||||
AI: 🤔 [REQUIREMENTS ANALYSIS]
|
||||
... asks clarifying questions, works through phases ...
|
||||
|
||||
AI: 🤔 [TRANSITION DECISION]
|
||||
Confidence: 92%. Creating specs/task-api/PLAN-DRAFT-20250204.md...
|
||||
```
|
||||
|
||||
**Session 2 - Documentation (New Chat):**
|
||||
|
||||
```
|
||||
User: [Paste or invoke Step 2 prompt]
|
||||
[Attach: specs/task-api/PLAN-DRAFT-20250204.md]
|
||||
|
||||
AI: 📝 [DOCUMENTATION]
|
||||
Creating specs/task-api/overview.md...
|
||||
Creating specs/task-api/Phase 1.md...
|
||||
Creating specs/task-api/Phase 2.md...
|
||||
...
|
||||
```
|
||||
|
||||
**Session 3 - Implementation Phase 1 (New Chat):**
|
||||
|
||||
```
|
||||
User: [Paste or invoke Step 3 prompt]
|
||||
[Provide: specs/task-api/overview.md]
|
||||
|
||||
AI: ⚡ [PHASE 1: Project Setup]
|
||||
(Auto-detected Phase 1 as next uncompleted phase)
|
||||
Implementing tasks...
|
||||
✓ Phase 1 complete. Updated checkboxes in Phase 1.md and overview.md.
|
||||
```
|
||||
|
||||
**Session 4 - Implementation Phase 2 (New Chat):**
|
||||
|
||||
```
|
||||
User: [Paste or invoke Step 3 prompt]
|
||||
[Provide: specs/task-api/overview.md]
|
||||
|
||||
AI: ⚡ [PHASE 2: Database Models]
|
||||
(Auto-detected Phase 2 as next uncompleted phase)
|
||||
Implementing tasks...
|
||||
✓ Phase 2 complete. Updated checkboxes in Phase 2.md and overview.md.
|
||||
```
|
||||
|
||||
**Sessions 5-N - Continue Implementation (New Chat for each phase):**
|
||||
|
||||
```
|
||||
... repeat for each remaining phase ...
|
||||
```
|
||||
|
||||
**Final Session - Finalization (New Chat):**
|
||||
|
||||
```
|
||||
User: [Paste or invoke Step 4 prompt]
|
||||
[Provide: specs/task-api/overview.md]
|
||||
|
||||
AI: 🧹 [VALIDATION]
|
||||
Verifying implementation...
|
||||
|
||||
AI: 🧹 [SPEC CLEANUP]
|
||||
Moving to specs--completed/task-api/
|
||||
|
||||
Implementation complete!
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Progress Tracking
|
||||
|
||||
The checkbox system enables seamless progress tracking across multiple sessions:
|
||||
|
||||
**Phase Status (overview.md):**
|
||||
|
||||
| Checkbox | Status | Meaning |
|
||||
|----------|--------|---------|
|
||||
| `[ ]` | Pending | Not yet started |
|
||||
| `[/]` | In Progress | Agent actively working (or paused/aborted) |
|
||||
| `[x]` | Complete | Finished and approved |
|
||||
|
||||
```markdown
|
||||
## Phases
|
||||
|
||||
- [x] Phase 1: Project Setup
|
||||
- [x] Phase 2: Database Models
|
||||
- [/] Phase 3: API Endpoints <- In progress (agent working)
|
||||
- [ ] Phase 4: Authentication <- Next available
|
||||
```
|
||||
|
||||
**Task Status (phase-X.md):**
|
||||
|
||||
```markdown
|
||||
## Tasks
|
||||
|
||||
- [x] Create routes file
|
||||
- [x] Implement GET /tasks
|
||||
- [ ] Implement POST /tasks <- Current task
|
||||
- [ ] Implement PUT /tasks/:id
|
||||
- [ ] Implement DELETE /tasks/:id
|
||||
```
|
||||
|
||||
The `[/]` status enables parallel execution - multiple agents can work on different phases simultaneously, and you can see which phases are actively being worked on.
|
||||
|
||||
---
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Start fresh conversations** - New chat for each step and each implementation phase
|
||||
2. **Always attach specs** - The AI needs the spec files to understand the current state
|
||||
3. **Don't skip planning** - The upfront investment prevents costly rework later
|
||||
4. **Confirm tech stack** - Ensure AI gets explicit approval before architecture design
|
||||
5. **One phase at a time** - Keeps conversations focused and manageable
|
||||
6. **Update checkboxes immediately** - Maintains accurate progress state
|
||||
7. **Review phase output** - Verify each phase before moving to the next
|
||||
8. **Keep spec files** - The completed folder serves as project documentation
|
||||
|
||||
---
|
||||
|
||||
## What to Attach at Each Step
|
||||
|
||||
| Step | Required Input |
|
||||
| ------------------ | ---------------------------------------------------- |
|
||||
| Step 1 (Plan) | None (describe your feature/project) |
|
||||
| Step 2 (Document) | `specs/<feature>/PLAN-DRAFT-<date>.md` or planning conversation |
|
||||
| Step 3 (Implement) | `specs/<feature>/overview.md` (auto-detects phase) |
|
||||
| Step 4 (Finalize) | `specs/<feature>/overview.md` |
|
||||
|
||||
---
|
||||
|
||||
## Autonomous Loop (Alternative to Step 3)
|
||||
|
||||
For hands-off implementation, Plan2Code includes an optional autonomous loop CLI that iterates through your spec tasks automatically.
|
||||
|
||||
> **Note:** The loop is an **alternative** to `/plan2code-3--implement`, not a replacement. Use the manual Step 3 workflow when you want direct control over each phase, or use the loop when you prefer autonomous execution.
|
||||
|
||||
### When to Use Each
|
||||
|
||||
| Approach | Best For |
|
||||
|----------|----------|
|
||||
| `/plan2code-3--implement` | Interactive control, reviewing each phase, complex logic requiring human judgment |
|
||||
| `plan2code-loop` | Straightforward implementations, batch processing, overnight runs |
|
||||
|
||||
### Installing the Loop
|
||||
|
||||
```bash
|
||||
# From the plan2code root directory:
|
||||
|
||||
# Option 1: Install everything (recommended)
|
||||
node install.js # Select I at the menu
|
||||
|
||||
# Option 2: Install loop only
|
||||
node install.js # Select C, then O at the menu
|
||||
```
|
||||
|
||||
### Using the Loop
|
||||
|
||||
```bash
|
||||
# Run the loop - fully interactive
|
||||
plan2code-loop
|
||||
```
|
||||
|
||||
The CLI will:
|
||||
1. Auto-detect specs in `./specs/` directory
|
||||
2. Let you select a spec if multiple are found
|
||||
3. Prompt to continue if an existing session is found
|
||||
4. Ask for JIRA ticket ID, agent selection, loop mode, and max iterations
|
||||
|
||||
### Loop Modes
|
||||
|
||||
| Mode | Behavior | Git Commits | Best For |
|
||||
|------|----------|-------------|----------|
|
||||
| **One task per loop** (default) | Each agent call implements one task | Node controller commits after each task | Smaller models, cautious execution |
|
||||
| **One phase per loop** | Each agent call implements all tasks in a phase | LLM commits after each task (with JIRA ID) | Smart models with larger context windows, related tasks |
|
||||
|
||||
Session state is stored per-spec in `specs/<feature>/.plan2code-loop/`, keeping each feature's progress isolated.
|
||||
|
||||
The loop will:
|
||||
1. Read your `overview.md` and phase files
|
||||
2. Find the first unchecked task (or phase, in phase mode)
|
||||
3. Implement it and mark the checkbox complete
|
||||
4. Repeat until all tasks are done or max iterations reached
|
||||
|
||||
See [plan2code-loop/](plan2code-loop/) for full documentation.
|
||||
|
||||
---
|
||||
|
||||
## File Structure After Complete Implementation
|
||||
|
||||
```
|
||||
your-project/
|
||||
├── specs/
|
||||
│ └── another-feature/ # In-progress feature
|
||||
│ ├── overview.md
|
||||
│ └── Phase 1.md
|
||||
├── specs--completed/
|
||||
│ └── feature-name/
|
||||
│ ├── overview.md # Archived with completion summary
|
||||
│ ├── Phase 1.md # All checkboxes marked [x]
|
||||
│ ├── Phase 2.md
|
||||
│ └── ...
|
||||
├── your project files...
|
||||
└── README.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Customization
|
||||
|
||||
Feel free to modify these prompts to fit your workflow:
|
||||
|
||||
- **Add testing phases** - Uncomment/add testing requirements in Step 2
|
||||
- **Adjust confidence threshold** - Change the 90% threshold in Step 1
|
||||
- **Modify output structure** - Customize the specs folder organization
|
||||
- **Add code review steps** - Enhance Step 3 with additional review gates
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**Slash commands/workflows not recognized:**
|
||||
|
||||
- Ensure you ran `node install.js` and selected the appropriate platform
|
||||
- Restart your AI tool after installation
|
||||
- For per-project installation, ensure the directory isn't in `.gitignore`
|
||||
|
||||
**AI jumps ahead to implementation during planning:**
|
||||
|
||||
- The prompts explicitly forbid this, but if it happens, remind the AI: "Stay in planning mode. Do not write code yet."
|
||||
|
||||
**AI doesn't know what to implement:**
|
||||
|
||||
- Make sure you provided the path to `overview.md`
|
||||
- The AI will auto-detect the next phase and read the corresponding `Phase X.md` file
|
||||
|
||||
**Lost progress between sessions:**
|
||||
|
||||
- Check `overview.md` for phase status
|
||||
- Review individual phase files for task completion status
|
||||
|
||||
**AI not following spec exactly:**
|
||||
|
||||
- Reference the specific phase document and ask it to re-read the requirements
|
||||
|
||||
**Too many/few phases:**
|
||||
|
||||
- Adjust during Step 2 (Documentation) - phases should represent logical groupings of work
|
||||
|
||||
# Plan2Code
|
||||
|
||||
<img src="docs/banner.png" alt="Plan2Code — send the plan, the build follows" style="max-width:1024px;">
|
||||
|
||||
**A spec-driven workflow for AI coding agents. Send the plan — the build follows.**
|
||||
|
||||
An AI agent is a fine builder and a terrible client. Plan2Code stops making it both: you approve a
|
||||
plan, the plan becomes a set of phase documents in your repo, and the agent builds to those documents
|
||||
one phase at a time. Progress lives in files instead of chat history — so the next session, the next
|
||||
agent, and the next engineer all start from the same specs.
|
||||
|
||||
Six commands, each posted separately. Two of them are optional.
|
||||
|
||||
Version 2.0.0 · MIT · 📖 [plan2code.jparkerweb.com](https://plan2code.jparkerweb.com)
|
||||
|
||||
---
|
||||
|
||||
## Install
|
||||
|
||||
Requires [Node.js](https://nodejs.org/) 14 or later. Re-run any time to update.
|
||||
|
||||
```bash
|
||||
npx --allow-git=all git+https://github.com/jparkerweb/plan2code.git
|
||||
```
|
||||
|
||||
This fetches the installer to a temp directory, runs it, writes the slash commands for whichever
|
||||
tools you pick, and cleans up after itself. The installed commands work independently from then on.
|
||||
|
||||
Either route lands you on the same menu:
|
||||
|
||||
```
|
||||
╔═════════════════════════════════════════════════════════╗
|
||||
║ INSTALL PLAN2CODE ║
|
||||
╠═════════════════════════════════════════════════════════╣
|
||||
║ I. INSTALL Install Plan2Code for all platforms ║
|
||||
║ A. ALL Install Plan2Code + dev tools ║
|
||||
║ U. UNINSTALL Remove Plan2Code files ║
|
||||
║ C. CUSTOM Advanced options ║
|
||||
║ Q. QUIT Exit ║
|
||||
╚═════════════════════════════════════════════════════════╝
|
||||
```
|
||||
|
||||
**Supported tools:** Claude Code · Cursor · Windsurf · Continue · Codeium (IntelliJ) ·
|
||||
GitHub Copilot CLI · VS Code Copilot · Crush · Pi · Amp · OpenCode · Devin · Zed
|
||||
|
||||
<details>
|
||||
<summary>Prefer to clone?</summary>
|
||||
|
||||
```bash
|
||||
git clone https://github.com/jparkerweb/plan2code.git
|
||||
cd plan2code
|
||||
node install.js
|
||||
|
||||
# Only if you plan to modify or contribute to Plan2Code itself
|
||||
npm install && npx husky
|
||||
```
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
## The workflow
|
||||
|
||||
```
|
||||
┌╴╴╴╴╴╴╴╴╴╴╴╴┐
|
||||
╎0 PATHFINDER╎ optional · new in 2.0 · for an idea too big or unclear to plan
|
||||
└╴╴╴╴╴╴┬╴╴╴╴╴┘
|
||||
▼
|
||||
┌────────────┐ ┌────────────┐ ┌────────────┐ ┌╴╴╴╴╴╴╴╴╴╴╴╴┐ ┌────────────┐
|
||||
│ 1 PLAN │─>│ 2 DOCUMENT │─>│3 IMPLEMENT │─>╎ REVIEW ╎─>│ 4 FINALIZE │
|
||||
│decide what │ │ draw it as │ │build to the│ ╎ optional ╎ │verify, sum,│
|
||||
│ to build │ │phase specs │ │ drawing │ ╎ any time ╎ │ archive │
|
||||
└────────────┘ └────────────┘ └────────────┘ └╴╴╴╴╴╴╴╴╴╴╴╴┘ └────────────┘
|
||||
new chat new chat new chat/phase new chat new chat
|
||||
├◀─────────────── one feature, start to archive ──────────────────▶┤
|
||||
```
|
||||
|
||||
Every box is its own conversation. That is not a style preference — planning context leaking into
|
||||
implementation is where most agent drift starts.
|
||||
|
||||
| Command | Use it when |
|
||||
|---------|-------------|
|
||||
| `/plan2code-0-pathfinder` | The idea is too big and unclear to plan. Charts it as decisions, clears one per session, hands a hot plan draft to Step 1 |
|
||||
| `/plan2code-1-plan` | Starting a feature. Full requirements → architecture pass |
|
||||
| `/plan2code-2-document` | Planning is done. Turn the plan into phase specs |
|
||||
| `/plan2code-3-implement` | Build the next phase (one per conversation) |
|
||||
| `/plan2code-review` | Independent second opinion on local changes, then optional fixes |
|
||||
| `/plan2code-4-finalize` | All phases done. Validate, summarize, archive |
|
||||
| `/plan2code-init` | Generate this repo's `AGENTS.md` so every agent starts informed |
|
||||
| `/plan2code-init-update` | Fold what you learned this session back into `AGENTS.md` |
|
||||
| `/plan2code-quick-task` | A small change that doesn't warrant the full sequence |
|
||||
| `/plan2code-1b-revise-plan` | Requirements moved mid-build. Revise the specs, not the code |
|
||||
| `/plan2code-handoff` | Compact this conversation into a doc the next one resumes from |
|
||||
|
||||
---
|
||||
|
||||
## The four rules that do most of the work
|
||||
|
||||
**1 · A fresh conversation for each step, and each implementation phase.**
|
||||
Step 3 gets a new chat per phase, not one chat for all of them.
|
||||
|
||||
**2 · No code until the plan hits 90% confidence.**
|
||||
Step 1 will not finalize below the threshold. Under it, the agent keeps asking and keeps reading your
|
||||
code — and writes every assumption down where you can argue with it.
|
||||
|
||||
**3 · Checkboxes are the state, not the chat.**
|
||||
Progress lives in the spec files. Any agent, any session, resumes cold from them.
|
||||
|
||||
**4 · Reply `approved` to close a phase.**
|
||||
Nothing advances on a guess about what you meant.
|
||||
|
||||
---
|
||||
|
||||
## What lands in your repo
|
||||
|
||||
```
|
||||
your-project/
|
||||
├── specs/
|
||||
│ └── task-api/ ← in progress
|
||||
│ ├── pathfinder/ ← only if you charted it in Step 0
|
||||
│ │ ├── map.md the destination, the decisions, the fog
|
||||
│ │ └── questions/NN-<slug>.md one decision per file
|
||||
│ ├── PLAN-DRAFT-20260804.md ← Step 1: the verified plan
|
||||
│ ├── PLAN-CONVERSATION-*.md ← Step 1: how you got there
|
||||
│ ├── overview.md ← Step 2: phase list + parallel groups
|
||||
│ └── Phase 1.md … Phase N.md ← Step 2: one-point tasks, self-contained
|
||||
├── specs--completed/
|
||||
│ └── auth-refresh/ ← Step 4 files finished work here
|
||||
└── ...your code
|
||||
```
|
||||
|
||||
`specs/` is gitignored by default — it's your working drawing, not a deliverable. Share a folder
|
||||
deliberately with `git add -f` when you want to.
|
||||
|
||||
### Progress marks
|
||||
|
||||
| Mark | Status | Meaning |
|
||||
|------|--------|---------|
|
||||
| `[ ]` | Open | Unclaimed. Any agent picks it up cold. |
|
||||
| `[/]` | In progress | Claimed right now — which is how two agents run parallel phases without colliding. |
|
||||
| `[x]` | Done | Built, self-reviewed against the spec, approved by you. |
|
||||
|
||||
```markdown
|
||||
## Phases
|
||||
|
||||
- [x] Phase 1: Project setup
|
||||
- [x] Phase 2: Data model
|
||||
- [/] Phase 3: API endpoints ← an agent is on this now
|
||||
- [ ] Phase 4: Authentication ← next available
|
||||
```
|
||||
|
||||
Step 2 marks which phases don't share files. Open a second agent on one of those, and the `[/]` marks
|
||||
keep the two out of each other's way.
|
||||
|
||||
---
|
||||
|
||||
## The six steps in detail
|
||||
|
||||
Each one travels on its own — a fresh conversation, opened and closed, with the specs on disk as the
|
||||
only thing carried between them.
|
||||
|
||||
### 0 · Pathfinder 🧭 — optional, new in 2.0
|
||||
|
||||
Some ideas are too big and unclear to plan: you can feel the shape of the work but you can't write
|
||||
it as requirements, so planning would just invent the answers. Pathfinder finds the *way* to the
|
||||
destination; Step 1 then walks it.
|
||||
|
||||
1. **Name the destination** — one or two lines fixing what this effort is finding its way to. Settled
|
||||
first, because it fixes scope. It also asks where the map should live: **local files** under
|
||||
gitignored `specs/` (private, solo — the default), or **GitHub Issues** (a map issue with one
|
||||
sub-issue per decision, native blocking, so your team can see and work the frontier in the tracker).
|
||||
2. **Chart the map** — a breadth-first grilling surfaces the open decisions. Anything you can phrase
|
||||
*sharply* becomes a question file; anything you can only sense stays listed as fog.
|
||||
3. **Clear one question per session** — resolving a question burns off the fog behind it, graduating
|
||||
whatever just became sharp into new questions.
|
||||
4. **Hand off** — when nothing is left to decide, it writes a `PLAN-DRAFT` that
|
||||
`/plan2code-1-plan` resumes from at Phase 4, with requirements, context, and scope already
|
||||
answered.
|
||||
|
||||
**Question types:** `grill` (a decision only you can make — the default) · `research` (a fact gates
|
||||
it; background agents resolve these, several in parallel) · `sketch` (you need something concrete to
|
||||
react to) · `legwork` (manual work that has to happen before a decision is possible).
|
||||
|
||||
It never answers its own questions, and it **plans, it never builds.** When the urge to just build it
|
||||
arrives, the map is done. Skip Step 0 entirely when you already know what you're building.
|
||||
|
||||
**Out:** `specs/<feature>/pathfinder/map.md` + `questions/` (or a `pathfinder:map` issue and its
|
||||
sub-issues) → `PLAN-DRAFT-<date>.md`. The draft is always a local file — that is what Step 1 reads.
|
||||
|
||||
### 1 · Plan 🤔
|
||||
|
||||
The agent works as a senior architect through six phases, stopping for you after each: requirements
|
||||
analysis · system context (reading your actual codebase) · tech stack (needs your explicit sign-off) ·
|
||||
architecture design · technical specification · transition decision.
|
||||
|
||||
It won't finalize below **90% confidence**, and every assumption it makes is written into the draft.
|
||||
|
||||
**In:** a description of the feature. **Out:** `PLAN-DRAFT-<date>.md` + `PLAN-CONVERSATION-<date>.md`
|
||||
|
||||
### 2 · Document 📝
|
||||
|
||||
The plan becomes the drawing. One `overview.md` with the phase checklist, plus one file per phase of
|
||||
one-story-point tasks. Each phase is **self-contained** — an agent opening `Phase 3.md` cold needs
|
||||
nothing else to build it. Unit and E2E tests are excluded unless you ask for them.
|
||||
|
||||
The overview also identifies the **parallel execution groups**: phases with no shared files or
|
||||
dependencies, safe to run in separate agents at once.
|
||||
|
||||
**In:** the `PLAN-DRAFT`. **Out:** `overview.md` + `Phase 1…N.md`
|
||||
|
||||
### 3 · Implement ⚡
|
||||
|
||||
Point it at `overview.md` and it does the rest: finds the next unchecked phase, implements every task
|
||||
exactly as specified, ticks tasks off as they land, then reviews its own work against the spec and
|
||||
writes a completion summary.
|
||||
|
||||
One phase per conversation. It won't run tests unless the phase says to.
|
||||
|
||||
**In:** `specs/<feature>/overview.md`. **Out:** working code, and updated checkboxes.
|
||||
|
||||
### Review 🔬 — optional, any time
|
||||
|
||||
An independent second opinion, not a rubber stamp. It figures out its own scope (conversation
|
||||
context, your instruction, or the git diff as a fallback), analyses across 11 dimensions, and ranks
|
||||
findings Critical / Warning / Suggestion. Every finding cites a file and a line, or it gets dropped —
|
||||
and the review pass is read-only. It fixes things only if you ask, and verifies each fix afterwards.
|
||||
|
||||
Spec-aware when `specs/` exists, and works fine without it. Most useful right after a planning or
|
||||
implementation step, but there's no wrong time to run it.
|
||||
|
||||
### 4 · Finalize 🧹
|
||||
|
||||
Validates every task against its phase spec, writes the summary and the list of files touched, flags
|
||||
the docs that drifted (`README`, `CHANGELOG`, `AGENTS.md`), then archives the whole spec folder —
|
||||
`pathfinder/` included — to `specs--completed/`. That folder is the record of *why* the code looks
|
||||
like this.
|
||||
|
||||
**In:** `specs/<feature>/overview.md`. **Out:** archived specs.
|
||||
|
||||
---
|
||||
|
||||
## What to bring to each step
|
||||
|
||||
| Step | Required input |
|
||||
|------|----------------|
|
||||
| 0 · Pathfinder | Nothing to start — just describe the idea. To continue: the feature name; it finds its own map |
|
||||
| 1 · Plan | Nothing — describe the feature |
|
||||
| 2 · Document | `specs/<feature>/PLAN-DRAFT-<date>.md`, or the planning conversation |
|
||||
| 3 · Implement | `specs/<feature>/overview.md` — it detects the phase itself |
|
||||
| Review | Scope guidance, e.g. "the last two phases", "just the auth module", "the whole PR". Auto-detects if you give none |
|
||||
| 4 · Finalize | `specs/<feature>/overview.md` |
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**Slash commands aren't recognised.** Re-run `node install.js` for that platform and restart your AI
|
||||
tool. For a per-project install, check the directory isn't gitignored.
|
||||
|
||||
**The agent starts coding during planning.** The prompts forbid it, but models drift. Say: "Stay in
|
||||
planning mode. Do not write code yet."
|
||||
|
||||
**The agent doesn't know what to implement.** Give it the path to `overview.md` — it reads the phase
|
||||
file itself from there.
|
||||
|
||||
**You lost track between sessions.** `overview.md` has the phase status; the phase files have the
|
||||
task status. That's the whole state.
|
||||
|
||||
**The agent isn't following the spec.** Point at the specific phase document and tell it to re-read
|
||||
the requirements.
|
||||
|
||||
**Too many or too few phases.** Fix it in Step 2 — a phase should be a logical grouping of work, not
|
||||
a fixed size.
|
||||
|
||||
---
|
||||
|
||||
## Customizing
|
||||
|
||||
The prompts are yours to edit. Common changes: add testing requirements in Step 2, move the 90%
|
||||
confidence threshold in Step 1, restructure the `specs/` layout, or add review gates to Step 3.
|
||||
Source files live in `src/`; re-run `node install.js` to push your edits out to every platform.
|
||||
|
||||
---
|
||||
|
||||
## Dive deeper
|
||||
|
||||
Core reference:
|
||||
|
||||
- **[QUICK-REFERENCE.md](QUICK-REFERENCE.md)** — the one-page card: commands, inputs, outputs, decision tree
|
||||
- **[.readme/walkthrough.md](.readme/walkthrough.md)** — one feature from a sentence to archived specs, session by session
|
||||
- **[AGENTS.md](AGENTS.md)** — architecture and contributor guide for this repo
|
||||
- **[CHANGELOG.md](CHANGELOG.md)** — what changed, and why
|
||||
|
||||
Optional tooling — none of it is required to use the workflow:
|
||||
|
||||
- **[.readme/autonomous-loop.md](.readme/autonomous-loop.md)** — `plan2code-loop`, a hands-off alternative to Step 3
|
||||
- **[.readme/status-line.md](.readme/status-line.md)** — three-line Claude Code status bar: model, context, quota, diff
|
||||
- **[.readme/metrics.md](.readme/metrics.md)** — `plan2code-metrics`, measuring and improving the prompts themselves
|
||||
- **[.readme/test-bot.md](.readme/test-bot.md)** — `plan2code-bot`, maintainer harness that runs the whole workflow unattended
|
||||
|
||||
|
After Width: | Height: | Size: 6.1 KiB |
|
After Width: | Height: | Size: 127 KiB |
|
Before Width: | Height: | Size: 185 KiB |
|
Before Width: | Height: | Size: 24 KiB After Width: | Height: | Size: 1.9 KiB |
|
Before Width: | Height: | Size: 2.4 KiB After Width: | Height: | Size: 584 B |
@@ -0,0 +1,33 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 64 64" width="64" height="64" role="img" aria-label="Plan2Code postage stamp">
|
||||
<defs>
|
||||
<!-- Perforated stamp silhouette: white keeps, black bites out.
|
||||
Few, deep notches so the scalloped edge still reads at 16px. -->
|
||||
<mask id="perf">
|
||||
<rect width="64" height="64" fill="#000"/>
|
||||
<rect x="2" y="2" width="60" height="60" fill="#fff"/>
|
||||
<g fill="#000">
|
||||
<circle cx="2" cy="2" r="5"/><circle cx="14" cy="2" r="5"/><circle cx="26" cy="2" r="5"/>
|
||||
<circle cx="38" cy="2" r="5"/><circle cx="50" cy="2" r="5"/><circle cx="62" cy="2" r="5"/>
|
||||
<circle cx="2" cy="62" r="5"/><circle cx="14" cy="62" r="5"/><circle cx="26" cy="62" r="5"/>
|
||||
<circle cx="38" cy="62" r="5"/><circle cx="50" cy="62" r="5"/><circle cx="62" cy="62" r="5"/>
|
||||
<circle cx="2" cy="14" r="5"/><circle cx="2" cy="26" r="5"/>
|
||||
<circle cx="2" cy="38" r="5"/><circle cx="2" cy="50" r="5"/>
|
||||
<circle cx="62" cy="14" r="5"/><circle cx="62" cy="26" r="5"/>
|
||||
<circle cx="62" cy="38" r="5"/><circle cx="62" cy="50" r="5"/>
|
||||
</g>
|
||||
</mask>
|
||||
</defs>
|
||||
|
||||
<g mask="url(#perf)">
|
||||
<!-- printed stamp -->
|
||||
<rect x="2" y="2" width="60" height="60" fill="#D33A38"/>
|
||||
<!-- the 2, set as a franking-machine numeral -->
|
||||
<g fill="#FBFAF6">
|
||||
<rect x="16" y="13" width="32" height="8"/>
|
||||
<rect x="40" y="13" width="8" height="22"/>
|
||||
<rect x="16" y="27" width="32" height="8"/>
|
||||
<rect x="16" y="27" width="8" height="24"/>
|
||||
<rect x="16" y="43" width="32" height="8"/>
|
||||
</g>
|
||||
</g>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 1.6 KiB |
|
Before Width: | Height: | Size: 147 KiB |
|
Before Width: | Height: | Size: 95 KiB |
@@ -1,12 +1,13 @@
|
||||
{
|
||||
"name": "plan2code",
|
||||
"version": "1.8.1",
|
||||
"version": "2.1.1",
|
||||
"private": true,
|
||||
"bin": {
|
||||
"plan2code": "./install.js"
|
||||
},
|
||||
"scripts": {
|
||||
"prepare": "husky"
|
||||
"prepare": "husky",
|
||||
"test": "node scripts/validate-char-count.js"
|
||||
},
|
||||
"devDependencies": {
|
||||
"husky": "^9.0.0"
|
||||
|
||||
@@ -0,0 +1,49 @@
|
||||
# Changelog
|
||||
|
||||
## 1.1.0
|
||||
|
||||
### Resume support
|
||||
- Add `--resume` flag to continue incomplete runs from saved state
|
||||
- Skip previously succeeded steps when resuming (init, plan, document, implement, finalize)
|
||||
- Restore idea name, description, project directory, and implement pass counter from state
|
||||
- Auto-detect state files in current directory (enhancement mode) or subdirectories (new-project mode)
|
||||
- Delete state file automatically after a fully successful run
|
||||
- Preserve state file on failure for later resume
|
||||
- Add `deleteState()` and `findExistingState()` utilities to bot-state module
|
||||
- Add bot-state unit tests (saveState, loadState, deleteState, findExistingState)
|
||||
|
||||
### Idea generation improvements
|
||||
- Expand idea categories from binary CLI/web-app coin flip to 12 diverse categories (games, dashboards, browser extensions, desktop utilities, etc.)
|
||||
- Add guidance to avoid defaulting to developer-centric tools (git analyzers, code formatters)
|
||||
- Strengthen `--idea` seed clause so the LLM stays aligned with the user's theme instead of ignoring it
|
||||
- Update system prompt to encourage creative, cross-domain ideas
|
||||
|
||||
### Init step overhaul (new projects)
|
||||
- Init now creates a minimal AGENTS.md stub (name, description, status) instead of running the full `/plan2code-init` skill
|
||||
- Prevents hallucinated architecture, commands, and `.agents-docs/` files before the plan step runs
|
||||
- Init evaluation criteria updated to reward minimalism and penalize premature detail
|
||||
|
||||
### Implement step overhaul
|
||||
- Implement step now works directly with Read/Write/Edit/Glob/Grep tools instead of delegating to Skill sub-session
|
||||
- Inlined step-by-step process: find specs, pick phase, implement tasks, mark checkboxes
|
||||
- Fixes issue where Skill sub-sessions did all work invisibly, causing zero tool observations
|
||||
|
||||
### Observation tracking fix
|
||||
- Capture `tool_use` blocks from the assistant message stream in session-runner as a fallback when `canUseTool` callback doesn't fire
|
||||
- Add deduplication in ObservationCollector to prevent double-counting from both sources
|
||||
- Fixes all steps reporting 0 tools used / 0 files created in BOT-NOTES and evaluations
|
||||
|
||||
### Evaluator improvements
|
||||
- Increase evaluator `maxTurns` from 3 to 30 so it has room for tool calls before producing the scored response
|
||||
- Add warning log when evaluation parser can't find SCORE in output (was silently defaulting to 50)
|
||||
|
||||
## 1.0.0
|
||||
|
||||
- Initial release
|
||||
- Two auto-detected modes: new-project and enhancement
|
||||
- `--idea` flag to seed the idea generator
|
||||
- Full workflow execution: init → plan → document → implement → finalize
|
||||
- Artifact validation after each step
|
||||
- State persistence to `.plan2code-bot-state.json`
|
||||
- Auto-responder for autonomous Claude Agent SDK sessions
|
||||
- Bot-friendly skill installation (strips `disable-model-invocation`)
|
||||
@@ -0,0 +1,234 @@
|
||||
# LLM-as-Judge Evaluation System
|
||||
|
||||
## Overview
|
||||
|
||||
The plan2code-bot now includes an **always-on LLM-as-judge evaluation system** that transforms it from a "yes-man" into an authentic QA agent. This provides realistic quality signals for `plan2code-metrics` to analyze and drive recursive self-improvement.
|
||||
|
||||
## Key Features
|
||||
|
||||
### 1. Intelligent Decision Making (Real-Time)
|
||||
|
||||
**What:** During execution, when `AskUserQuestion` is called, the bot uses an LLM to make thoughtful decisions based on current observations.
|
||||
|
||||
**How it works:**
|
||||
- Collects observations up to the current point (tools used, files created, errors)
|
||||
- Queries LLM with context: "Given what you've seen, should you approve this plan?"
|
||||
- LLM inspects current artifacts using Read/Glob/Grep
|
||||
- Returns evidence-based answer with reasoning
|
||||
- All decisions are recorded for metrics analysis
|
||||
|
||||
**Example:**
|
||||
```
|
||||
Question: "Approve this plan?"
|
||||
Observations: Created PLAN-DRAFT.md, 3 phases, 42s duration, no errors
|
||||
LLM reads PLAN-DRAFT.md, evaluates quality
|
||||
LLM decides: "Yes, approve - phases are well-scoped and realistic"
|
||||
```
|
||||
|
||||
### 2. Post-Step Evaluation
|
||||
|
||||
**What:** After each step completes, the bot evaluates quality using step-specific criteria.
|
||||
|
||||
**How it works:**
|
||||
- Collects complete execution observations
|
||||
- Queries LLM with evaluation criteria for the step
|
||||
- LLM inspects final artifacts
|
||||
- Returns structured evaluation (score, strengths, weaknesses, suggestions)
|
||||
- Writes `specs/<feature>/BOT-EVALUATION.md` and `specs/<feature>/BOT-NOTES.md` (falls back to project root if no spec folder exists yet, e.g. during `init`)
|
||||
|
||||
**Example output:**
|
||||
```markdown
|
||||
# Evaluation: plan Step
|
||||
|
||||
**Score:** 78/100
|
||||
|
||||
## Strengths
|
||||
- Clear phase breakdown with realistic scope
|
||||
- Tech stack choices appropriate
|
||||
|
||||
## Weaknesses
|
||||
- Phase 3 description too vague
|
||||
- No testing strategy mentioned
|
||||
|
||||
## Suggestions
|
||||
- Expand Phase 3 with concrete tasks
|
||||
- Add explicit testing phase
|
||||
```
|
||||
|
||||
### 3. Quality Gate
|
||||
|
||||
**What:** Before finalize, checks that average quality score is acceptable.
|
||||
|
||||
**How it works:**
|
||||
- Calculates average score across all evaluated steps
|
||||
- If average < 60, blocks finalization
|
||||
- Displays clear message about quality issues
|
||||
- User must review `specs/<feature>/BOT-EVALUATION.md` and fix problems
|
||||
|
||||
## Files Created
|
||||
|
||||
### New Files
|
||||
|
||||
1. **`src/observation-collector.ts`**
|
||||
- Tracks execution details (tools, files, messages, errors, questions)
|
||||
- Provides snapshots for real-time decisions
|
||||
- Captures complete history for evaluation
|
||||
|
||||
2. **`src/intelligent-responder.ts`**
|
||||
- Replaces hardcoded auto-responder
|
||||
- Uses LLM to answer AskUserQuestion prompts
|
||||
- Provides reasoning for all decisions
|
||||
- Falls back gracefully if LLM unavailable
|
||||
|
||||
3. **`src/prompts/evaluation-criteria.ts`**
|
||||
- Step-specific evaluation criteria (init, plan, document, implement, finalize)
|
||||
- Quality checks, common pitfalls, scoring guidance
|
||||
- Emphasizes honest scoring (most work should score 70-85)
|
||||
|
||||
4. **`src/evaluator.ts`**
|
||||
- Post-step evaluation using LLM-as-judge
|
||||
- Queries LLM with observations and criteria
|
||||
- Parses structured evaluation output
|
||||
- Writes `specs/<feature>/BOT-EVALUATION.md` and `specs/<feature>/BOT-NOTES.md`
|
||||
|
||||
### Modified Files
|
||||
|
||||
1. **`src/types.ts`**
|
||||
- Added interfaces: `ToolObservation`, `QuestionContext`, `ExecutionObservation`, `EvaluationResult`
|
||||
- Extended `StepResult` with `evaluation` and `observations` fields
|
||||
|
||||
2. **`src/session-runner.ts`**
|
||||
- Added `collector` parameter to `SessionOptions`
|
||||
- Returns `observations` in `SessionResult`
|
||||
- Records all messages for observation tracking
|
||||
- Uses intelligent responder instead of auto-responder
|
||||
|
||||
3. **`src/cli.ts`**
|
||||
- Creates `ObservationCollector` for each step
|
||||
- Always runs evaluation after successful steps
|
||||
- Displays scores with color coding (green/yellow/red)
|
||||
- Implements quality gate before finalize
|
||||
- Shows evaluation summary in step output
|
||||
|
||||
4. **`src/bin/plan2code-bot.ts`**
|
||||
- Updated help text to mention LLM-as-judge evaluation
|
||||
- No new CLI flags (evaluation is always on)
|
||||
|
||||
### Deleted Files
|
||||
|
||||
1. **`src/auto-responder.ts`** - Replaced by intelligent-responder.ts
|
||||
2. **`src/auto-responder.test.ts`** - No longer needed
|
||||
|
||||
## Output Files (Created During Execution)
|
||||
|
||||
Both files are written to `specs/<feature>/` so they stay co-located with the feature they describe. If no spec folder exists yet (e.g. during `init`), they fall back to the project root.
|
||||
|
||||
### BOT-EVALUATION.md
|
||||
|
||||
Contains evaluation results for each step:
|
||||
- Score (0-100)
|
||||
- Strengths identified
|
||||
- Weaknesses found
|
||||
- Suggestions for improvement
|
||||
- Critical issues (if any)
|
||||
- Full reasoning from LLM
|
||||
|
||||
### BOT-NOTES.md
|
||||
|
||||
Contains execution observations:
|
||||
- Duration, tool counts, file changes
|
||||
- Questions asked and LLM reasoning for answers
|
||||
- Tool usage timeline
|
||||
- Files created/modified
|
||||
- Assistant output summary
|
||||
|
||||
## Data Structure for Metrics
|
||||
|
||||
All evaluation data is structured in `StepResult`:
|
||||
|
||||
```typescript
|
||||
{
|
||||
step: 'plan',
|
||||
success: true,
|
||||
duration: 42000,
|
||||
evaluation: {
|
||||
score: 78,
|
||||
strengths: ["Clear phases", "Realistic scope"],
|
||||
weaknesses: ["Phase 3 too vague"],
|
||||
suggestions: ["Add specific tasks to Phase 3"],
|
||||
criticalIssues: [],
|
||||
reasoning: "...",
|
||||
timestamp: 1234567890,
|
||||
evaluatorModel: 'claude-sonnet-4-5'
|
||||
},
|
||||
observations: {
|
||||
tools: [{ toolName, input, output, timestamp }, ...],
|
||||
questionsAsked: [
|
||||
{
|
||||
question: "Approve plan?",
|
||||
selectedAnswer: "Yes, approve",
|
||||
llmReasoning: "Phases are well-scoped...",
|
||||
timestamp: 1234567890
|
||||
}
|
||||
],
|
||||
filesCreated: [...],
|
||||
filesModified: [...],
|
||||
errors: []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Benefits for Recursive Improvement
|
||||
|
||||
1. **Authentic Signals:** Real quality scores identify actual problem areas
|
||||
2. **Detailed Context:** Observations + reasoning explain WHY failures happen
|
||||
3. **Correlation Analysis:** Link patterns (tool usage, duration, errors) to quality
|
||||
4. **Continuous Loop:** Better metrics → improved workflows → higher scores → repeat
|
||||
|
||||
## Usage
|
||||
|
||||
No special flags needed - evaluation is always on:
|
||||
|
||||
```bash
|
||||
# New project
|
||||
plan2code-bot --idea "todo app"
|
||||
|
||||
# Enhancement
|
||||
cd my-project && plan2code-bot
|
||||
|
||||
# Resume with evaluation data preserved
|
||||
plan2code-bot --resume
|
||||
```
|
||||
|
||||
## Verification
|
||||
|
||||
After running the bot, check:
|
||||
|
||||
1. **`specs/<feature>/BOT-EVALUATION.md`** - Should show realistic scores (not all 100s)
|
||||
2. **`specs/<feature>/BOT-NOTES.md`** - Should show LLM reasoning for decisions
|
||||
3. **Console output** - Should display color-coded scores after each step
|
||||
4. **State file** (`.plan2code-bot-state.json`) - Should include evaluation data
|
||||
|
||||
## Trade-offs
|
||||
|
||||
### Latency
|
||||
- Adds ~2-3s per AskUserQuestion call (~15-20s total per run)
|
||||
- Worth it for authentic evaluation
|
||||
|
||||
### Token Cost
|
||||
- ~20-26K tokens per run (~$0.60 with Opus 4.6)
|
||||
- Investment pays off through metrics-driven improvement
|
||||
|
||||
### Determinism
|
||||
- LLM decisions vary between runs (non-deterministic)
|
||||
- Realistic - humans vary too
|
||||
- Metrics average over many runs
|
||||
|
||||
## Future Enhancements
|
||||
|
||||
Potential improvements:
|
||||
- Model selection per step (use Haiku for simple decisions)
|
||||
- Configurable quality gate threshold
|
||||
- Historical score tracking across runs
|
||||
- Comparison with previous evaluations
|
||||
- More sophisticated scoring (weighted by step importance)
|
||||
@@ -0,0 +1,111 @@
|
||||
# plan2code-bot
|
||||
|
||||
Autonomous workflow test runner for plan2code. Uses the Claude Agent SDK to simulate a human running through the entire plan2code workflow (init, plan, document, implement, finalize) end-to-end.
|
||||
|
||||
## Two Modes (Auto-Detected)
|
||||
|
||||
1. **New Project Mode** — No `AGENTS.md` in cwd: generates an app idea, creates a subdirectory, writes IDEA.md, runs init, then all 4 steps.
|
||||
2. **Enhancement Mode** — `AGENTS.md` exists in cwd: scans the existing codebase and proposes a realistic enhancement, writes IDEA.md, then runs plan through finalize.
|
||||
|
||||
## Installation
|
||||
|
||||
From the plan2code root:
|
||||
|
||||
```bash
|
||||
node install.js
|
||||
# Select C > B to install bot only, or I to install everything
|
||||
```
|
||||
|
||||
Or manually:
|
||||
|
||||
```bash
|
||||
cd plan2code-bot
|
||||
npm install
|
||||
npm run build
|
||||
npm link
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# New project mode (run from an empty directory)
|
||||
mkdir /tmp/test-bot && cd /tmp/test-bot
|
||||
plan2code-bot
|
||||
|
||||
# Enhancement mode (run from an existing project with AGENTS.md)
|
||||
cd my-project
|
||||
plan2code-bot
|
||||
|
||||
# Seed the idea generator with a specific concept
|
||||
plan2code-bot --idea "web app that displays the current weather as vector images"
|
||||
|
||||
# Resume a previous incomplete run
|
||||
plan2code-bot --resume
|
||||
```
|
||||
|
||||
### `--idea`
|
||||
|
||||
Pass a quoted string after `--idea` to seed the idea generator with a specific concept. The AI will use it as inspiration rather than generating a completely random idea. Wrap the value in double quotes so the shell treats it as a single argument.
|
||||
|
||||
```bash
|
||||
# Specific app concept
|
||||
plan2code-bot --idea "web app that displays the current weather as vector images"
|
||||
|
||||
# Short keyword to nudge the category
|
||||
plan2code-bot --idea "markdown editor"
|
||||
|
||||
# Detailed constraint
|
||||
plan2code-bot --idea "CLI tool that converts CSV files to SQLite databases with type inference"
|
||||
|
||||
# Works in enhancement mode too — guides what kind of enhancement to propose
|
||||
cd my-existing-project
|
||||
plan2code-bot --idea "add dark mode support"
|
||||
```
|
||||
|
||||
Without `--idea`, the bot picks a random category (CLI tool or web app) and invents something on its own.
|
||||
|
||||
### `--resume`
|
||||
|
||||
Resume a previous incomplete run. The bot searches for a `.plan2code-bot-state.json` file in the current directory (enhancement mode) or in immediate subdirectories (new-project mode). If found, it restores the idea, config, and progress — skipping steps that already succeeded and continuing from where it left off.
|
||||
|
||||
```bash
|
||||
# A run failed at the implement step — resume it
|
||||
plan2code-bot --resume
|
||||
|
||||
# Can combine with --idea (idea is ignored when resuming since it's restored from state)
|
||||
plan2code-bot --resume --idea "ignored when state exists"
|
||||
```
|
||||
|
||||
If no state file is found, the bot starts a fresh run.
|
||||
|
||||
## How It Works
|
||||
|
||||
1. Detects mode based on presence of `AGENTS.md`
|
||||
2. Generates an idea (new app or enhancement) via Claude, optionally guided by `--idea` seed
|
||||
3. Writes `IDEA.md` to the project directory
|
||||
4. Installs bot-friendly copies of plan2code skills (strips `disable-model-invocation` so sessions can invoke them)
|
||||
5. Runs each workflow step as a separate Claude Agent SDK session:
|
||||
- **init** — generates `AGENTS.md` (new-project mode only)
|
||||
- **plan** — creates plan draft in `specs/<feature>/`
|
||||
- **document** — produces `overview.md` and `phase-*.md` files
|
||||
- **implement** — loops until all phases are complete (max 10 passes)
|
||||
- **finalize** — validates and archives to `specs--completed/`
|
||||
6. Validates expected artifacts after each step (aborts on missing artifacts)
|
||||
7. Moves `IDEA.md` into `specs/<feature>/` after the plan step so it stays with its feature
|
||||
8. Auto-responds to `AskUserQuestion` prompts (approvals, testing gates, name questions)
|
||||
9. Saves state to `.plan2code-bot-state.json` after each step
|
||||
|
||||
## State File
|
||||
|
||||
After each step, the bot saves its state to `.plan2code-bot-state.json` in the project directory. This includes the config, all step results, and progress tracking.
|
||||
|
||||
- **On full success** — the state file is automatically deleted (clean finish)
|
||||
- **On failure/incomplete** — the state file is preserved so you can `--resume` later
|
||||
|
||||
## Development
|
||||
|
||||
```bash
|
||||
npm run build # Build with tsup
|
||||
npm run dev # Watch mode
|
||||
npm test # Run tests (vitest)
|
||||
```
|
||||
@@ -0,0 +1,36 @@
|
||||
{
|
||||
"name": "plan2code-bot",
|
||||
"version": "1.1.0",
|
||||
"description": "Plan2Code Bot - Autonomous workflow runner for testing plan2code end-to-end",
|
||||
"type": "module",
|
||||
"main": "dist/index.js",
|
||||
"bin": {
|
||||
"plan2code-bot": "./dist/bin/plan2code-bot.js"
|
||||
},
|
||||
"files": [
|
||||
"dist"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=18.0.0"
|
||||
},
|
||||
"scripts": {
|
||||
"build": "tsup",
|
||||
"dev": "tsup --watch",
|
||||
"start": "node dist/bin/plan2code-bot.js",
|
||||
"test": "vitest run",
|
||||
"prepublishOnly": "npm run build"
|
||||
},
|
||||
"dependencies": {
|
||||
"@anthropic-ai/claude-agent-sdk": "^0.2.63",
|
||||
"chalk": "^5.6.2",
|
||||
"fs-extra": "^11.3.3",
|
||||
"ora": "^9.0.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
"@types/fs-extra": "^11.0.4",
|
||||
"@types/node": "^25.0.3",
|
||||
"tsup": "^8.5.1",
|
||||
"typescript": "^5.9.3",
|
||||
"vitest": "^4.0.18"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,116 @@
|
||||
import { runCLI } from '../cli.js';
|
||||
|
||||
function showHelp(): void {
|
||||
console.log(`
|
||||
+----------------------------------------------------------------+
|
||||
— PLAN2CODEDE-BOT —
|
||||
—----------------------------------------------------------------—
|
||||
— Autonomous workflow test runner foplan2codede —
|
||||
— Features LLM-as-judge for honest quality evaluation —
|
||||
+----------------------------------------------------------------+
|
||||
|
||||
Usage:
|
||||
plan2code-bot [options]
|
||||
|
||||
Options:
|
||||
--help Show this help message
|
||||
--idea <string> Seed the idea generator with a specific concept
|
||||
Example: --idea "web app for weather"
|
||||
Example: --idea="CLI tool for CSV conversion"
|
||||
--resume Resume a previous incomplete run
|
||||
|
||||
Modes:
|
||||
— New Project Mode - Run from empty directory
|
||||
The bot generates an app idea, creates a subdirectory, writes
|
||||
IDEA.md, runs init, then all 4 workflow steps.
|
||||
|
||||
— Enhancement Mode - Run from directory with AGENTS.md
|
||||
The bot scans the existing codebase, proposes an enhancement,
|
||||
writes IDEA.md, then runs plan through finalize.
|
||||
|
||||
Evaluation:
|
||||
The bot acts as an authentic QA agent, using LLM-based decision
|
||||
making during execution and providing honest quality assessments
|
||||
after each step. Results are written to specs/<feature>/BOT-EVALUATION.md
|
||||
and specs/<feature>/BOT-NOTES.md for metrics analysis.
|
||||
|
||||
Examples:
|
||||
# New project (from empty directory)
|
||||
plan2code-bot
|
||||
|
||||
# Enhancement (from existing project)
|
||||
cd my-project && plan2code-bot
|
||||
|
||||
# With specific idea
|
||||
plan2code-bot --idea "markdown editor with live preview"
|
||||
|
||||
# Resume incomplete run
|
||||
plan2code-bot --resume
|
||||
|
||||
Documentation:
|
||||
https://github.com/jparkerweb/plan2code
|
||||
`);
|
||||
}
|
||||
|
||||
function stripQuotes(str: string): string {
|
||||
// Remove surrounding quotes if present (both single and double)
|
||||
if ((str.startsWith('"') && str.endsWith('"')) ||
|
||||
(str.startsWith("'") && str.endsWith("'"))) {
|
||||
return str.slice(1, -1);
|
||||
}
|
||||
return str;
|
||||
}
|
||||
|
||||
function parseArgs(): { idea?: string; resume?: boolean; help?: boolean } {
|
||||
const args = process.argv.slice(2);
|
||||
|
||||
// Check for --help
|
||||
if (args.includes('--help') || args.includes('-h')) {
|
||||
return { help: true };
|
||||
}
|
||||
|
||||
// Parse --idea (supports both --idea="value" and --idea "value")
|
||||
let idea: string | undefined;
|
||||
for (let i = 0; i < args.length; i++) {
|
||||
const arg = args[i];
|
||||
|
||||
// Format: --idea="value"
|
||||
if (arg.startsWith('--idea=')) {
|
||||
idea = stripQuotes(arg.substring('--idea='.length));
|
||||
break;
|
||||
}
|
||||
|
||||
// Format: --idea "value"
|
||||
if (arg === '--idea' && i + 1 < args.length) {
|
||||
idea = stripQuotes(args[i + 1]);
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
// Parse --resume
|
||||
const resume = args.includes('--resume');
|
||||
|
||||
return { idea, resume: resume || undefined };
|
||||
}
|
||||
|
||||
async function main() {
|
||||
try {
|
||||
const { idea, resume, help } = parseArgs();
|
||||
|
||||
if (help) {
|
||||
showHelp();
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
await runCLI({ idea, resume });
|
||||
process.exit(0);
|
||||
} catch (err) {
|
||||
if (err instanceof Error && err.message.includes('User force closed')) {
|
||||
process.exit(0);
|
||||
}
|
||||
console.error(err instanceof Error ? err.message : String(err));
|
||||
process.exit(1);
|
||||
}
|
||||
}
|
||||
|
||||
main();
|
||||
@@ -0,0 +1,121 @@
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
|
||||
import fs from 'fs-extra';
|
||||
import path from 'path';
|
||||
import os from 'os';
|
||||
import { saveState, loadState, deleteState, findExistingState } from './bot-state.js';
|
||||
import type { BotState } from './types.js';
|
||||
|
||||
function makeTmpDir(): string {
|
||||
return fs.mkdtempSync(path.join(os.tmpdir(), 'bot-state-test-'));
|
||||
}
|
||||
|
||||
function makeState(projectDir: string): BotState {
|
||||
return {
|
||||
config: {
|
||||
workDir: path.dirname(projectDir),
|
||||
projectDir,
|
||||
ideaName: 'test-idea',
|
||||
ideaDescription: 'A test idea',
|
||||
mode: 'new-project',
|
||||
},
|
||||
steps: [],
|
||||
currentStep: null,
|
||||
implementPasses: 0,
|
||||
allPhasesComplete: false,
|
||||
};
|
||||
}
|
||||
|
||||
describe('bot-state', () => {
|
||||
let tmpDir: string;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = makeTmpDir();
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
fs.removeSync(tmpDir);
|
||||
});
|
||||
|
||||
describe('saveState', () => {
|
||||
it('writes valid JSON', () => {
|
||||
const projectDir = path.join(tmpDir, 'project');
|
||||
fs.ensureDirSync(projectDir);
|
||||
const state = makeState(projectDir);
|
||||
|
||||
saveState(state);
|
||||
|
||||
const filePath = path.join(projectDir, '.plan2code-bot-state.json');
|
||||
expect(fs.existsSync(filePath)).toBe(true);
|
||||
const parsed = fs.readJsonSync(filePath);
|
||||
expect(parsed.config.ideaName).toBe('test-idea');
|
||||
});
|
||||
});
|
||||
|
||||
describe('loadState', () => {
|
||||
it('returns state from file', () => {
|
||||
const projectDir = path.join(tmpDir, 'project');
|
||||
fs.ensureDirSync(projectDir);
|
||||
const state = makeState(projectDir);
|
||||
saveState(state);
|
||||
|
||||
const loaded = loadState(projectDir);
|
||||
expect(loaded).not.toBeNull();
|
||||
expect(loaded!.config.ideaName).toBe('test-idea');
|
||||
expect(loaded!.implementPasses).toBe(0);
|
||||
});
|
||||
|
||||
it('returns null when file does not exist', () => {
|
||||
const result = loadState(path.join(tmpDir, 'nonexistent'));
|
||||
expect(result).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
describe('deleteState', () => {
|
||||
it('removes the file', () => {
|
||||
const projectDir = path.join(tmpDir, 'project');
|
||||
fs.ensureDirSync(projectDir);
|
||||
const state = makeState(projectDir);
|
||||
saveState(state);
|
||||
|
||||
const filePath = path.join(projectDir, '.plan2code-bot-state.json');
|
||||
expect(fs.existsSync(filePath)).toBe(true);
|
||||
|
||||
deleteState(projectDir);
|
||||
expect(fs.existsSync(filePath)).toBe(false);
|
||||
});
|
||||
|
||||
it('is a no-op when file does not exist', () => {
|
||||
// Should not throw
|
||||
deleteState(path.join(tmpDir, 'nonexistent'));
|
||||
});
|
||||
});
|
||||
|
||||
describe('findExistingState', () => {
|
||||
it('finds state in workDir (enhancement mode)', () => {
|
||||
const state = makeState(tmpDir);
|
||||
state.config.projectDir = tmpDir;
|
||||
state.config.mode = 'enhancement';
|
||||
saveState(state);
|
||||
|
||||
const found = findExistingState(tmpDir);
|
||||
expect(found).not.toBeNull();
|
||||
expect(found!.config.ideaName).toBe('test-idea');
|
||||
});
|
||||
|
||||
it('finds state in a subdirectory (new-project mode)', () => {
|
||||
const projectDir = path.join(tmpDir, 'my-app');
|
||||
fs.ensureDirSync(projectDir);
|
||||
const state = makeState(projectDir);
|
||||
saveState(state);
|
||||
|
||||
const found = findExistingState(tmpDir);
|
||||
expect(found).not.toBeNull();
|
||||
expect(found!.config.projectDir).toBe(projectDir);
|
||||
});
|
||||
|
||||
it('returns null when no state exists', () => {
|
||||
const found = findExistingState(tmpDir);
|
||||
expect(found).toBeNull();
|
||||
});
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,48 @@
|
||||
import fs from 'fs-extra';
|
||||
import path from 'path';
|
||||
import type { BotState } from './types.js';
|
||||
|
||||
const STATE_FILE = '.plan2code-bot-state.json';
|
||||
|
||||
export function saveState(state: BotState): void {
|
||||
const filePath = path.join(state.config.projectDir, STATE_FILE);
|
||||
fs.writeJsonSync(filePath, state, { spaces: 2 });
|
||||
}
|
||||
|
||||
export function loadState(projectDir: string): BotState | null {
|
||||
const filePath = path.join(projectDir, STATE_FILE);
|
||||
if (!fs.existsSync(filePath)) {
|
||||
return null;
|
||||
}
|
||||
return fs.readJsonSync(filePath) as BotState;
|
||||
}
|
||||
|
||||
export function deleteState(projectDir: string): void {
|
||||
const filePath = path.join(projectDir, STATE_FILE);
|
||||
if (fs.existsSync(filePath)) {
|
||||
fs.removeSync(filePath);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Search for an existing state file in workDir (enhancement mode)
|
||||
* or in immediate subdirectories (new-project mode).
|
||||
*/
|
||||
export function findExistingState(workDir: string): BotState | null {
|
||||
// Enhancement mode: state is in workDir directly
|
||||
const direct = loadState(workDir);
|
||||
if (direct) return direct;
|
||||
|
||||
// New-project mode: state is in a subdirectory
|
||||
try {
|
||||
for (const entry of fs.readdirSync(workDir, { withFileTypes: true })) {
|
||||
if (entry.isDirectory()) {
|
||||
const sub = loadState(path.join(workDir, entry.name));
|
||||
if (sub) return sub;
|
||||
}
|
||||
}
|
||||
} catch {
|
||||
// workDir not readable — ignore
|
||||
}
|
||||
return null;
|
||||
}
|
||||
@@ -0,0 +1,195 @@
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
|
||||
import fs from 'fs-extra';
|
||||
import os from 'os';
|
||||
import path from 'path';
|
||||
import { validateStepArtifacts, installSkillsForBot } from './cli.js';
|
||||
|
||||
let tmpDir: string;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'p2c-cli-test-'));
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
fs.removeSync(tmpDir);
|
||||
});
|
||||
|
||||
// ── validateStepArtifacts ───────────────────────────────────────────
|
||||
|
||||
describe('validateStepArtifacts', () => {
|
||||
describe('init', () => {
|
||||
it('valid when AGENTS.md exists', () => {
|
||||
fs.writeFileSync(path.join(tmpDir, 'AGENTS.md'), '# Agents');
|
||||
const result = validateStepArtifacts('init', tmpDir);
|
||||
expect(result.valid).toBe(true);
|
||||
expect(result.missing).toEqual([]);
|
||||
});
|
||||
|
||||
it('missing when AGENTS.md does not exist', () => {
|
||||
const result = validateStepArtifacts('init', tmpDir);
|
||||
expect(result.valid).toBe(false);
|
||||
expect(result.missing).toContain('AGENTS.md');
|
||||
});
|
||||
});
|
||||
|
||||
describe('plan', () => {
|
||||
it('valid when specs/feature/PLAN-DRAFT-*.md exists', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(path.join(specDir, 'PLAN-DRAFT-v1.md'), '# Plan');
|
||||
const result = validateStepArtifacts('plan', tmpDir);
|
||||
expect(result.valid).toBe(true);
|
||||
});
|
||||
|
||||
it('missing when specs/ is absent', () => {
|
||||
const result = validateStepArtifacts('plan', tmpDir);
|
||||
expect(result.valid).toBe(false);
|
||||
expect(result.missing).toContain('specs/ directory');
|
||||
});
|
||||
|
||||
it('missing when specs/ has dirs but no plan draft files', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(path.join(specDir, 'notes.md'), '# Notes');
|
||||
const result = validateStepArtifacts('plan', tmpDir);
|
||||
expect(result.valid).toBe(false);
|
||||
expect(result.missing).toContain('specs/*/PLAN-DRAFT-*.md');
|
||||
});
|
||||
});
|
||||
|
||||
describe('document', () => {
|
||||
it('valid when overview.md and phase-*.md exist', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(path.join(specDir, 'overview.md'), '# Overview');
|
||||
fs.writeFileSync(path.join(specDir, 'phase-1.md'), '# Phase 1');
|
||||
const result = validateStepArtifacts('document', tmpDir);
|
||||
expect(result.valid).toBe(true);
|
||||
});
|
||||
|
||||
it('missing overview.md when absent', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(path.join(specDir, 'phase-1.md'), '# Phase 1');
|
||||
const result = validateStepArtifacts('document', tmpDir);
|
||||
expect(result.valid).toBe(false);
|
||||
expect(result.missing).toContain('specs/*/overview.md');
|
||||
});
|
||||
|
||||
it('missing phase files when absent', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(path.join(specDir, 'overview.md'), '# Overview');
|
||||
const result = validateStepArtifacts('document', tmpDir);
|
||||
expect(result.valid).toBe(false);
|
||||
expect(result.missing).toContain('specs/*/phase-*.md files');
|
||||
});
|
||||
|
||||
it('accepts phase files inside phases/ subdirectory', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
const phasesDir = path.join(specDir, 'phases');
|
||||
fs.ensureDirSync(phasesDir);
|
||||
fs.writeFileSync(path.join(specDir, 'overview.md'), '# Overview');
|
||||
fs.writeFileSync(path.join(phasesDir, 'phase-1.md'), '# Phase 1');
|
||||
const result = validateStepArtifacts('document', tmpDir);
|
||||
expect(result.valid).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('finalize', () => {
|
||||
it('valid when specs--completed/ has a subdirectory', () => {
|
||||
fs.ensureDirSync(path.join(tmpDir, 'specs--completed', 'my-feature'));
|
||||
const result = validateStepArtifacts('finalize', tmpDir);
|
||||
expect(result.valid).toBe(true);
|
||||
});
|
||||
|
||||
it('missing when specs--completed/ is absent', () => {
|
||||
const result = validateStepArtifacts('finalize', tmpDir);
|
||||
expect(result.valid).toBe(false);
|
||||
expect(result.missing).toContain('specs--completed/ directory');
|
||||
});
|
||||
|
||||
it('missing when specs--completed/ is empty', () => {
|
||||
fs.ensureDirSync(path.join(tmpDir, 'specs--completed'));
|
||||
const result = validateStepArtifacts('finalize', tmpDir);
|
||||
expect(result.valid).toBe(false);
|
||||
expect(result.missing).toContain('archived spec in specs--completed/');
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
// ── installSkillsForBot ─────────────────────────────────────────────
|
||||
|
||||
describe('installSkillsForBot', () => {
|
||||
let fakeHome: string;
|
||||
let projectDir: string;
|
||||
let origHome: string | undefined;
|
||||
let origUserProfile: string | undefined;
|
||||
|
||||
beforeEach(() => {
|
||||
fakeHome = fs.mkdtempSync(path.join(os.tmpdir(), 'p2c-home-'));
|
||||
projectDir = fs.mkdtempSync(path.join(os.tmpdir(), 'p2c-proj-'));
|
||||
origHome = process.env.HOME;
|
||||
origUserProfile = process.env.USERPROFILE;
|
||||
process.env.HOME = fakeHome;
|
||||
process.env.USERPROFILE = fakeHome;
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
process.env.HOME = origHome;
|
||||
process.env.USERPROFILE = origUserProfile;
|
||||
fs.removeSync(fakeHome);
|
||||
fs.removeSync(projectDir);
|
||||
});
|
||||
|
||||
it('copies skills from user dir to project .claude/skills/', () => {
|
||||
const srcDir = path.join(fakeHome, '.claude', 'skills', 'plan2code-init');
|
||||
fs.ensureDirSync(srcDir);
|
||||
fs.writeFileSync(path.join(srcDir, 'SKILL.md'), '---\nname: init\n---\nSome content');
|
||||
|
||||
installSkillsForBot(projectDir);
|
||||
|
||||
const dest = path.join(projectDir, '.claude', 'skills', 'plan2code-init', 'SKILL.md');
|
||||
expect(fs.existsSync(dest)).toBe(true);
|
||||
expect(fs.readFileSync(dest, 'utf-8')).toContain('Some content');
|
||||
});
|
||||
|
||||
it('strips disable-model-invocation: true from copied SKILL.md', () => {
|
||||
const srcDir = path.join(fakeHome, '.claude', 'skills', 'plan2code-1-plan');
|
||||
fs.ensureDirSync(srcDir);
|
||||
fs.writeFileSync(
|
||||
path.join(srcDir, 'SKILL.md'),
|
||||
'disable-model-invocation: true\n---\nname: plan\n---\nPlan content',
|
||||
);
|
||||
|
||||
installSkillsForBot(projectDir);
|
||||
|
||||
const dest = path.join(projectDir, '.claude', 'skills', 'plan2code-1-plan', 'SKILL.md');
|
||||
const content = fs.readFileSync(dest, 'utf-8');
|
||||
expect(content).not.toContain('disable-model-invocation');
|
||||
expect(content).toContain('Plan content');
|
||||
});
|
||||
|
||||
it('preserves rest of skill content', () => {
|
||||
const srcDir = path.join(fakeHome, '.claude', 'skills', 'plan2code-2-document');
|
||||
fs.ensureDirSync(srcDir);
|
||||
const original = 'disable-model-invocation: true\n---\nname: doc\n---\nLine 1\nLine 2\nLine 3';
|
||||
fs.writeFileSync(path.join(srcDir, 'SKILL.md'), original);
|
||||
|
||||
installSkillsForBot(projectDir);
|
||||
|
||||
const dest = path.join(projectDir, '.claude', 'skills', 'plan2code-2-document', 'SKILL.md');
|
||||
const content = fs.readFileSync(dest, 'utf-8');
|
||||
expect(content).toContain('Line 1');
|
||||
expect(content).toContain('Line 2');
|
||||
expect(content).toContain('Line 3');
|
||||
expect(content).toContain('name: doc');
|
||||
});
|
||||
|
||||
it('skips gracefully when source skill does not exist', () => {
|
||||
// No skills in fakeHome — should not throw
|
||||
expect(() => installSkillsForBot(projectDir)).not.toThrow();
|
||||
// No .claude/skills/ created in project
|
||||
expect(fs.existsSync(path.join(projectDir, '.claude', 'skills'))).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,587 @@
|
||||
import fs from 'fs-extra';
|
||||
import path from 'path';
|
||||
import chalk from 'chalk';
|
||||
import ora from 'ora';
|
||||
import { generateNewAppIdea, generateEnhancementIdea } from './idea-generator.js';
|
||||
import { runSession } from './session-runner.js';
|
||||
import { buildStepPrompt } from './prompts/step-instructions.js';
|
||||
import { checkAllPhasesComplete } from './step-detector.js';
|
||||
import { saveState, loadState, deleteState, findExistingState } from './bot-state.js';
|
||||
import { ObservationCollector } from './observation-collector.js';
|
||||
import { evaluateStep } from './evaluator.js';
|
||||
import type { BotConfig, BotMode, BotState, StepName, StepResult } from './types.js';
|
||||
|
||||
const BANNER = `
|
||||
╔══════════════════════════════════════╗
|
||||
║ plan2code-bot v1.1.0 ║
|
||||
║ Autonomous Workflow Test Runner ║
|
||||
╚══════════════════════════════════════╝
|
||||
`;
|
||||
|
||||
const MAX_IMPLEMENT_PASSES = 10;
|
||||
|
||||
/** Skill name mapping: bot step name → installed skill directory name */
|
||||
const SKILL_MAP: Record<StepName, string> = {
|
||||
init: 'plan2code-init',
|
||||
plan: 'plan2code-1-plan',
|
||||
document: 'plan2code-2-document',
|
||||
implement: 'plan2code-3-implement',
|
||||
finalize: 'plan2code-4-finalize',
|
||||
};
|
||||
|
||||
/**
|
||||
* Install bot-friendly copies of plan2code skills into the project's .claude/skills/.
|
||||
* The global skills have `disable-model-invocation: true` which prevents the autonomous
|
||||
* session from invoking them via the Skill tool. We copy them with that flag removed
|
||||
* so the session can invoke the real workflow prompts instead of guessing.
|
||||
*/
|
||||
export function installSkillsForBot(projectDir: string): void {
|
||||
const userSkillsDir = path.join(
|
||||
process.env.HOME || process.env.USERPROFILE || '',
|
||||
'.claude',
|
||||
'skills',
|
||||
);
|
||||
const projectSkillsDir = path.join(projectDir, '.claude', 'skills');
|
||||
|
||||
for (const skillName of Object.values(SKILL_MAP)) {
|
||||
const srcFile = path.join(userSkillsDir, skillName, 'SKILL.md');
|
||||
if (!fs.existsSync(srcFile)) continue;
|
||||
|
||||
const content = fs.readFileSync(srcFile, 'utf-8');
|
||||
// Remove the disable-model-invocation line so the bot session can invoke the skill
|
||||
const patched = content.replace(/^disable-model-invocation:\s*true\n?/m, '');
|
||||
|
||||
const destDir = path.join(projectSkillsDir, skillName);
|
||||
fs.ensureDirSync(destDir);
|
||||
fs.writeFileSync(path.join(destDir, 'SKILL.md'), patched);
|
||||
}
|
||||
}
|
||||
|
||||
export function validateStepArtifacts(step: StepName, projectDir: string): { valid: boolean; missing: string[] } {
|
||||
const missing: string[] = [];
|
||||
|
||||
switch (step) {
|
||||
case 'init': {
|
||||
const agentsPath = path.join(projectDir, 'AGENTS.md');
|
||||
if (!fs.existsSync(agentsPath)) missing.push('AGENTS.md');
|
||||
break;
|
||||
}
|
||||
case 'plan': {
|
||||
const specsDir = path.join(projectDir, 'specs');
|
||||
if (!fs.existsSync(specsDir)) {
|
||||
missing.push('specs/ directory');
|
||||
} else {
|
||||
const entries = fs.readdirSync(specsDir, { withFileTypes: true });
|
||||
const specDirs = entries.filter((e) => e.isDirectory());
|
||||
const hasPlanDraft = specDirs.some((d) => {
|
||||
const files = fs.readdirSync(path.join(specsDir, d.name));
|
||||
return files.some((f) => /plan[-_]?draft/i.test(f));
|
||||
});
|
||||
if (!hasPlanDraft) missing.push('specs/*/PLAN-DRAFT-*.md');
|
||||
}
|
||||
break;
|
||||
}
|
||||
case 'document': {
|
||||
const specsDir = path.join(projectDir, 'specs');
|
||||
if (!fs.existsSync(specsDir)) {
|
||||
missing.push('specs/ directory');
|
||||
} else {
|
||||
const entries = fs.readdirSync(specsDir, { withFileTypes: true });
|
||||
const specDirs = entries.filter((e) => e.isDirectory());
|
||||
let foundOverview = false;
|
||||
let foundPhaseFile = false;
|
||||
for (const d of specDirs) {
|
||||
const specPath = path.join(specsDir, d.name);
|
||||
const files = fs.readdirSync(specPath);
|
||||
if (files.some((f) => f === 'overview.md')) foundOverview = true;
|
||||
if (files.some((f) => /^phase[-_]?\d+.*\.md$/i.test(f))) foundPhaseFile = true;
|
||||
// Also check phases/ subdirectory
|
||||
const phasesSubdir = path.join(specPath, 'phases');
|
||||
if (fs.existsSync(phasesSubdir)) {
|
||||
const subFiles = fs.readdirSync(phasesSubdir);
|
||||
if (subFiles.some((f) => /phase/i.test(f))) foundPhaseFile = true;
|
||||
}
|
||||
}
|
||||
if (!foundOverview) missing.push('specs/*/overview.md');
|
||||
if (!foundPhaseFile) missing.push('specs/*/phase-*.md files');
|
||||
}
|
||||
break;
|
||||
}
|
||||
case 'finalize': {
|
||||
const completedDir = path.join(projectDir, 'specs--completed');
|
||||
if (!fs.existsSync(completedDir)) {
|
||||
missing.push('specs--completed/ directory');
|
||||
} else {
|
||||
const entries = fs.readdirSync(completedDir, { withFileTypes: true });
|
||||
const specDirs = entries.filter((e) => e.isDirectory());
|
||||
if (specDirs.length === 0) missing.push('archived spec in specs--completed/');
|
||||
}
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
return { valid: missing.length === 0, missing };
|
||||
}
|
||||
|
||||
function detectMode(workDir: string): BotMode {
|
||||
const agentsPath = path.join(workDir, 'AGENTS.md');
|
||||
return fs.existsSync(agentsPath) ? 'enhancement' : 'new-project';
|
||||
}
|
||||
|
||||
function formatDuration(ms: number): string {
|
||||
const seconds = Math.floor(ms / 1000);
|
||||
const minutes = Math.floor(seconds / 60);
|
||||
const remainingSeconds = seconds % 60;
|
||||
if (minutes > 0) {
|
||||
return `${minutes}m ${remainingSeconds}s`;
|
||||
}
|
||||
return `${seconds}s`;
|
||||
}
|
||||
|
||||
async function runStep(
|
||||
step: StepName,
|
||||
config: BotConfig,
|
||||
state: BotState,
|
||||
): Promise<StepResult> {
|
||||
const spinner = ora({
|
||||
text: chalk.cyan(`Running ${step} step...`),
|
||||
spinner: 'dots',
|
||||
}).start();
|
||||
|
||||
const prompt = buildStepPrompt(step, config);
|
||||
|
||||
try {
|
||||
// Create observation collector
|
||||
const collector = new ObservationCollector(step);
|
||||
|
||||
const result = await runSession({
|
||||
prompt,
|
||||
config,
|
||||
step,
|
||||
maxTurns: step === 'implement' ? 80 : 50,
|
||||
collector,
|
||||
});
|
||||
|
||||
if (result.success) {
|
||||
spinner.succeed(
|
||||
chalk.green(`${step} completed in ${formatDuration(result.duration)}`)
|
||||
);
|
||||
|
||||
// Run evaluation
|
||||
spinner.text = chalk.cyan('Evaluating step quality...');
|
||||
spinner.start();
|
||||
|
||||
const evaluation = await evaluateStep(step, result.observations, config.projectDir);
|
||||
|
||||
spinner.succeed(
|
||||
chalk.cyan(`Evaluation complete: ${formatScore(evaluation.score)}`)
|
||||
);
|
||||
|
||||
// Display evaluation summary
|
||||
console.log(chalk.dim(` Score: ${formatScore(evaluation.score)}`));
|
||||
if (evaluation.strengths.length > 0) {
|
||||
console.log(chalk.green(` ✓ ${evaluation.strengths[0]}`));
|
||||
}
|
||||
if (evaluation.weaknesses.length > 0) {
|
||||
console.log(chalk.yellow(` ⚠ ${evaluation.weaknesses[0]}`));
|
||||
}
|
||||
|
||||
const stepResult: StepResult = {
|
||||
step,
|
||||
success: result.success,
|
||||
sessionId: result.sessionId,
|
||||
duration: result.duration,
|
||||
error: null,
|
||||
evaluation,
|
||||
observations: result.observations,
|
||||
};
|
||||
|
||||
return stepResult;
|
||||
} else {
|
||||
spinner.fail(chalk.red(`${step} failed after ${formatDuration(result.duration)}`));
|
||||
|
||||
return {
|
||||
step,
|
||||
success: false,
|
||||
sessionId: result.sessionId,
|
||||
duration: result.duration,
|
||||
error: 'Session failed',
|
||||
observations: result.observations,
|
||||
};
|
||||
}
|
||||
} catch (error) {
|
||||
const errorMsg = error instanceof Error ? error.message : String(error);
|
||||
spinner.fail(chalk.red(`${step} error: ${errorMsg}`));
|
||||
return {
|
||||
step,
|
||||
success: false,
|
||||
sessionId: null,
|
||||
duration: 0,
|
||||
error: errorMsg,
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
function formatScore(score: number): string {
|
||||
if (score >= 85) return chalk.green(`${score}/100`);
|
||||
if (score >= 70) return chalk.yellow(`${score}/100`);
|
||||
return chalk.red(`${score}/100`);
|
||||
}
|
||||
|
||||
export interface CLIOptions {
|
||||
idea?: string;
|
||||
resume?: boolean;
|
||||
}
|
||||
|
||||
/** Check if a step completed successfully in a saved state */
|
||||
function stepSucceeded(state: BotState, step: StepName): boolean {
|
||||
return state.steps.some((s) => s.step === step && s.success);
|
||||
}
|
||||
|
||||
export async function runCLI(options: CLIOptions = {}): Promise<void> {
|
||||
console.log(chalk.cyan(BANNER));
|
||||
|
||||
const workDir = process.cwd();
|
||||
|
||||
// Check for resumable state
|
||||
let resuming = false;
|
||||
let savedState: BotState | null = null;
|
||||
|
||||
if (options.resume) {
|
||||
savedState = findExistingState(workDir);
|
||||
if (savedState) {
|
||||
resuming = true;
|
||||
console.log(chalk.yellow('Resuming previous incomplete run...'));
|
||||
} else {
|
||||
console.log(chalk.dim('No previous state found, starting fresh.'));
|
||||
}
|
||||
}
|
||||
|
||||
const mode = resuming ? savedState!.config.mode : detectMode(workDir);
|
||||
|
||||
console.log(chalk.dim(`Working directory: ${workDir}`));
|
||||
console.log(chalk.dim(`Mode: ${mode === 'new-project' ? 'New Project' : 'Enhancement'}`));
|
||||
if (options.idea) {
|
||||
console.log(chalk.dim(`Idea seed: ${options.idea}`));
|
||||
}
|
||||
console.log('');
|
||||
|
||||
let ideaName: string;
|
||||
let ideaDescription: string;
|
||||
let projectDir: string;
|
||||
|
||||
if (resuming) {
|
||||
// Restore from saved state
|
||||
ideaName = savedState!.config.ideaName;
|
||||
ideaDescription = savedState!.config.ideaDescription;
|
||||
projectDir = savedState!.config.projectDir;
|
||||
console.log(chalk.green(`Restored idea: ${ideaName}`));
|
||||
console.log(chalk.dim(` ${ideaDescription}`));
|
||||
console.log('');
|
||||
} else {
|
||||
// Step 1: Generate idea
|
||||
const ideaSpinner = ora({
|
||||
text: chalk.cyan('Generating idea...'),
|
||||
spinner: 'dots',
|
||||
}).start();
|
||||
|
||||
try {
|
||||
if (mode === 'new-project') {
|
||||
const idea = await generateNewAppIdea(options.idea);
|
||||
ideaName = idea.name;
|
||||
ideaDescription = idea.description;
|
||||
} else {
|
||||
const idea = await generateEnhancementIdea(workDir, options.idea);
|
||||
ideaName = idea.name;
|
||||
ideaDescription = idea.description;
|
||||
}
|
||||
ideaSpinner.succeed(chalk.green(`Idea generated: ${ideaName}`));
|
||||
} catch (error) {
|
||||
ideaSpinner.fail(chalk.red('Failed to generate idea'));
|
||||
throw error;
|
||||
}
|
||||
|
||||
console.log(chalk.dim(` ${ideaDescription}`));
|
||||
console.log('');
|
||||
|
||||
// Determine project directory
|
||||
projectDir = mode === 'new-project'
|
||||
? path.join(workDir, ideaName)
|
||||
: workDir;
|
||||
|
||||
// Create project directory for new projects
|
||||
if (mode === 'new-project') {
|
||||
fs.ensureDirSync(projectDir);
|
||||
}
|
||||
}
|
||||
|
||||
// Install bot-friendly skills (without disable-model-invocation)
|
||||
installSkillsForBot(projectDir);
|
||||
console.log(chalk.dim('Installed plan2code skills for bot sessions'));
|
||||
|
||||
// Write IDEA.md (only if not resuming past plan step, since it gets moved)
|
||||
if (!resuming || !stepSucceeded(savedState!, 'plan')) {
|
||||
const ideaContent = `# ${ideaName}\n\n${ideaDescription}\n`;
|
||||
fs.writeFileSync(path.join(projectDir, 'IDEA.md'), ideaContent);
|
||||
console.log(chalk.dim(`Wrote IDEA.md to ${projectDir}`));
|
||||
}
|
||||
console.log('');
|
||||
|
||||
// Build config
|
||||
const config: BotConfig = {
|
||||
workDir,
|
||||
projectDir,
|
||||
ideaDescription,
|
||||
ideaName,
|
||||
mode,
|
||||
};
|
||||
|
||||
// Initialize or restore state
|
||||
const state: BotState = resuming
|
||||
? { ...savedState!, config }
|
||||
: {
|
||||
config,
|
||||
steps: [],
|
||||
currentStep: null,
|
||||
implementPasses: 0,
|
||||
allPhasesComplete: false,
|
||||
};
|
||||
|
||||
// Step 2: Run init (new project only)
|
||||
if (mode === 'new-project') {
|
||||
if (resuming && stepSucceeded(savedState!, 'init')) {
|
||||
console.log(chalk.dim('--- Init (skipped — previously succeeded) ---'));
|
||||
} else {
|
||||
console.log(chalk.bold('--- Init ---'));
|
||||
state.currentStep = 'init';
|
||||
const initResult = await runStep('init', config, state);
|
||||
state.steps.push(initResult);
|
||||
saveState(state);
|
||||
|
||||
if (!initResult.success) {
|
||||
console.log(chalk.red('\nInit failed. Aborting.'));
|
||||
printSummary(state, workDir, false);
|
||||
return;
|
||||
}
|
||||
const initValidation = validateStepArtifacts('init', projectDir);
|
||||
if (!initValidation.valid) {
|
||||
console.log(chalk.yellow(` ⚠ Missing artifacts: ${initValidation.missing.join(', ')}`));
|
||||
initResult.success = false;
|
||||
initResult.error = `Missing artifacts: ${initValidation.missing.join(', ')}`;
|
||||
console.log(chalk.red('\nInit artifacts missing. Aborting.'));
|
||||
printSummary(state, workDir, false);
|
||||
return;
|
||||
}
|
||||
}
|
||||
console.log('');
|
||||
}
|
||||
|
||||
// Step 3: Plan
|
||||
if (resuming && stepSucceeded(savedState!, 'plan')) {
|
||||
console.log(chalk.dim('--- Plan (skipped — previously succeeded) ---'));
|
||||
} else {
|
||||
console.log(chalk.bold('--- Plan ---'));
|
||||
state.currentStep = 'plan';
|
||||
const planResult = await runStep('plan', config, state);
|
||||
state.steps.push(planResult);
|
||||
saveState(state);
|
||||
|
||||
if (!planResult.success) {
|
||||
console.log(chalk.red('\nPlan step failed. Aborting.'));
|
||||
printSummary(state, workDir, false);
|
||||
return;
|
||||
}
|
||||
const planValidation = validateStepArtifacts('plan', projectDir);
|
||||
if (!planValidation.valid) {
|
||||
console.log(chalk.yellow(` ⚠ Missing artifacts: ${planValidation.missing.join(', ')}`));
|
||||
planResult.success = false;
|
||||
planResult.error = `Missing artifacts: ${planValidation.missing.join(', ')}`;
|
||||
console.log(chalk.red('\nPlan artifacts missing. Aborting.'));
|
||||
printSummary(state, workDir, false);
|
||||
return;
|
||||
}
|
||||
|
||||
// Move IDEA.md into the spec directory so it stays with its feature
|
||||
const ideaPath = path.join(projectDir, 'IDEA.md');
|
||||
if (fs.existsSync(ideaPath)) {
|
||||
const specEntries = fs.readdirSync(path.join(projectDir, 'specs'), { withFileTypes: true });
|
||||
const firstSpecDir = specEntries.find((e) => e.isDirectory());
|
||||
if (firstSpecDir) {
|
||||
const dest = path.join(projectDir, 'specs', firstSpecDir.name, 'IDEA.md');
|
||||
fs.moveSync(ideaPath, dest, { overwrite: true });
|
||||
console.log(chalk.dim(`Moved IDEA.md → specs/${firstSpecDir.name}/IDEA.md`));
|
||||
}
|
||||
}
|
||||
}
|
||||
console.log('');
|
||||
|
||||
// Step 4: Document
|
||||
if (resuming && stepSucceeded(savedState!, 'document')) {
|
||||
console.log(chalk.dim('--- Document (skipped — previously succeeded) ---'));
|
||||
} else {
|
||||
console.log(chalk.bold('--- Document ---'));
|
||||
state.currentStep = 'document';
|
||||
const docResult = await runStep('document', config, state);
|
||||
state.steps.push(docResult);
|
||||
saveState(state);
|
||||
|
||||
if (!docResult.success) {
|
||||
console.log(chalk.red('\nDocument step failed. Aborting.'));
|
||||
printSummary(state, workDir, false);
|
||||
return;
|
||||
}
|
||||
const docValidation = validateStepArtifacts('document', projectDir);
|
||||
if (!docValidation.valid) {
|
||||
console.log(chalk.yellow(` ⚠ Missing artifacts: ${docValidation.missing.join(', ')}`));
|
||||
docResult.success = false;
|
||||
docResult.error = `Missing artifacts: ${docValidation.missing.join(', ')}`;
|
||||
console.log(chalk.red('\nDocument artifacts missing. Aborting.'));
|
||||
printSummary(state, workDir, false);
|
||||
return;
|
||||
}
|
||||
}
|
||||
console.log('');
|
||||
|
||||
// Step 5: Implement (loop until all phases complete)
|
||||
if (resuming && savedState!.allPhasesComplete) {
|
||||
console.log(chalk.dim('--- Implement (skipped — all phases already complete) ---'));
|
||||
} else {
|
||||
console.log(chalk.bold('--- Implement ---'));
|
||||
while (state.implementPasses < MAX_IMPLEMENT_PASSES) {
|
||||
state.implementPasses++;
|
||||
state.currentStep = 'implement';
|
||||
|
||||
console.log(chalk.dim(` Pass ${state.implementPasses}/${MAX_IMPLEMENT_PASSES}`));
|
||||
const implResult = await runStep('implement', config, state);
|
||||
state.steps.push(implResult);
|
||||
saveState(state);
|
||||
|
||||
if (!implResult.success) {
|
||||
console.log(chalk.yellow(`\nImplement pass ${state.implementPasses} failed. Continuing...`));
|
||||
}
|
||||
|
||||
// Check if all phases are complete
|
||||
if (checkAllPhasesComplete(projectDir)) {
|
||||
state.allPhasesComplete = true;
|
||||
saveState(state);
|
||||
console.log(chalk.green(' All phases complete!'));
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
if (!state.allPhasesComplete && state.implementPasses >= MAX_IMPLEMENT_PASSES) {
|
||||
console.log(chalk.yellow(`\nMax implement passes (${MAX_IMPLEMENT_PASSES}) reached.`));
|
||||
}
|
||||
}
|
||||
console.log('');
|
||||
|
||||
// Step 6: Finalize
|
||||
if (resuming && stepSucceeded(savedState!, 'finalize')) {
|
||||
console.log(chalk.dim('--- Finalize (skipped — previously succeeded) ---'));
|
||||
} else {
|
||||
console.log(chalk.bold('--- Finalize ---'));
|
||||
|
||||
// Check quality gate: average score must be >= 60
|
||||
const evaluatedSteps = state.steps.filter((s) => s.evaluation);
|
||||
if (evaluatedSteps.length > 0) {
|
||||
const avgScore =
|
||||
evaluatedSteps.reduce((sum, s) => sum + (s.evaluation?.score ?? 0), 0) /
|
||||
evaluatedSteps.length;
|
||||
|
||||
console.log(chalk.dim(` Average quality score: ${formatScore(Math.round(avgScore))}`));
|
||||
|
||||
if (avgScore < 60) {
|
||||
console.log(
|
||||
chalk.red(
|
||||
'\n⚠ Quality gate failed: Average score is below 60. Please review and fix issues before finalizing.'
|
||||
)
|
||||
);
|
||||
console.log(chalk.dim(' Check specs/<feature>/BOT-EVALUATION.md for detailed feedback.'));
|
||||
printSummary(state, workDir, false);
|
||||
return;
|
||||
}
|
||||
}
|
||||
|
||||
state.currentStep = 'finalize';
|
||||
const finalizeResult = await runStep('finalize', config, state);
|
||||
state.steps.push(finalizeResult);
|
||||
saveState(state);
|
||||
|
||||
if (finalizeResult.success) {
|
||||
const finalValidation = validateStepArtifacts('finalize', projectDir);
|
||||
if (!finalValidation.valid) {
|
||||
console.log(chalk.yellow(` ⚠ Missing artifacts: ${finalValidation.missing.join(', ')}`));
|
||||
finalizeResult.success = false;
|
||||
finalizeResult.error = `Missing artifacts: ${finalValidation.missing.join(', ')}`;
|
||||
saveState(state);
|
||||
}
|
||||
}
|
||||
}
|
||||
console.log('');
|
||||
|
||||
// Determine overall success
|
||||
const allSucceeded = state.steps.length > 0 && state.steps.every((s) => s.success);
|
||||
printSummary(state, workDir, allSucceeded);
|
||||
|
||||
// Clean up state file on full success
|
||||
if (allSucceeded) {
|
||||
deleteState(projectDir);
|
||||
console.log(chalk.dim('Cleaned up state file (run succeeded)'));
|
||||
} else {
|
||||
console.log(chalk.dim('State file preserved for resume (run incomplete)'));
|
||||
}
|
||||
console.log('');
|
||||
}
|
||||
|
||||
function printSummary(state: BotState, workDir: string, allSucceeded: boolean): void {
|
||||
const { ideaName, mode, projectDir } = state.config;
|
||||
|
||||
console.log(chalk.cyan('═══════════════════════════════════════'));
|
||||
console.log(chalk.bold(' Bot Run Summary'));
|
||||
console.log(chalk.cyan('═══════════════════════════════════════'));
|
||||
console.log(chalk.dim(` Project: ${ideaName}`));
|
||||
console.log(chalk.dim(` Mode: ${mode}`));
|
||||
console.log(chalk.dim(` Directory: ${projectDir}`));
|
||||
console.log('');
|
||||
|
||||
const totalDuration = state.steps.reduce((sum, s) => sum + s.duration, 0);
|
||||
const successCount = state.steps.filter((s) => s.success).length;
|
||||
|
||||
for (const step of state.steps) {
|
||||
const icon = step.success ? chalk.green('✓') : chalk.red('✗');
|
||||
const scoreText = step.evaluation
|
||||
? ` [${formatScore(step.evaluation.score)}]`
|
||||
: '';
|
||||
console.log(
|
||||
` ${icon} ${step.step.padEnd(12)} ${formatDuration(step.duration)}${scoreText}`
|
||||
);
|
||||
}
|
||||
|
||||
console.log('');
|
||||
|
||||
// Display average quality score if available
|
||||
const evaluatedSteps = state.steps.filter((s) => s.evaluation);
|
||||
if (evaluatedSteps.length > 0) {
|
||||
const avgScore =
|
||||
evaluatedSteps.reduce((sum, s) => sum + (s.evaluation?.score ?? 0), 0) /
|
||||
evaluatedSteps.length;
|
||||
console.log(
|
||||
chalk.dim(` Average quality: ${formatScore(Math.round(avgScore))}`)
|
||||
);
|
||||
}
|
||||
|
||||
console.log(
|
||||
chalk.dim(
|
||||
` Total: ${successCount}/${state.steps.length} steps succeeded in ${formatDuration(totalDuration)}`
|
||||
)
|
||||
);
|
||||
console.log(chalk.dim(` Implement passes: ${state.implementPasses}`));
|
||||
if (!allSucceeded) {
|
||||
console.log(
|
||||
chalk.dim(
|
||||
` State saved to: ${path.relative(workDir, projectDir)}/.plan2code-bot-state.json`
|
||||
)
|
||||
);
|
||||
}
|
||||
console.log('');
|
||||
}
|
||||
@@ -0,0 +1,311 @@
|
||||
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||
import fs from 'fs-extra';
|
||||
import path from 'path';
|
||||
import type { EvaluationResult, ExecutionObservation, StepName } from './types.js';
|
||||
import { getCriteriaForStep } from './prompts/evaluation-criteria.js';
|
||||
|
||||
/**
|
||||
* Evaluates a completed step using LLM-as-judge.
|
||||
* Returns honest quality assessment based on observations and artifacts.
|
||||
*/
|
||||
export async function evaluateStep(
|
||||
step: StepName,
|
||||
observations: ExecutionObservation,
|
||||
projectDir: string
|
||||
): Promise<EvaluationResult> {
|
||||
const criteria = getCriteriaForStep(step);
|
||||
const prompt = buildEvaluationPrompt(step, observations, criteria, projectDir);
|
||||
|
||||
try {
|
||||
// Query LLM with ability to inspect artifacts
|
||||
const session = query({
|
||||
prompt,
|
||||
options: {
|
||||
maxTurns: 30,
|
||||
cwd: projectDir,
|
||||
permissionMode: 'bypassPermissions',
|
||||
allowDangerouslySkipPermissions: true,
|
||||
allowedTools: ['Read', 'Glob', 'Grep'],
|
||||
systemPrompt:
|
||||
'You are a QA engineer evaluating completed work. Be thorough, honest, and constructive.',
|
||||
},
|
||||
});
|
||||
|
||||
let output = '';
|
||||
for await (const message of session) {
|
||||
if (message.type === 'assistant') {
|
||||
for (const block of message.message.content) {
|
||||
if (block.type === 'text') output += block.text;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const evaluation = parseEvaluationOutput(output, step);
|
||||
|
||||
// Write evaluation files
|
||||
await writeEvaluationFiles(evaluation, observations, projectDir);
|
||||
|
||||
return evaluation;
|
||||
} catch (error) {
|
||||
console.warn('Evaluation failed, using fallback:', error);
|
||||
return createFallbackEvaluation(step, observations);
|
||||
}
|
||||
}
|
||||
|
||||
function buildEvaluationPrompt(
|
||||
step: StepName,
|
||||
observations: ExecutionObservation,
|
||||
criteria: ReturnType<typeof getCriteriaForStep>,
|
||||
projectDir: string
|
||||
): string {
|
||||
const duration = observations.endTime - observations.startTime;
|
||||
const durationSec = Math.floor(duration / 1000);
|
||||
|
||||
const toolsSummary = observations.tools
|
||||
.map((t) => `- ${t.toolName} (${new Date(t.timestamp).toISOString()})`)
|
||||
.join('\n');
|
||||
|
||||
const filesSummary = [
|
||||
...observations.filesCreated.map((f) => `CREATED: ${f}`),
|
||||
...observations.filesModified.map((f) => `MODIFIED: ${f}`),
|
||||
].join('\n');
|
||||
|
||||
const questionsSummary = observations.questionsAsked
|
||||
.map(
|
||||
(q) =>
|
||||
`Q: ${q.question}\nA: ${q.selectedAnswer}\nReasoning: ${q.llmReasoning}`
|
||||
)
|
||||
.join('\n\n');
|
||||
|
||||
return `You are a QA engineer evaluating the quality of a completed workflow step.
|
||||
|
||||
## Step Information
|
||||
- Step: ${step}
|
||||
- Duration: ${durationSec}s
|
||||
- Tools used: ${observations.tools.length}
|
||||
- Files created: ${observations.filesCreated.length}
|
||||
- Files modified: ${observations.filesModified.length}
|
||||
- Errors: ${observations.errors.length}
|
||||
|
||||
## Execution Details
|
||||
|
||||
### Tools Used
|
||||
${toolsSummary || '(none)'}
|
||||
|
||||
### Files Changed
|
||||
${filesSummary || '(none)'}
|
||||
|
||||
${observations.questionsAsked.length > 0 ? `### Questions & Decisions\n${questionsSummary}` : ''}
|
||||
|
||||
${observations.errors.length > 0 ? `### Errors Encountered\n${observations.errors.join('\n')}` : ''}
|
||||
|
||||
## Evaluation Criteria
|
||||
|
||||
**Key Artifacts Expected:**
|
||||
${criteria.keyArtifacts.map((a) => `- ${a}`).join('\n')}
|
||||
|
||||
**Quality Checks:**
|
||||
${criteria.qualityChecks.map((c) => `- ${c}`).join('\n')}
|
||||
|
||||
**Common Pitfalls to Watch For:**
|
||||
${criteria.commonPitfalls.map((p) => `- ${p}`).join('\n')}
|
||||
|
||||
**Scoring Guidance:**
|
||||
${criteria.scoringGuidance}
|
||||
|
||||
## Your Task
|
||||
|
||||
Evaluate this step honestly and thoroughly:
|
||||
|
||||
1. **Inspect the artifacts** using Read, Glob, and Grep tools
|
||||
2. **Check against quality criteria** listed above
|
||||
3. **Identify strengths and weaknesses** based on actual evidence
|
||||
4. **Provide constructive suggestions** for improvement
|
||||
5. **Assign an honest score** (0-100) following the guidance
|
||||
|
||||
Be critical but fair. Most work scores 70-85. Don't inflate scores.
|
||||
|
||||
## Response Format
|
||||
|
||||
SCORE: <number 0-100>
|
||||
|
||||
STRENGTHS:
|
||||
- <strength 1>
|
||||
- <strength 2>
|
||||
- <strength 3>
|
||||
|
||||
WEAKNESSES:
|
||||
- <weakness 1>
|
||||
- <weakness 2>
|
||||
|
||||
SUGGESTIONS:
|
||||
- <suggestion 1>
|
||||
- <suggestion 2>
|
||||
|
||||
CRITICAL_ISSUES:
|
||||
- <critical issue 1 (or "None")>
|
||||
|
||||
REASONING:
|
||||
<1-2 paragraphs explaining your evaluation, referencing specific files/evidence>
|
||||
|
||||
Provide honest, evidence-based evaluation.`;
|
||||
}
|
||||
|
||||
function parseEvaluationOutput(
|
||||
output: string,
|
||||
step: StepName
|
||||
): EvaluationResult {
|
||||
// Extract score
|
||||
const scoreMatch = output.match(/SCORE:\s*(\d+)/i);
|
||||
if (!scoreMatch) {
|
||||
console.warn(`Evaluation parser: no SCORE found in output (${output.length} chars). Defaulting to 50.`);
|
||||
}
|
||||
const score = scoreMatch ? parseInt(scoreMatch[1], 10) : 50;
|
||||
|
||||
// Extract sections
|
||||
const strengthsMatch = output.match(
|
||||
/STRENGTHS:\s*((?:- .+\n?)+)/i
|
||||
);
|
||||
const weaknessesMatch = output.match(
|
||||
/WEAKNESSES:\s*((?:- .+\n?)+)/i
|
||||
);
|
||||
const suggestionsMatch = output.match(
|
||||
/SUGGESTIONS:\s*((?:- .+\n?)+)/i
|
||||
);
|
||||
const criticalMatch = output.match(
|
||||
/CRITICAL_ISSUES:\s*((?:- .+\n?)+)/i
|
||||
);
|
||||
const reasoningMatch = output.match(/REASONING:\s*(.+?)(?=\n\n|$)/is);
|
||||
|
||||
const parseList = (text: string | undefined): string[] => {
|
||||
if (!text) return [];
|
||||
return text
|
||||
.split('\n')
|
||||
.map((line) => line.replace(/^-\s*/, '').trim())
|
||||
.filter((line) => line.length > 0 && !line.toLowerCase().includes('none'));
|
||||
};
|
||||
|
||||
return {
|
||||
step,
|
||||
score: Math.max(0, Math.min(100, score)),
|
||||
strengths: parseList(strengthsMatch?.[1]),
|
||||
weaknesses: parseList(weaknessesMatch?.[1]),
|
||||
suggestions: parseList(suggestionsMatch?.[1]),
|
||||
criticalIssues: parseList(criticalMatch?.[1]),
|
||||
timestamp: Date.now(),
|
||||
evaluatorModel: 'claude-sonnet-4-5',
|
||||
reasoning: reasoningMatch?.[1]?.trim() || 'No reasoning provided',
|
||||
};
|
||||
}
|
||||
|
||||
function createFallbackEvaluation(
|
||||
step: StepName,
|
||||
observations: ExecutionObservation
|
||||
): EvaluationResult {
|
||||
return {
|
||||
step,
|
||||
score: 50,
|
||||
strengths: ['Step completed'],
|
||||
weaknesses: ['Evaluation failed - using fallback'],
|
||||
suggestions: ['Re-run with evaluation enabled'],
|
||||
criticalIssues: ['Evaluation system unavailable'],
|
||||
timestamp: Date.now(),
|
||||
evaluatorModel: 'fallback',
|
||||
reasoning: 'Evaluation failed, using fallback. Cannot provide detailed assessment.',
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves the active spec subdirectory (e.g. specs/<feature>/).
|
||||
* Falls back to projectDir if no spec folder exists yet (e.g. during init).
|
||||
*/
|
||||
function resolveOutputDir(projectDir: string): string {
|
||||
const specsDir = path.join(projectDir, 'specs');
|
||||
if (fs.existsSync(specsDir)) {
|
||||
const entries = fs.readdirSync(specsDir, { withFileTypes: true });
|
||||
const firstSpec = entries.find((e) => e.isDirectory());
|
||||
if (firstSpec) {
|
||||
return path.join(specsDir, firstSpec.name);
|
||||
}
|
||||
}
|
||||
return projectDir;
|
||||
}
|
||||
|
||||
async function writeEvaluationFiles(
|
||||
evaluation: EvaluationResult,
|
||||
observations: ExecutionObservation,
|
||||
projectDir: string
|
||||
): Promise<void> {
|
||||
const outputDir = resolveOutputDir(projectDir);
|
||||
|
||||
// Write BOT-EVALUATION.md
|
||||
const evalContent = formatEvaluationMarkdown(evaluation);
|
||||
const evalPath = path.join(outputDir, 'BOT-EVALUATION.md');
|
||||
|
||||
if (fs.existsSync(evalPath)) {
|
||||
// Append to existing file
|
||||
const existing = fs.readFileSync(evalPath, 'utf-8');
|
||||
fs.writeFileSync(evalPath, existing + '\n\n---\n\n' + evalContent);
|
||||
} else {
|
||||
fs.writeFileSync(evalPath, evalContent);
|
||||
}
|
||||
|
||||
// Write BOT-NOTES.md
|
||||
const notesContent = formatObservationsMarkdown(observations);
|
||||
const notesPath = path.join(outputDir, 'BOT-NOTES.md');
|
||||
|
||||
if (fs.existsSync(notesPath)) {
|
||||
const existing = fs.readFileSync(notesPath, 'utf-8');
|
||||
fs.writeFileSync(notesPath, existing + '\n\n---\n\n' + notesContent);
|
||||
} else {
|
||||
fs.writeFileSync(notesPath, notesContent);
|
||||
}
|
||||
}
|
||||
|
||||
function formatEvaluationMarkdown(evaluation: EvaluationResult): string {
|
||||
return `# Evaluation: ${evaluation.step} Step
|
||||
|
||||
**Score:** ${evaluation.score}/100
|
||||
**Timestamp:** ${new Date(evaluation.timestamp).toISOString()}
|
||||
**Evaluator:** ${evaluation.evaluatorModel}
|
||||
|
||||
## Strengths
|
||||
${evaluation.strengths.map((s) => `- ${s}`).join('\n') || '(none)'}
|
||||
|
||||
## Weaknesses
|
||||
${evaluation.weaknesses.map((w) => `- ${w}`).join('\n') || '(none)'}
|
||||
|
||||
## Suggestions for Improvement
|
||||
${evaluation.suggestions.map((s) => `- ${s}`).join('\n') || '(none)'}
|
||||
|
||||
${evaluation.criticalIssues.length > 0 ? `## Critical Issues\n${evaluation.criticalIssues.map((i) => `- ${i}`).join('\n')}\n` : ''}
|
||||
|
||||
## Reasoning
|
||||
${evaluation.reasoning}`;
|
||||
}
|
||||
|
||||
function formatObservationsMarkdown(observations: ExecutionObservation): string {
|
||||
const duration = observations.endTime - observations.startTime;
|
||||
const durationSec = Math.floor(duration / 1000);
|
||||
|
||||
return `# Observations: ${observations.step} Step
|
||||
|
||||
**Duration:** ${durationSec}s
|
||||
**Tools Used:** ${observations.tools.length}
|
||||
**Files Created:** ${observations.filesCreated.length}
|
||||
**Files Modified:** ${observations.filesModified.length}
|
||||
**Errors:** ${observations.errors.length}
|
||||
|
||||
${observations.questionsAsked.length > 0 ? `## Questions Asked & Answers\n\n${observations.questionsAsked.map((q) => `### Question: "${q.question}"\n**Selected:** "${q.selectedAnswer}"\n**Reasoning:** ${q.llmReasoning}`).join('\n\n')}\n` : ''}
|
||||
|
||||
## Tool Usage
|
||||
${observations.tools.map((t) => `- ${t.toolName}`).join('\n')}
|
||||
|
||||
## Files Created
|
||||
${observations.filesCreated.map((f) => `- ${f}`).join('\n') || '(none)'}
|
||||
|
||||
## Files Modified
|
||||
${observations.filesModified.map((f) => `- ${f}`).join('\n') || '(none)'}
|
||||
|
||||
${observations.assistantMessages.length > 0 ? `## Assistant Output Summary\n${observations.assistantMessages.slice(0, 3).map((m) => `> ${m.substring(0, 100)}...`).join('\n')}\n` : ''}`;
|
||||
}
|
||||
@@ -0,0 +1,114 @@
|
||||
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||
|
||||
interface IdeaResult {
|
||||
name: string;
|
||||
description: string;
|
||||
}
|
||||
|
||||
function parseIdea(text: string): IdeaResult {
|
||||
const nameMatch = text.match(/NAME:\s*(.+)/i);
|
||||
const descMatch = text.match(/DESCRIPTION:\s*([\s\S]+?)(?:\n\n|$)/i);
|
||||
|
||||
const name = nameMatch?.[1]?.trim() ?? 'auto-project';
|
||||
const description = descMatch?.[1]?.trim() ?? text.trim();
|
||||
|
||||
// Ensure kebab-case
|
||||
const kebabName = name
|
||||
.toLowerCase()
|
||||
.replace(/[^a-z0-9]+/g, '-')
|
||||
.replace(/^-|-$/g, '');
|
||||
|
||||
return { name: kebabName, description };
|
||||
}
|
||||
|
||||
export async function generateNewAppIdea(seed?: string): Promise<IdeaResult> {
|
||||
const categories = [
|
||||
'CLI tool',
|
||||
'single-page web app',
|
||||
'REST API service',
|
||||
'browser extension',
|
||||
'interactive data visualization dashboard',
|
||||
'terminal-based game',
|
||||
'real-time web app (using WebSockets)',
|
||||
'static site generator or theme',
|
||||
'browser-based game',
|
||||
'desktop utility (using Electron or Tauri)',
|
||||
'chat bot or conversational tool',
|
||||
'automation script or workflow tool',
|
||||
];
|
||||
const category = categories[Math.floor(Math.random() * categories.length)];
|
||||
|
||||
const seedClause = seed
|
||||
? `\n\nThe user provided this seed for inspiration. Stay closely aligned with the theme and intent of the seed — build on it, don't ignore it:\n"${seed}"`
|
||||
: '';
|
||||
|
||||
const prompt = `Generate a random, creative idea for a ${category}. The project should be achievable in a single coding session (1-2 hours) and should be interesting but not overly complex.
|
||||
|
||||
IMPORTANT: Be creative and diverse with your ideas. Avoid defaulting to developer-centric tools (git analyzers, code formatters, repo scanners, etc.) unless the category specifically calls for it. Think about ideas that would appeal to a broad audience — productivity, entertainment, education, health, finance, art, music, social, cooking, travel, fitness, etc.${seedClause}
|
||||
|
||||
Respond in EXACTLY this format (no other text):
|
||||
NAME: <kebab-case-project-name>
|
||||
DESCRIPTION: <2-3 sentence description of what the app does, its key features, and the tech stack to use>`;
|
||||
|
||||
const session = query({
|
||||
prompt,
|
||||
options: {
|
||||
maxTurns: 1,
|
||||
systemPrompt: 'You are a wildly creative project idea generator. You come up with surprising, fun, and diverse software project ideas spanning many domains — not just developer tools. Respond only in the exact format requested.',
|
||||
},
|
||||
});
|
||||
|
||||
let output = '';
|
||||
for await (const message of session) {
|
||||
if (message.type === 'assistant') {
|
||||
for (const block of message.message.content) {
|
||||
if (block.type === 'text') {
|
||||
output += block.text;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return parseIdea(output);
|
||||
}
|
||||
|
||||
export async function generateEnhancementIdea(projectDir: string, seed?: string): Promise<IdeaResult> {
|
||||
const seedClause = seed
|
||||
? `\n\nUse this as inspiration for the enhancement: "${seed}"`
|
||||
: '';
|
||||
|
||||
const prompt = `You are in a project directory. Scan the existing codebase to understand what it does, then propose a realistic enhancement (new feature, refactor, improvement, or extension).
|
||||
|
||||
Use the Read, Glob, and Grep tools to explore the project. Look at:
|
||||
- Package.json or similar config files for project info
|
||||
- Source files for current functionality
|
||||
- README or docs for context${seedClause}
|
||||
|
||||
Then respond in EXACTLY this format (no other text):
|
||||
NAME: <kebab-case-enhancement-name>
|
||||
DESCRIPTION: <2-3 sentence description of the enhancement, what it adds/changes, and why it would be valuable>`;
|
||||
|
||||
const session = query({
|
||||
prompt,
|
||||
options: {
|
||||
maxTurns: 8,
|
||||
cwd: projectDir,
|
||||
tools: { type: 'preset', preset: 'claude_code' },
|
||||
allowedTools: ['Read', 'Glob', 'Grep'],
|
||||
systemPrompt: { type: 'preset', preset: 'claude_code' },
|
||||
},
|
||||
});
|
||||
|
||||
let output = '';
|
||||
for await (const message of session) {
|
||||
if (message.type === 'assistant') {
|
||||
for (const block of message.message.content) {
|
||||
if (block.type === 'text') {
|
||||
output += block.text;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return parseIdea(output);
|
||||
}
|
||||
@@ -0,0 +1,2 @@
|
||||
export { runCLI } from './cli.js';
|
||||
export type { BotConfig, BotMode, BotState, StepName, StepResult } from './types.js';
|
||||
@@ -0,0 +1,243 @@
|
||||
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||
import type { BotConfig, ExecutionObservation, StepName } from './types.js';
|
||||
import type { ObservationCollector } from './observation-collector.js';
|
||||
|
||||
interface AskUserQuestionInput {
|
||||
questions: Array<{
|
||||
question: string;
|
||||
options: Array<{ label: string; description: string }>;
|
||||
multiSelect?: boolean;
|
||||
}>;
|
||||
}
|
||||
|
||||
interface IntelligentAnswers {
|
||||
answers: Record<string, string>;
|
||||
reasoning: Record<string, string>;
|
||||
}
|
||||
|
||||
/**
|
||||
* Creates an intelligent responder that uses LLM-as-judge for ALL decisions.
|
||||
* Replaces the old hardcoded auto-responder logic.
|
||||
*/
|
||||
export function createIntelligentResponder(
|
||||
config: BotConfig,
|
||||
step: StepName,
|
||||
collector: ObservationCollector
|
||||
) {
|
||||
return async (
|
||||
toolName: string,
|
||||
input: Record<string, unknown>
|
||||
): Promise<{ behavior: 'allow'; updatedInput?: Record<string, unknown> } | { behavior: 'deny'; message: string }> => {
|
||||
// Record every tool invocation for observations
|
||||
collector.recordToolUse(toolName, input, undefined, true);
|
||||
|
||||
// Handle AskUserQuestion with LLM-generated answers
|
||||
if (toolName === 'AskUserQuestion') {
|
||||
const askInput = input as unknown as AskUserQuestionInput;
|
||||
|
||||
// Get current observations to provide context to LLM
|
||||
const observations = collector.getSnapshot();
|
||||
|
||||
// Generate answers using LLM-as-judge
|
||||
const answers = await generateIntelligentAnswers(
|
||||
askInput,
|
||||
observations,
|
||||
config,
|
||||
step
|
||||
);
|
||||
|
||||
// Record each question/answer pair for metrics
|
||||
for (const q of askInput.questions) {
|
||||
const answer = answers.answers[q.question];
|
||||
const reasoning = answers.reasoning[q.question] || 'No reasoning provided';
|
||||
collector.recordQuestion(q.question, q.options, answer, reasoning);
|
||||
}
|
||||
|
||||
return {
|
||||
behavior: 'allow',
|
||||
updatedInput: { ...input, answers: answers.answers },
|
||||
};
|
||||
}
|
||||
|
||||
// Allow all other tools
|
||||
return { behavior: 'allow' };
|
||||
};
|
||||
}
|
||||
|
||||
async function generateIntelligentAnswers(
|
||||
askInput: AskUserQuestionInput,
|
||||
observations: ExecutionObservation,
|
||||
config: BotConfig,
|
||||
step: StepName
|
||||
): Promise<IntelligentAnswers> {
|
||||
const prompt = buildDecisionPrompt(askInput, observations, config, step);
|
||||
|
||||
try {
|
||||
// Query LLM for decision (single turn, read-only tools)
|
||||
const session = query({
|
||||
prompt,
|
||||
options: {
|
||||
maxTurns: 1,
|
||||
cwd: config.projectDir,
|
||||
permissionMode: 'bypassPermissions',
|
||||
allowDangerouslySkipPermissions: true,
|
||||
allowedTools: ['Read', 'Glob', 'Grep'],
|
||||
systemPrompt: 'You are a QA engineer reviewing work-in-progress. Be thoughtful and honest.',
|
||||
},
|
||||
});
|
||||
|
||||
let output = '';
|
||||
for await (const message of session) {
|
||||
if (message.type === 'assistant') {
|
||||
for (const block of message.message.content) {
|
||||
if (block.type === 'text') output += block.text;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return parseDecisionOutput(output, askInput);
|
||||
} catch (error) {
|
||||
console.warn('LLM decision failed, using fallback logic:', error);
|
||||
// Fallback to reasonable defaults if LLM fails
|
||||
return generateFallbackAnswers(askInput, step);
|
||||
}
|
||||
}
|
||||
|
||||
function buildDecisionPrompt(
|
||||
askInput: AskUserQuestionInput,
|
||||
observations: ExecutionObservation,
|
||||
config: BotConfig,
|
||||
step: StepName
|
||||
): string {
|
||||
const duration = observations.endTime - observations.startTime;
|
||||
const toolSummary = observations.tools
|
||||
.map((t) => `- ${t.toolName}`)
|
||||
.join('\n');
|
||||
const filesSummary = [
|
||||
...observations.filesCreated.map((f) => `CREATED: ${f}`),
|
||||
...observations.filesModified.map((f) => `MODIFIED: ${f}`),
|
||||
].join('\n');
|
||||
|
||||
const questionsText = askInput.questions
|
||||
.map((q, i) => {
|
||||
const optionsText = q.options
|
||||
.map((o, j) => ` ${j + 1}. ${o.label} - ${o.description}`)
|
||||
.join('\n');
|
||||
return `QUESTION ${i + 1}: ${q.question}\nOptions:\n${optionsText}`;
|
||||
})
|
||||
.join('\n\n');
|
||||
|
||||
return `You are a QA engineer reviewing a workflow step in progress.
|
||||
|
||||
## Context
|
||||
- Step: ${step}
|
||||
- Mode: ${config.mode}
|
||||
- Duration so far: ${Math.floor(duration / 1000)}s
|
||||
- Tools used: ${observations.tools.length}
|
||||
- Errors encountered: ${observations.errors.length}
|
||||
|
||||
## What's Happened So Far
|
||||
|
||||
### Tools Used
|
||||
${toolSummary || '(none yet)'}
|
||||
|
||||
### Files Changed
|
||||
${filesSummary || '(none yet)'}
|
||||
|
||||
${observations.errors.length > 0 ? `### Errors\n${observations.errors.join('\n')}` : ''}
|
||||
|
||||
## Questions to Answer
|
||||
|
||||
${questionsText}
|
||||
|
||||
## Your Task
|
||||
|
||||
You need to answer these questions as a thoughtful QA engineer would:
|
||||
1. Use Read, Glob, and Grep tools to inspect the current state of artifacts if needed
|
||||
2. Consider what you've observed (tools used, files created, errors)
|
||||
3. For each question, select the most appropriate answer
|
||||
4. Provide brief reasoning for your choice
|
||||
|
||||
Respond in this format:
|
||||
|
||||
QUESTION 1:
|
||||
ANSWER: <option label>
|
||||
REASONING: <1-2 sentences explaining your choice>
|
||||
|
||||
QUESTION 2:
|
||||
ANSWER: <option label>
|
||||
REASONING: <1-2 sentences>
|
||||
|
||||
Be honest. If work looks incomplete or problematic, don't approve it.
|
||||
If tests should be run but haven't been, don't skip them without good reason.
|
||||
Act like a real developer who cares about quality.`;
|
||||
}
|
||||
|
||||
function parseDecisionOutput(
|
||||
output: string,
|
||||
askInput: AskUserQuestionInput
|
||||
): IntelligentAnswers {
|
||||
const answers: Record<string, string> = {};
|
||||
const reasoning: Record<string, string> = {};
|
||||
|
||||
// Parse structured output
|
||||
const questionBlocks = output.split(/QUESTION \d+:/i).slice(1);
|
||||
|
||||
askInput.questions.forEach((q, i) => {
|
||||
const block = questionBlocks[i] || '';
|
||||
|
||||
const answerMatch = block.match(/ANSWER:\s*(.+?)(?=\n|$)/i);
|
||||
const reasoningMatch = block.match(/REASONING:\s*(.+?)(?=\n\n|$)/is);
|
||||
|
||||
const selectedLabel = answerMatch?.[1]?.trim() || '';
|
||||
|
||||
// Find matching option by label (case-insensitive partial match)
|
||||
const matchedOption = q.options.find(
|
||||
(opt) =>
|
||||
opt.label.toLowerCase().includes(selectedLabel.toLowerCase()) ||
|
||||
selectedLabel.toLowerCase().includes(opt.label.toLowerCase())
|
||||
);
|
||||
|
||||
answers[q.question] = matchedOption?.label || q.options[0].label;
|
||||
reasoning[q.question] = reasoningMatch?.[1]?.trim() || 'No reasoning provided';
|
||||
});
|
||||
|
||||
return { answers, reasoning };
|
||||
}
|
||||
|
||||
function generateFallbackAnswers(
|
||||
askInput: AskUserQuestionInput,
|
||||
step: StepName
|
||||
): IntelligentAnswers {
|
||||
// Simple fallback: pick first option for most questions
|
||||
// For approval questions, approve; for testing, skip
|
||||
const answers: Record<string, string> = {};
|
||||
const reasoning: Record<string, string> = {};
|
||||
|
||||
for (const q of askInput.questions) {
|
||||
const questionLower = q.question.toLowerCase();
|
||||
|
||||
if (questionLower.includes('approve') || questionLower.includes('proceed')) {
|
||||
const approveOption = q.options.find(
|
||||
(o) =>
|
||||
o.label.toLowerCase().includes('approve') ||
|
||||
o.label.toLowerCase().includes('yes')
|
||||
);
|
||||
answers[q.question] = approveOption?.label || q.options[0].label;
|
||||
reasoning[q.question] = 'Fallback approval (LLM unavailable)';
|
||||
} else if (questionLower.includes('test')) {
|
||||
const skipOption = q.options.find(
|
||||
(o) =>
|
||||
o.label.toLowerCase().includes('skip') ||
|
||||
o.label.toLowerCase().includes('none')
|
||||
);
|
||||
answers[q.question] = skipOption?.label || q.options[0].label;
|
||||
reasoning[q.question] = 'Fallback skip (LLM unavailable)';
|
||||
} else {
|
||||
answers[q.question] = q.options[0].label;
|
||||
reasoning[q.question] = 'Fallback first option (LLM unavailable)';
|
||||
}
|
||||
}
|
||||
|
||||
return { answers, reasoning };
|
||||
}
|
||||
@@ -0,0 +1,113 @@
|
||||
import type { ExecutionObservation, StepName } from './types.js';
|
||||
|
||||
/**
|
||||
* Collects detailed observations during step execution.
|
||||
* Tracks tool usage, messages, errors, file changes, and questions asked.
|
||||
*/
|
||||
export class ObservationCollector {
|
||||
private observations: ExecutionObservation;
|
||||
private recentToolKeys: Set<string> = new Set();
|
||||
|
||||
constructor(step: StepName) {
|
||||
this.observations = {
|
||||
step,
|
||||
startTime: Date.now(),
|
||||
endTime: 0,
|
||||
tools: [],
|
||||
assistantMessages: [],
|
||||
errors: [],
|
||||
filesCreated: [],
|
||||
filesModified: [],
|
||||
questionsAsked: [],
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Records a tool invocation with its input and output.
|
||||
* Deduplicates if the same tool+input is recorded from both canUseTool and the message stream.
|
||||
*/
|
||||
recordToolUse(
|
||||
toolName: string,
|
||||
input: Record<string, unknown>,
|
||||
output: unknown,
|
||||
allowed: boolean
|
||||
): void {
|
||||
// Deduplicate based on tool name + serialized input (within a short time window)
|
||||
const key = `${toolName}:${JSON.stringify(input)}`;
|
||||
if (this.recentToolKeys.has(key)) {
|
||||
return;
|
||||
}
|
||||
this.recentToolKeys.add(key);
|
||||
// Clean up old keys periodically to avoid unbounded growth
|
||||
if (this.recentToolKeys.size > 500) {
|
||||
const entries = [...this.recentToolKeys];
|
||||
this.recentToolKeys = new Set(entries.slice(entries.length - 250));
|
||||
}
|
||||
|
||||
this.observations.tools.push({
|
||||
toolName,
|
||||
input,
|
||||
output,
|
||||
timestamp: Date.now(),
|
||||
allowed,
|
||||
autoAnswered: toolName === 'AskUserQuestion',
|
||||
});
|
||||
|
||||
// Extract file paths from common tools
|
||||
if (toolName === 'Write' && input.file_path) {
|
||||
this.observations.filesCreated.push(input.file_path as string);
|
||||
}
|
||||
if (toolName === 'Edit' && input.file_path) {
|
||||
this.observations.filesModified.push(input.file_path as string);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Records messages from the session (assistant text, errors).
|
||||
*/
|
||||
recordMessage(message: any): void {
|
||||
if (message.type === 'assistant') {
|
||||
for (const block of message.message.content) {
|
||||
if (block.type === 'text') {
|
||||
this.observations.assistantMessages.push(block.text);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (message.type === 'error') {
|
||||
this.observations.errors.push(message.error?.message ?? 'Unknown error');
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Records a question that was asked and the LLM-generated answer.
|
||||
*/
|
||||
recordQuestion(
|
||||
question: string,
|
||||
options: Array<{ label: string; description: string }>,
|
||||
selectedAnswer: string,
|
||||
llmReasoning: string
|
||||
): void {
|
||||
this.observations.questionsAsked.push({
|
||||
question,
|
||||
options,
|
||||
llmReasoning,
|
||||
selectedAnswer,
|
||||
timestamp: Date.now(),
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Gets a snapshot of current observations (for real-time decision making).
|
||||
*/
|
||||
getSnapshot(): ExecutionObservation {
|
||||
return { ...this.observations, endTime: Date.now() };
|
||||
}
|
||||
|
||||
/**
|
||||
* Finalizes observations and returns the complete record.
|
||||
*/
|
||||
finalize(): ExecutionObservation {
|
||||
this.observations.endTime = Date.now();
|
||||
return { ...this.observations };
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,158 @@
|
||||
import type { StepName } from '../types.js';
|
||||
|
||||
export interface StepEvaluationCriteria {
|
||||
step: StepName;
|
||||
keyArtifacts: string[];
|
||||
qualityChecks: string[];
|
||||
commonPitfalls: string[];
|
||||
scoringGuidance: string;
|
||||
}
|
||||
|
||||
export const EVALUATION_CRITERIA: Record<StepName, StepEvaluationCriteria> = {
|
||||
init: {
|
||||
step: 'init',
|
||||
keyArtifacts: ['AGENTS.md', 'IDEA.md'],
|
||||
qualityChecks: [
|
||||
'AGENTS.md exists with project name and description',
|
||||
'AGENTS.md includes a brief intro line and a Status section indicating the project is in planning phase',
|
||||
'AGENTS.md does NOT contain hallucinated architecture, commands, file structures, or tech stack details',
|
||||
'No .agents-docs/ directory was created (too early for detail files)',
|
||||
'No project scaffolding (package.json, dependencies, src/) was created',
|
||||
'IDEA.md exists with the project idea',
|
||||
],
|
||||
commonPitfalls: [
|
||||
'Hallucinating architecture or tech stack details before the plan step',
|
||||
'Creating .agents-docs/ detail files with invented content',
|
||||
'Scaffolding project files or installing dependencies prematurely',
|
||||
],
|
||||
scoringGuidance: `
|
||||
100 = Perfect: AGENTS.md stub with intro line, project name/description, and status section. IDEA.md present. Nothing else created.
|
||||
85-95 = Good but minor extra content beyond the expected stub format (e.g., an extra placeholder heading)
|
||||
70-84 = AGENTS.md exists but includes some hallucinated details (e.g., assumed tech stack or commands)
|
||||
50-69 = Significant hallucination (e.g., .agents-docs/ created with invented content, project scaffolded)
|
||||
<50 = Major problems (e.g., AGENTS.md missing, full project structure hallucinated)
|
||||
|
||||
The expected AGENTS.md format is: intro line, Project Overview (name + description), and a Status section. This is the target for a 100 score.
|
||||
`,
|
||||
},
|
||||
|
||||
plan: {
|
||||
step: 'plan',
|
||||
keyArtifacts: ['specs/*/PLAN-DRAFT-*.md', 'IDEA.md'],
|
||||
qualityChecks: [
|
||||
'Plan breaks work into clear, achievable phases',
|
||||
'Each phase has specific goals and deliverables',
|
||||
'Technical approach is appropriate',
|
||||
'Scope is realistic for the idea',
|
||||
'Dependencies between phases are identified',
|
||||
],
|
||||
commonPitfalls: [
|
||||
'Phases too vague ("polish the app")',
|
||||
'Missing specific tasks within phases',
|
||||
'No testing strategy mentioned',
|
||||
'Overly ambitious scope',
|
||||
'Missing file paths or specific actions',
|
||||
],
|
||||
scoringGuidance: `
|
||||
100 = Exceptional plan: detailed phases, realistic scope, clear tasks, testing included
|
||||
85-95 = Good plan with minor improvements possible (e.g., one phase could be more specific)
|
||||
70-84 = Acceptable but has vague sections or missing testing strategy
|
||||
50-69 = Significant issues (e.g., multiple vague phases, unrealistic scope)
|
||||
<50 = Major problems (e.g., no clear phases, plan doesn't match idea)
|
||||
|
||||
Most plans should score 70-85. Be critical of vague language.
|
||||
`,
|
||||
},
|
||||
|
||||
document: {
|
||||
step: 'document',
|
||||
keyArtifacts: ['specs/*/overview.md', 'specs/*/phase-*.md files'],
|
||||
qualityChecks: [
|
||||
'overview.md provides clear project summary',
|
||||
'Each phase file has specific tasks with checkboxes',
|
||||
'File paths are explicit (not generic)',
|
||||
'Dependencies between tasks are identified',
|
||||
'Technical details are specific',
|
||||
'Acceptance criteria are clear',
|
||||
],
|
||||
commonPitfalls: [
|
||||
'Tasks too generic ("implement feature X")',
|
||||
'Missing file paths',
|
||||
'No checkboxes or unclear task structure',
|
||||
'Missing dependencies',
|
||||
'Overly verbose or lacking specifics',
|
||||
],
|
||||
scoringGuidance: `
|
||||
100 = Exceptional documentation: specific tasks, explicit file paths, clear dependencies
|
||||
85-95 = Good documentation with minor vagueness in one or two tasks
|
||||
70-84 = Acceptable but multiple tasks lack specifics or file paths
|
||||
50-69 = Significant issues (e.g., many generic tasks, missing file paths)
|
||||
<50 = Major problems (e.g., tasks don't match plan, fundamentally vague)
|
||||
|
||||
Most documentation should score 70-85. Penalize generic language heavily.
|
||||
`,
|
||||
},
|
||||
|
||||
implement: {
|
||||
step: 'implement',
|
||||
keyArtifacts: ['actual code files', 'checked-off tasks in phase files'],
|
||||
qualityChecks: [
|
||||
'Phase tasks are being completed',
|
||||
'Code files are actually created/modified',
|
||||
'Implementation follows the documented plan',
|
||||
'No major errors blocking progress',
|
||||
'Tests are written (if applicable)',
|
||||
],
|
||||
commonPitfalls: [
|
||||
'Tasks marked complete but files not actually changed',
|
||||
'Implementation deviates significantly from plan',
|
||||
'Errors not addressed',
|
||||
'Skipping tests without justification',
|
||||
'Working on wrong phase',
|
||||
],
|
||||
scoringGuidance: `
|
||||
100 = Exceptional implementation: all tasks complete, code works, tests pass
|
||||
85-95 = Good implementation with minor issues or incomplete tests
|
||||
70-84 = Acceptable but some tasks incomplete or code has issues
|
||||
50-69 = Significant issues (e.g., many tasks incomplete, code doesn't work)
|
||||
<50 = Major problems (e.g., wrong phase, no actual work done)
|
||||
|
||||
Implementation scoring depends heavily on actual progress. Be realistic.
|
||||
`,
|
||||
},
|
||||
|
||||
finalize: {
|
||||
step: 'finalize',
|
||||
keyArtifacts: ['specs--completed/', 'README or docs', 'final code state'],
|
||||
qualityChecks: [
|
||||
'All phases are marked complete',
|
||||
'Spec moved to specs--completed/',
|
||||
'Documentation is updated',
|
||||
'Code is in working state',
|
||||
'No obvious loose ends',
|
||||
],
|
||||
commonPitfalls: [
|
||||
'Incomplete phases',
|
||||
'Missing specs--completed/ move',
|
||||
'Documentation not updated',
|
||||
'Code broken or incomplete',
|
||||
'Unrealistic self-assessment',
|
||||
],
|
||||
scoringGuidance: `
|
||||
100 = Exceptional finalization: everything complete, polished, documented
|
||||
85-95 = Good finalization with minor issues
|
||||
70-84 = Acceptable but some loose ends or documentation gaps
|
||||
50-69 = Significant issues (e.g., incomplete phases, broken code)
|
||||
<50 = Major problems (e.g., work not actually done, fundamentally incomplete)
|
||||
|
||||
Finalize scores should reflect overall project quality. Be honest.
|
||||
`,
|
||||
},
|
||||
};
|
||||
|
||||
/**
|
||||
* Gets evaluation criteria for a specific step.
|
||||
*/
|
||||
export function getCriteriaForStep(step: StepName): StepEvaluationCriteria {
|
||||
return EVALUATION_CRITERIA[step];
|
||||
}
|
||||
@@ -0,0 +1,58 @@
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import { buildStepPrompt } from './step-instructions.js';
|
||||
import type { BotConfig } from '../types.js';
|
||||
|
||||
const config: BotConfig = {
|
||||
workDir: '/tmp/work',
|
||||
projectDir: '/tmp/work/my-app',
|
||||
ideaDescription: 'A todo app with drag-and-drop',
|
||||
ideaName: 'drag-todo',
|
||||
mode: 'new-project',
|
||||
};
|
||||
|
||||
describe('buildStepPrompt', () => {
|
||||
it('each step includes the autonomous preamble', () => {
|
||||
const steps = ['init', 'plan', 'document', 'implement', 'finalize'] as const;
|
||||
for (const step of steps) {
|
||||
const prompt = buildStepPrompt(step, config);
|
||||
expect(prompt).toContain('running autonomously');
|
||||
}
|
||||
});
|
||||
|
||||
it('init prompt includes /plan2code-init skill invocation', () => {
|
||||
const prompt = buildStepPrompt('init', config);
|
||||
expect(prompt).toContain('/plan2code-init');
|
||||
});
|
||||
|
||||
it('plan prompt includes /plan2code-1-plan skill invocation', () => {
|
||||
const prompt = buildStepPrompt('plan', config);
|
||||
expect(prompt).toContain('/plan2code-1-plan');
|
||||
});
|
||||
|
||||
it('document prompt includes /plan2code-2-document skill invocation', () => {
|
||||
const prompt = buildStepPrompt('document', config);
|
||||
expect(prompt).toContain('/plan2code-2-document');
|
||||
});
|
||||
|
||||
it('implement prompt instructs direct tool usage without skill invocation', () => {
|
||||
const prompt = buildStepPrompt('implement', config);
|
||||
expect(prompt).toContain('Do NOT use the Skill tool');
|
||||
});
|
||||
|
||||
it('finalize prompt includes /plan2code-4-finalize skill invocation', () => {
|
||||
const prompt = buildStepPrompt('finalize', config);
|
||||
expect(prompt).toContain('/plan2code-4-finalize');
|
||||
});
|
||||
|
||||
it('plan prompt includes project name and description', () => {
|
||||
const prompt = buildStepPrompt('plan', config);
|
||||
expect(prompt).toContain('drag-todo');
|
||||
expect(prompt).toContain('A todo app with drag-and-drop');
|
||||
});
|
||||
|
||||
it('init prompt includes project name and description', () => {
|
||||
const prompt = buildStepPrompt('init', config);
|
||||
expect(prompt).toContain('drag-todo');
|
||||
expect(prompt).toContain('A todo app with drag-and-drop');
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,131 @@
|
||||
import type { BotConfig, StepName } from '../types.js';
|
||||
|
||||
const AUTONOMOUS_PREAMBLE = `You are running autonomously as part of an automated test pipeline. Do NOT pause for human input. When you encounter questions or approval gates, make reasonable decisions and proceed. If asked for confirmation, approve. If asked to choose, pick the most reasonable option. Complete the entire step without stopping.`;
|
||||
|
||||
export function buildStepPrompt(step: StepName, config: BotConfig): string {
|
||||
switch (step) {
|
||||
case 'init':
|
||||
return buildInitPrompt(config);
|
||||
case 'plan':
|
||||
return buildPlanPrompt(config);
|
||||
case 'document':
|
||||
return buildDocumentPrompt(config);
|
||||
case 'implement':
|
||||
return buildImplementPrompt(config);
|
||||
case 'finalize':
|
||||
return buildFinalizePrompt(config);
|
||||
}
|
||||
}
|
||||
|
||||
function buildInitPrompt(config: BotConfig): string {
|
||||
return `${AUTONOMOUS_PREAMBLE}
|
||||
|
||||
This is a brand new project with no existing code. Create a minimal stub AGENTS.md file and an IDEA.md file. Do NOT run the /plan2code-init skill — there is no codebase to analyze yet.
|
||||
|
||||
The project idea is: ${config.ideaDescription}
|
||||
The project name is: ${config.ideaName}
|
||||
|
||||
## What to create
|
||||
|
||||
### AGENTS.md
|
||||
Create a minimal stub with ONLY the following — do NOT invent architecture, tech stack details, commands, or file structures:
|
||||
|
||||
\`\`\`markdown
|
||||
# AGENTS.md
|
||||
|
||||
This file provides guidance to AI coding agents when working with code in this repository.
|
||||
|
||||
## Project Overview
|
||||
**Name:** ${config.ideaName}
|
||||
**Description:** ${config.ideaDescription}
|
||||
|
||||
## Status
|
||||
This project is in the planning phase. Architecture, commands, and detailed documentation will be added after the plan and document steps are complete.
|
||||
\`\`\`
|
||||
|
||||
### IDEA.md
|
||||
If IDEA.md does not already exist, create it with the project name and description.
|
||||
|
||||
## Rules
|
||||
- Do NOT create .agents-docs/ or any detail files — there is nothing to document yet
|
||||
- Do NOT hallucinate architecture, dependencies, file structures, or tech stack choices
|
||||
- Do NOT install dependencies or scaffold project files
|
||||
- ONLY create the two files above`;
|
||||
}
|
||||
|
||||
function buildPlanPrompt(config: BotConfig): string {
|
||||
return `${AUTONOMOUS_PREAMBLE}
|
||||
|
||||
Run the /plan2code-1-plan skill to create a plan for this project.
|
||||
|
||||
Read the IDEA.md file first to understand the project. The project is: ${config.ideaDescription}
|
||||
|
||||
When making decisions during planning:
|
||||
- Set confidence levels reasonably high (85-95%)
|
||||
- Accept the generated tech stack without revision
|
||||
- Do not request additional reference files
|
||||
- Approve the plan when asked for sign-off
|
||||
- Keep scope small and achievable (3-4 phases max)
|
||||
- Use the project name: ${config.ideaName}`;
|
||||
}
|
||||
|
||||
function buildDocumentPrompt(config: BotConfig): string {
|
||||
return `${AUTONOMOUS_PREAMBLE}
|
||||
|
||||
Run the /plan2code-2-document skill to transform the plan into implementation specs.
|
||||
|
||||
Read the existing plan output in specs/ first to understand what was planned.
|
||||
|
||||
When making decisions:
|
||||
- Accept all generated documentation
|
||||
- Approve phase breakdowns and task lists
|
||||
- Do not request changes to the generated docs`;
|
||||
}
|
||||
|
||||
function buildImplementPrompt(config: BotConfig): string {
|
||||
return `${AUTONOMOUS_PREAMBLE}
|
||||
|
||||
You are a senior software engineer implementing a project phase. Do NOT use the Skill tool — implement directly using Read, Write, Edit, Glob, and Grep tools.
|
||||
|
||||
## Process
|
||||
|
||||
1. **Find the spec**: Read \`specs/*/overview.md\` to find the phase checklist
|
||||
2. **Pick the next phase**: Find the first phase marked \`[ ]\` (pending) or \`[/]\` (in-progress)
|
||||
3. **Read the phase file**: Read the corresponding \`phase-X.md\` from the same directory
|
||||
4. **Mark phase in-progress**: Update \`[ ]\` to \`[/]\` in overview.md
|
||||
5. **Implement each task sequentially**:
|
||||
- Read the task specification completely
|
||||
- Write the code using Write or Edit tools — create real files, not code blocks
|
||||
- Mark the task \`[x]\` in the phase file immediately after completing it
|
||||
6. **Complete the phase**: After all tasks, fill in the "Phase Completion Summary" in the phase file
|
||||
7. **Mark phase complete**: Update \`[/]\` to \`[x]\` in overview.md
|
||||
|
||||
## Rules
|
||||
|
||||
- Follow AGENTS.md if it exists
|
||||
- Implement specs EXACTLY — no creative additions or unsolicited improvements
|
||||
- Write task completion status (\`[x]\`) to disk immediately after each task — never batch
|
||||
- Only create files mentioned in the spec tasks
|
||||
- Use the specified file paths, function names, and structures from the spec
|
||||
- No placeholder code — fully implement every function
|
||||
- Match existing codebase conventions
|
||||
- Do NOT run git commands
|
||||
- Skip running tests unless explicitly listed as a phase task
|
||||
|
||||
## Project info
|
||||
- Project: ${config.ideaName}
|
||||
- Description: ${config.ideaDescription}
|
||||
- Project directory: ${config.projectDir}`;
|
||||
}
|
||||
|
||||
function buildFinalizePrompt(config: BotConfig): string {
|
||||
return `${AUTONOMOUS_PREAMBLE}
|
||||
|
||||
Run the /plan2code-4-finalize skill to validate and archive the completed project.
|
||||
|
||||
When making decisions:
|
||||
- Approve all documentation updates
|
||||
- Accept the completion summary
|
||||
- If asked for a rating or feedback, give 8/10 and positive feedback
|
||||
- Complete the archival process fully`;
|
||||
}
|
||||
@@ -0,0 +1,82 @@
|
||||
import { describe, it, expect, vi } from 'vitest';
|
||||
|
||||
// Mock the SDK before importing session-runner
|
||||
vi.mock('@anthropic-ai/claude-agent-sdk', () => ({
|
||||
query: vi.fn(),
|
||||
}));
|
||||
|
||||
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||
import { runSession } from './session-runner.js';
|
||||
import { ObservationCollector } from './observation-collector.js';
|
||||
import type { BotConfig } from './types.js';
|
||||
|
||||
const config: BotConfig = {
|
||||
workDir: '/tmp/work',
|
||||
projectDir: '/tmp/work/my-app',
|
||||
ideaDescription: 'test app',
|
||||
ideaName: 'test-app',
|
||||
mode: 'new-project',
|
||||
};
|
||||
|
||||
describe('runSession', () => {
|
||||
it('returns success: false when output is empty', async () => {
|
||||
// Simulate a session that yields an assistant message with empty text
|
||||
const mockQuery = vi.mocked(query);
|
||||
mockQuery.mockReturnValue(
|
||||
(async function* () {
|
||||
yield {
|
||||
type: 'assistant' as const,
|
||||
session_id: 'sess-1',
|
||||
message: { content: [{ type: 'text' as const, text: '' }] },
|
||||
};
|
||||
yield {
|
||||
type: 'result' as const,
|
||||
session_id: 'sess-1',
|
||||
};
|
||||
})() as any,
|
||||
);
|
||||
|
||||
const collector = new ObservationCollector('init');
|
||||
const result = await runSession({
|
||||
prompt: 'do something',
|
||||
config,
|
||||
step: 'init',
|
||||
collector,
|
||||
});
|
||||
|
||||
expect(result.success).toBe(false);
|
||||
expect(result.output.trim()).toBe('');
|
||||
expect(result.observations).toBeDefined();
|
||||
});
|
||||
|
||||
it('returns success: true when output has content', async () => {
|
||||
const mockQuery = vi.mocked(query);
|
||||
mockQuery.mockReturnValue(
|
||||
(async function* () {
|
||||
yield {
|
||||
type: 'assistant' as const,
|
||||
session_id: 'sess-2',
|
||||
message: { content: [{ type: 'text' as const, text: 'AGENTS.md has been created successfully' }] },
|
||||
};
|
||||
yield {
|
||||
type: 'result' as const,
|
||||
session_id: 'sess-2',
|
||||
};
|
||||
})() as any,
|
||||
);
|
||||
|
||||
const collector = new ObservationCollector('init');
|
||||
const result = await runSession({
|
||||
prompt: 'do something',
|
||||
config,
|
||||
step: 'init',
|
||||
collector,
|
||||
});
|
||||
|
||||
expect(result.success).toBe(true);
|
||||
expect(result.output).toContain('AGENTS.md');
|
||||
expect(result.sessionId).toBe('sess-2');
|
||||
expect(result.observations).toBeDefined();
|
||||
expect(result.observations.step).toBe('init');
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,77 @@
|
||||
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||
import { createIntelligentResponder } from './intelligent-responder.js';
|
||||
import type { BotConfig, ExecutionObservation, StepName } from './types.js';
|
||||
import type { ObservationCollector } from './observation-collector.js';
|
||||
|
||||
export interface SessionOptions {
|
||||
prompt: string;
|
||||
config: BotConfig;
|
||||
step: StepName;
|
||||
maxTurns?: number;
|
||||
collector: ObservationCollector;
|
||||
}
|
||||
|
||||
export interface SessionResult {
|
||||
sessionId: string | null;
|
||||
output: string;
|
||||
success: boolean;
|
||||
duration: number;
|
||||
observations: ExecutionObservation;
|
||||
}
|
||||
|
||||
export async function runSession(options: SessionOptions): Promise<SessionResult> {
|
||||
const { prompt, config, step, maxTurns = 50, collector } = options;
|
||||
const startTime = Date.now();
|
||||
let output = '';
|
||||
let sessionId: string | null = null;
|
||||
|
||||
try {
|
||||
const session = query({
|
||||
prompt,
|
||||
options: {
|
||||
cwd: config.projectDir,
|
||||
maxTurns,
|
||||
permissionMode: 'bypassPermissions',
|
||||
allowDangerouslySkipPermissions: true,
|
||||
canUseTool: createIntelligentResponder(config, step, collector),
|
||||
systemPrompt: { type: 'preset', preset: 'claude_code' },
|
||||
settingSources: ['project'],
|
||||
},
|
||||
});
|
||||
|
||||
for await (const message of session) {
|
||||
// Record all messages for observations
|
||||
collector.recordMessage(message);
|
||||
|
||||
if (message.type === 'assistant') {
|
||||
sessionId = message.session_id ?? sessionId;
|
||||
for (const block of message.message.content) {
|
||||
if (block.type === 'text') {
|
||||
output += block.text + '\n';
|
||||
} else if (block.type === 'tool_use') {
|
||||
// Capture tool invocations from the message stream as a fallback
|
||||
// in case canUseTool doesn't fire (e.g., Skill sub-sessions)
|
||||
collector.recordToolUse(
|
||||
block.name,
|
||||
block.input as Record<string, unknown>,
|
||||
undefined,
|
||||
true
|
||||
);
|
||||
}
|
||||
}
|
||||
} else if (message.type === 'result') {
|
||||
sessionId = message.session_id ?? sessionId;
|
||||
}
|
||||
}
|
||||
|
||||
const duration = Date.now() - startTime;
|
||||
const hasOutput = output.trim().length > 0;
|
||||
const observations = collector.finalize();
|
||||
return { sessionId, output, success: hasOutput, duration, observations };
|
||||
} catch (error) {
|
||||
const duration = Date.now() - startTime;
|
||||
const errorMsg = error instanceof Error ? error.message : String(error);
|
||||
const observations = collector.finalize();
|
||||
return { sessionId, output, success: false, duration, observations };
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,164 @@
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
|
||||
import fs from 'fs-extra';
|
||||
import os from 'os';
|
||||
import path from 'path';
|
||||
import { checkAllPhasesComplete, detectStepCompletion } from './step-detector.js';
|
||||
|
||||
let tmpDir: string;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'p2c-test-'));
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
fs.removeSync(tmpDir);
|
||||
});
|
||||
|
||||
// ── checkAllPhasesComplete ──────────────────────────────────────────
|
||||
|
||||
describe('checkAllPhasesComplete', () => {
|
||||
it('returns false when no specs/ directory exists', () => {
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(false);
|
||||
});
|
||||
|
||||
it('returns false when specs/ has no subdirectories', () => {
|
||||
fs.ensureDirSync(path.join(tmpDir, 'specs'));
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(false);
|
||||
});
|
||||
|
||||
it('returns false when spec dir exists but has no overview.md and no phase files', () => {
|
||||
// This is the false-positive bug we fixed — an empty spec dir should NOT be "complete"
|
||||
fs.ensureDirSync(path.join(tmpDir, 'specs', 'my-feature'));
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(false);
|
||||
});
|
||||
|
||||
it('returns true when overview.md has all phases checked [x]', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(
|
||||
path.join(specDir, 'overview.md'),
|
||||
`# Overview\n- [x] Phase 1: Setup\n- [x] Phase 2: Core\n- [x] Phase 3: Polish\n`,
|
||||
);
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(true);
|
||||
});
|
||||
|
||||
it('returns false when overview.md has unchecked [ ] phases', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(
|
||||
path.join(specDir, 'overview.md'),
|
||||
`# Overview\n- [x] Phase 1: Setup\n- [ ] Phase 2: Core\n- [ ] Phase 3: Polish\n`,
|
||||
);
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(false);
|
||||
});
|
||||
|
||||
it('returns true via fallback: phase-*.md files with all checked, no overview.md', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(
|
||||
path.join(specDir, 'phase-1.md'),
|
||||
`# Phase 1\n- [x] Task A\n- [x] Task B\n`,
|
||||
);
|
||||
fs.writeFileSync(
|
||||
path.join(specDir, 'phase-2.md'),
|
||||
`# Phase 2\n- [x] Task C\n`,
|
||||
);
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(true);
|
||||
});
|
||||
|
||||
it('returns false via fallback: phase-*.md with unchecked items', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
fs.ensureDirSync(specDir);
|
||||
fs.writeFileSync(
|
||||
path.join(specDir, 'phase-1.md'),
|
||||
`# Phase 1\n- [x] Task A\n- [ ] Task B\n`,
|
||||
);
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(false);
|
||||
});
|
||||
|
||||
it('returns true via nested phases/ subdirectory fallback with all checked', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
const phasesDir = path.join(specDir, 'phases');
|
||||
fs.ensureDirSync(phasesDir);
|
||||
fs.writeFileSync(
|
||||
path.join(phasesDir, 'phase-1.md'),
|
||||
`# Phase 1\n- [x] Task A\n- [x] Task B\n`,
|
||||
);
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(true);
|
||||
});
|
||||
|
||||
it('returns false via nested phases/ with unchecked items', () => {
|
||||
const specDir = path.join(tmpDir, 'specs', 'my-feature');
|
||||
const phasesDir = path.join(specDir, 'phases');
|
||||
fs.ensureDirSync(phasesDir);
|
||||
fs.writeFileSync(
|
||||
path.join(phasesDir, 'phase-1.md'),
|
||||
`# Phase 1\n- [x] Task A\n- [ ] Task B\n`,
|
||||
);
|
||||
expect(checkAllPhasesComplete(tmpDir)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
// ── detectStepCompletion ────────────────────────────────────────────
|
||||
|
||||
describe('detectStepCompletion', () => {
|
||||
it('init: completed when output mentions agents.md created', () => {
|
||||
const result = detectStepCompletion('AGENTS.md has been created successfully', 'init');
|
||||
expect(result.completed).toBe(true);
|
||||
expect(result.nextStep).toBe('plan');
|
||||
});
|
||||
|
||||
it('init: not completed for unrelated output', () => {
|
||||
const result = detectStepCompletion('Hello world, nothing happened', 'init');
|
||||
expect(result.completed).toBe(false);
|
||||
expect(result.nextStep).toBeNull();
|
||||
});
|
||||
|
||||
it('plan: completed when plan is saved', () => {
|
||||
const result = detectStepCompletion('The plan has been saved and finalized', 'plan');
|
||||
expect(result.completed).toBe(true);
|
||||
expect(result.nextStep).toBe('document');
|
||||
});
|
||||
|
||||
it('plan: not completed for unrelated output', () => {
|
||||
const result = detectStepCompletion('Reading the codebase...', 'plan');
|
||||
expect(result.completed).toBe(false);
|
||||
expect(result.nextStep).toBeNull();
|
||||
});
|
||||
|
||||
it('document: completed when overview.md is mentioned', () => {
|
||||
const result = detectStepCompletion('Created overview.md with all phases', 'document');
|
||||
expect(result.completed).toBe(true);
|
||||
expect(result.nextStep).toBe('implement');
|
||||
});
|
||||
|
||||
it('document: not completed for unrelated output', () => {
|
||||
const result = detectStepCompletion('Thinking about the design...', 'document');
|
||||
expect(result.completed).toBe(false);
|
||||
expect(result.nextStep).toBeNull();
|
||||
});
|
||||
|
||||
it('implement: completed when phase is done', () => {
|
||||
const result = detectStepCompletion('Phase 1 is now complete!', 'implement');
|
||||
expect(result.completed).toBe(true);
|
||||
expect(result.nextStep).toBe('finalize');
|
||||
});
|
||||
|
||||
it('implement: not completed for unrelated output', () => {
|
||||
const result = detectStepCompletion('Working on some files', 'implement');
|
||||
expect(result.completed).toBe(false);
|
||||
expect(result.nextStep).toBeNull();
|
||||
});
|
||||
|
||||
it('finalize: completed when finalize is done', () => {
|
||||
const result = detectStepCompletion('Finalize step is complete and archived', 'finalize');
|
||||
expect(result.completed).toBe(true);
|
||||
expect(result.nextStep).toBeNull(); // last step
|
||||
});
|
||||
|
||||
it('finalize: not completed for unrelated output', () => {
|
||||
const result = detectStepCompletion('Just starting...', 'finalize');
|
||||
expect(result.completed).toBe(false);
|
||||
expect(result.nextStep).toBeNull();
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,151 @@
|
||||
import fs from 'fs-extra';
|
||||
import path from 'path';
|
||||
import type { StepName } from './types.js';
|
||||
|
||||
export interface DetectionResult {
|
||||
completed: boolean;
|
||||
nextStep: StepName | null;
|
||||
needsAnotherImplementPass: boolean;
|
||||
}
|
||||
|
||||
const STEP_ORDER: StepName[] = ['init', 'plan', 'document', 'implement', 'finalize'];
|
||||
|
||||
export function detectStepCompletion(output: string, step: StepName): DetectionResult {
|
||||
const lowerOutput = output.toLowerCase();
|
||||
|
||||
let completed = false;
|
||||
|
||||
switch (step) {
|
||||
case 'init':
|
||||
completed = lowerOutput.includes('agents.md') && (
|
||||
lowerOutput.includes('created') ||
|
||||
lowerOutput.includes('generated') ||
|
||||
lowerOutput.includes('written')
|
||||
);
|
||||
break;
|
||||
|
||||
case 'plan':
|
||||
completed = lowerOutput.includes('plan') && (
|
||||
lowerOutput.includes('complete') ||
|
||||
lowerOutput.includes('approved') ||
|
||||
lowerOutput.includes('finalized') ||
|
||||
lowerOutput.includes('saved')
|
||||
);
|
||||
break;
|
||||
|
||||
case 'document':
|
||||
completed = lowerOutput.includes('overview.md') || (
|
||||
lowerOutput.includes('document') && lowerOutput.includes('complete')
|
||||
);
|
||||
break;
|
||||
|
||||
case 'implement':
|
||||
completed = lowerOutput.includes('phase') && (
|
||||
lowerOutput.includes('complete') ||
|
||||
lowerOutput.includes('done') ||
|
||||
lowerOutput.includes('finished')
|
||||
);
|
||||
break;
|
||||
|
||||
case 'finalize':
|
||||
completed = lowerOutput.includes('finalize') && (
|
||||
lowerOutput.includes('complete') ||
|
||||
lowerOutput.includes('archived') ||
|
||||
lowerOutput.includes('done')
|
||||
);
|
||||
break;
|
||||
}
|
||||
|
||||
// Determine next step
|
||||
const currentIdx = STEP_ORDER.indexOf(step);
|
||||
const nextStep = currentIdx < STEP_ORDER.length - 1 ? STEP_ORDER[currentIdx + 1] : null;
|
||||
|
||||
return {
|
||||
completed,
|
||||
nextStep: completed ? nextStep : null,
|
||||
needsAnotherImplementPass: false,
|
||||
};
|
||||
}
|
||||
|
||||
export function checkAllPhasesComplete(projectDir: string): boolean {
|
||||
const specsDir = path.join(projectDir, 'specs');
|
||||
|
||||
if (!fs.existsSync(specsDir)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
const entries = fs.readdirSync(specsDir, { withFileTypes: true });
|
||||
const specDirs = entries.filter((e) => e.isDirectory());
|
||||
|
||||
if (specDirs.length === 0) {
|
||||
return false;
|
||||
}
|
||||
|
||||
let foundPhaseTracking = false;
|
||||
|
||||
for (const dir of specDirs) {
|
||||
const specPath = path.join(specsDir, dir.name);
|
||||
|
||||
// Try overview.md first
|
||||
const overviewPath = path.join(specPath, 'overview.md');
|
||||
if (fs.existsSync(overviewPath)) {
|
||||
foundPhaseTracking = true;
|
||||
if (hasUncheckedPhases(fs.readFileSync(overviewPath, 'utf-8'))) {
|
||||
return false;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Fallback: look for phase-*.md or PHASE-*.md in the spec dir
|
||||
const phaseFiles = findPhaseFiles(specPath);
|
||||
if (phaseFiles.length > 0) {
|
||||
foundPhaseTracking = true;
|
||||
for (const pf of phaseFiles) {
|
||||
if (hasUncheckedPhases(fs.readFileSync(pf, 'utf-8'))) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Fallback: look inside a phases/ subdirectory
|
||||
const phasesSubdir = path.join(specPath, 'phases');
|
||||
if (fs.existsSync(phasesSubdir)) {
|
||||
const subPhaseFiles = findPhaseFiles(phasesSubdir);
|
||||
if (subPhaseFiles.length > 0) {
|
||||
foundPhaseTracking = true;
|
||||
for (const pf of subPhaseFiles) {
|
||||
if (hasUncheckedPhases(fs.readFileSync(pf, 'utf-8'))) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Only return true if we positively confirmed all phases are checked off
|
||||
return foundPhaseTracking;
|
||||
}
|
||||
|
||||
function hasUncheckedPhases(content: string): boolean {
|
||||
const lines = content.split('\n');
|
||||
|
||||
const phaseLines = lines.filter((line) =>
|
||||
line.match(/^[-*]\s*\[[ x]\]/i) && line.toLowerCase().includes('phase')
|
||||
);
|
||||
|
||||
if (phaseLines.length > 0) {
|
||||
return phaseLines.some((line) => line.includes('[ ]'));
|
||||
}
|
||||
|
||||
// No phase-specific checkboxes — check for any unchecked boxes
|
||||
return lines.some((line) => /^[-*]\s*\[ \]/.test(line));
|
||||
}
|
||||
|
||||
function findPhaseFiles(dir: string): string[] {
|
||||
if (!fs.existsSync(dir)) return [];
|
||||
const entries = fs.readdirSync(dir);
|
||||
return entries
|
||||
.filter((name) => /^phase[-_]?\d+.*\.md$/i.test(name))
|
||||
.map((name) => path.join(dir, name));
|
||||
}
|
||||
@@ -0,0 +1,70 @@
|
||||
export type BotMode = 'new-project' | 'enhancement';
|
||||
|
||||
export interface BotConfig {
|
||||
workDir: string;
|
||||
projectDir: string;
|
||||
ideaDescription: string;
|
||||
ideaName: string;
|
||||
mode: BotMode;
|
||||
}
|
||||
|
||||
export type StepName = 'init' | 'plan' | 'document' | 'implement' | 'finalize';
|
||||
|
||||
export interface ToolObservation {
|
||||
toolName: string;
|
||||
input: Record<string, unknown>;
|
||||
output?: unknown;
|
||||
timestamp: number;
|
||||
allowed: boolean;
|
||||
autoAnswered?: boolean;
|
||||
}
|
||||
|
||||
export interface QuestionContext {
|
||||
question: string;
|
||||
options: Array<{ label: string; description: string }>;
|
||||
llmReasoning: string;
|
||||
selectedAnswer: string;
|
||||
timestamp: number;
|
||||
}
|
||||
|
||||
export interface ExecutionObservation {
|
||||
step: StepName;
|
||||
startTime: number;
|
||||
endTime: number;
|
||||
tools: ToolObservation[];
|
||||
assistantMessages: string[];
|
||||
errors: string[];
|
||||
filesCreated: string[];
|
||||
filesModified: string[];
|
||||
questionsAsked: QuestionContext[];
|
||||
}
|
||||
|
||||
export interface EvaluationResult {
|
||||
step: StepName;
|
||||
score: number;
|
||||
strengths: string[];
|
||||
weaknesses: string[];
|
||||
suggestions: string[];
|
||||
criticalIssues: string[];
|
||||
timestamp: number;
|
||||
evaluatorModel: string;
|
||||
reasoning: string;
|
||||
}
|
||||
|
||||
export interface StepResult {
|
||||
step: StepName;
|
||||
success: boolean;
|
||||
sessionId: string | null;
|
||||
duration: number;
|
||||
error: string | null;
|
||||
evaluation?: EvaluationResult;
|
||||
observations?: ExecutionObservation;
|
||||
}
|
||||
|
||||
export interface BotState {
|
||||
config: BotConfig;
|
||||
steps: StepResult[];
|
||||
currentStep: StepName | null;
|
||||
implementPasses: number;
|
||||
allPhasesComplete: boolean;
|
||||
}
|
||||
@@ -0,0 +1,20 @@
|
||||
{
|
||||
"compilerOptions": {
|
||||
"target": "ES2022",
|
||||
"module": "ESNext",
|
||||
"moduleResolution": "bundler",
|
||||
"lib": ["ES2022"],
|
||||
"outDir": "dist",
|
||||
"rootDir": ".",
|
||||
"strict": true,
|
||||
"esModuleInterop": true,
|
||||
"skipLibCheck": true,
|
||||
"forceConsistentCasingInFileNames": true,
|
||||
"resolveJsonModule": true,
|
||||
"declaration": true,
|
||||
"declarationMap": true,
|
||||
"sourceMap": true
|
||||
},
|
||||
"include": ["src/**/*"],
|
||||
"exclude": ["node_modules", "dist"]
|
||||
}
|
||||
@@ -0,0 +1,15 @@
|
||||
import { defineConfig } from 'tsup';
|
||||
|
||||
export default defineConfig({
|
||||
entry: {
|
||||
'bin/plan2code-bot': 'src/bin/plan2code-bot.ts',
|
||||
index: 'src/index.ts',
|
||||
},
|
||||
format: ['esm'],
|
||||
dts: false,
|
||||
clean: true,
|
||||
sourcemap: true,
|
||||
banner: {
|
||||
js: '#!/usr/bin/env node',
|
||||
},
|
||||
});
|
||||
@@ -0,0 +1,7 @@
|
||||
import { defineConfig } from 'vitest/config';
|
||||
|
||||
export default defineConfig({
|
||||
test: {
|
||||
include: ['src/**/*.test.ts'],
|
||||
},
|
||||
});
|
||||
@@ -1,210 +1,212 @@
|
||||
# Plan2Code Loop
|
||||
|
||||
An autonomous CLI tool that implements Plan2Code specs by looping through tasks automatically.
|
||||
|
||||
> **Note:** This is an **alternative** to `/plan2code-3--implement`, not a replacement. Use the manual Step 3 workflow when you want interactive control over each phase, or use this loop when you prefer hands-off autonomous execution.
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
# From the plan2code root directory:
|
||||
|
||||
# Option 1: Install everything (recommended)
|
||||
node install.js # Select option A
|
||||
|
||||
# Option 2: Install loop only
|
||||
node install.js # Select option O
|
||||
|
||||
# Option 3: Manual build and link
|
||||
cd plan2code-loop
|
||||
npm install
|
||||
npm run build
|
||||
npm link
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Simply run the command - everything is interactive:
|
||||
|
||||
```bash
|
||||
plan2code-loop
|
||||
```
|
||||
|
||||
The CLI will:
|
||||
1. Auto-detect specs in `./specs/` directory
|
||||
2. Let you select a spec if multiple are found
|
||||
3. Prompt to continue if an existing session is found
|
||||
4. Ask for JIRA ticket ID, agent selection, loop mode, and max iterations
|
||||
|
||||
## How It Works
|
||||
|
||||
The loop uses an **LLM-driven discovery** approach:
|
||||
|
||||
1. **Spec Selection** - Interactive menu to select from discovered specs
|
||||
2. **Task Discovery** - The AI reads `overview.md` and phase files to find unchecked tasks
|
||||
3. **Implementation** - The AI implements tasks (one per iteration in task mode, or all in a phase in phase mode)
|
||||
4. **Checkbox Update** - The AI marks tasks complete in the markdown file
|
||||
5. **Scratchpad Update** - The AI appends notes to the per-spec scratchpad
|
||||
6. **Completion Marker** - The AI outputs structured markers (e.g., `TASK_COMPLETE: 1.1 - description`)
|
||||
7. **Loop** - Repeat until all tasks done or max iterations reached
|
||||
|
||||
### Loop Modes
|
||||
|
||||
The CLI asks you to choose a loop mode:
|
||||
|
||||
| Mode | Behavior | Git Commits | Best For |
|
||||
|------|----------|-------------|----------|
|
||||
| **One task per loop** (default) | Each agent invocation implements exactly one task | Node controller commits after each task | Smaller models, careful step-by-step execution |
|
||||
| **One phase per loop** | Each agent invocation implements all remaining tasks in the current phase | LLM commits after each task (with JIRA ID if provided) | Smart models with larger context windows, keeping related tasks together |
|
||||
|
||||
### Why LLM-Driven?
|
||||
|
||||
The Node app does NOT parse markdown to find tasks. Instead, the AI reads the spec files directly and decides what to work on. This is:
|
||||
|
||||
- **More flexible** - Works with any reasonable markdown format
|
||||
- **Smarter** - AI can handle edge cases and ambiguity
|
||||
- **Simpler** - Node code is just orchestration, not parsing
|
||||
|
||||
## Completion Markers
|
||||
|
||||
The AI must output one of these markers at the end of each iteration:
|
||||
|
||||
```
|
||||
TASK_COMPLETE: 1.1 - Initialize project structure
|
||||
TASK_BLOCKED: 2.3 - Missing API credentials
|
||||
PHASE_COMPLETE
|
||||
LOOP_COMPLETE
|
||||
```
|
||||
|
||||
| Marker | Meaning |
|
||||
|--------|---------|
|
||||
| `TASK_COMPLETE: X.X - desc` | Task implemented and marked complete |
|
||||
| `TASK_BLOCKED: X.X - reason` | Cannot complete task (explains why) |
|
||||
| `PHASE_COMPLETE` | Current phase finished (phase mode only) |
|
||||
| `LOOP_COMPLETE` | All phases finished |
|
||||
|
||||
In **phase mode**, the AI outputs multiple `TASK_COMPLETE` markers (one per task) within a single iteration, followed by `PHASE_COMPLETE` or `LOOP_COMPLETE`.
|
||||
|
||||
## Session Files
|
||||
|
||||
Session state is stored **per-spec** inside the spec directory:
|
||||
|
||||
```
|
||||
specs/my-feature/
|
||||
├── overview.md
|
||||
├── phase-1.md
|
||||
├── phase-2.md
|
||||
└── .plan2code-loop/ # Per-spec session state
|
||||
├── config.json # Session configuration
|
||||
├── scratchpad.md # LLM-managed progress notes
|
||||
├── iteration.log # JSON log of each iteration
|
||||
└── spec.hash # Hash for detecting spec changes
|
||||
```
|
||||
|
||||
The scratchpad is managed by the LLM itself - after each task, the AI appends notes about what was done, decisions made, and files changed. This helps future iterations skip exploration.
|
||||
|
||||
## Supported Agents
|
||||
|
||||
| Agent | Status |
|
||||
|-------|--------|
|
||||
| Claude Code | Supported |
|
||||
| GitHub Copilot CLI | Supported |
|
||||
|
||||
The loop uses your configured default model for each agent.
|
||||
|
||||
## Example Session
|
||||
|
||||
```
|
||||
$ plan2code-loop
|
||||
|
||||
╭──────────────────────────────────────╮
|
||||
│ │
|
||||
│ 🔮 Plany's Loop │
|
||||
│ Autonomous Implementation │
|
||||
│ │
|
||||
╰──────────────────────────────────────╯
|
||||
|
||||
Found spec: specs/todo-app
|
||||
Feature: Todo App
|
||||
Phases: 0/3
|
||||
Tasks: 0/15
|
||||
|
||||
══════════════════════════════════════════════
|
||||
Spec: Todo App
|
||||
══════════════════════════════════════════════
|
||||
Phases: 0/3
|
||||
Tasks: 0/15
|
||||
|
||||
? JIRA Ticket ID (optional): PROJ-123
|
||||
? Select AI agent: Claude Code
|
||||
? Tasks per loop iteration: One task per loop (default)
|
||||
? Maximum iterations: 100
|
||||
|
||||
Starting Plan2Code Loop
|
||||
══════════════════════════════════════════════
|
||||
Agent: Claude Code
|
||||
Model: default
|
||||
Spec: C:\projects\my-app\specs\todo-app
|
||||
Loop mode: One task per loop
|
||||
Max iterations: 100
|
||||
|
||||
Iteration 1/100
|
||||
[1/100] Task 1.1: Initialize project with Vite (45s)
|
||||
✓ Completed: Task 1.1: Initialize project with Vite
|
||||
|
||||
Iteration 2/100
|
||||
[2/100] Task 1.2: Configure TypeScript (32s)
|
||||
✓ Completed: Task 1.2: Configure TypeScript
|
||||
|
||||
...
|
||||
|
||||
Session Summary
|
||||
══════════════════════════════════════════════
|
||||
Total iterations: 12
|
||||
Tasks completed: 12
|
||||
Exit reason: all_complete
|
||||
|
||||
Session files saved to specs/todo-app/.plan2code-loop:
|
||||
- config.json (session configuration)
|
||||
- scratchpad.md (LLM-managed notes)
|
||||
- iteration.log (history)
|
||||
```
|
||||
|
||||
## Development
|
||||
|
||||
```bash
|
||||
# Install dependencies
|
||||
npm install
|
||||
|
||||
# Build
|
||||
npm run build
|
||||
|
||||
# Link for global usage
|
||||
npm link
|
||||
|
||||
# Unlink
|
||||
npm unlink
|
||||
```
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
src/
|
||||
├── bin/plan2code-loop.ts # CLI entry point
|
||||
├── index.ts # Main exports
|
||||
├── controller.ts # Loop orchestration
|
||||
├── cli.ts # Interactive prompts
|
||||
├── agents/ # Agent implementations
|
||||
│ ├── claude-code.ts
|
||||
│ └── copilot-cli.ts
|
||||
├── prompt/ # Prompt building
|
||||
│ ├── templates.ts
|
||||
│ └── builder.ts
|
||||
├── state/ # Session state management
|
||||
│ └── manager.ts
|
||||
├── spec/ # Spec utilities (for CLI display)
|
||||
│ └── utils.ts
|
||||
└── utils/ # Utilities
|
||||
├── completion.ts # Marker parsing
|
||||
└── logger.ts
|
||||
```
|
||||
# Plan2Code Loop
|
||||
|
||||
An autonomous CLI tool that implements Plan2Code specs by looping through tasks automatically.
|
||||
|
||||
> **Note:** This is an **alternative** to `/plan2code-3-implement`, not a replacement. Use the manual Step 3 workflow when you want interactive control over each phase, or use this loop when you prefer hands-off autonomous execution.
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
# From the plan2code root directory:
|
||||
|
||||
# Option 1: Install everything (recommended)
|
||||
node install.js # Select option A
|
||||
|
||||
# Option 2: Install loop only
|
||||
node install.js # Select option O
|
||||
|
||||
# Option 3: Manual build and link
|
||||
cd plan2code-loop
|
||||
npm install
|
||||
npm run build
|
||||
npm link
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Simply run the command - everything is interactive:
|
||||
|
||||
```bash
|
||||
plan2code-loop
|
||||
```
|
||||
|
||||
The CLI will:
|
||||
1. Auto-detect specs in `./specs/` directory
|
||||
2. Let you select a spec if multiple are found
|
||||
3. Prompt to continue if an existing session is found
|
||||
4. Ask for JIRA ticket ID, agent selection, loop mode, and max iterations
|
||||
|
||||
## How It Works
|
||||
|
||||
The loop uses an **LLM-driven discovery** approach:
|
||||
|
||||
1. **Spec Selection** - Interactive menu to select from discovered specs
|
||||
2. **Task Discovery** - The AI reads `overview.md` and phase files to find unchecked tasks
|
||||
3. **Implementation** - The AI implements tasks (one per iteration in task mode, or all in a phase in phase mode)
|
||||
4. **Checkbox Update** - The AI marks tasks complete in the markdown file
|
||||
5. **Scratchpad Update** - The AI appends notes to the per-spec scratchpad
|
||||
6. **Completion Marker** - The AI outputs structured markers (e.g., `TASK_COMPLETE: 1.1 - description`)
|
||||
7. **Loop** - Repeat until all tasks done or max iterations reached
|
||||
|
||||
### Loop Modes
|
||||
|
||||
The CLI asks you to choose a loop mode:
|
||||
|
||||
| Mode | Behavior | Git Commits | Best For |
|
||||
|------|----------|-------------|----------|
|
||||
| **One task per loop** (default) | Each agent invocation implements exactly one task | Node controller commits after each task | Smaller models, careful step-by-step execution |
|
||||
| **One phase per loop** | Each agent invocation implements all remaining tasks in the current phase | LLM commits after each task (with JIRA ID if provided) | Smart models with larger context windows, keeping related tasks together |
|
||||
|
||||
### Why LLM-Driven?
|
||||
|
||||
The Node app does NOT parse markdown to find tasks. Instead, the AI reads the spec files directly and decides what to work on. This is:
|
||||
|
||||
- **More flexible** - Works with any reasonable markdown format
|
||||
- **Smarter** - AI can handle edge cases and ambiguity
|
||||
- **Simpler** - Node code is just orchestration, not parsing
|
||||
|
||||
## Completion Markers
|
||||
|
||||
The AI must output one of these markers at the end of each iteration:
|
||||
|
||||
```
|
||||
TASK_COMPLETE: 1.1 - Initialize project structure
|
||||
TASK_BLOCKED: 2.3 - Missing API credentials
|
||||
PHASE_COMPLETE
|
||||
LOOP_COMPLETE
|
||||
```
|
||||
|
||||
| Marker | Meaning |
|
||||
|--------|---------|
|
||||
| `TASK_COMPLETE: X.X - desc` | Task implemented and marked complete |
|
||||
| `TASK_BLOCKED: X.X - reason` | Cannot complete task (explains why) |
|
||||
| `PHASE_COMPLETE` | Current phase finished (phase mode only) |
|
||||
| `LOOP_COMPLETE` | All phases finished |
|
||||
|
||||
In **phase mode**, the AI outputs multiple `TASK_COMPLETE` markers (one per task) within a single iteration, followed by `PHASE_COMPLETE` or `LOOP_COMPLETE`.
|
||||
|
||||
## Session Files
|
||||
|
||||
Session state is stored **per-spec** inside the spec directory:
|
||||
|
||||
```
|
||||
specs/my-feature/
|
||||
├── overview.md
|
||||
├── phase-1.md
|
||||
├── phase-2.md
|
||||
└── .plan2code-loop/ # Per-spec session state
|
||||
├── config.json # Session configuration
|
||||
├── scratchpad.md # LLM-managed progress notes
|
||||
├── iteration.log # JSON log of each iteration
|
||||
└── spec.hash # Hash for detecting spec changes
|
||||
```
|
||||
|
||||
The scratchpad is managed by the LLM itself - after each task, the AI appends notes about what was done, decisions made, and files changed. This helps future iterations skip exploration.
|
||||
|
||||
## Supported Agents
|
||||
|
||||
| Agent | Status |
|
||||
|-------|--------|
|
||||
| Claude Code | Supported |
|
||||
| GitHub Copilot CLI | Supported |
|
||||
| Devin CLI | Supported |
|
||||
|
||||
The loop uses your configured default model for each agent.
|
||||
|
||||
## Example Session
|
||||
|
||||
```
|
||||
$ plan2code-loop
|
||||
|
||||
╭──────────────────────────────────────╮
|
||||
│ │
|
||||
│ 🔮 Planny's Loop │
|
||||
│ Autonomous Implementation │
|
||||
│ │
|
||||
╰──────────────────────────────────────╯
|
||||
|
||||
Found spec: specs/todo-app
|
||||
Feature: Todo App
|
||||
Phases: 0/3
|
||||
Tasks: 0/15
|
||||
|
||||
══════════════════════════════════════════════
|
||||
Spec: Todo App
|
||||
══════════════════════════════════════════════
|
||||
Phases: 0/3
|
||||
Tasks: 0/15
|
||||
|
||||
? JIRA Ticket ID (optional): PROJ-123
|
||||
? Select AI agent: Claude Code
|
||||
? Tasks per loop iteration: One task per loop (default)
|
||||
? Maximum iterations: 100
|
||||
|
||||
Starting Plan2Code Loop
|
||||
══════════════════════════════════════════════
|
||||
Agent: Claude Code
|
||||
Model: default
|
||||
Spec: C:\projects\my-app\specs\todo-app
|
||||
Loop mode: One task per loop
|
||||
Max iterations: 100
|
||||
|
||||
Iteration 1/100
|
||||
[1/100] Task 1.1: Initialize project with Vite (45s)
|
||||
✓ Completed: Task 1.1: Initialize project with Vite
|
||||
|
||||
Iteration 2/100
|
||||
[2/100] Task 1.2: Configure TypeScript (32s)
|
||||
✓ Completed: Task 1.2: Configure TypeScript
|
||||
|
||||
...
|
||||
|
||||
Session Summary
|
||||
══════════════════════════════════════════════
|
||||
Total iterations: 12
|
||||
Tasks completed: 12
|
||||
Exit reason: all_complete
|
||||
|
||||
Session files saved to specs/todo-app/.plan2code-loop:
|
||||
- config.json (session configuration)
|
||||
- scratchpad.md (LLM-managed notes)
|
||||
- iteration.log (history)
|
||||
```
|
||||
|
||||
## Development
|
||||
|
||||
```bash
|
||||
# Install dependencies
|
||||
npm install
|
||||
|
||||
# Build
|
||||
npm run build
|
||||
|
||||
# Link for global usage
|
||||
npm link
|
||||
|
||||
# Unlink
|
||||
npm unlink
|
||||
```
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
src/
|
||||
├── bin/plan2code-loop.ts # CLI entry point
|
||||
├── index.ts # Main exports
|
||||
├── controller.ts # Loop orchestration
|
||||
├── cli.ts # Interactive prompts
|
||||
├── agents/ # Agent implementations
|
||||
│ ├── claude-code.ts
|
||||
│ ├── copilot-cli.ts
|
||||
│ └── devin-cli.ts
|
||||
├── prompt/ # Prompt building
|
||||
│ ├── templates.ts
|
||||
│ └── builder.ts
|
||||
├── state/ # Session state management
|
||||
│ └── manager.ts
|
||||
├── spec/ # Spec utilities (for CLI display)
|
||||
│ └── utils.ts
|
||||
└── utils/ # Utilities
|
||||
├── completion.ts # Marker parsing
|
||||
└── logger.ts
|
||||
```
|
||||
|
||||
@@ -1,46 +1,47 @@
|
||||
{
|
||||
"name": "plan2code-loop",
|
||||
"version": "1.6.2",
|
||||
"description": "Plan2Code Loop - Autonomous spec-driven implementation CLI",
|
||||
"type": "module",
|
||||
"main": "dist/index.js",
|
||||
"bin": {
|
||||
"plan2code-loop": "./dist/bin/plan2code-loop.js"
|
||||
},
|
||||
"files": [
|
||||
"dist"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=18.0.0"
|
||||
},
|
||||
"keywords": [
|
||||
"cli",
|
||||
"ai",
|
||||
"agent",
|
||||
"automation",
|
||||
"claude",
|
||||
"copilot",
|
||||
"spec-driven"
|
||||
],
|
||||
"license": "MIT",
|
||||
"scripts": {
|
||||
"build": "tsup",
|
||||
"dev": "tsup --watch",
|
||||
"start": "node dist/bin/plan2code-loop.js",
|
||||
"prepublishOnly": "npm run build"
|
||||
},
|
||||
"dependencies": {
|
||||
"@inquirer/prompts": "^8.1.0",
|
||||
"chalk": "^5.6.2",
|
||||
"execa": "^9.6.1",
|
||||
"fs-extra": "^11.3.3",
|
||||
"ora": "^9.0.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
"@types/fs-extra": "^11.0.4",
|
||||
"@types/node": "^25.0.3",
|
||||
"tsup": "^8.5.1",
|
||||
"typescript": "^5.9.3"
|
||||
}
|
||||
}
|
||||
|
||||
{
|
||||
"name": "plan2code-loop",
|
||||
"version": "1.6.2",
|
||||
"description": "Plan2Code Loop - Autonomous spec-driven implementation CLI",
|
||||
"type": "module",
|
||||
"main": "dist/index.js",
|
||||
"bin": {
|
||||
"plan2code-loop": "./dist/bin/plan2code-loop.js"
|
||||
},
|
||||
"files": [
|
||||
"dist"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=18.0.0"
|
||||
},
|
||||
"keywords": [
|
||||
"cli",
|
||||
"ai",
|
||||
"agent",
|
||||
"automation",
|
||||
"claude",
|
||||
"copilot",
|
||||
"devin",
|
||||
"spec-driven"
|
||||
],
|
||||
"license": "MIT",
|
||||
"scripts": {
|
||||
"build": "tsup",
|
||||
"dev": "tsup --watch",
|
||||
"start": "node dist/bin/plan2code-loop.js",
|
||||
"prepublishOnly": "npm run build"
|
||||
},
|
||||
"dependencies": {
|
||||
"@inquirer/prompts": "^8.1.0",
|
||||
"chalk": "^5.6.2",
|
||||
"execa": "^9.6.1",
|
||||
"fs-extra": "^11.3.3",
|
||||
"ora": "^9.0.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
"@types/fs-extra": "^11.0.4",
|
||||
"@types/node": "^25.0.3",
|
||||
"tsup": "^8.5.1",
|
||||
"typescript": "^5.9.3"
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -1,82 +1,82 @@
|
||||
import type { Agent, AgentConfig, AgentExecutionOptions, AgentExecutionResult } from './types.js';
|
||||
import { executeCommand } from '../utils/process.js';
|
||||
import { writeFileSync, unlinkSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { tmpdir } from 'os';
|
||||
|
||||
const claudeCodeConfig: AgentConfig = {
|
||||
name: 'claude-code',
|
||||
displayName: 'Claude Code',
|
||||
command: 'claude',
|
||||
models: [
|
||||
{ value: 'default', label: 'Default (use Claude config)' },
|
||||
],
|
||||
defaultModel: 'default',
|
||||
flags: {
|
||||
prompt: '--print',
|
||||
model: '--model',
|
||||
skipPermissions: '--dangerously-skip-permissions',
|
||||
},
|
||||
};
|
||||
|
||||
class ClaudeCodeAgent implements Agent {
|
||||
readonly config = claudeCodeConfig;
|
||||
|
||||
async execute(options: AgentExecutionOptions): Promise<AgentExecutionResult> {
|
||||
// Write prompt to temp file - more reliable than stdin on Windows
|
||||
const tempFile = join(tmpdir(), `plan2code-prompt-${Date.now()}.txt`);
|
||||
writeFileSync(tempFile, options.prompt, 'utf-8');
|
||||
|
||||
try {
|
||||
// Build args: flags first, then read prompt from temp file via shell
|
||||
const args: string[] = [
|
||||
this.config.flags.prompt, // --print for non-interactive mode
|
||||
this.config.flags.skipPermissions,
|
||||
];
|
||||
|
||||
// Only add --model if not using default
|
||||
if (options.model && options.model !== 'default') {
|
||||
args.push(this.config.flags.model, options.model);
|
||||
}
|
||||
|
||||
// Use stdin from the temp file
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args,
|
||||
cwd: options.cwd,
|
||||
timeout: options.timeout,
|
||||
signal: options.signal,
|
||||
stdinFile: tempFile,
|
||||
});
|
||||
|
||||
return {
|
||||
stdout: result.stdout,
|
||||
stderr: result.stderr,
|
||||
exitCode: result.exitCode,
|
||||
timedOut: result.timedOut,
|
||||
cancelled: result.cancelled,
|
||||
duration: result.duration,
|
||||
};
|
||||
} finally {
|
||||
// Clean up temp file
|
||||
try {
|
||||
unlinkSync(tempFile);
|
||||
} catch {
|
||||
// Ignore cleanup errors
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
async isAvailable(): Promise<boolean> {
|
||||
// Run claude --version to verify it's actually installed and working
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args: ['--version'],
|
||||
cwd: process.cwd(),
|
||||
timeout: 5000,
|
||||
});
|
||||
return result.exitCode === 0;
|
||||
}
|
||||
}
|
||||
|
||||
export const claudeCodeAgent = new ClaudeCodeAgent();
|
||||
import type { Agent, AgentConfig, AgentExecutionOptions, AgentExecutionResult } from './types.js';
|
||||
import { executeCommand } from '../utils/process.js';
|
||||
import { writeFileSync, unlinkSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { tmpdir } from 'os';
|
||||
|
||||
const claudeCodeConfig: AgentConfig = {
|
||||
name: 'claude-code',
|
||||
displayName: 'Claude Code',
|
||||
command: 'claude',
|
||||
models: [
|
||||
{ value: 'default', label: 'Default (use Claude config)' },
|
||||
],
|
||||
defaultModel: 'default',
|
||||
flags: {
|
||||
prompt: '--print',
|
||||
model: '--model',
|
||||
skipPermissions: '--dangerously-skip-permissions',
|
||||
},
|
||||
};
|
||||
|
||||
class ClaudeCodeAgent implements Agent {
|
||||
readonly config = claudeCodeConfig;
|
||||
|
||||
async execute(options: AgentExecutionOptions): Promise<AgentExecutionResult> {
|
||||
// Write prompt to temp file - more reliable than stdin on Windows
|
||||
const tempFile = join(tmpdir(), `plan2code-prompt-${Date.now()}.txt`);
|
||||
writeFileSync(tempFile, options.prompt, 'utf-8');
|
||||
|
||||
try {
|
||||
// Build args: flags first, then read prompt from temp file via shell
|
||||
const args: string[] = [
|
||||
this.config.flags.prompt, // --print for non-interactive mode
|
||||
this.config.flags.skipPermissions,
|
||||
];
|
||||
|
||||
// Only add --model if not using default
|
||||
if (options.model && options.model !== 'default') {
|
||||
args.push(this.config.flags.model, options.model);
|
||||
}
|
||||
|
||||
// Use stdin from the temp file
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args,
|
||||
cwd: options.cwd,
|
||||
timeout: options.timeout,
|
||||
signal: options.signal,
|
||||
stdinFile: tempFile,
|
||||
});
|
||||
|
||||
return {
|
||||
stdout: result.stdout,
|
||||
stderr: result.stderr,
|
||||
exitCode: result.exitCode,
|
||||
timedOut: result.timedOut,
|
||||
cancelled: result.cancelled,
|
||||
duration: result.duration,
|
||||
};
|
||||
} finally {
|
||||
// Clean up temp file
|
||||
try {
|
||||
unlinkSync(tempFile);
|
||||
} catch {
|
||||
// Ignore cleanup errors
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
async isAvailable(): Promise<boolean> {
|
||||
// Run claude --version to verify it's actually installed and working
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args: ['--version'],
|
||||
cwd: process.cwd(),
|
||||
timeout: 5000,
|
||||
});
|
||||
return result.exitCode === 0;
|
||||
}
|
||||
}
|
||||
|
||||
export const claudeCodeAgent = new ClaudeCodeAgent();
|
||||
|
||||
@@ -1,70 +1,70 @@
|
||||
import type { Agent, AgentConfig, AgentExecutionOptions, AgentExecutionResult } from './types.js';
|
||||
import { executeCommand } from '../utils/process.js';
|
||||
|
||||
const copilotCliConfig: AgentConfig = {
|
||||
name: 'copilot-cli',
|
||||
displayName: 'GitHub Copilot CLI',
|
||||
command: 'copilot',
|
||||
models: [
|
||||
{ value: 'claude-sonnet-4', label: 'Claude Sonnet 4 (Default)' },
|
||||
{ value: 'claude-sonnet-4.5', label: 'Claude Sonnet 4.5' },
|
||||
{ value: 'claude-opus-4.5', label: 'Claude Opus 4.5' },
|
||||
{ value: 'gpt-5', label: 'GPT-5' },
|
||||
{ value: 'gpt-5-mini', label: 'GPT-5 Mini' },
|
||||
{ value: 'gemini-3-pro-preview', label: 'Gemini 3 Pro' },
|
||||
],
|
||||
defaultModel: 'claude-sonnet-4',
|
||||
flags: {
|
||||
prompt: '-p',
|
||||
model: '--model',
|
||||
skipPermissions: '--allow-all-tools',
|
||||
silent: '-s',
|
||||
},
|
||||
};
|
||||
|
||||
class CopilotCliAgent implements Agent {
|
||||
readonly config = copilotCliConfig;
|
||||
|
||||
async execute(options: AgentExecutionOptions): Promise<AgentExecutionResult> {
|
||||
// Use stdin for prompt to handle multi-line text properly
|
||||
const args: string[] = [];
|
||||
|
||||
// Only add --model if not using default
|
||||
if (options.model && options.model !== 'default') {
|
||||
args.push(this.config.flags.model, options.model);
|
||||
}
|
||||
|
||||
args.push(this.config.flags.skipPermissions, this.config.flags.silent!);
|
||||
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args,
|
||||
cwd: options.cwd,
|
||||
timeout: options.timeout,
|
||||
signal: options.signal,
|
||||
stdin: options.prompt,
|
||||
});
|
||||
|
||||
return {
|
||||
stdout: result.stdout,
|
||||
stderr: result.stderr,
|
||||
exitCode: result.exitCode,
|
||||
timedOut: result.timedOut,
|
||||
cancelled: result.cancelled,
|
||||
duration: result.duration,
|
||||
};
|
||||
}
|
||||
|
||||
async isAvailable(): Promise<boolean> {
|
||||
// Run copilot --version to verify it's installed
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args: ['--version'],
|
||||
cwd: process.cwd(),
|
||||
timeout: 5000,
|
||||
});
|
||||
return result.exitCode === 0;
|
||||
}
|
||||
}
|
||||
|
||||
export const copilotCliAgent = new CopilotCliAgent();
|
||||
import type { Agent, AgentConfig, AgentExecutionOptions, AgentExecutionResult } from './types.js';
|
||||
import { executeCommand } from '../utils/process.js';
|
||||
|
||||
const copilotCliConfig: AgentConfig = {
|
||||
name: 'copilot-cli',
|
||||
displayName: 'GitHub Copilot CLI',
|
||||
command: 'copilot',
|
||||
models: [
|
||||
{ value: 'claude-sonnet-4', label: 'Claude Sonnet 4 (Default)' },
|
||||
{ value: 'claude-sonnet-4.5', label: 'Claude Sonnet 4.5' },
|
||||
{ value: 'claude-opus-4.5', label: 'Claude Opus 4.5' },
|
||||
{ value: 'gpt-5', label: 'GPT-5' },
|
||||
{ value: 'gpt-5-mini', label: 'GPT-5 Mini' },
|
||||
{ value: 'gemini-3-pro-preview', label: 'Gemini 3 Pro' },
|
||||
],
|
||||
defaultModel: 'claude-sonnet-4',
|
||||
flags: {
|
||||
prompt: '-p',
|
||||
model: '--model',
|
||||
skipPermissions: '--allow-all-tools',
|
||||
silent: '-s',
|
||||
},
|
||||
};
|
||||
|
||||
class CopilotCliAgent implements Agent {
|
||||
readonly config = copilotCliConfig;
|
||||
|
||||
async execute(options: AgentExecutionOptions): Promise<AgentExecutionResult> {
|
||||
// Use stdin for prompt to handle multi-line text properly
|
||||
const args: string[] = [];
|
||||
|
||||
// Only add --model if not using default
|
||||
if (options.model && options.model !== 'default') {
|
||||
args.push(this.config.flags.model, options.model);
|
||||
}
|
||||
|
||||
args.push(this.config.flags.skipPermissions, this.config.flags.silent!);
|
||||
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args,
|
||||
cwd: options.cwd,
|
||||
timeout: options.timeout,
|
||||
signal: options.signal,
|
||||
stdin: options.prompt,
|
||||
});
|
||||
|
||||
return {
|
||||
stdout: result.stdout,
|
||||
stderr: result.stderr,
|
||||
exitCode: result.exitCode,
|
||||
timedOut: result.timedOut,
|
||||
cancelled: result.cancelled,
|
||||
duration: result.duration,
|
||||
};
|
||||
}
|
||||
|
||||
async isAvailable(): Promise<boolean> {
|
||||
// Run copilot --version to verify it's installed
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args: ['--version'],
|
||||
cwd: process.cwd(),
|
||||
timeout: 5000,
|
||||
});
|
||||
return result.exitCode === 0;
|
||||
}
|
||||
}
|
||||
|
||||
export const copilotCliAgent = new CopilotCliAgent();
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
import type { Agent, AgentConfig, AgentExecutionOptions, AgentExecutionResult } from './types.js';
|
||||
import { executeCommand } from '../utils/process.js';
|
||||
import { writeFileSync, unlinkSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { tmpdir } from 'os';
|
||||
|
||||
const devinCliConfig: AgentConfig = {
|
||||
name: 'devin-cli',
|
||||
displayName: 'Devin CLI',
|
||||
command: 'devin',
|
||||
models: [
|
||||
{ value: 'default', label: 'Default (use Devin config)' },
|
||||
],
|
||||
defaultModel: 'default',
|
||||
flags: {
|
||||
prompt: '--print',
|
||||
promptFile: '--prompt-file',
|
||||
model: '--model',
|
||||
skipPermissions: '--permission-mode',
|
||||
},
|
||||
};
|
||||
|
||||
class DevinCliAgent implements Agent {
|
||||
readonly config = devinCliConfig;
|
||||
|
||||
async execute(options: AgentExecutionOptions): Promise<AgentExecutionResult> {
|
||||
// Devin CLI takes the prompt via --prompt-file rather than stdin
|
||||
const tempFile = join(tmpdir(), `plan2code-prompt-${Date.now()}.txt`);
|
||||
writeFileSync(tempFile, options.prompt, 'utf-8');
|
||||
|
||||
try {
|
||||
const args: string[] = [
|
||||
this.config.flags.prompt, // --print for non-interactive mode
|
||||
this.config.flags.promptFile!, tempFile, // --prompt-file <path>
|
||||
this.config.flags.skipPermissions, 'dangerous', // --permission-mode dangerous (auto-approve all tools)
|
||||
];
|
||||
|
||||
// Only add --model if not using default
|
||||
if (options.model && options.model !== 'default') {
|
||||
args.push(this.config.flags.model, options.model);
|
||||
}
|
||||
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args,
|
||||
cwd: options.cwd,
|
||||
timeout: options.timeout,
|
||||
signal: options.signal,
|
||||
});
|
||||
|
||||
return {
|
||||
stdout: result.stdout,
|
||||
stderr: result.stderr,
|
||||
exitCode: result.exitCode,
|
||||
timedOut: result.timedOut,
|
||||
cancelled: result.cancelled,
|
||||
duration: result.duration,
|
||||
};
|
||||
} finally {
|
||||
// Clean up temp file
|
||||
try {
|
||||
unlinkSync(tempFile);
|
||||
} catch {
|
||||
// Ignore cleanup errors
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
async isAvailable(): Promise<boolean> {
|
||||
// Run devin --version to verify it's actually installed and working
|
||||
const result = await executeCommand({
|
||||
command: this.config.command,
|
||||
args: ['--version'],
|
||||
cwd: process.cwd(),
|
||||
timeout: 5000,
|
||||
});
|
||||
return result.exitCode === 0;
|
||||
}
|
||||
}
|
||||
|
||||
export const devinCliAgent = new DevinCliAgent();
|
||||
@@ -1,19 +1,22 @@
|
||||
export type {
|
||||
Agent,
|
||||
AgentConfig,
|
||||
AgentExecutionOptions,
|
||||
AgentExecutionResult,
|
||||
ModelOption,
|
||||
} from './types.js';
|
||||
|
||||
export { agentRegistry } from './registry.js';
|
||||
export { claudeCodeAgent } from './claude-code.js';
|
||||
export { copilotCliAgent } from './copilot-cli.js';
|
||||
|
||||
// Register all agents
|
||||
import { agentRegistry } from './registry.js';
|
||||
import { claudeCodeAgent } from './claude-code.js';
|
||||
import { copilotCliAgent } from './copilot-cli.js';
|
||||
|
||||
agentRegistry.register(claudeCodeAgent);
|
||||
agentRegistry.register(copilotCliAgent);
|
||||
export type {
|
||||
Agent,
|
||||
AgentConfig,
|
||||
AgentExecutionOptions,
|
||||
AgentExecutionResult,
|
||||
ModelOption,
|
||||
} from './types.js';
|
||||
|
||||
export { agentRegistry } from './registry.js';
|
||||
export { claudeCodeAgent } from './claude-code.js';
|
||||
export { copilotCliAgent } from './copilot-cli.js';
|
||||
export { devinCliAgent } from './devin-cli.js';
|
||||
|
||||
// Register all agents
|
||||
import { agentRegistry } from './registry.js';
|
||||
import { claudeCodeAgent } from './claude-code.js';
|
||||
import { copilotCliAgent } from './copilot-cli.js';
|
||||
import { devinCliAgent } from './devin-cli.js';
|
||||
|
||||
agentRegistry.register(claudeCodeAgent);
|
||||
agentRegistry.register(copilotCliAgent);
|
||||
agentRegistry.register(devinCliAgent);
|
||||
|
||||
@@ -1,34 +1,34 @@
|
||||
import type { Agent } from './types.js';
|
||||
|
||||
class AgentRegistry {
|
||||
private agents: Map<string, Agent> = new Map();
|
||||
|
||||
register(agent: Agent): void {
|
||||
this.agents.set(agent.config.name, agent);
|
||||
}
|
||||
|
||||
get(name: string): Agent | undefined {
|
||||
return this.agents.get(name);
|
||||
}
|
||||
|
||||
getAll(): Agent[] {
|
||||
return Array.from(this.agents.values());
|
||||
}
|
||||
|
||||
getAvailable(): Promise<Agent[]> {
|
||||
return Promise.all(
|
||||
this.getAll().map(async (agent) => ({
|
||||
agent,
|
||||
available: await agent.isAvailable(),
|
||||
}))
|
||||
).then((results) =>
|
||||
results.filter((r) => r.available).map((r) => r.agent)
|
||||
);
|
||||
}
|
||||
|
||||
getNames(): string[] {
|
||||
return Array.from(this.agents.keys());
|
||||
}
|
||||
}
|
||||
|
||||
export const agentRegistry = new AgentRegistry();
|
||||
import type { Agent } from './types.js';
|
||||
|
||||
class AgentRegistry {
|
||||
private agents: Map<string, Agent> = new Map();
|
||||
|
||||
register(agent: Agent): void {
|
||||
this.agents.set(agent.config.name, agent);
|
||||
}
|
||||
|
||||
get(name: string): Agent | undefined {
|
||||
return this.agents.get(name);
|
||||
}
|
||||
|
||||
getAll(): Agent[] {
|
||||
return Array.from(this.agents.values());
|
||||
}
|
||||
|
||||
getAvailable(): Promise<Agent[]> {
|
||||
return Promise.all(
|
||||
this.getAll().map(async (agent) => ({
|
||||
agent,
|
||||
available: await agent.isAvailable(),
|
||||
}))
|
||||
).then((results) =>
|
||||
results.filter((r) => r.available).map((r) => r.agent)
|
||||
);
|
||||
}
|
||||
|
||||
getNames(): string[] {
|
||||
return Array.from(this.agents.keys());
|
||||
}
|
||||
}
|
||||
|
||||
export const agentRegistry = new AgentRegistry();
|
||||
|
||||
@@ -1,42 +1,43 @@
|
||||
export interface ModelOption {
|
||||
value: string;
|
||||
label: string;
|
||||
}
|
||||
|
||||
export interface AgentConfig {
|
||||
name: string;
|
||||
displayName: string;
|
||||
command: string;
|
||||
models: ModelOption[];
|
||||
defaultModel: string;
|
||||
flags: {
|
||||
prompt: string;
|
||||
model: string;
|
||||
skipPermissions: string;
|
||||
silent?: string;
|
||||
};
|
||||
}
|
||||
|
||||
export interface AgentExecutionOptions {
|
||||
prompt: string;
|
||||
model: string;
|
||||
timeout: number; // milliseconds
|
||||
verbose: boolean;
|
||||
cwd: string;
|
||||
signal?: AbortSignal; // For cancellation
|
||||
}
|
||||
|
||||
export interface AgentExecutionResult {
|
||||
stdout: string;
|
||||
stderr: string;
|
||||
exitCode: number;
|
||||
timedOut: boolean;
|
||||
cancelled: boolean;
|
||||
duration: number; // milliseconds
|
||||
}
|
||||
|
||||
export interface Agent {
|
||||
config: AgentConfig;
|
||||
execute(options: AgentExecutionOptions): Promise<AgentExecutionResult>;
|
||||
isAvailable(): Promise<boolean>;
|
||||
}
|
||||
export interface ModelOption {
|
||||
value: string;
|
||||
label: string;
|
||||
}
|
||||
|
||||
export interface AgentConfig {
|
||||
name: string;
|
||||
displayName: string;
|
||||
command: string;
|
||||
models: ModelOption[];
|
||||
defaultModel: string;
|
||||
flags: {
|
||||
prompt: string;
|
||||
model: string;
|
||||
skipPermissions: string;
|
||||
silent?: string;
|
||||
promptFile?: string;
|
||||
};
|
||||
}
|
||||
|
||||
export interface AgentExecutionOptions {
|
||||
prompt: string;
|
||||
model: string;
|
||||
timeout: number; // milliseconds
|
||||
verbose: boolean;
|
||||
cwd: string;
|
||||
signal?: AbortSignal; // For cancellation
|
||||
}
|
||||
|
||||
export interface AgentExecutionResult {
|
||||
stdout: string;
|
||||
stderr: string;
|
||||
exitCode: number;
|
||||
timedOut: boolean;
|
||||
cancelled: boolean;
|
||||
duration: number; // milliseconds
|
||||
}
|
||||
|
||||
export interface Agent {
|
||||
config: AgentConfig;
|
||||
execute(options: AgentExecutionOptions): Promise<AgentExecutionResult>;
|
||||
isAvailable(): Promise<boolean>;
|
||||
}
|
||||
|
||||
@@ -1,34 +1,34 @@
|
||||
import { run } from '../index.js';
|
||||
import { logger } from '../utils/index.js';
|
||||
|
||||
async function main() {
|
||||
try {
|
||||
const result = await run();
|
||||
|
||||
if (!result) {
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
// Exit codes per spec
|
||||
switch (result.exitReason) {
|
||||
case 'all_complete':
|
||||
logger.success('Loop completed successfully - all tasks done!');
|
||||
process.exit(0);
|
||||
case 'max_iterations':
|
||||
logger.warning('Loop ended: max iterations reached');
|
||||
process.exit(1);
|
||||
case 'interrupted':
|
||||
logger.info('Loop interrupted by user');
|
||||
process.exit(2);
|
||||
case 'error':
|
||||
logger.error('Loop ended with error');
|
||||
process.exit(3);
|
||||
}
|
||||
|
||||
} catch (error) {
|
||||
logger.error(error instanceof Error ? error.message : String(error));
|
||||
process.exit(3);
|
||||
}
|
||||
}
|
||||
|
||||
main();
|
||||
import { run } from '../index.js';
|
||||
import { logger } from '../utils/index.js';
|
||||
|
||||
async function main() {
|
||||
try {
|
||||
const result = await run();
|
||||
|
||||
if (!result) {
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
// Exit codes per spec
|
||||
switch (result.exitReason) {
|
||||
case 'all_complete':
|
||||
logger.success('Loop completed successfully - all tasks done!');
|
||||
process.exit(0);
|
||||
case 'max_iterations':
|
||||
logger.warning('Loop ended: max iterations reached');
|
||||
process.exit(1);
|
||||
case 'interrupted':
|
||||
logger.info('Loop interrupted by user');
|
||||
process.exit(2);
|
||||
case 'error':
|
||||
logger.error('Loop ended with error');
|
||||
process.exit(3);
|
||||
}
|
||||
|
||||
} catch (error) {
|
||||
logger.error(error instanceof Error ? error.message : String(error));
|
||||
process.exit(3);
|
||||
}
|
||||
}
|
||||
|
||||
main();
|
||||
|
||||
@@ -1,239 +1,240 @@
|
||||
import path from 'path';
|
||||
import { confirm, input, select } from '@inquirer/prompts';
|
||||
import { agentRegistry } from './agents/index.js';
|
||||
import { StateManager, type SessionConfig, type LoopMode } from './state/index.js';
|
||||
import { detectSpecDirectories, getSpecProgress } from './spec/utils.js';
|
||||
import { logger } from './utils/index.js';
|
||||
|
||||
export interface SessionSetupResult {
|
||||
config: SessionConfig;
|
||||
isResume: boolean;
|
||||
}
|
||||
|
||||
/**
|
||||
* Detect and select a spec directory
|
||||
*/
|
||||
async function selectSpec(cwd: string = process.cwd()): Promise<string | null> {
|
||||
// Auto-detect spec directories
|
||||
const specDirs = await detectSpecDirectories(cwd);
|
||||
|
||||
if (specDirs.length === 0) {
|
||||
logger.error('No spec directories found!');
|
||||
logger.info('Expected: specs/<feature>/overview.md');
|
||||
logger.info('');
|
||||
logger.info('To get started:');
|
||||
logger.info('');
|
||||
logger.info('1. Create a spec using `plan2code-1--plan`');
|
||||
logger.info(' command in our AI Agent');
|
||||
logger.info('');
|
||||
logger.info('2. Come back here and run `plan2code-loop`');
|
||||
logger.info(' as an alternative to `plan2code-3--implement`');
|
||||
logger.info('');
|
||||
return null;
|
||||
}
|
||||
|
||||
if (specDirs.length === 1) {
|
||||
const spec = specDirs[0];
|
||||
const progress = await getSpecProgress(spec);
|
||||
logger.info(`Found spec: ${path.relative(cwd, spec)}`);
|
||||
logger.dim(` Feature: ${progress.featureName}`);
|
||||
logger.dim(` Phases: ${progress.totalPhases}`);
|
||||
return spec;
|
||||
}
|
||||
|
||||
// Multiple specs - let user choose
|
||||
const choices = await Promise.all(
|
||||
specDirs.map(async (spec) => {
|
||||
const progress = await getSpecProgress(spec);
|
||||
const relativePath = path.relative(cwd, spec);
|
||||
return {
|
||||
name: `${progress.featureName} (${progress.totalPhases} phases) - ${relativePath}`,
|
||||
value: spec,
|
||||
};
|
||||
})
|
||||
);
|
||||
|
||||
const selectedSpec = await select({
|
||||
message: 'Select spec to implement:',
|
||||
choices,
|
||||
});
|
||||
|
||||
return selectedSpec;
|
||||
}
|
||||
|
||||
/**
|
||||
* Select AI agent
|
||||
*/
|
||||
async function selectAgent(): Promise<string> {
|
||||
const allAgents = agentRegistry.getAll();
|
||||
|
||||
const agentName = await select({
|
||||
message: 'Select AI agent:',
|
||||
choices: allAgents.map((agent) => ({
|
||||
name: agent.config.displayName,
|
||||
value: agent.config.name,
|
||||
})),
|
||||
});
|
||||
|
||||
return agentName;
|
||||
}
|
||||
|
||||
/**
|
||||
* Prompt for JIRA ticket ID
|
||||
*/
|
||||
async function promptJiraTicketId(): Promise<string | undefined> {
|
||||
const ticketId = await input({
|
||||
message: 'JIRA Ticket ID (optional, for commit messages):',
|
||||
});
|
||||
|
||||
return ticketId.trim() || undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Select maximum iterations
|
||||
*/
|
||||
async function selectMaxIterations(): Promise<number> {
|
||||
const choices = [15, 30, 50, 75, 100, 125, 150, 200];
|
||||
|
||||
const max = await select({
|
||||
message: 'Maximum iterations:',
|
||||
choices: choices.map((n) => ({
|
||||
name: n.toString(),
|
||||
value: n,
|
||||
})),
|
||||
default: 100,
|
||||
});
|
||||
|
||||
return max;
|
||||
}
|
||||
|
||||
/**
|
||||
* Select loop mode: one task per loop or one phase per loop
|
||||
*/
|
||||
async function selectLoopMode(): Promise<LoopMode> {
|
||||
const mode = await select<LoopMode>({
|
||||
message: 'Tasks per loop iteration:',
|
||||
choices: [
|
||||
{
|
||||
name: 'One task per loop (default)',
|
||||
value: 'task' as LoopMode,
|
||||
},
|
||||
{
|
||||
name: 'One phase per loop (related tasks together)',
|
||||
value: 'phase' as LoopMode,
|
||||
},
|
||||
],
|
||||
default: 'task',
|
||||
});
|
||||
|
||||
return mode;
|
||||
}
|
||||
|
||||
/**
|
||||
* Handle existing session - returns action to take
|
||||
*/
|
||||
async function handleExistingSession(
|
||||
stateManager: StateManager,
|
||||
specPath: string
|
||||
): Promise<'continue' | 'fresh' | 'new'> {
|
||||
const state = await stateManager.detectSessionState(specPath);
|
||||
|
||||
if (state === 'new') {
|
||||
return 'new';
|
||||
}
|
||||
|
||||
if (state === 'continue') {
|
||||
const config = await stateManager.readConfig();
|
||||
const lastIter = config?.currentIteration ?? 0;
|
||||
|
||||
logger.info(`Found existing session at iteration ${lastIter}`);
|
||||
|
||||
const continueSession = await confirm({
|
||||
message: 'Continue previous session?',
|
||||
default: true,
|
||||
});
|
||||
|
||||
return continueSession ? 'continue' : 'fresh';
|
||||
}
|
||||
|
||||
if (state === 'changed') {
|
||||
logger.warning('Spec has changed since last session.');
|
||||
|
||||
const startFresh = await confirm({
|
||||
message: 'Start fresh? (This will clear previous progress)',
|
||||
default: false,
|
||||
});
|
||||
|
||||
return startFresh ? 'fresh' : 'continue';
|
||||
}
|
||||
|
||||
return 'new';
|
||||
}
|
||||
|
||||
/**
|
||||
* Setup session via interactive prompts
|
||||
*/
|
||||
export async function setupSession(
|
||||
stateManager: StateManager
|
||||
): Promise<SessionSetupResult | null> {
|
||||
// Show Planny welcome
|
||||
logger.welcome();
|
||||
|
||||
// Select spec
|
||||
const specPath = await selectSpec(process.cwd());
|
||||
if (!specPath) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// Set spec path on state manager for per-spec state directory
|
||||
stateManager.setSpecPath(specPath);
|
||||
|
||||
// Show spec progress
|
||||
const progress = await getSpecProgress(specPath);
|
||||
console.log();
|
||||
logger.header(`Spec: ${progress.featureName}`);
|
||||
logger.info(`Phases: ${progress.totalPhases}`);
|
||||
console.log();
|
||||
|
||||
// Check for existing session
|
||||
const sessionAction = await handleExistingSession(stateManager, specPath);
|
||||
|
||||
if (sessionAction === 'continue') {
|
||||
const existingConfig = await stateManager.readConfig();
|
||||
if (existingConfig) {
|
||||
const agent = await selectAgent();
|
||||
existingConfig.agent = agent;
|
||||
await stateManager.writeConfig(existingConfig);
|
||||
logger.info('Resuming previous session...');
|
||||
return { config: existingConfig, isResume: true };
|
||||
}
|
||||
}
|
||||
|
||||
if (sessionAction === 'fresh') {
|
||||
await stateManager.clearState();
|
||||
}
|
||||
|
||||
// Collect new session configuration
|
||||
const jiraTicketId = await promptJiraTicketId();
|
||||
const agent = await selectAgent();
|
||||
const loopMode = await selectLoopMode();
|
||||
const maxIterations = await selectMaxIterations();
|
||||
|
||||
const config: SessionConfig = {
|
||||
agent,
|
||||
model: 'default',
|
||||
maxIterations,
|
||||
specPath,
|
||||
timeout: 30,
|
||||
verbose: false,
|
||||
startedAt: new Date().toISOString(),
|
||||
currentIteration: 0,
|
||||
jiraTicketId,
|
||||
loopMode,
|
||||
};
|
||||
|
||||
// Initialize session
|
||||
await stateManager.initializeNewSession(config);
|
||||
|
||||
return { config, isResume: false };
|
||||
}
|
||||
import path from 'path';
|
||||
import { confirm, input, select } from '@inquirer/prompts';
|
||||
import { agentRegistry } from './agents/index.js';
|
||||
import { StateManager, type SessionConfig, type LoopMode } from './state/index.js';
|
||||
import { detectSpecDirectories, getSpecProgress } from './spec/utils.js';
|
||||
import { logger } from './utils/index.js';
|
||||
|
||||
export interface SessionSetupResult {
|
||||
config: SessionConfig;
|
||||
isResume: boolean;
|
||||
}
|
||||
|
||||
/**
|
||||
* Detect and select a spec directory
|
||||
*/
|
||||
async function selectSpec(cwd: string = process.cwd()): Promise<string | null> {
|
||||
// Auto-detect spec directories
|
||||
const specDirs = await detectSpecDirectories(cwd);
|
||||
|
||||
if (specDirs.length === 0) {
|
||||
logger.error('No spec directories found!');
|
||||
logger.info('Expected: specs/<feature>/overview.md');
|
||||
logger.info('');
|
||||
logger.info('To get started:');
|
||||
logger.info('');
|
||||
logger.info('1. Create a spec using `plan2code-1-plan`');
|
||||
logger.info(' command in our AI Agent');
|
||||
logger.info('');
|
||||
logger.info('2. Come back here and run `plan2code-loop`');
|
||||
logger.info(' as an alternative to `plan2code-3-implement`');
|
||||
logger.info('');
|
||||
return null;
|
||||
}
|
||||
|
||||
if (specDirs.length === 1) {
|
||||
const spec = specDirs[0];
|
||||
const progress = await getSpecProgress(spec);
|
||||
logger.info(`Found spec: ${path.relative(cwd, spec)}`);
|
||||
logger.dim(` Feature: ${progress.featureName}`);
|
||||
logger.dim(` Phases: ${progress.totalPhases}`);
|
||||
return spec;
|
||||
}
|
||||
|
||||
// Multiple specs - let user choose
|
||||
const choices = await Promise.all(
|
||||
specDirs.map(async (spec) => {
|
||||
const progress = await getSpecProgress(spec);
|
||||
const relativePath = path.relative(cwd, spec);
|
||||
return {
|
||||
name: `${progress.featureName} (${progress.totalPhases} phases) - ${relativePath}`,
|
||||
value: spec,
|
||||
};
|
||||
})
|
||||
);
|
||||
|
||||
const selectedSpec = await select({
|
||||
message: 'Select spec to implement:',
|
||||
choices,
|
||||
});
|
||||
|
||||
return selectedSpec;
|
||||
}
|
||||
|
||||
/**
|
||||
* Select AI agent
|
||||
*/
|
||||
async function selectAgent(): Promise<string> {
|
||||
const allAgents = agentRegistry.getAll();
|
||||
|
||||
const agentName = await select({
|
||||
message: 'Select AI agent:',
|
||||
choices: allAgents.map((agent) => ({
|
||||
name: agent.config.displayName,
|
||||
value: agent.config.name,
|
||||
})),
|
||||
});
|
||||
|
||||
return agentName;
|
||||
}
|
||||
|
||||
/**
|
||||
* Prompt for JIRA ticket ID
|
||||
*/
|
||||
async function promptJiraTicketId(): Promise<string | undefined> {
|
||||
const ticketId = await input({
|
||||
message: 'JIRA Ticket ID (optional, for commit messages):',
|
||||
});
|
||||
|
||||
return ticketId.trim() || undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Select maximum iterations
|
||||
*/
|
||||
async function selectMaxIterations(): Promise<number> {
|
||||
const choices = [15, 30, 50, 75, 100, 125, 150, 200];
|
||||
|
||||
const max = await select({
|
||||
message: 'Maximum iterations:',
|
||||
choices: choices.map((n) => ({
|
||||
name: n.toString(),
|
||||
value: n,
|
||||
})),
|
||||
default: 100,
|
||||
});
|
||||
|
||||
return max;
|
||||
}
|
||||
|
||||
/**
|
||||
* Select loop mode: one task per loop or one phase per loop
|
||||
*/
|
||||
async function selectLoopMode(): Promise<LoopMode> {
|
||||
const mode = await select<LoopMode>({
|
||||
message: 'Tasks per loop iteration:',
|
||||
choices: [
|
||||
{
|
||||
name: 'One task per loop (default)',
|
||||
value: 'task' as LoopMode,
|
||||
},
|
||||
{
|
||||
name: 'One phase per loop (related tasks together)',
|
||||
value: 'phase' as LoopMode,
|
||||
},
|
||||
],
|
||||
default: 'task',
|
||||
});
|
||||
|
||||
return mode;
|
||||
}
|
||||
|
||||
/**
|
||||
* Handle existing session - returns action to take
|
||||
*/
|
||||
async function handleExistingSession(
|
||||
stateManager: StateManager,
|
||||
specPath: string
|
||||
): Promise<'continue' | 'fresh' | 'new'> {
|
||||
const state = await stateManager.detectSessionState(specPath);
|
||||
|
||||
if (state === 'new') {
|
||||
return 'new';
|
||||
}
|
||||
|
||||
if (state === 'continue') {
|
||||
const config = await stateManager.readConfig();
|
||||
const lastIter = config?.currentIteration ?? 0;
|
||||
|
||||
logger.info(`Found existing session at iteration ${lastIter}`);
|
||||
|
||||
const continueSession = await confirm({
|
||||
message: 'Continue previous session?',
|
||||
default: true,
|
||||
});
|
||||
|
||||
return continueSession ? 'continue' : 'fresh';
|
||||
}
|
||||
|
||||
if (state === 'changed') {
|
||||
logger.warning('Spec has changed since last session.');
|
||||
|
||||
const startFresh = await confirm({
|
||||
message: 'Start fresh? (This will clear previous progress)',
|
||||
default: false,
|
||||
});
|
||||
|
||||
return startFresh ? 'fresh' : 'continue';
|
||||
}
|
||||
|
||||
return 'new';
|
||||
}
|
||||
|
||||
/**
|
||||
* Setup session via interactive prompts
|
||||
*/
|
||||
export async function setupSession(
|
||||
stateManager: StateManager
|
||||
): Promise<SessionSetupResult | null> {
|
||||
// Show Planny welcome
|
||||
logger.welcome();
|
||||
|
||||
// Select spec
|
||||
const specPath = await selectSpec(process.cwd());
|
||||
if (!specPath) {
|
||||
return null;
|
||||
}
|
||||
|
||||
// Set spec path on state manager for per-spec state directory
|
||||
stateManager.setSpecPath(specPath);
|
||||
|
||||
// Show spec progress
|
||||
const progress = await getSpecProgress(specPath);
|
||||
console.log();
|
||||
logger.header(`Spec: ${progress.featureName}`);
|
||||
logger.info(`Phases: ${progress.totalPhases}`);
|
||||
console.log();
|
||||
|
||||
// Check for existing session
|
||||
const sessionAction = await handleExistingSession(stateManager, specPath);
|
||||
|
||||
if (sessionAction === 'continue') {
|
||||
const existingConfig = await stateManager.readConfig();
|
||||
if (existingConfig) {
|
||||
const agent = await selectAgent();
|
||||
existingConfig.agent = agent;
|
||||
await stateManager.writeConfig(existingConfig);
|
||||
logger.info('Resuming previous session...');
|
||||
return { config: existingConfig, isResume: true };
|
||||
}
|
||||
}
|
||||
|
||||
if (sessionAction === 'fresh') {
|
||||
await stateManager.clearState();
|
||||
}
|
||||
|
||||
// Collect new session configuration
|
||||
const jiraTicketId = await promptJiraTicketId();
|
||||
const agent = await selectAgent();
|
||||
const loopMode = await selectLoopMode();
|
||||
const maxIterations = await selectMaxIterations();
|
||||
|
||||
const config: SessionConfig = {
|
||||
agent,
|
||||
model: 'default',
|
||||
maxIterations,
|
||||
specPath,
|
||||
timeout: 3,
|
||||
maxRetries: 5,
|
||||
verbose: false,
|
||||
startedAt: new Date().toISOString(),
|
||||
currentIteration: 0,
|
||||
jiraTicketId,
|
||||
loopMode,
|
||||
};
|
||||
|
||||
// Initialize session
|
||||
await stateManager.initializeNewSession(config);
|
||||
|
||||
return { config, isResume: false };
|
||||
}
|
||||
|
||||
@@ -1,452 +1,522 @@
|
||||
import { agentRegistry, type Agent, type AgentExecutionResult } from './agents/index.js';
|
||||
import { StateManager, type SessionConfig, type IterationLogEntry } from './state/index.js';
|
||||
import { buildLoopPrompt } from './prompt/index.js';
|
||||
import { checkForCompletion, checkForAllCompletions, logger, ensureGitRepo, ensureGitignore, type CompletionCheckResult } from './utils/index.js';
|
||||
|
||||
export interface TaskCompleteInfo {
|
||||
marker: string;
|
||||
taskId?: string;
|
||||
taskName?: string;
|
||||
}
|
||||
|
||||
export interface ControllerOptions {
|
||||
config: SessionConfig;
|
||||
stateManager: StateManager;
|
||||
onIteration?: (iteration: number, max: number) => void;
|
||||
onTaskComplete?: (info: TaskCompleteInfo) => void | Promise<void>;
|
||||
onLoopComplete?: () => void;
|
||||
}
|
||||
|
||||
export interface LoopResult {
|
||||
completed: boolean;
|
||||
iterations: number;
|
||||
finalMarker?: string;
|
||||
exitReason: 'all_complete' | 'max_iterations' | 'interrupted' | 'error';
|
||||
tasksCompleted: number;
|
||||
prereqsCompleted: number;
|
||||
error?: Error;
|
||||
}
|
||||
|
||||
export class Controller {
|
||||
private readonly config: SessionConfig;
|
||||
private readonly stateManager: StateManager;
|
||||
private readonly agent: Agent;
|
||||
private readonly onIteration?: (iteration: number, max: number) => void;
|
||||
private readonly onTaskComplete?: (info: TaskCompleteInfo) => void | Promise<void>;
|
||||
private readonly onLoopComplete?: () => void;
|
||||
private interrupted = false;
|
||||
private abortController: AbortController | null = null;
|
||||
private tasksCompleted = 0;
|
||||
private prereqsCompleted = 0;
|
||||
|
||||
constructor(options: ControllerOptions) {
|
||||
this.config = options.config;
|
||||
this.stateManager = options.stateManager;
|
||||
this.onIteration = options.onIteration;
|
||||
this.onTaskComplete = options.onTaskComplete;
|
||||
this.onLoopComplete = options.onLoopComplete;
|
||||
|
||||
const agent = agentRegistry.get(this.config.agent);
|
||||
if (!agent) {
|
||||
throw new Error(`Agent not found: ${this.config.agent}`);
|
||||
}
|
||||
this.agent = agent;
|
||||
}
|
||||
|
||||
private async buildPrompt(): Promise<string> {
|
||||
return buildLoopPrompt({
|
||||
specPath: this.config.specPath,
|
||||
iteration: this.config.currentIteration + 1,
|
||||
maxIterations: this.config.maxIterations,
|
||||
stateManager: this.stateManager,
|
||||
loopMode: this.config.loopMode || 'task',
|
||||
jiraTicketId: this.config.jiraTicketId,
|
||||
});
|
||||
}
|
||||
|
||||
private async executeIteration(prompt: string): Promise<AgentExecutionResult> {
|
||||
const timeoutMs = this.config.timeout * 60 * 1000;
|
||||
|
||||
// Create new AbortController for this iteration
|
||||
this.abortController = new AbortController();
|
||||
|
||||
const result = await this.agent.execute({
|
||||
prompt,
|
||||
model: this.config.model,
|
||||
timeout: timeoutMs,
|
||||
verbose: this.config.verbose,
|
||||
cwd: process.cwd(),
|
||||
signal: this.abortController.signal,
|
||||
});
|
||||
|
||||
this.abortController = null;
|
||||
return result;
|
||||
}
|
||||
|
||||
private createLogEntry(
|
||||
result: AgentExecutionResult,
|
||||
status: IterationLogEntry['status'],
|
||||
marker?: string
|
||||
): IterationLogEntry {
|
||||
return {
|
||||
iteration: this.config.currentIteration + 1,
|
||||
timestamp: new Date().toISOString(),
|
||||
duration: result.duration,
|
||||
exitCode: result.exitCode,
|
||||
status,
|
||||
completionMarker: marker,
|
||||
};
|
||||
}
|
||||
|
||||
private formatTaskDisplay(completion: CompletionCheckResult): string {
|
||||
if (completion.taskId && completion.taskName) {
|
||||
return `Task ${completion.taskId}: ${completion.taskName}`;
|
||||
} else if (completion.taskId) {
|
||||
return `Task ${completion.taskId}`;
|
||||
}
|
||||
return '';
|
||||
}
|
||||
|
||||
private displayIterationResult(result: AgentExecutionResult, iterNum: number, completion: CompletionCheckResult): void {
|
||||
const duration = Math.round(result.duration / 1000);
|
||||
const taskDisplay = this.formatTaskDisplay(completion);
|
||||
|
||||
// Show iteration completion with task info if available
|
||||
if (taskDisplay) {
|
||||
logger.iteration(iterNum, this.config.maxIterations, `${taskDisplay} (${duration}s)`);
|
||||
} else {
|
||||
logger.iteration(iterNum, this.config.maxIterations, `completed in ${duration}s`);
|
||||
}
|
||||
|
||||
// Verbose mode: show full output
|
||||
if (this.config.verbose) {
|
||||
console.log();
|
||||
logger.dim('--- Agent Output ---');
|
||||
console.log(result.stdout);
|
||||
if (result.stderr) {
|
||||
logger.dim('--- Agent Stderr ---');
|
||||
console.log(result.stderr);
|
||||
}
|
||||
logger.dim('--- End Output ---');
|
||||
console.log();
|
||||
} else if (result.exitCode !== 0 && result.stderr.trim()) {
|
||||
logger.error(` ${result.stderr.trim().split('\n')[0]}`);
|
||||
}
|
||||
}
|
||||
|
||||
async run(): Promise<LoopResult> {
|
||||
// Pre-flight: verify agent CLI is available
|
||||
const isAvailable = await this.agent.isAvailable();
|
||||
if (!isAvailable) {
|
||||
logger.error(`"${this.agent.config.displayName}" is not available!`);
|
||||
logger.info(`Please ensure the "${this.agent.config.command}" command is installed and available in your PATH.`);
|
||||
throw new Error(`Agent "${this.agent.config.displayName}" is not available. Please install it and try again.`);
|
||||
}
|
||||
|
||||
// Ensure git repo and .gitignore are set up before any iterations
|
||||
const gitReady = await ensureGitRepo(process.cwd());
|
||||
if (!gitReady) {
|
||||
throw new Error('Failed to initialize a git repository in the current working directory. Cannot start Plan2Code Loop.');
|
||||
}
|
||||
ensureGitignore(process.cwd());
|
||||
|
||||
const loopModeLabel = (this.config.loopMode || 'task') === 'phase' ? 'One phase per loop' : 'One task per loop';
|
||||
logger.header('Starting Plan2Code Loop');
|
||||
logger.info(`Agent: ${this.agent.config.displayName}`);
|
||||
logger.info(`Model: ${this.config.model}`);
|
||||
logger.info(`Spec: ${this.config.specPath}`);
|
||||
logger.info(`Loop mode: ${loopModeLabel}`);
|
||||
logger.info(`Max iterations: ${this.config.maxIterations}`);
|
||||
console.log();
|
||||
|
||||
while (this.config.currentIteration < this.config.maxIterations) {
|
||||
if (this.interrupted) {
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'interrupted',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
|
||||
const iterNum = this.config.currentIteration + 1;
|
||||
this.onIteration?.(iterNum, this.config.maxIterations);
|
||||
|
||||
console.log();
|
||||
logger.info(`Iteration ${iterNum}/${this.config.maxIterations}`);
|
||||
|
||||
// Build prompt - simple, just spec path and iteration info
|
||||
const prompt = await this.buildPrompt();
|
||||
|
||||
const spinner = logger.spinner('Waiting for AI Agent response (please be patient)');
|
||||
const startTime = Date.now();
|
||||
|
||||
// Update spinner with elapsed time every second
|
||||
const elapsedInterval = setInterval(() => {
|
||||
const elapsed = Math.round((Date.now() - startTime) / 1000);
|
||||
spinner.text = `Waiting for AI Agent response (please be patient) ... (${elapsed}s)`;
|
||||
}, 1000);
|
||||
|
||||
try {
|
||||
const result = await this.executeIteration(prompt);
|
||||
clearInterval(elapsedInterval);
|
||||
spinner.stop();
|
||||
|
||||
// Check if cancelled
|
||||
if (result.cancelled || this.interrupted) {
|
||||
logger.info('Agent process cancelled');
|
||||
|
||||
const entry: IterationLogEntry = {
|
||||
iteration: iterNum,
|
||||
timestamp: new Date().toISOString(),
|
||||
duration: result.duration,
|
||||
exitCode: -1,
|
||||
status: 'interrupted',
|
||||
};
|
||||
await this.stateManager.appendIterationLog(entry);
|
||||
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'interrupted',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
|
||||
// Branch completion handling based on loop mode
|
||||
const isPhaseMode = (this.config.loopMode || 'task') === 'phase';
|
||||
|
||||
if (isPhaseMode) {
|
||||
// Phase mode: parse ALL completion markers from output
|
||||
const allCompletions = checkForAllCompletions(result.stdout + result.stderr);
|
||||
|
||||
// Determine status
|
||||
let status: IterationLogEntry['status'] = 'running';
|
||||
if (allCompletions.tasks.length > 0 || allCompletions.loopComplete || allCompletions.phaseComplete) {
|
||||
const hasBlocked = allCompletions.tasks.some(t => t.marker === 'TASK_BLOCKED');
|
||||
const hasCompleted = allCompletions.tasks.some(t => t.marker === 'TASK_COMPLETE' || t.marker === 'PREREQ_COMPLETE' || t.marker === 'PREREQ_ASSUMED');
|
||||
status = hasCompleted || allCompletions.phaseComplete || allCompletions.loopComplete ? 'completed' : hasBlocked ? 'blocked' : 'running';
|
||||
} else if (result.timedOut) {
|
||||
status = 'timeout';
|
||||
} else if (result.exitCode !== 0) {
|
||||
status = 'error';
|
||||
}
|
||||
|
||||
// Log iteration with count of tasks
|
||||
const markerSummary = allCompletions.tasks.map(t => `${t.marker}: ${t.taskId}`).join(', ');
|
||||
const logEntry = this.createLogEntry(result, status, markerSummary || undefined);
|
||||
await this.stateManager.appendIterationLog(logEntry);
|
||||
|
||||
// Display each completed task
|
||||
const duration = Math.round(result.duration / 1000);
|
||||
for (const task of allCompletions.tasks) {
|
||||
if (task.marker === 'TASK_COMPLETE') {
|
||||
this.tasksCompleted++;
|
||||
const taskDisplay = this.formatTaskDisplay(task);
|
||||
logger.success(taskDisplay ? `Completed: ${taskDisplay}` : 'Task completed!');
|
||||
} else if (task.marker === 'PREREQ_COMPLETE') {
|
||||
this.prereqsCompleted++;
|
||||
const prereqDisplay = task.taskId && task.taskName
|
||||
? `Prereq ${task.taskId}: ${task.taskName}`
|
||||
: task.taskId ? `Prereq ${task.taskId}` : 'Prerequisite';
|
||||
logger.success(`Verified: ${prereqDisplay}`);
|
||||
} else if (task.marker === 'PREREQ_ASSUMED') {
|
||||
this.prereqsCompleted++;
|
||||
const prereqDisplay = task.taskId && task.taskName
|
||||
? `Prereq ${task.taskId}: ${task.taskName}`
|
||||
: task.taskId ? `Prereq ${task.taskId}` : 'Prerequisite';
|
||||
logger.success(`Assumed: ${prereqDisplay}`);
|
||||
} else if (task.marker === 'TASK_BLOCKED') {
|
||||
const blockInfo = task.taskId
|
||||
? `Task ${task.taskId} blocked: ${task.reason || 'Unknown reason'}`
|
||||
: `Task blocked: ${task.reason || 'Unknown reason'}`;
|
||||
logger.warning(blockInfo);
|
||||
}
|
||||
}
|
||||
|
||||
// Show phase-level summary
|
||||
if (allCompletions.tasks.length > 0) {
|
||||
const completedCount = allCompletions.tasks.filter(t => t.marker === 'TASK_COMPLETE').length;
|
||||
const prereqCount = allCompletions.tasks.filter(t => t.marker === 'PREREQ_COMPLETE' || t.marker === 'PREREQ_ASSUMED').length;
|
||||
const blockedCount = allCompletions.tasks.filter(t => t.marker === 'TASK_BLOCKED').length;
|
||||
const parts: string[] = [];
|
||||
if (completedCount > 0) parts.push(`${completedCount} task(s) completed`);
|
||||
if (prereqCount > 0) parts.push(`${prereqCount} prereq(s) verified`);
|
||||
if (blockedCount > 0) parts.push(`${blockedCount} blocked`);
|
||||
logger.iteration(iterNum, this.config.maxIterations,
|
||||
`Phase done: ${parts.join(', ')} (${duration}s)`);
|
||||
} else {
|
||||
logger.iteration(iterNum, this.config.maxIterations, `completed in ${duration}s`);
|
||||
}
|
||||
|
||||
// Verbose output
|
||||
if (this.config.verbose) {
|
||||
console.log();
|
||||
logger.dim('--- Agent Output ---');
|
||||
console.log(result.stdout);
|
||||
if (result.stderr) {
|
||||
logger.dim('--- Agent Stderr ---');
|
||||
console.log(result.stderr);
|
||||
}
|
||||
logger.dim('--- End Output ---');
|
||||
console.log();
|
||||
} else if (result.exitCode !== 0 && result.stderr.trim()) {
|
||||
logger.error(` ${result.stderr.trim().split('\n')[0]}`);
|
||||
}
|
||||
|
||||
// Handle LOOP_COMPLETE
|
||||
if (allCompletions.loopComplete) {
|
||||
this.onLoopComplete?.();
|
||||
logger.success('All tasks complete!');
|
||||
return {
|
||||
completed: true,
|
||||
iterations: iterNum,
|
||||
finalMarker: 'LOOP_COMPLETE',
|
||||
exitReason: 'all_complete',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
} else {
|
||||
// Task mode (default): existing single-marker logic
|
||||
const completion = checkForCompletion(result.stdout + result.stderr);
|
||||
|
||||
// Determine status
|
||||
let status: IterationLogEntry['status'] = 'running';
|
||||
if (completion.completed) {
|
||||
if (completion.marker === 'TASK_BLOCKED') {
|
||||
status = 'blocked';
|
||||
} else {
|
||||
status = 'completed';
|
||||
}
|
||||
} else if (result.timedOut) {
|
||||
status = 'timeout';
|
||||
} else if (result.exitCode !== 0) {
|
||||
status = 'error';
|
||||
}
|
||||
|
||||
// Log iteration
|
||||
const logEntry = this.createLogEntry(result, status, completion.marker);
|
||||
await this.stateManager.appendIterationLog(logEntry);
|
||||
|
||||
// Display iteration result with task info from completion marker
|
||||
this.displayIterationResult(result, iterNum, completion);
|
||||
|
||||
// Handle completion markers
|
||||
if (completion.completed) {
|
||||
const taskDisplay = this.formatTaskDisplay(completion);
|
||||
|
||||
if (completion.marker === 'TASK_COMPLETE') {
|
||||
this.tasksCompleted++;
|
||||
await this.onTaskComplete?.({
|
||||
marker: completion.marker,
|
||||
taskId: completion.taskId,
|
||||
taskName: completion.taskName,
|
||||
});
|
||||
logger.success(taskDisplay ? `Completed: ${taskDisplay}` : 'Task completed!');
|
||||
} else if (completion.marker === 'PREREQ_COMPLETE' || completion.marker === 'PREREQ_ASSUMED') {
|
||||
this.prereqsCompleted++;
|
||||
await this.onTaskComplete?.({
|
||||
marker: completion.marker,
|
||||
taskId: completion.taskId,
|
||||
taskName: completion.taskName,
|
||||
});
|
||||
const prereqDisplay = completion.taskId && completion.taskName
|
||||
? `Prereq ${completion.taskId}: ${completion.taskName}`
|
||||
: completion.taskId ? `Prereq ${completion.taskId}` : 'Prerequisite';
|
||||
const verb = completion.marker === 'PREREQ_COMPLETE' ? 'Verified' : 'Assumed';
|
||||
logger.success(`${verb}: ${prereqDisplay}`);
|
||||
} else if (completion.marker === 'TASK_BLOCKED') {
|
||||
const blockInfo = completion.taskId
|
||||
? `Task ${completion.taskId} blocked: ${completion.reason || 'Unknown reason'}`
|
||||
: `Task blocked: ${completion.reason || 'Unknown reason'}`;
|
||||
logger.warning(blockInfo);
|
||||
} else if (completion.marker === 'LOOP_COMPLETE') {
|
||||
// Commit any final changes before completing
|
||||
this.tasksCompleted++;
|
||||
await this.onTaskComplete?.({
|
||||
marker: completion.marker,
|
||||
taskId: completion.taskId,
|
||||
taskName: completion.taskName || 'Final implementation complete',
|
||||
});
|
||||
this.onLoopComplete?.();
|
||||
logger.success('All tasks complete!');
|
||||
return {
|
||||
completed: true,
|
||||
iterations: iterNum,
|
||||
finalMarker: 'LOOP_COMPLETE',
|
||||
exitReason: 'all_complete',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Increment iteration and update spec hash (so resume doesn't see false changes)
|
||||
await this.stateManager.incrementIteration();
|
||||
await this.stateManager.updateSpecHash(this.config.specPath);
|
||||
this.config.currentIteration++;
|
||||
|
||||
// Handle timeout
|
||||
if (result.timedOut) {
|
||||
logger.warning(`Iteration ${iterNum} timed out, continuing...`);
|
||||
}
|
||||
|
||||
// Handle error (but continue - LLM might recover)
|
||||
if (result.exitCode !== 0 && !result.timedOut) {
|
||||
// In phase mode, check if any tasks completed despite error exit code
|
||||
const hasCompletions = isPhaseMode
|
||||
? checkForAllCompletions(result.stdout + result.stderr).tasks.length > 0
|
||||
: checkForCompletion(result.stdout + result.stderr).completed;
|
||||
if (!hasCompletions) {
|
||||
logger.warning(`Iteration ${iterNum} exited with code ${result.exitCode}, continuing...`);
|
||||
}
|
||||
}
|
||||
|
||||
} catch (error) {
|
||||
clearInterval(elapsedInterval);
|
||||
spinner.stop();
|
||||
logger.error(`Iteration ${iterNum} failed: ${error}`);
|
||||
|
||||
// Log the error
|
||||
const entry: IterationLogEntry = {
|
||||
iteration: iterNum,
|
||||
timestamp: new Date().toISOString(),
|
||||
duration: 0,
|
||||
exitCode: -1,
|
||||
status: 'error',
|
||||
};
|
||||
await this.stateManager.appendIterationLog(entry);
|
||||
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'error',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
error: error instanceof Error ? error : new Error(String(error)),
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
// Max iterations reached
|
||||
logger.warning('Max iterations reached');
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'max_iterations',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
|
||||
interrupt(): void {
|
||||
this.interrupted = true;
|
||||
if (this.abortController) {
|
||||
this.abortController.abort();
|
||||
}
|
||||
}
|
||||
}
|
||||
import { agentRegistry, type Agent, type AgentExecutionResult } from './agents/index.js';
|
||||
import { StateManager, type SessionConfig, type IterationLogEntry } from './state/index.js';
|
||||
import { buildLoopPrompt } from './prompt/index.js';
|
||||
import { checkForCompletion, checkForAllCompletions, logger, ensureGitRepo, ensureGitignore, type CompletionCheckResult } from './utils/index.js';
|
||||
|
||||
export interface TaskCompleteInfo {
|
||||
marker: string;
|
||||
taskId?: string;
|
||||
taskName?: string;
|
||||
}
|
||||
|
||||
export interface ControllerOptions {
|
||||
config: SessionConfig;
|
||||
stateManager: StateManager;
|
||||
onIteration?: (iteration: number, max: number) => void;
|
||||
onTaskComplete?: (info: TaskCompleteInfo) => void | Promise<void>;
|
||||
onLoopComplete?: () => void;
|
||||
}
|
||||
|
||||
export interface LoopResult {
|
||||
completed: boolean;
|
||||
iterations: number;
|
||||
finalMarker?: string;
|
||||
exitReason: 'all_complete' | 'max_iterations' | 'interrupted' | 'error';
|
||||
tasksCompleted: number;
|
||||
prereqsCompleted: number;
|
||||
error?: Error;
|
||||
}
|
||||
|
||||
export class Controller {
|
||||
private readonly config: SessionConfig;
|
||||
private readonly stateManager: StateManager;
|
||||
private readonly agent: Agent;
|
||||
private readonly onIteration?: (iteration: number, max: number) => void;
|
||||
private readonly onTaskComplete?: (info: TaskCompleteInfo) => void | Promise<void>;
|
||||
private readonly onLoopComplete?: () => void;
|
||||
private interrupted = false;
|
||||
private abortController: AbortController | null = null;
|
||||
private tasksCompleted = 0;
|
||||
private prereqsCompleted = 0;
|
||||
|
||||
constructor(options: ControllerOptions) {
|
||||
this.config = options.config;
|
||||
this.stateManager = options.stateManager;
|
||||
this.onIteration = options.onIteration;
|
||||
this.onTaskComplete = options.onTaskComplete;
|
||||
this.onLoopComplete = options.onLoopComplete;
|
||||
|
||||
const agent = agentRegistry.get(this.config.agent);
|
||||
if (!agent) {
|
||||
throw new Error(`Agent not found: ${this.config.agent}`);
|
||||
}
|
||||
this.agent = agent;
|
||||
}
|
||||
|
||||
private async buildPrompt(): Promise<string> {
|
||||
return buildLoopPrompt({
|
||||
specPath: this.config.specPath,
|
||||
iteration: this.config.currentIteration + 1,
|
||||
maxIterations: this.config.maxIterations,
|
||||
stateManager: this.stateManager,
|
||||
loopMode: this.config.loopMode || 'task',
|
||||
jiraTicketId: this.config.jiraTicketId,
|
||||
});
|
||||
}
|
||||
|
||||
private computeTimeoutMs(attempt: number): number {
|
||||
const baseMs = this.config.timeout * 60 * 1000;
|
||||
return baseMs + (attempt * 30 * 1000);
|
||||
}
|
||||
|
||||
private formatDuration(ms: number): string {
|
||||
const totalSec = Math.round(ms / 1000);
|
||||
const min = Math.floor(totalSec / 60);
|
||||
const sec = totalSec % 60;
|
||||
if (min === 0) return `${sec}s`;
|
||||
if (sec === 0) return `${min}m`;
|
||||
return `${min}m ${sec}s`;
|
||||
}
|
||||
|
||||
private async executeWithRetry(prompt: string, iterNum: number): Promise<AgentExecutionResult | 'fatal_timeout'> {
|
||||
const maxAttempts = (this.config.maxRetries ?? 5) + 1;
|
||||
|
||||
for (let attempt = 0; attempt < maxAttempts; attempt++) {
|
||||
const timeoutMs = this.computeTimeoutMs(attempt);
|
||||
|
||||
if (attempt > 0) {
|
||||
const prevTimeoutMs = this.computeTimeoutMs(attempt - 1);
|
||||
logger.warning(
|
||||
`Timed out after ${this.formatDuration(prevTimeoutMs)}, retrying (${attempt}/${maxAttempts - 1})...`
|
||||
);
|
||||
}
|
||||
|
||||
const spinnerBase = attempt > 0
|
||||
? `Waiting for AI Agent response (retry ${attempt}/${maxAttempts - 1})`
|
||||
: 'Waiting for AI Agent response (please be patient)';
|
||||
const spinner = logger.spinner(spinnerBase);
|
||||
const startTime = Date.now();
|
||||
|
||||
const elapsedInterval = setInterval(() => {
|
||||
const elapsed = Math.round((Date.now() - startTime) / 1000);
|
||||
spinner.text = `${spinnerBase} ... (${elapsed}s)`;
|
||||
}, 1000);
|
||||
|
||||
const result = await this.executeIteration(prompt, timeoutMs);
|
||||
clearInterval(elapsedInterval);
|
||||
spinner.stop();
|
||||
|
||||
// Cancelled or interrupted — return immediately, don't retry
|
||||
if (result.cancelled || this.interrupted) {
|
||||
return result;
|
||||
}
|
||||
|
||||
// Completed (success or error exit code) — return to caller
|
||||
if (!result.timedOut) {
|
||||
return result;
|
||||
}
|
||||
|
||||
// Timed out — retry if attempts remain, otherwise fatal
|
||||
if (attempt < maxAttempts - 1) {
|
||||
continue;
|
||||
}
|
||||
|
||||
// All attempts exhausted
|
||||
logger.error(
|
||||
`Iteration ${iterNum} timed out on all ${maxAttempts} attempt${maxAttempts === 1 ? '' : 's'}. Stopping loop.`
|
||||
);
|
||||
const entry: IterationLogEntry = {
|
||||
iteration: iterNum,
|
||||
timestamp: new Date().toISOString(),
|
||||
duration: result.duration,
|
||||
exitCode: -1,
|
||||
status: 'timeout',
|
||||
};
|
||||
await this.stateManager.appendIterationLog(entry);
|
||||
return 'fatal_timeout';
|
||||
}
|
||||
|
||||
return 'fatal_timeout'; // unreachable, satisfies TS
|
||||
}
|
||||
|
||||
private async executeIteration(prompt: string, timeoutMs: number): Promise<AgentExecutionResult> {
|
||||
// Create new AbortController for this iteration
|
||||
this.abortController = new AbortController();
|
||||
|
||||
const result = await this.agent.execute({
|
||||
prompt,
|
||||
model: this.config.model,
|
||||
timeout: timeoutMs,
|
||||
verbose: this.config.verbose,
|
||||
cwd: process.cwd(),
|
||||
signal: this.abortController.signal,
|
||||
});
|
||||
|
||||
this.abortController = null;
|
||||
return result;
|
||||
}
|
||||
|
||||
private createLogEntry(
|
||||
result: AgentExecutionResult,
|
||||
status: IterationLogEntry['status'],
|
||||
marker?: string
|
||||
): IterationLogEntry {
|
||||
return {
|
||||
iteration: this.config.currentIteration + 1,
|
||||
timestamp: new Date().toISOString(),
|
||||
duration: result.duration,
|
||||
exitCode: result.exitCode,
|
||||
status,
|
||||
completionMarker: marker,
|
||||
};
|
||||
}
|
||||
|
||||
private formatTaskDisplay(completion: CompletionCheckResult): string {
|
||||
if (completion.taskId && completion.taskName) {
|
||||
return `Task ${completion.taskId}: ${completion.taskName}`;
|
||||
} else if (completion.taskId) {
|
||||
return `Task ${completion.taskId}`;
|
||||
}
|
||||
return '';
|
||||
}
|
||||
|
||||
private displayIterationResult(result: AgentExecutionResult, iterNum: number, completion: CompletionCheckResult): void {
|
||||
const duration = Math.round(result.duration / 1000);
|
||||
const taskDisplay = this.formatTaskDisplay(completion);
|
||||
|
||||
// Show iteration completion with task info if available
|
||||
if (taskDisplay) {
|
||||
logger.iteration(iterNum, this.config.maxIterations, `${taskDisplay} (${duration}s)`);
|
||||
} else {
|
||||
logger.iteration(iterNum, this.config.maxIterations, `completed in ${duration}s`);
|
||||
}
|
||||
|
||||
// Verbose mode: show full output
|
||||
if (this.config.verbose) {
|
||||
console.log();
|
||||
logger.dim('--- Agent Output ---');
|
||||
console.log(result.stdout);
|
||||
if (result.stderr) {
|
||||
logger.dim('--- Agent Stderr ---');
|
||||
console.log(result.stderr);
|
||||
}
|
||||
logger.dim('--- End Output ---');
|
||||
console.log();
|
||||
} else if (result.exitCode !== 0 && result.stderr.trim()) {
|
||||
logger.error(` ${result.stderr.trim().split('\n')[0]}`);
|
||||
}
|
||||
}
|
||||
|
||||
async run(): Promise<LoopResult> {
|
||||
// Pre-flight: verify agent CLI is available
|
||||
const isAvailable = await this.agent.isAvailable();
|
||||
if (!isAvailable) {
|
||||
logger.error(`"${this.agent.config.displayName}" is not available!`);
|
||||
logger.info(`Please ensure the "${this.agent.config.command}" command is installed and available in your PATH.`);
|
||||
throw new Error(`Agent "${this.agent.config.displayName}" is not available. Please install it and try again.`);
|
||||
}
|
||||
|
||||
// Ensure git repo and .gitignore are set up before any iterations
|
||||
const gitReady = await ensureGitRepo(process.cwd());
|
||||
if (!gitReady) {
|
||||
throw new Error('Failed to initialize a git repository in the current working directory. Cannot start Plan2Code Loop.');
|
||||
}
|
||||
ensureGitignore(process.cwd());
|
||||
|
||||
const loopModeLabel = (this.config.loopMode || 'task') === 'phase' ? 'One phase per loop' : 'One task per loop';
|
||||
logger.header('Starting Plan2Code Loop');
|
||||
logger.info(`Agent: ${this.agent.config.displayName}`);
|
||||
logger.info(`Model: ${this.config.model}`);
|
||||
logger.info(`Spec: ${this.config.specPath}`);
|
||||
logger.info(`Loop mode: ${loopModeLabel}`);
|
||||
logger.info(`Max iterations: ${this.config.maxIterations}`);
|
||||
console.log();
|
||||
|
||||
while (this.config.currentIteration < this.config.maxIterations) {
|
||||
if (this.interrupted) {
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'interrupted',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
|
||||
const iterNum = this.config.currentIteration + 1;
|
||||
this.onIteration?.(iterNum, this.config.maxIterations);
|
||||
|
||||
console.log();
|
||||
logger.info(`Iteration ${iterNum}/${this.config.maxIterations}`);
|
||||
|
||||
// Build prompt - simple, just spec path and iteration info
|
||||
const prompt = await this.buildPrompt();
|
||||
|
||||
try {
|
||||
const retryResult = await this.executeWithRetry(prompt, iterNum);
|
||||
|
||||
if (retryResult === 'fatal_timeout') {
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'error',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
error: new Error(`Iteration ${iterNum} failed after all retry attempts`),
|
||||
};
|
||||
}
|
||||
|
||||
const result = retryResult;
|
||||
|
||||
// Check if cancelled
|
||||
if (result.cancelled || this.interrupted) {
|
||||
logger.info('Agent process cancelled');
|
||||
|
||||
const entry: IterationLogEntry = {
|
||||
iteration: iterNum,
|
||||
timestamp: new Date().toISOString(),
|
||||
duration: result.duration,
|
||||
exitCode: -1,
|
||||
status: 'interrupted',
|
||||
};
|
||||
await this.stateManager.appendIterationLog(entry);
|
||||
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'interrupted',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
|
||||
// Branch completion handling based on loop mode
|
||||
const isPhaseMode = (this.config.loopMode || 'task') === 'phase';
|
||||
|
||||
if (isPhaseMode) {
|
||||
// Phase mode: parse ALL completion markers from output
|
||||
const allCompletions = checkForAllCompletions(result.stdout + result.stderr);
|
||||
|
||||
// Determine status
|
||||
let status: IterationLogEntry['status'] = 'running';
|
||||
if (allCompletions.tasks.length > 0 || allCompletions.loopComplete || allCompletions.phaseComplete) {
|
||||
const hasBlocked = allCompletions.tasks.some(t => t.marker === 'TASK_BLOCKED');
|
||||
const hasCompleted = allCompletions.tasks.some(t => t.marker === 'TASK_COMPLETE' || t.marker === 'PREREQ_COMPLETE' || t.marker === 'PREREQ_ASSUMED');
|
||||
status = hasCompleted || allCompletions.phaseComplete || allCompletions.loopComplete ? 'completed' : hasBlocked ? 'blocked' : 'running';
|
||||
} else if (result.timedOut) {
|
||||
status = 'timeout';
|
||||
} else if (result.exitCode !== 0) {
|
||||
status = 'error';
|
||||
}
|
||||
|
||||
// Log iteration with count of tasks
|
||||
const markerSummary = allCompletions.tasks.map(t => `${t.marker}: ${t.taskId}`).join(', ');
|
||||
const logEntry = this.createLogEntry(result, status, markerSummary || undefined);
|
||||
await this.stateManager.appendIterationLog(logEntry);
|
||||
|
||||
// Display each completed task
|
||||
const duration = Math.round(result.duration / 1000);
|
||||
for (const task of allCompletions.tasks) {
|
||||
if (task.marker === 'TASK_COMPLETE') {
|
||||
this.tasksCompleted++;
|
||||
const taskDisplay = this.formatTaskDisplay(task);
|
||||
logger.success(taskDisplay ? `Completed: ${taskDisplay}` : 'Task completed!');
|
||||
} else if (task.marker === 'PREREQ_COMPLETE') {
|
||||
this.prereqsCompleted++;
|
||||
const prereqDisplay = task.taskId && task.taskName
|
||||
? `Prereq ${task.taskId}: ${task.taskName}`
|
||||
: task.taskId ? `Prereq ${task.taskId}` : 'Prerequisite';
|
||||
logger.success(`Verified: ${prereqDisplay}`);
|
||||
} else if (task.marker === 'PREREQ_ASSUMED') {
|
||||
this.prereqsCompleted++;
|
||||
const prereqDisplay = task.taskId && task.taskName
|
||||
? `Prereq ${task.taskId}: ${task.taskName}`
|
||||
: task.taskId ? `Prereq ${task.taskId}` : 'Prerequisite';
|
||||
logger.success(`Assumed: ${prereqDisplay}`);
|
||||
} else if (task.marker === 'TASK_BLOCKED') {
|
||||
const blockInfo = task.taskId
|
||||
? `Task ${task.taskId} blocked: ${task.reason || 'Unknown reason'}`
|
||||
: `Task blocked: ${task.reason || 'Unknown reason'}`;
|
||||
logger.warning(blockInfo);
|
||||
}
|
||||
}
|
||||
|
||||
// Show phase-level summary
|
||||
if (allCompletions.tasks.length > 0) {
|
||||
const completedCount = allCompletions.tasks.filter(t => t.marker === 'TASK_COMPLETE').length;
|
||||
const prereqCount = allCompletions.tasks.filter(t => t.marker === 'PREREQ_COMPLETE' || t.marker === 'PREREQ_ASSUMED').length;
|
||||
const blockedCount = allCompletions.tasks.filter(t => t.marker === 'TASK_BLOCKED').length;
|
||||
const parts: string[] = [];
|
||||
if (completedCount > 0) parts.push(`${completedCount} task(s) completed`);
|
||||
if (prereqCount > 0) parts.push(`${prereqCount} prereq(s) verified`);
|
||||
if (blockedCount > 0) parts.push(`${blockedCount} blocked`);
|
||||
logger.iteration(iterNum, this.config.maxIterations,
|
||||
`Phase done: ${parts.join(', ')} (${duration}s)`);
|
||||
} else {
|
||||
logger.iteration(iterNum, this.config.maxIterations, `completed in ${duration}s`);
|
||||
}
|
||||
|
||||
// Verbose output
|
||||
if (this.config.verbose) {
|
||||
console.log();
|
||||
logger.dim('--- Agent Output ---');
|
||||
console.log(result.stdout);
|
||||
if (result.stderr) {
|
||||
logger.dim('--- Agent Stderr ---');
|
||||
console.log(result.stderr);
|
||||
}
|
||||
logger.dim('--- End Output ---');
|
||||
console.log();
|
||||
} else if (result.exitCode !== 0 && result.stderr.trim()) {
|
||||
logger.error(` ${result.stderr.trim().split('\n')[0]}`);
|
||||
}
|
||||
|
||||
// Handle LOOP_COMPLETE
|
||||
if (allCompletions.loopComplete) {
|
||||
this.onLoopComplete?.();
|
||||
logger.success('All tasks complete!');
|
||||
return {
|
||||
completed: true,
|
||||
iterations: iterNum,
|
||||
finalMarker: 'LOOP_COMPLETE',
|
||||
exitReason: 'all_complete',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
} else {
|
||||
// Task mode (default): existing single-marker logic
|
||||
const completion = checkForCompletion(result.stdout + result.stderr);
|
||||
|
||||
// Determine status
|
||||
let status: IterationLogEntry['status'] = 'running';
|
||||
if (completion.completed) {
|
||||
if (completion.marker === 'TASK_BLOCKED') {
|
||||
status = 'blocked';
|
||||
} else {
|
||||
status = 'completed';
|
||||
}
|
||||
} else if (result.timedOut) {
|
||||
status = 'timeout';
|
||||
} else if (result.exitCode !== 0) {
|
||||
status = 'error';
|
||||
}
|
||||
|
||||
// Log iteration
|
||||
const logEntry = this.createLogEntry(result, status, completion.marker);
|
||||
await this.stateManager.appendIterationLog(logEntry);
|
||||
|
||||
// Display iteration result with task info from completion marker
|
||||
this.displayIterationResult(result, iterNum, completion);
|
||||
|
||||
// Handle completion markers
|
||||
if (completion.completed) {
|
||||
const taskDisplay = this.formatTaskDisplay(completion);
|
||||
|
||||
if (completion.marker === 'TASK_COMPLETE') {
|
||||
this.tasksCompleted++;
|
||||
await this.onTaskComplete?.({
|
||||
marker: completion.marker,
|
||||
taskId: completion.taskId,
|
||||
taskName: completion.taskName,
|
||||
});
|
||||
logger.success(taskDisplay ? `Completed: ${taskDisplay}` : 'Task completed!');
|
||||
} else if (completion.marker === 'PREREQ_COMPLETE' || completion.marker === 'PREREQ_ASSUMED') {
|
||||
this.prereqsCompleted++;
|
||||
await this.onTaskComplete?.({
|
||||
marker: completion.marker,
|
||||
taskId: completion.taskId,
|
||||
taskName: completion.taskName,
|
||||
});
|
||||
const prereqDisplay = completion.taskId && completion.taskName
|
||||
? `Prereq ${completion.taskId}: ${completion.taskName}`
|
||||
: completion.taskId ? `Prereq ${completion.taskId}` : 'Prerequisite';
|
||||
const verb = completion.marker === 'PREREQ_COMPLETE' ? 'Verified' : 'Assumed';
|
||||
logger.success(`${verb}: ${prereqDisplay}`);
|
||||
} else if (completion.marker === 'TASK_BLOCKED') {
|
||||
const blockInfo = completion.taskId
|
||||
? `Task ${completion.taskId} blocked: ${completion.reason || 'Unknown reason'}`
|
||||
: `Task blocked: ${completion.reason || 'Unknown reason'}`;
|
||||
logger.warning(blockInfo);
|
||||
} else if (completion.marker === 'LOOP_COMPLETE') {
|
||||
// Commit any final changes before completing
|
||||
this.tasksCompleted++;
|
||||
await this.onTaskComplete?.({
|
||||
marker: completion.marker,
|
||||
taskId: completion.taskId,
|
||||
taskName: completion.taskName || 'Final implementation complete',
|
||||
});
|
||||
this.onLoopComplete?.();
|
||||
logger.success('All tasks complete!');
|
||||
return {
|
||||
completed: true,
|
||||
iterations: iterNum,
|
||||
finalMarker: 'LOOP_COMPLETE',
|
||||
exitReason: 'all_complete',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Increment iteration and update spec hash (so resume doesn't see false changes)
|
||||
await this.stateManager.incrementIteration();
|
||||
await this.stateManager.updateSpecHash(this.config.specPath);
|
||||
this.config.currentIteration++;
|
||||
|
||||
// Handle error (but continue - LLM might recover)
|
||||
if (result.exitCode !== 0) {
|
||||
// In phase mode, check if any tasks completed despite error exit code
|
||||
const hasCompletions = isPhaseMode
|
||||
? checkForAllCompletions(result.stdout + result.stderr).tasks.length > 0
|
||||
: checkForCompletion(result.stdout + result.stderr).completed;
|
||||
if (!hasCompletions) {
|
||||
logger.warning(`Iteration ${iterNum} exited with code ${result.exitCode}, continuing...`);
|
||||
}
|
||||
}
|
||||
|
||||
} catch (error) {
|
||||
clearInterval(elapsedInterval);
|
||||
spinner.stop();
|
||||
logger.error(`Iteration ${iterNum} failed: ${error}`);
|
||||
|
||||
// Log the error
|
||||
const entry: IterationLogEntry = {
|
||||
iteration: iterNum,
|
||||
timestamp: new Date().toISOString(),
|
||||
duration: 0,
|
||||
exitCode: -1,
|
||||
status: 'error',
|
||||
};
|
||||
await this.stateManager.appendIterationLog(entry);
|
||||
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'error',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
error: error instanceof Error ? error : new Error(String(error)),
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
// Max iterations reached
|
||||
logger.warning('Max iterations reached');
|
||||
return {
|
||||
completed: false,
|
||||
iterations: this.config.currentIteration,
|
||||
exitReason: 'max_iterations',
|
||||
tasksCompleted: this.tasksCompleted,
|
||||
prereqsCompleted: this.prereqsCompleted,
|
||||
};
|
||||
}
|
||||
|
||||
interrupt(): void {
|
||||
this.interrupted = true;
|
||||
if (this.abortController) {
|
||||
this.abortController.abort();
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,97 +1,96 @@
|
||||
import path from 'path';
|
||||
import { StateManager } from './state/index.js';
|
||||
import { Controller, type LoopResult, type TaskCompleteInfo } from './controller.js';
|
||||
import { setupSession } from './cli.js';
|
||||
import { logger, createTaskCommit } from './utils/index.js';
|
||||
|
||||
export async function run(): Promise<LoopResult | null> {
|
||||
// Ensure agents are registered
|
||||
await import('./agents/index.js');
|
||||
|
||||
const stateManager = new StateManager();
|
||||
|
||||
const result = await setupSession(stateManager);
|
||||
if (!result) {
|
||||
return null;
|
||||
}
|
||||
|
||||
const { config, isResume } = result;
|
||||
|
||||
if (isResume) {
|
||||
logger.info(`Resuming from iteration ${config.currentIteration}`);
|
||||
}
|
||||
|
||||
const controller = new Controller({
|
||||
config,
|
||||
stateManager,
|
||||
onIteration: (iter, max) => {
|
||||
// Could add git checkpoint logic here if needed
|
||||
},
|
||||
onTaskComplete: async (info: TaskCompleteInfo) => {
|
||||
// Create git commit for completed task
|
||||
const taskName = info.taskName || info.taskId || 'Task completed';
|
||||
await createTaskCommit({
|
||||
taskName,
|
||||
jiraTicketId: config.jiraTicketId,
|
||||
cwd: process.cwd(),
|
||||
});
|
||||
},
|
||||
onLoopComplete: () => {
|
||||
// All tasks completed callback
|
||||
},
|
||||
});
|
||||
|
||||
// Setup interrupt handler
|
||||
const handleInterrupt = () => {
|
||||
logger.warning('\nInterrupt received, saving state...');
|
||||
controller.interrupt();
|
||||
};
|
||||
|
||||
process.on('SIGINT', handleInterrupt);
|
||||
process.on('SIGTERM', handleInterrupt);
|
||||
|
||||
try {
|
||||
const loopResult = await controller.run();
|
||||
|
||||
// Display summary
|
||||
console.log();
|
||||
logger.header('Session Summary');
|
||||
logger.info(`Total iterations: ${loopResult.iterations}`);
|
||||
logger.info(`Tasks completed: ${loopResult.tasksCompleted}`);
|
||||
if (loopResult.prereqsCompleted > 0) {
|
||||
logger.info(`Prerequisites verified: ${loopResult.prereqsCompleted}`);
|
||||
}
|
||||
logger.info(`Exit reason: ${loopResult.exitReason}`);
|
||||
if (loopResult.finalMarker) {
|
||||
logger.info(`Completion marker: ${loopResult.finalMarker}`);
|
||||
}
|
||||
if (loopResult.error) {
|
||||
logger.error(`Error: ${loopResult.error.message}`);
|
||||
}
|
||||
|
||||
// Show completion celebration and finalize reminder when all phases complete
|
||||
if (loopResult.exitReason === 'all_complete') {
|
||||
logger.allPhasesComplete();
|
||||
}
|
||||
|
||||
// Show state file locations (now per-spec)
|
||||
console.log();
|
||||
logger.dim(`Session files saved to ${path.relative(process.cwd(), stateManager.getStateDir())}:`);
|
||||
logger.dim(' - config.json (session configuration)');
|
||||
logger.dim(' - scratchpad.md (LLM-managed notes)');
|
||||
logger.dim(' - iteration.log (history)');
|
||||
logger.dim('Tip: run `plan2code-metrics` to capture run data for prompt improvement.');
|
||||
|
||||
return loopResult;
|
||||
} finally {
|
||||
process.off('SIGINT', handleInterrupt);
|
||||
process.off('SIGTERM', handleInterrupt);
|
||||
}
|
||||
}
|
||||
|
||||
// Re-export types and classes
|
||||
export { Controller, type ControllerOptions, type LoopResult, type TaskCompleteInfo } from './controller.js';
|
||||
export { StateManager } from './state/index.js';
|
||||
export { setupSession } from './cli.js';
|
||||
export { agentRegistry, type Agent, type AgentConfig } from './agents/index.js';
|
||||
export { detectSpecDirectories, getSpecProgress } from './spec/index.js';
|
||||
import path from 'path';
|
||||
import { StateManager } from './state/index.js';
|
||||
import { Controller, type LoopResult, type TaskCompleteInfo } from './controller.js';
|
||||
import { setupSession } from './cli.js';
|
||||
import { logger, createTaskCommit } from './utils/index.js';
|
||||
|
||||
export async function run(): Promise<LoopResult | null> {
|
||||
// Ensure agents are registered
|
||||
await import('./agents/index.js');
|
||||
|
||||
const stateManager = new StateManager();
|
||||
|
||||
const result = await setupSession(stateManager);
|
||||
if (!result) {
|
||||
return null;
|
||||
}
|
||||
|
||||
const { config, isResume } = result;
|
||||
|
||||
if (isResume) {
|
||||
logger.info(`Resuming from iteration ${config.currentIteration}`);
|
||||
}
|
||||
|
||||
const controller = new Controller({
|
||||
config,
|
||||
stateManager,
|
||||
onIteration: (iter, max) => {
|
||||
// Could add git checkpoint logic here if needed
|
||||
},
|
||||
onTaskComplete: async (info: TaskCompleteInfo) => {
|
||||
// Create git commit for completed task
|
||||
const taskName = info.taskName || info.taskId || 'Task completed';
|
||||
await createTaskCommit({
|
||||
taskName,
|
||||
jiraTicketId: config.jiraTicketId,
|
||||
cwd: process.cwd(),
|
||||
});
|
||||
},
|
||||
onLoopComplete: () => {
|
||||
// All tasks completed callback
|
||||
},
|
||||
});
|
||||
|
||||
// Setup interrupt handler
|
||||
const handleInterrupt = () => {
|
||||
logger.warning('\nInterrupt received, saving state...');
|
||||
controller.interrupt();
|
||||
};
|
||||
|
||||
process.on('SIGINT', handleInterrupt);
|
||||
process.on('SIGTERM', handleInterrupt);
|
||||
|
||||
try {
|
||||
const loopResult = await controller.run();
|
||||
|
||||
// Display summary
|
||||
console.log();
|
||||
logger.header('Session Summary');
|
||||
logger.info(`Total iterations: ${loopResult.iterations}`);
|
||||
logger.info(`Tasks completed: ${loopResult.tasksCompleted}`);
|
||||
if (loopResult.prereqsCompleted > 0) {
|
||||
logger.info(`Prerequisites verified: ${loopResult.prereqsCompleted}`);
|
||||
}
|
||||
logger.info(`Exit reason: ${loopResult.exitReason}`);
|
||||
if (loopResult.finalMarker) {
|
||||
logger.info(`Completion marker: ${loopResult.finalMarker}`);
|
||||
}
|
||||
if (loopResult.error) {
|
||||
logger.error(`Error: ${loopResult.error.message}`);
|
||||
}
|
||||
|
||||
// Show completion celebration and finalize reminder when all phases complete
|
||||
if (loopResult.exitReason === 'all_complete') {
|
||||
logger.allPhasesComplete();
|
||||
}
|
||||
|
||||
// Show state file locations (now per-spec)
|
||||
console.log();
|
||||
logger.dim(`Session files saved to ${path.relative(process.cwd(), stateManager.getStateDir())}:`);
|
||||
logger.dim(' - config.json (session configuration)');
|
||||
logger.dim(' - scratchpad.md (LLM-managed notes)');
|
||||
logger.dim(' - iteration.log (history)');
|
||||
|
||||
return loopResult;
|
||||
} finally {
|
||||
process.off('SIGINT', handleInterrupt);
|
||||
process.off('SIGTERM', handleInterrupt);
|
||||
}
|
||||
}
|
||||
|
||||
// Re-export types and classes
|
||||
export { Controller, type ControllerOptions, type LoopResult, type TaskCompleteInfo } from './controller.js';
|
||||
export { StateManager } from './state/index.js';
|
||||
export { setupSession } from './cli.js';
|
||||
export { agentRegistry, type Agent, type AgentConfig } from './agents/index.js';
|
||||
export { detectSpecDirectories, getSpecProgress } from './spec/index.js';
|
||||
|
||||
@@ -1,39 +1,39 @@
|
||||
import type { StateManager, LoopMode } from '../state/index.js';
|
||||
import { LOOP_PROMPT_TEMPLATE, LOOP_PROMPT_TEMPLATE_PHASE } from './templates.js';
|
||||
|
||||
export interface PromptContext {
|
||||
specPath: string;
|
||||
iteration: number;
|
||||
maxIterations: number;
|
||||
stateManager: StateManager;
|
||||
loopMode: LoopMode;
|
||||
jiraTicketId?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the prompt for the AI agent
|
||||
* Selects template based on loop mode (task vs phase)
|
||||
*/
|
||||
export async function buildLoopPrompt(context: PromptContext): Promise<string> {
|
||||
const { specPath, iteration, maxIterations, stateManager, loopMode, jiraTicketId } = context;
|
||||
|
||||
// Read scratchpad content for session continuity (LLM writes to this)
|
||||
const scratchpadContent = await stateManager.readScratchpad();
|
||||
|
||||
// Project root is where plan2code-loop was invoked from
|
||||
const projectRoot = process.cwd();
|
||||
|
||||
// Select template based on loop mode
|
||||
const template = loopMode === 'phase' ? LOOP_PROMPT_TEMPLATE_PHASE : LOOP_PROMPT_TEMPLATE;
|
||||
|
||||
// Template substitution
|
||||
const prompt = template
|
||||
.replace(/{{projectRoot}}/g, projectRoot)
|
||||
.replace(/{{specPath}}/g, specPath)
|
||||
.replace(/{{iteration}}/g, iteration.toString())
|
||||
.replace(/{{maxIterations}}/g, maxIterations.toString())
|
||||
.replace(/{{scratchpadContent}}/g, scratchpadContent || '(First iteration - no previous progress)')
|
||||
.replace(/{{jiraTicketId}}/g, jiraTicketId || '');
|
||||
|
||||
return prompt;
|
||||
}
|
||||
import type { StateManager, LoopMode } from '../state/index.js';
|
||||
import { LOOP_PROMPT_TEMPLATE, LOOP_PROMPT_TEMPLATE_PHASE } from './templates.js';
|
||||
|
||||
export interface PromptContext {
|
||||
specPath: string;
|
||||
iteration: number;
|
||||
maxIterations: number;
|
||||
stateManager: StateManager;
|
||||
loopMode: LoopMode;
|
||||
jiraTicketId?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the prompt for the AI agent
|
||||
* Selects template based on loop mode (task vs phase)
|
||||
*/
|
||||
export async function buildLoopPrompt(context: PromptContext): Promise<string> {
|
||||
const { specPath, iteration, maxIterations, stateManager, loopMode, jiraTicketId } = context;
|
||||
|
||||
// Read scratchpad content for session continuity (LLM writes to this)
|
||||
const scratchpadContent = await stateManager.readScratchpad();
|
||||
|
||||
// Project root is where plan2code-loop was invoked from
|
||||
const projectRoot = process.cwd();
|
||||
|
||||
// Select template based on loop mode
|
||||
const template = loopMode === 'phase' ? LOOP_PROMPT_TEMPLATE_PHASE : LOOP_PROMPT_TEMPLATE;
|
||||
|
||||
// Template substitution
|
||||
const prompt = template
|
||||
.replace(/{{projectRoot}}/g, projectRoot)
|
||||
.replace(/{{specPath}}/g, specPath)
|
||||
.replace(/{{iteration}}/g, iteration.toString())
|
||||
.replace(/{{maxIterations}}/g, maxIterations.toString())
|
||||
.replace(/{{scratchpadContent}}/g, scratchpadContent || '(First iteration - no previous progress)')
|
||||
.replace(/{{jiraTicketId}}/g, jiraTicketId || '');
|
||||
|
||||
return prompt;
|
||||
}
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
export {
|
||||
buildLoopPrompt,
|
||||
type PromptContext,
|
||||
} from './builder.js';
|
||||
|
||||
export { LOOP_PROMPT_TEMPLATE, LOOP_PROMPT_TEMPLATE_PHASE } from './templates.js';
|
||||
export {
|
||||
buildLoopPrompt,
|
||||
type PromptContext,
|
||||
} from './builder.js';
|
||||
|
||||
export { LOOP_PROMPT_TEMPLATE, LOOP_PROMPT_TEMPLATE_PHASE } from './templates.js';
|
||||
|
||||
@@ -1,186 +1,200 @@
|
||||
export const LOOP_PROMPT_TEMPLATE = `# PLAN2CODE-LOOP: Autonomous Task Implementation
|
||||
|
||||
## CRITICAL CONSTRAINT
|
||||
**IMPLEMENT EXACTLY ONE TASK PER ITERATION.**
|
||||
Do NOT implement multiple tasks. Do NOT complete an entire phase.
|
||||
Find the FIRST unchecked task, implement ONLY that task, then STOP and report.
|
||||
|
||||
## Project Information
|
||||
- **Project Root:** \`{{projectRoot}}\`
|
||||
- **Spec Location:** \`{{specPath}}\`
|
||||
- Read \`AGENTS.md\` for project-specific guidance if available
|
||||
|
||||
## IMPORTANT: File Locations
|
||||
- Write ALL code files relative to the **project root** (\`{{projectRoot}}\`)
|
||||
- The spec directory (\`{{specPath}}\`) is for documentation ONLY - never write code there
|
||||
- Example: Create \`{{projectRoot}}/src/index.ts\`, NOT \`{{specPath}}/src/index.ts\`
|
||||
|
||||
## Iteration
|
||||
{{iteration}} of {{maxIterations}}
|
||||
|
||||
## Task Discovery Process
|
||||
1. Read \`{{specPath}}/overview.md\` to see all phases
|
||||
2. Find the FIRST phase with an unchecked checkbox (\`- [ ]\` or \`- [/]\`)
|
||||
3. Read that phase's file (e.g., \`phase-1.md\`)
|
||||
4. Check the \`## Prerequisites\` section FIRST
|
||||
5. Find the FIRST unchecked prerequisite (\`- [ ]\`)
|
||||
- If found, that is your task for this iteration
|
||||
- Verify/complete it, then mark \`[x]\` or \`[?]\`
|
||||
6. Only if ALL prerequisites are complete (\`[x]\` or \`[?]\`), find the FIRST unchecked task
|
||||
7. That is your ONE task - implement ONLY that task
|
||||
|
||||
## Checkbox States
|
||||
- \`[ ]\` = incomplete/pending (do the FIRST one you find)
|
||||
- \`[x]\` = complete (skip)
|
||||
- \`[?]\` = assumed complete, couldn't verify (skip)
|
||||
- \`[!]\` = blocked (skip)
|
||||
|
||||
## Implementation Steps
|
||||
1. Read and understand the single task
|
||||
2. Implement it completely
|
||||
3. Validate it works (run tests if applicable and double-check code)
|
||||
4. Mark ONLY that task's checkbox as \`[x]\` in the phase file
|
||||
5. If that was the LAST task in the phase, also mark the phase \`[x]\` in overview.md
|
||||
6. Output your completion marker and STOP
|
||||
|
||||
## Git Policy
|
||||
**DO NOT create git commits.** The orchestration system handles commits automatically after each task completion. Just implement the code and leave changes uncommitted.
|
||||
|
||||
## Completion Markers (REQUIRED FORMAT)
|
||||
Output exactly ONE of these at the end, including the task ID and description:
|
||||
|
||||
**PREREQ_COMPLETE: [prereq_id] - [description]**
|
||||
Example: \`PREREQ_COMPLETE: P1.1 - Verified Phase 1 complete\`
|
||||
|
||||
**PREREQ_ASSUMED: [prereq_id] - [description]**
|
||||
Example: \`PREREQ_ASSUMED: P2.1 - Design approval (cannot verify)\`
|
||||
|
||||
**TASK_COMPLETE: [task_id] - [task_description]**
|
||||
Example: \`TASK_COMPLETE: 1.1 - Initialize project structure\`
|
||||
|
||||
**TASK_BLOCKED: [task_id] - [reason]**
|
||||
Example: \`TASK_BLOCKED: 2.3 - Missing API credentials\`
|
||||
|
||||
**LOOP_COMPLETE**
|
||||
Use only when ALL phases in overview.md are marked complete.
|
||||
|
||||
## Scratchpad Management
|
||||
|
||||
After completing each task, append to \`{{specPath}}/.plan2code-loop/scratchpad.md\`:
|
||||
- Task completed and Phase item reference
|
||||
- Key decisions made and reasoning
|
||||
- Files changed
|
||||
- Any blockers or notes for next iteration
|
||||
|
||||
Keep entries concise. Sacrifice grammar for concision. This file helps future iterations skip exploration.
|
||||
|
||||
If key patterns or learnings were discovered, update \`./AGENTS.md\` if it exists.
|
||||
|
||||
## Previous Session Context
|
||||
{{scratchpadContent}}
|
||||
|
||||
---
|
||||
|
||||
Remember: ONE TASK ONLY. Find it, implement it, mark it done, output TASK_COMPLETE with the task ID and description, then stop.
|
||||
`;
|
||||
|
||||
export const LOOP_PROMPT_TEMPLATE_PHASE = `# PLAN2CODE-LOOP: Autonomous Phase Implementation
|
||||
|
||||
## CRITICAL CONSTRAINT
|
||||
**IMPLEMENT ALL REMAINING TASKS IN THE CURRENT PHASE.**
|
||||
Find the first incomplete phase, then implement every remaining task in that phase before stopping.
|
||||
Complete each task fully before moving to the next task within the phase.
|
||||
## Project Information
|
||||
- **Project Root:** \`{{projectRoot}}\`
|
||||
- **Spec Location:** \`{{specPath}}\`
|
||||
- Read \`AGENTS.md\` for project-specific guidance if available
|
||||
|
||||
## IMPORTANT: File Locations
|
||||
- Write ALL code files relative to the **project root** (\`{{projectRoot}}\`)
|
||||
- The spec directory (\`{{specPath}}\`) is for documentation ONLY - never write code there
|
||||
- Example: Create \`{{projectRoot}}/src/index.ts\`, NOT \`{{specPath}}/src/index.ts\`
|
||||
|
||||
## Iteration
|
||||
{{iteration}} of {{maxIterations}}
|
||||
|
||||
## Phase Discovery Process
|
||||
1. Read \`{{specPath}}/overview.md\` to see all phases
|
||||
2. Find the FIRST phase with an unchecked checkbox (\`- [ ]\` or \`- [/]\`)
|
||||
3. Read that phase's file (e.g., \`phase-1.md\`)
|
||||
4. Check the \`## Prerequisites\` section FIRST
|
||||
5. Complete ALL unchecked prerequisites (\`- [ ]\`) first, in order
|
||||
- Verify/complete each, then mark \`[x]\` or \`[?]\`
|
||||
6. Once ALL prerequisites are complete, implement ALL unchecked tasks in order
|
||||
7. Continue until every task in the phase is marked \`[x]\`
|
||||
|
||||
## Checkbox States
|
||||
- \`[ ]\` = incomplete/pending
|
||||
- \`[x]\` = complete (skip)
|
||||
- \`[?]\` = assumed complete, couldn't verify (skip)
|
||||
- \`[!]\` = blocked (skip, note in scratchpad)
|
||||
|
||||
## Implementation Steps (repeat for EACH task in the phase)
|
||||
1. Read and understand the task
|
||||
2. Implement it completely
|
||||
3. Validate it works (run tests if applicable and double-check code)
|
||||
4. Mark that task's checkbox as \`[x]\` in the phase file
|
||||
5. **Create a git commit** for this task (see Git Policy below)
|
||||
6. Output a TASK_COMPLETE marker for this task
|
||||
7. Move to the next unchecked task in the same phase
|
||||
8. When ALL tasks in the phase are done, mark the phase \`[x]\` in overview.md
|
||||
|
||||
## Git Policy
|
||||
**YOU are responsible for creating git commits after each task.** The orchestration system does NOT handle commits in phase mode.
|
||||
|
||||
After completing each task:
|
||||
\`\`\`bash
|
||||
git add -A
|
||||
git commit -m "<commit message>"
|
||||
\`\`\`
|
||||
|
||||
**Commit message format:**
|
||||
- With JIRA ticket: Use \`-m "Task X.Y: description" -m "{{jiraTicketId}}" -m "AI Assisted"\` (three \`-m\` flags)
|
||||
- Without JIRA ticket: \`-m "Task X.Y: description" -m "AI Assisted"\` (two \`-m\` flags)
|
||||
- ALWAYS include the "AI Assisted" footer as the final \`-m\` flag
|
||||
|
||||
Replace X.Y with the actual task ID and description with a concise summary of what was implemented.
|
||||
|
||||
## Completion Markers (REQUIRED FORMAT)
|
||||
Output one of these **after each task** you complete:
|
||||
|
||||
**PREREQ_COMPLETE: [prereq_id] - [description]**
|
||||
Example: \`PREREQ_COMPLETE: P1.1 - Verified Phase 1 complete\`
|
||||
|
||||
**PREREQ_ASSUMED: [prereq_id] - [description]**
|
||||
Example: \`PREREQ_ASSUMED: P2.1 - Design approval (cannot verify)\`
|
||||
|
||||
**TASK_COMPLETE: [task_id] - [task_description]**
|
||||
Example: \`TASK_COMPLETE: 1.1 - Initialize project structure\`
|
||||
|
||||
**TASK_BLOCKED: [task_id] - [reason]**
|
||||
Example: \`TASK_BLOCKED: 2.3 - Missing API credentials\`
|
||||
If a task is blocked, skip it and continue to the next task.
|
||||
|
||||
After ALL tasks in the phase are complete (or blocked), output:
|
||||
**PHASE_COMPLETE** - if only this phase is done
|
||||
**LOOP_COMPLETE** - if ALL phases in overview.md are now marked complete
|
||||
|
||||
## Scratchpad Management
|
||||
|
||||
After completing each task, append to \`{{specPath}}/.plan2code-loop/scratchpad.md\`:
|
||||
- Task completed and Phase item reference
|
||||
- Key decisions made and reasoning
|
||||
- Files changed
|
||||
- Any blockers or notes for next iteration
|
||||
|
||||
Keep entries concise. Sacrifice grammar for concision. This file helps future iterations skip exploration.
|
||||
|
||||
If key patterns or learnings were discovered, update \`./AGENTS.md\` if it exists.
|
||||
|
||||
## Previous Session Context
|
||||
{{scratchpadContent}}
|
||||
|
||||
---
|
||||
|
||||
Remember: Complete ALL tasks in the current phase. Implement each task, commit it, output TASK_COMPLETE, then continue to the next. Stop only when the phase is done.
|
||||
`;
|
||||
export const LOOP_PROMPT_TEMPLATE = `# PLAN2CODE-LOOP: Autonomous Task Implementation
|
||||
|
||||
## CRITICAL CONSTRAINT
|
||||
**IMPLEMENT EXACTLY ONE TASK PER ITERATION.**
|
||||
Do NOT implement multiple tasks. Do NOT complete an entire phase.
|
||||
Find the FIRST unchecked task, implement ONLY that task, then STOP and report.
|
||||
|
||||
## Project Information
|
||||
- **Project Root:** \`{{projectRoot}}\`
|
||||
- **Spec Location:** \`{{specPath}}\`
|
||||
- Read \`AGENTS.md\` for project-specific guidance if available
|
||||
|
||||
## IMPORTANT: File Locations
|
||||
- Write ALL code files relative to the **project root** (\`{{projectRoot}}\`)
|
||||
- The spec directory (\`{{specPath}}\`) is for documentation ONLY - never write code there
|
||||
- Example: Create \`{{projectRoot}}/src/index.ts\`, NOT \`{{specPath}}/src/index.ts\`
|
||||
|
||||
## Iteration
|
||||
{{iteration}} of {{maxIterations}}
|
||||
|
||||
## Task Discovery Process
|
||||
1. Read \`{{specPath}}/overview.md\` to see all phases
|
||||
2. Find the FIRST phase with an unchecked checkbox (\`- [ ]\` or \`- [/]\`)
|
||||
3. Read that phase's file (e.g., \`phase-1.md\`)
|
||||
4. Check the \`## Prerequisites\` section FIRST
|
||||
5. Find the FIRST unverified prerequisite (no "VERIFIED" or "ASSUMED" annotation)
|
||||
- If found, verify/complete it, then annotate "VERIFIED" or "ASSUMED: [reason]" inline
|
||||
6. Only if ALL prerequisites are verified or assumed, find the FIRST unchecked task (\`- [ ]\`)
|
||||
7. That is your ONE task - implement ONLY that task
|
||||
|
||||
## Checkbox States (Task items only)
|
||||
- \`[ ]\` = incomplete/pending (do the FIRST one you find)
|
||||
- \`[x]\` = complete (skip)
|
||||
- \`[?]\` = assumed complete, couldn't verify (skip)
|
||||
- \`[!]\` = blocked (skip)
|
||||
|
||||
Prerequisites use plain bullets with inline annotations, not checkboxes.
|
||||
|
||||
## Implementation Steps
|
||||
1. Read and understand the single task
|
||||
2. Implement it completely
|
||||
3. Validate it works (run tests if applicable and double-check code)
|
||||
4. Mark ONLY that task's checkbox as \`[x]\` in the phase file
|
||||
5. If that was the LAST task in the phase, also mark the phase \`[x]\` in overview.md
|
||||
6. Output your completion marker and STOP
|
||||
|
||||
## Git Policy
|
||||
**DO NOT create git commits.** The orchestration system handles commits automatically after each task completion. Just implement the code and leave changes uncommitted.
|
||||
|
||||
## Completion Markers (REQUIRED FORMAT)
|
||||
Output exactly ONE of these at the end, including the task ID and description:
|
||||
|
||||
**PREREQ_COMPLETE: [prereq_id] - [description]**
|
||||
Example: \`PREREQ_COMPLETE: P1.1 - Verified Phase 1 complete\`
|
||||
|
||||
**PREREQ_ASSUMED: [prereq_id] - [description]**
|
||||
Example: \`PREREQ_ASSUMED: P2.1 - Design approval (cannot verify)\`
|
||||
|
||||
**TASK_COMPLETE: [task_id] - [task_description]**
|
||||
Example: \`TASK_COMPLETE: 1.1 - Initialize project structure\`
|
||||
|
||||
**TASK_BLOCKED: [task_id] - [reason]**
|
||||
Example: \`TASK_BLOCKED: 2.3 - Missing API credentials\`
|
||||
|
||||
**LOOP_COMPLETE**
|
||||
Use only when ALL phases in overview.md are marked complete.
|
||||
|
||||
## Scratchpad Management
|
||||
|
||||
After completing each task, add a new entry at the **bottom** of \`{{specPath}}/.plan2code-loop/scratchpad.md\`.
|
||||
Never edit, reorganize, or insert into existing content — only append new entries to the end of the file.
|
||||
|
||||
Each entry should include:
|
||||
- Task completed and Phase item reference
|
||||
- Key decisions made and reasoning
|
||||
- Files changed
|
||||
- Any blockers or notes for next iteration
|
||||
|
||||
Keep entries concise. Sacrifice grammar for concision. This file helps future iterations skip exploration.
|
||||
|
||||
If key patterns or learnings were discovered, update \`./AGENTS.md\` if it exists.
|
||||
|
||||
## Previous Session Context
|
||||
{{scratchpadContent}}
|
||||
|
||||
---
|
||||
|
||||
Remember: ONE TASK ONLY. Find it, implement it, mark it done, output TASK_COMPLETE with the task ID and description, then stop.
|
||||
`;
|
||||
|
||||
export const LOOP_PROMPT_TEMPLATE_PHASE = `# PLAN2CODE-LOOP: Autonomous Phase Implementation
|
||||
|
||||
## CRITICAL CONSTRAINT
|
||||
**IMPLEMENT ALL REMAINING TASKS IN THE CURRENT PHASE.**
|
||||
Find the first incomplete phase, then implement every remaining task in that phase before stopping.
|
||||
Complete each task fully before moving to the next task within the phase.
|
||||
|
||||
## Project Information
|
||||
- **Project Root:** \`{{projectRoot}}\`
|
||||
- **Spec Location:** \`{{specPath}}\`
|
||||
- Read \`AGENTS.md\` for project-specific guidance if available
|
||||
|
||||
## IMPORTANT: File Locations
|
||||
- Write ALL code files relative to the **project root** (\`{{projectRoot}}\`)
|
||||
- The spec directory (\`{{specPath}}\`) is for documentation ONLY - never write code there
|
||||
- Example: Create \`{{projectRoot}}/src/index.ts\`, NOT \`{{specPath}}/src/index.ts\`
|
||||
|
||||
## Iteration
|
||||
{{iteration}} of {{maxIterations}}
|
||||
|
||||
## Phase Discovery Process
|
||||
1. Read \`{{specPath}}/overview.md\` to see all phases
|
||||
2. Find the FIRST phase with an unchecked checkbox (\`- [ ]\` or \`- [/]\`)
|
||||
3. Read that phase's file (e.g., \`phase-1.md\`)
|
||||
4. Check the \`## Prerequisites\` section FIRST
|
||||
5. Verify ALL unverified prerequisites first, in order
|
||||
- Annotate each "VERIFIED" or "ASSUMED: [reason]" inline
|
||||
6. Once ALL prerequisites are verified, implement ALL unchecked tasks in order
|
||||
7. Continue until every task in the phase is marked \`[x]\`
|
||||
|
||||
## Checkbox States (Task items only)
|
||||
- \`[ ]\` = incomplete/pending
|
||||
- \`[x]\` = complete (skip)
|
||||
- \`[?]\` = assumed complete, couldn't verify (skip)
|
||||
- \`[!]\` = blocked (skip, note in scratchpad)
|
||||
|
||||
Prerequisites use plain bullets with inline annotations, not checkboxes.
|
||||
|
||||
## Implementation Steps (repeat for EACH task in the phase)
|
||||
1. Read and understand the task
|
||||
2. Implement it completely
|
||||
3. Validate it works (run tests if applicable and double-check code)
|
||||
4. Mark that task's checkbox as \`[x]\` in the phase file
|
||||
5. **Create a git commit** for this task (see Git Policy below)
|
||||
6. Output a TASK_COMPLETE marker for this task
|
||||
7. Move to the next unchecked task in the same phase
|
||||
8. When ALL tasks in the phase are done, mark the phase \`[x]\` in overview.md
|
||||
|
||||
## Git Policy
|
||||
**YOU are responsible for creating git commits after each task.** The orchestration system does NOT handle commits in phase mode.
|
||||
|
||||
After completing each task:
|
||||
\`\`\`bash
|
||||
git add -A
|
||||
git commit -m "<commit message>"
|
||||
\`\`\`
|
||||
|
||||
**Commit message format:**
|
||||
\`\`\`
|
||||
git add -A
|
||||
git commit -m "Task X.Y: description" -m "{{jiraTicketId}}" -m "AI Assisted"
|
||||
\`\`\`
|
||||
- With JIRA ticket: three \`-m\` flags (description, ticket ID, AI Assisted)
|
||||
- Without JIRA ticket: two \`-m\` flags (description, AI Assisted)
|
||||
- ALWAYS include "AI Assisted" as the final \`-m\` flag
|
||||
|
||||
Replace X.Y with the actual task ID and description with a concise summary of what was implemented.
|
||||
|
||||
## Completion Markers (REQUIRED FORMAT)
|
||||
Output one of these **after each task** you complete:
|
||||
|
||||
**PREREQ_COMPLETE: [prereq_id] - [description]**
|
||||
Example: \`PREREQ_COMPLETE: P1.1 - Verified Phase 1 complete\`
|
||||
|
||||
**PREREQ_ASSUMED: [prereq_id] - [description]**
|
||||
Example: \`PREREQ_ASSUMED: P2.1 - Design approval (cannot verify)\`
|
||||
|
||||
**TASK_COMPLETE: [task_id] - [task_description]**
|
||||
Example: \`TASK_COMPLETE: 1.1 - Initialize project structure\`
|
||||
|
||||
**TASK_BLOCKED: [task_id] - [reason]**
|
||||
Example: \`TASK_BLOCKED: 2.3 - Missing API credentials\`
|
||||
If a task is blocked, skip it and continue to the next task.
|
||||
|
||||
After ALL tasks in the phase are complete (or blocked), output:
|
||||
**PHASE_COMPLETE** - if only this phase is done
|
||||
**LOOP_COMPLETE** - if ALL phases in overview.md are now marked complete
|
||||
|
||||
## Scratchpad Management
|
||||
|
||||
After completing each task, add a new entry at the **bottom** of \`{{specPath}}/.plan2code-loop/scratchpad.md\`.
|
||||
Never edit, reorganize, or insert into existing content — only append new entries to the end of the file.
|
||||
|
||||
Each entry should include:
|
||||
- Task completed and Phase item reference
|
||||
- Key decisions made and reasoning
|
||||
- Files changed
|
||||
- Any blockers or notes for next iteration
|
||||
|
||||
Keep entries concise. Sacrifice grammar for concision. This file helps future iterations skip exploration.
|
||||
|
||||
If key patterns or learnings were discovered, update \`./AGENTS.md\` if it exists.
|
||||
|
||||
## Previous Session Context
|
||||
{{scratchpadContent}}
|
||||
|
||||
---
|
||||
|
||||
Remember: Complete ALL tasks in the current phase. Implement each task, commit it, output TASK_COMPLETE, then continue to the next. Stop only when the phase is done.
|
||||
`;
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
export {
|
||||
detectSpecDirectories,
|
||||
getSpecProgress,
|
||||
} from './utils.js';
|
||||
export {
|
||||
detectSpecDirectories,
|
||||
getSpecProgress,
|
||||
} from './utils.js';
|
||||
|
||||
@@ -1,70 +1,70 @@
|
||||
import path from 'path';
|
||||
import fs from 'fs-extra';
|
||||
|
||||
/**
|
||||
* Auto-detect spec directories in the project
|
||||
* Looks for directories containing overview.md
|
||||
*/
|
||||
export async function detectSpecDirectories(cwd: string = process.cwd()): Promise<string[]> {
|
||||
const specsDir = path.join(cwd, 'specs');
|
||||
const specsDirs: string[] = [];
|
||||
|
||||
if (await fs.pathExists(specsDir)) {
|
||||
// Look for overview.md files in subdirectories
|
||||
const entries = await fs.readdir(specsDir, { withFileTypes: true });
|
||||
|
||||
for (const entry of entries) {
|
||||
if (entry.isDirectory()) {
|
||||
const overviewPath = path.join(specsDir, entry.name, 'overview.md');
|
||||
if (await fs.pathExists(overviewPath)) {
|
||||
specsDirs.push(path.join(specsDir, entry.name));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Also check if specs/ itself contains overview.md
|
||||
const rootOverview = path.join(specsDir, 'overview.md');
|
||||
if (await fs.pathExists(rootOverview)) {
|
||||
specsDirs.push(specsDir);
|
||||
}
|
||||
}
|
||||
|
||||
return specsDirs;
|
||||
}
|
||||
|
||||
/**
|
||||
* Simple progress stats by counting phase-*.md files
|
||||
* Used for CLI display only - LLM handles actual task discovery
|
||||
*/
|
||||
export async function getSpecProgress(specPath: string): Promise<{
|
||||
featureName: string;
|
||||
totalPhases: number;
|
||||
}> {
|
||||
const overviewPath = path.join(specPath, 'overview.md');
|
||||
|
||||
// Extract feature name from overview.md
|
||||
let featureName = path.basename(specPath);
|
||||
try {
|
||||
const overviewContent = await fs.readFile(overviewPath, 'utf8');
|
||||
const h1Match = overviewContent.match(/^#\s+(.+)$/m);
|
||||
if (h1Match) {
|
||||
featureName = h1Match[1].trim();
|
||||
}
|
||||
} catch {
|
||||
// Use directory name as fallback
|
||||
}
|
||||
|
||||
// Count phase-*.md files
|
||||
let totalPhases = 0;
|
||||
try {
|
||||
const entries = await fs.readdir(specPath);
|
||||
totalPhases = entries.filter(name => /^phase-\d+\.md$/i.test(name)).length;
|
||||
} catch {
|
||||
// Directory read failed
|
||||
}
|
||||
|
||||
return {
|
||||
featureName,
|
||||
totalPhases,
|
||||
};
|
||||
}
|
||||
import path from 'path';
|
||||
import fs from 'fs-extra';
|
||||
|
||||
/**
|
||||
* Auto-detect spec directories in the project
|
||||
* Looks for directories containing overview.md
|
||||
*/
|
||||
export async function detectSpecDirectories(cwd: string = process.cwd()): Promise<string[]> {
|
||||
const specsDir = path.join(cwd, 'specs');
|
||||
const specsDirs: string[] = [];
|
||||
|
||||
if (await fs.pathExists(specsDir)) {
|
||||
// Look for overview.md files in subdirectories
|
||||
const entries = await fs.readdir(specsDir, { withFileTypes: true });
|
||||
|
||||
for (const entry of entries) {
|
||||
if (entry.isDirectory()) {
|
||||
const overviewPath = path.join(specsDir, entry.name, 'overview.md');
|
||||
if (await fs.pathExists(overviewPath)) {
|
||||
specsDirs.push(path.join(specsDir, entry.name));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Also check if specs/ itself contains overview.md
|
||||
const rootOverview = path.join(specsDir, 'overview.md');
|
||||
if (await fs.pathExists(rootOverview)) {
|
||||
specsDirs.push(specsDir);
|
||||
}
|
||||
}
|
||||
|
||||
return specsDirs;
|
||||
}
|
||||
|
||||
/**
|
||||
* Simple progress stats by counting phase-*.md files
|
||||
* Used for CLI display only - LLM handles actual task discovery
|
||||
*/
|
||||
export async function getSpecProgress(specPath: string): Promise<{
|
||||
featureName: string;
|
||||
totalPhases: number;
|
||||
}> {
|
||||
const overviewPath = path.join(specPath, 'overview.md');
|
||||
|
||||
// Extract feature name from overview.md
|
||||
let featureName = path.basename(specPath);
|
||||
try {
|
||||
const overviewContent = await fs.readFile(overviewPath, 'utf8');
|
||||
const h1Match = overviewContent.match(/^#\s+(.+)$/m);
|
||||
if (h1Match) {
|
||||
featureName = h1Match[1].trim();
|
||||
}
|
||||
} catch {
|
||||
// Use directory name as fallback
|
||||
}
|
||||
|
||||
// Count phase-*.md files
|
||||
let totalPhases = 0;
|
||||
try {
|
||||
const entries = await fs.readdir(specPath);
|
||||
totalPhases = entries.filter(name => /^phase-\d+\.md$/i.test(name)).length;
|
||||
} catch {
|
||||
// Directory read failed
|
||||
}
|
||||
|
||||
return {
|
||||
featureName,
|
||||
totalPhases,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -1,33 +1,35 @@
|
||||
export type LoopMode = 'task' | 'phase';
|
||||
|
||||
export interface SessionConfig {
|
||||
agent: string; // "claude-code" | "copilot-cli"
|
||||
model: string; // Selected model
|
||||
maxIterations: number; // 5-50
|
||||
timeout: number; // Minutes per iteration
|
||||
verbose: boolean;
|
||||
specPath: string; // Path to the spec directory
|
||||
startedAt: string; // ISO timestamp
|
||||
currentIteration: number;
|
||||
jiraTicketId?: string; // JIRA ticket ID for commit messages
|
||||
loopMode: LoopMode; // "task" = one task per loop, "phase" = one phase per loop
|
||||
}
|
||||
|
||||
export interface IterationLogEntry {
|
||||
iteration: number;
|
||||
timestamp: string; // ISO timestamp
|
||||
duration: number; // milliseconds
|
||||
exitCode: number;
|
||||
status: 'running' | 'completed' | 'error' | 'timeout' | 'interrupted' | 'blocked';
|
||||
completionMarker?: string;
|
||||
}
|
||||
|
||||
export type SessionState = 'new' | 'continue' | 'changed';
|
||||
|
||||
export const DEFAULT_CONFIG: Partial<SessionConfig> = {
|
||||
maxIterations: 100,
|
||||
timeout: 30,
|
||||
verbose: false,
|
||||
currentIteration: 0,
|
||||
loopMode: 'task',
|
||||
};
|
||||
export type LoopMode = 'task' | 'phase';
|
||||
|
||||
export interface SessionConfig {
|
||||
agent: string; // "claude-code" | "copilot-cli" | "devin-cli"
|
||||
model: string; // Selected model
|
||||
maxIterations: number; // 5-50
|
||||
timeout: number; // Base timeout in minutes per iteration attempt
|
||||
maxRetries: number; // Max retry attempts per iteration (timeout increments by 30s each retry)
|
||||
verbose: boolean;
|
||||
specPath: string; // Path to the spec directory
|
||||
startedAt: string; // ISO timestamp
|
||||
currentIteration: number;
|
||||
jiraTicketId?: string; // JIRA ticket ID for commit messages
|
||||
loopMode: LoopMode; // "task" = one task per loop, "phase" = one phase per loop
|
||||
}
|
||||
|
||||
export interface IterationLogEntry {
|
||||
iteration: number;
|
||||
timestamp: string; // ISO timestamp
|
||||
duration: number; // milliseconds
|
||||
exitCode: number;
|
||||
status: 'running' | 'completed' | 'error' | 'timeout' | 'interrupted' | 'blocked';
|
||||
completionMarker?: string;
|
||||
}
|
||||
|
||||
export type SessionState = 'new' | 'continue' | 'changed';
|
||||
|
||||
export const DEFAULT_CONFIG: Partial<SessionConfig> = {
|
||||
maxIterations: 100,
|
||||
timeout: 3,
|
||||
maxRetries: 5,
|
||||
verbose: false,
|
||||
currentIteration: 0,
|
||||
loopMode: 'task',
|
||||
};
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
import { createHash } from 'crypto';
|
||||
|
||||
export function computeHash(content: string): string {
|
||||
return createHash('sha256').update(content).digest('hex').slice(0, 16);
|
||||
}
|
||||
|
||||
export function hashesMatch(a: string, b: string): boolean {
|
||||
return a === b;
|
||||
}
|
||||
import { createHash } from 'crypto';
|
||||
|
||||
export function computeHash(content: string): string {
|
||||
return createHash('sha256').update(content).digest('hex').slice(0, 16);
|
||||
}
|
||||
|
||||
export function hashesMatch(a: string, b: string): boolean {
|
||||
return a === b;
|
||||
}
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
export {
|
||||
type SessionConfig,
|
||||
type IterationLogEntry,
|
||||
type SessionState,
|
||||
type LoopMode,
|
||||
DEFAULT_CONFIG,
|
||||
} from './config.js';
|
||||
|
||||
export { StateManager } from './manager.js';
|
||||
export { computeHash, hashesMatch } from './hash.js';
|
||||
export {
|
||||
type SessionConfig,
|
||||
type IterationLogEntry,
|
||||
type SessionState,
|
||||
type LoopMode,
|
||||
DEFAULT_CONFIG,
|
||||
} from './config.js';
|
||||
|
||||
export { StateManager } from './manager.js';
|
||||
export { computeHash, hashesMatch } from './hash.js';
|
||||
|
||||
@@ -1,231 +1,231 @@
|
||||
import path from 'path';
|
||||
import fs from 'fs-extra';
|
||||
import type { SessionConfig, IterationLogEntry, SessionState } from './config.js';
|
||||
import { DEFAULT_CONFIG } from './config.js';
|
||||
import { computeHash, hashesMatch } from './hash.js';
|
||||
|
||||
const SCRATCHPAD_TEMPLATE = `# Scratchpad
|
||||
|
||||
Session notes appended by LLM during implementation.
|
||||
|
||||
---
|
||||
|
||||
`;
|
||||
|
||||
export class StateManager {
|
||||
private readonly stateDir: string;
|
||||
private readonly configPath: string;
|
||||
private readonly scratchpadPath: string;
|
||||
private readonly logPath: string;
|
||||
private readonly hashPath: string;
|
||||
private specPath: string | null = null;
|
||||
|
||||
constructor(cwd: string = process.cwd()) {
|
||||
// Default to project root - will be updated when spec is selected
|
||||
this.stateDir = path.join(cwd, '.plan2code-loop');
|
||||
this.configPath = path.join(this.stateDir, 'config.json');
|
||||
this.scratchpadPath = path.join(this.stateDir, 'scratchpad.md');
|
||||
this.logPath = path.join(this.stateDir, 'iteration.log');
|
||||
this.hashPath = path.join(this.stateDir, 'spec.hash');
|
||||
}
|
||||
|
||||
/**
|
||||
* Set the spec path and update all state paths to be inside the spec directory
|
||||
*/
|
||||
setSpecPath(specPath: string): void {
|
||||
this.specPath = specPath;
|
||||
const stateDir = path.join(specPath, '.plan2code-loop');
|
||||
// Update all paths to be relative to spec directory
|
||||
(this as any).stateDir = stateDir;
|
||||
(this as any).configPath = path.join(stateDir, 'config.json');
|
||||
(this as any).scratchpadPath = path.join(stateDir, 'scratchpad.md');
|
||||
(this as any).logPath = path.join(stateDir, 'iteration.log');
|
||||
(this as any).hashPath = path.join(stateDir, 'spec.hash');
|
||||
}
|
||||
|
||||
// Directory operations
|
||||
|
||||
async ensureStateDir(): Promise<boolean> {
|
||||
const existed = await fs.pathExists(this.stateDir);
|
||||
await fs.ensureDir(this.stateDir);
|
||||
return existed;
|
||||
}
|
||||
|
||||
getStateDir(): string {
|
||||
return this.stateDir;
|
||||
}
|
||||
|
||||
// Config operations
|
||||
|
||||
async readConfig(): Promise<SessionConfig | null> {
|
||||
try {
|
||||
const content = await fs.readFile(this.configPath, 'utf8');
|
||||
return JSON.parse(content) as SessionConfig;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
async writeConfig(config: SessionConfig): Promise<void> {
|
||||
await fs.writeFile(
|
||||
this.configPath,
|
||||
JSON.stringify(config, null, 2),
|
||||
'utf8'
|
||||
);
|
||||
}
|
||||
|
||||
async updateConfig(updates: Partial<SessionConfig>): Promise<SessionConfig> {
|
||||
const existing = await this.readConfig();
|
||||
const updated = { ...DEFAULT_CONFIG, ...existing, ...updates } as SessionConfig;
|
||||
await this.writeConfig(updated);
|
||||
return updated;
|
||||
}
|
||||
|
||||
async incrementIteration(): Promise<number> {
|
||||
const config = await this.readConfig();
|
||||
if (!config) throw new Error('No session config found');
|
||||
config.currentIteration++;
|
||||
await this.writeConfig(config);
|
||||
return config.currentIteration;
|
||||
}
|
||||
|
||||
// Session state detection
|
||||
|
||||
async hasExistingSession(): Promise<boolean> {
|
||||
const [configExists, scratchpadExists] = await Promise.all([
|
||||
fs.pathExists(this.configPath),
|
||||
fs.pathExists(this.scratchpadPath),
|
||||
]);
|
||||
return configExists || scratchpadExists;
|
||||
}
|
||||
|
||||
async detectSessionState(specPath: string): Promise<SessionState> {
|
||||
// Ensure we're using the correct spec path
|
||||
this.setSpecPath(specPath);
|
||||
|
||||
const hasSession = await this.hasExistingSession();
|
||||
if (!hasSession) {
|
||||
return 'new';
|
||||
}
|
||||
|
||||
// Check if the spec path has changed
|
||||
const existingConfig = await this.readConfig();
|
||||
if (existingConfig && existingConfig.specPath !== specPath) {
|
||||
return 'changed';
|
||||
}
|
||||
|
||||
// Check if spec content has changed (using hash of overview.md)
|
||||
const specChanged = await this.hasSpecChanged(specPath);
|
||||
return specChanged ? 'changed' : 'continue';
|
||||
}
|
||||
|
||||
// Hash management for spec change detection
|
||||
|
||||
async computeSpecHash(specPath: string): Promise<string> {
|
||||
const overviewPath = path.join(specPath, 'overview.md');
|
||||
try {
|
||||
const content = await fs.readFile(overviewPath, 'utf8');
|
||||
return computeHash(content);
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
async readStoredHash(): Promise<string | null> {
|
||||
try {
|
||||
return await fs.readFile(this.hashPath, 'utf8');
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
async storeHash(hash: string): Promise<void> {
|
||||
await fs.writeFile(this.hashPath, hash, 'utf8');
|
||||
}
|
||||
|
||||
async hasSpecChanged(specPath: string): Promise<boolean> {
|
||||
const stored = await this.readStoredHash();
|
||||
if (!stored) return true;
|
||||
const current = await this.computeSpecHash(specPath);
|
||||
return !hashesMatch(stored, current);
|
||||
}
|
||||
|
||||
async updateSpecHash(specPath: string): Promise<void> {
|
||||
const hash = await this.computeSpecHash(specPath);
|
||||
await this.storeHash(hash);
|
||||
}
|
||||
|
||||
// Scratchpad operations
|
||||
|
||||
async initializeScratchpad(): Promise<void> {
|
||||
await fs.writeFile(this.scratchpadPath, SCRATCHPAD_TEMPLATE, 'utf8');
|
||||
}
|
||||
|
||||
async readScratchpad(): Promise<string> {
|
||||
try {
|
||||
return await fs.readFile(this.scratchpadPath, 'utf8');
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
async writeScratchpad(content: string): Promise<void> {
|
||||
await fs.writeFile(this.scratchpadPath, content, 'utf8');
|
||||
}
|
||||
|
||||
// Iteration log operations
|
||||
|
||||
async appendIterationLog(entry: IterationLogEntry): Promise<void> {
|
||||
const line = JSON.stringify(entry) + '\n';
|
||||
await fs.appendFile(this.logPath, line, 'utf8');
|
||||
}
|
||||
|
||||
async readIterationLog(): Promise<IterationLogEntry[]> {
|
||||
try {
|
||||
const content = await fs.readFile(this.logPath, 'utf8');
|
||||
return content
|
||||
.trim()
|
||||
.split('\n')
|
||||
.filter(Boolean)
|
||||
.map((line) => JSON.parse(line) as IterationLogEntry);
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
}
|
||||
|
||||
async getLastIteration(): Promise<IterationLogEntry | null> {
|
||||
const log = await this.readIterationLog();
|
||||
return log.length > 0 ? log[log.length - 1] : null;
|
||||
}
|
||||
|
||||
// State clearing and initialization
|
||||
|
||||
async clearState(): Promise<void> {
|
||||
const filesToDelete = [
|
||||
this.scratchpadPath,
|
||||
this.configPath,
|
||||
this.logPath,
|
||||
this.hashPath,
|
||||
];
|
||||
|
||||
await Promise.all(
|
||||
filesToDelete.map((file) => fs.remove(file).catch(() => {}))
|
||||
);
|
||||
}
|
||||
|
||||
async initializeNewSession(config: SessionConfig): Promise<void> {
|
||||
// Ensure spec path is set before initializing
|
||||
this.setSpecPath(config.specPath);
|
||||
|
||||
await this.clearState();
|
||||
await this.ensureStateDir();
|
||||
await this.initializeScratchpad();
|
||||
await this.writeConfig({
|
||||
...config,
|
||||
startedAt: new Date().toISOString(),
|
||||
currentIteration: 0,
|
||||
});
|
||||
const hash = await this.computeSpecHash(config.specPath);
|
||||
await this.storeHash(hash);
|
||||
}
|
||||
}
|
||||
import path from 'path';
|
||||
import fs from 'fs-extra';
|
||||
import type { SessionConfig, IterationLogEntry, SessionState } from './config.js';
|
||||
import { DEFAULT_CONFIG } from './config.js';
|
||||
import { computeHash, hashesMatch } from './hash.js';
|
||||
|
||||
const SCRATCHPAD_TEMPLATE = `# Scratchpad
|
||||
|
||||
Session notes appended by LLM during implementation.
|
||||
|
||||
---
|
||||
|
||||
`;
|
||||
|
||||
export class StateManager {
|
||||
private readonly stateDir: string;
|
||||
private readonly configPath: string;
|
||||
private readonly scratchpadPath: string;
|
||||
private readonly logPath: string;
|
||||
private readonly hashPath: string;
|
||||
private specPath: string | null = null;
|
||||
|
||||
constructor(cwd: string = process.cwd()) {
|
||||
// Default to project root - will be updated when spec is selected
|
||||
this.stateDir = path.join(cwd, '.plan2code-loop');
|
||||
this.configPath = path.join(this.stateDir, 'config.json');
|
||||
this.scratchpadPath = path.join(this.stateDir, 'scratchpad.md');
|
||||
this.logPath = path.join(this.stateDir, 'iteration.log');
|
||||
this.hashPath = path.join(this.stateDir, 'spec.hash');
|
||||
}
|
||||
|
||||
/**
|
||||
* Set the spec path and update all state paths to be inside the spec directory
|
||||
*/
|
||||
setSpecPath(specPath: string): void {
|
||||
this.specPath = specPath;
|
||||
const stateDir = path.join(specPath, '.plan2code-loop');
|
||||
// Update all paths to be relative to spec directory
|
||||
(this as any).stateDir = stateDir;
|
||||
(this as any).configPath = path.join(stateDir, 'config.json');
|
||||
(this as any).scratchpadPath = path.join(stateDir, 'scratchpad.md');
|
||||
(this as any).logPath = path.join(stateDir, 'iteration.log');
|
||||
(this as any).hashPath = path.join(stateDir, 'spec.hash');
|
||||
}
|
||||
|
||||
// Directory operations
|
||||
|
||||
async ensureStateDir(): Promise<boolean> {
|
||||
const existed = await fs.pathExists(this.stateDir);
|
||||
await fs.ensureDir(this.stateDir);
|
||||
return existed;
|
||||
}
|
||||
|
||||
getStateDir(): string {
|
||||
return this.stateDir;
|
||||
}
|
||||
|
||||
// Config operations
|
||||
|
||||
async readConfig(): Promise<SessionConfig | null> {
|
||||
try {
|
||||
const content = await fs.readFile(this.configPath, 'utf8');
|
||||
return JSON.parse(content) as SessionConfig;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
async writeConfig(config: SessionConfig): Promise<void> {
|
||||
await fs.writeFile(
|
||||
this.configPath,
|
||||
JSON.stringify(config, null, 2),
|
||||
'utf8'
|
||||
);
|
||||
}
|
||||
|
||||
async updateConfig(updates: Partial<SessionConfig>): Promise<SessionConfig> {
|
||||
const existing = await this.readConfig();
|
||||
const updated = { ...DEFAULT_CONFIG, ...existing, ...updates } as SessionConfig;
|
||||
await this.writeConfig(updated);
|
||||
return updated;
|
||||
}
|
||||
|
||||
async incrementIteration(): Promise<number> {
|
||||
const config = await this.readConfig();
|
||||
if (!config) throw new Error('No session config found');
|
||||
config.currentIteration++;
|
||||
await this.writeConfig(config);
|
||||
return config.currentIteration;
|
||||
}
|
||||
|
||||
// Session state detection
|
||||
|
||||
async hasExistingSession(): Promise<boolean> {
|
||||
const [configExists, scratchpadExists] = await Promise.all([
|
||||
fs.pathExists(this.configPath),
|
||||
fs.pathExists(this.scratchpadPath),
|
||||
]);
|
||||
return configExists || scratchpadExists;
|
||||
}
|
||||
|
||||
async detectSessionState(specPath: string): Promise<SessionState> {
|
||||
// Ensure we're using the correct spec path
|
||||
this.setSpecPath(specPath);
|
||||
|
||||
const hasSession = await this.hasExistingSession();
|
||||
if (!hasSession) {
|
||||
return 'new';
|
||||
}
|
||||
|
||||
// Check if the spec path has changed
|
||||
const existingConfig = await this.readConfig();
|
||||
if (existingConfig && existingConfig.specPath !== specPath) {
|
||||
return 'changed';
|
||||
}
|
||||
|
||||
// Check if spec content has changed (using hash of overview.md)
|
||||
const specChanged = await this.hasSpecChanged(specPath);
|
||||
return specChanged ? 'changed' : 'continue';
|
||||
}
|
||||
|
||||
// Hash management for spec change detection
|
||||
|
||||
async computeSpecHash(specPath: string): Promise<string> {
|
||||
const overviewPath = path.join(specPath, 'overview.md');
|
||||
try {
|
||||
const content = await fs.readFile(overviewPath, 'utf8');
|
||||
return computeHash(content);
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
async readStoredHash(): Promise<string | null> {
|
||||
try {
|
||||
return await fs.readFile(this.hashPath, 'utf8');
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
async storeHash(hash: string): Promise<void> {
|
||||
await fs.writeFile(this.hashPath, hash, 'utf8');
|
||||
}
|
||||
|
||||
async hasSpecChanged(specPath: string): Promise<boolean> {
|
||||
const stored = await this.readStoredHash();
|
||||
if (!stored) return true;
|
||||
const current = await this.computeSpecHash(specPath);
|
||||
return !hashesMatch(stored, current);
|
||||
}
|
||||
|
||||
async updateSpecHash(specPath: string): Promise<void> {
|
||||
const hash = await this.computeSpecHash(specPath);
|
||||
await this.storeHash(hash);
|
||||
}
|
||||
|
||||
// Scratchpad operations
|
||||
|
||||
async initializeScratchpad(): Promise<void> {
|
||||
await fs.writeFile(this.scratchpadPath, SCRATCHPAD_TEMPLATE, 'utf8');
|
||||
}
|
||||
|
||||
async readScratchpad(): Promise<string> {
|
||||
try {
|
||||
return await fs.readFile(this.scratchpadPath, 'utf8');
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
async writeScratchpad(content: string): Promise<void> {
|
||||
await fs.writeFile(this.scratchpadPath, content, 'utf8');
|
||||
}
|
||||
|
||||
// Iteration log operations
|
||||
|
||||
async appendIterationLog(entry: IterationLogEntry): Promise<void> {
|
||||
const line = JSON.stringify(entry) + '\n';
|
||||
await fs.appendFile(this.logPath, line, 'utf8');
|
||||
}
|
||||
|
||||
async readIterationLog(): Promise<IterationLogEntry[]> {
|
||||
try {
|
||||
const content = await fs.readFile(this.logPath, 'utf8');
|
||||
return content
|
||||
.trim()
|
||||
.split('\n')
|
||||
.filter(Boolean)
|
||||
.map((line) => JSON.parse(line) as IterationLogEntry);
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
}
|
||||
|
||||
async getLastIteration(): Promise<IterationLogEntry | null> {
|
||||
const log = await this.readIterationLog();
|
||||
return log.length > 0 ? log[log.length - 1] : null;
|
||||
}
|
||||
|
||||
// State clearing and initialization
|
||||
|
||||
async clearState(): Promise<void> {
|
||||
const filesToDelete = [
|
||||
this.scratchpadPath,
|
||||
this.configPath,
|
||||
this.logPath,
|
||||
this.hashPath,
|
||||
];
|
||||
|
||||
await Promise.all(
|
||||
filesToDelete.map((file) => fs.remove(file).catch(() => {}))
|
||||
);
|
||||
}
|
||||
|
||||
async initializeNewSession(config: SessionConfig): Promise<void> {
|
||||
// Ensure spec path is set before initializing
|
||||
this.setSpecPath(config.specPath);
|
||||
|
||||
await this.clearState();
|
||||
await this.ensureStateDir();
|
||||
await this.initializeScratchpad();
|
||||
await this.writeConfig({
|
||||
...config,
|
||||
startedAt: new Date().toISOString(),
|
||||
currentIteration: 0,
|
||||
});
|
||||
const hash = await this.computeSpecHash(config.specPath);
|
||||
await this.storeHash(hash);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,169 +1,169 @@
|
||||
// Completion markers for plan2code-loop
|
||||
const COMPLETION_MARKERS = ['TASK_COMPLETE', 'TASK_BLOCKED', 'LOOP_COMPLETE', 'PREREQ_COMPLETE', 'PREREQ_ASSUMED'] as const;
|
||||
|
||||
export type CompletionMarker = (typeof COMPLETION_MARKERS)[number];
|
||||
|
||||
export interface CompletionCheckResult {
|
||||
completed: boolean;
|
||||
marker?: CompletionMarker;
|
||||
taskId?: string; // e.g., "1.1", "2.3"
|
||||
taskName?: string; // e.g., "Initialize project structure"
|
||||
reason?: string; // For TASK_BLOCKED: reason
|
||||
}
|
||||
|
||||
export function checkForCompletion(output: string): CompletionCheckResult {
|
||||
// Check for LOOP_COMPLETE first (highest priority)
|
||||
if (output.includes('LOOP_COMPLETE')) {
|
||||
return { completed: true, marker: 'LOOP_COMPLETE' };
|
||||
}
|
||||
|
||||
// Check for TASK_COMPLETE with task info
|
||||
// Format: TASK_COMPLETE: 1.1 - Task description
|
||||
// Or: TASK_COMPLETE: 1.1: Task description
|
||||
// Or: TASK_COMPLETE[1.1]: Task description
|
||||
const completeMatch = output.match(/TASK_COMPLETE[:\[\s]+(\d+\.\d+)[\]:\-\s]+(.+?)(?:\n|$)/i);
|
||||
if (completeMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'TASK_COMPLETE',
|
||||
taskId: completeMatch[1],
|
||||
taskName: completeMatch[2].trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// Simple TASK_COMPLETE without structured info - try to extract task from context
|
||||
if (output.includes('TASK_COMPLETE')) {
|
||||
// Try to find task info nearby
|
||||
const taskMatch = output.match(/(?:task|completed?)\s+(\d+\.\d+)[:\s]+([^\n]{5,80})/i);
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'TASK_COMPLETE',
|
||||
taskId: taskMatch?.[1],
|
||||
taskName: taskMatch?.[2]?.trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// Check for PREREQ_COMPLETE with prereq info
|
||||
// Format: PREREQ_COMPLETE: P1.1 - Verified Phase 1 complete
|
||||
const prereqMatch = output.match(/PREREQ_COMPLETE[:\[\s]+([^\]:\-\s]+)[\]:\-\s]+(.+?)(?:\n|$)/i);
|
||||
if (prereqMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'PREREQ_COMPLETE',
|
||||
taskId: prereqMatch[1],
|
||||
taskName: prereqMatch[2].trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// Check for PREREQ_ASSUMED with prereq info
|
||||
// Format: PREREQ_ASSUMED: P2.1 - Design approval (cannot verify)
|
||||
const assumedMatch = output.match(/PREREQ_ASSUMED[:\[\s]+([^\]:\-\s]+)[\]:\-\s]+(.+?)(?:\n|$)/i);
|
||||
if (assumedMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'PREREQ_ASSUMED',
|
||||
taskId: assumedMatch[1],
|
||||
taskName: assumedMatch[2].trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// Check for TASK_BLOCKED with task info and reason
|
||||
// Format: TASK_BLOCKED: 1.1 - Reason why blocked
|
||||
const blockedWithTaskMatch = output.match(/TASK_BLOCKED[:\[\s]+(\d+\.\d+)[\]:\-\s]+(.+?)(?:\n|$)/i);
|
||||
if (blockedWithTaskMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'TASK_BLOCKED',
|
||||
taskId: blockedWithTaskMatch[1],
|
||||
reason: blockedWithTaskMatch[2].trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// TASK_BLOCKED with just reason (no task ID)
|
||||
const blockedMatch = output.match(/TASK_BLOCKED:\s*(.+?)(?:\n|$)/);
|
||||
if (blockedMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'TASK_BLOCKED',
|
||||
reason: blockedMatch[1].trim()
|
||||
};
|
||||
}
|
||||
|
||||
// Simple TASK_BLOCKED without any info
|
||||
if (output.includes('TASK_BLOCKED')) {
|
||||
return { completed: true, marker: 'TASK_BLOCKED', reason: 'No reason provided' };
|
||||
}
|
||||
|
||||
return { completed: false };
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract ALL completion markers from output (for phase mode).
|
||||
* Returns an array of all TASK_COMPLETE/TASK_BLOCKED markers found,
|
||||
* plus whether LOOP_COMPLETE or PHASE_COMPLETE was present.
|
||||
*/
|
||||
export function checkForAllCompletions(output: string): {
|
||||
tasks: CompletionCheckResult[];
|
||||
loopComplete: boolean;
|
||||
phaseComplete: boolean;
|
||||
} {
|
||||
const tasks: CompletionCheckResult[] = [];
|
||||
let loopComplete = false;
|
||||
let phaseComplete = false;
|
||||
|
||||
if (output.includes('LOOP_COMPLETE')) {
|
||||
loopComplete = true;
|
||||
}
|
||||
|
||||
if (output.includes('PHASE_COMPLETE')) {
|
||||
phaseComplete = true;
|
||||
}
|
||||
|
||||
// Find all TASK_COMPLETE markers with task info
|
||||
// Format: TASK_COMPLETE: 1.1 - Task description
|
||||
const completeRegex = /TASK_COMPLETE[:\[\s]+(\d+\.\d+)[\]:\-\s]+(.+?)(?:\n|$)/gi;
|
||||
let match: RegExpExecArray | null;
|
||||
while ((match = completeRegex.exec(output)) !== null) {
|
||||
tasks.push({
|
||||
completed: true,
|
||||
marker: 'TASK_COMPLETE',
|
||||
taskId: match[1],
|
||||
taskName: match[2].trim(),
|
||||
});
|
||||
}
|
||||
|
||||
// Find all TASK_BLOCKED markers with task info
|
||||
const blockedRegex = /TASK_BLOCKED[:\[\s]+(\d+\.\d+)[\]:\-\s]+(.+?)(?:\n|$)/gi;
|
||||
while ((match = blockedRegex.exec(output)) !== null) {
|
||||
tasks.push({
|
||||
completed: true,
|
||||
marker: 'TASK_BLOCKED',
|
||||
taskId: match[1],
|
||||
reason: match[2].trim(),
|
||||
});
|
||||
}
|
||||
|
||||
// Find PREREQ_COMPLETE markers
|
||||
const prereqRegex = /PREREQ_COMPLETE[:\[\s]+([^\]:\-\s]+)[\]:\-\s]+(.+?)(?:\n|$)/gi;
|
||||
while ((match = prereqRegex.exec(output)) !== null) {
|
||||
tasks.push({
|
||||
completed: true,
|
||||
marker: 'PREREQ_COMPLETE',
|
||||
taskId: match[1],
|
||||
taskName: match[2].trim(),
|
||||
});
|
||||
}
|
||||
|
||||
// Find PREREQ_ASSUMED markers
|
||||
const assumedRegex = /PREREQ_ASSUMED[:\[\s]+([^\]:\-\s]+)[\]:\-\s]+(.+?)(?:\n|$)/gi;
|
||||
while ((match = assumedRegex.exec(output)) !== null) {
|
||||
tasks.push({
|
||||
completed: true,
|
||||
marker: 'PREREQ_ASSUMED',
|
||||
taskId: match[1],
|
||||
taskName: match[2].trim(),
|
||||
});
|
||||
}
|
||||
|
||||
return { tasks, loopComplete, phaseComplete };
|
||||
}
|
||||
// Completion markers for plan2code-loop
|
||||
const COMPLETION_MARKERS = ['TASK_COMPLETE', 'TASK_BLOCKED', 'LOOP_COMPLETE', 'PREREQ_COMPLETE', 'PREREQ_ASSUMED'] as const;
|
||||
|
||||
export type CompletionMarker = (typeof COMPLETION_MARKERS)[number];
|
||||
|
||||
export interface CompletionCheckResult {
|
||||
completed: boolean;
|
||||
marker?: CompletionMarker;
|
||||
taskId?: string; // e.g., "1.1", "2.3"
|
||||
taskName?: string; // e.g., "Initialize project structure"
|
||||
reason?: string; // For TASK_BLOCKED: reason
|
||||
}
|
||||
|
||||
export function checkForCompletion(output: string): CompletionCheckResult {
|
||||
// Check for LOOP_COMPLETE first (highest priority)
|
||||
if (output.includes('LOOP_COMPLETE')) {
|
||||
return { completed: true, marker: 'LOOP_COMPLETE' };
|
||||
}
|
||||
|
||||
// Check for TASK_COMPLETE with task info
|
||||
// Format: TASK_COMPLETE: 1.1 - Task description
|
||||
// Or: TASK_COMPLETE: 1.1: Task description
|
||||
// Or: TASK_COMPLETE[1.1]: Task description
|
||||
const completeMatch = output.match(/TASK_COMPLETE[:\[\s]+(\d+\.\d+)[\]:\-\s]+(.+?)(?:\n|$)/i);
|
||||
if (completeMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'TASK_COMPLETE',
|
||||
taskId: completeMatch[1],
|
||||
taskName: completeMatch[2].trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// Simple TASK_COMPLETE without structured info - try to extract task from context
|
||||
if (output.includes('TASK_COMPLETE')) {
|
||||
// Try to find task info nearby
|
||||
const taskMatch = output.match(/(?:task|completed?)\s+(\d+\.\d+)[:\s]+([^\n]{5,80})/i);
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'TASK_COMPLETE',
|
||||
taskId: taskMatch?.[1],
|
||||
taskName: taskMatch?.[2]?.trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// Check for PREREQ_COMPLETE with prereq info
|
||||
// Format: PREREQ_COMPLETE: P1.1 - Verified Phase 1 complete
|
||||
const prereqMatch = output.match(/PREREQ_COMPLETE[:\[\s]+([^\]:\-\s]+)[\]:\-\s]+(.+?)(?:\n|$)/i);
|
||||
if (prereqMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'PREREQ_COMPLETE',
|
||||
taskId: prereqMatch[1],
|
||||
taskName: prereqMatch[2].trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// Check for PREREQ_ASSUMED with prereq info
|
||||
// Format: PREREQ_ASSUMED: P2.1 - Design approval (cannot verify)
|
||||
const assumedMatch = output.match(/PREREQ_ASSUMED[:\[\s]+([^\]:\-\s]+)[\]:\-\s]+(.+?)(?:\n|$)/i);
|
||||
if (assumedMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'PREREQ_ASSUMED',
|
||||
taskId: assumedMatch[1],
|
||||
taskName: assumedMatch[2].trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// Check for TASK_BLOCKED with task info and reason
|
||||
// Format: TASK_BLOCKED: 1.1 - Reason why blocked
|
||||
const blockedWithTaskMatch = output.match(/TASK_BLOCKED[:\[\s]+(\d+\.\d+)[\]:\-\s]+(.+?)(?:\n|$)/i);
|
||||
if (blockedWithTaskMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'TASK_BLOCKED',
|
||||
taskId: blockedWithTaskMatch[1],
|
||||
reason: blockedWithTaskMatch[2].trim(),
|
||||
};
|
||||
}
|
||||
|
||||
// TASK_BLOCKED with just reason (no task ID)
|
||||
const blockedMatch = output.match(/TASK_BLOCKED:\s*(.+?)(?:\n|$)/);
|
||||
if (blockedMatch) {
|
||||
return {
|
||||
completed: true,
|
||||
marker: 'TASK_BLOCKED',
|
||||
reason: blockedMatch[1].trim()
|
||||
};
|
||||
}
|
||||
|
||||
// Simple TASK_BLOCKED without any info
|
||||
if (output.includes('TASK_BLOCKED')) {
|
||||
return { completed: true, marker: 'TASK_BLOCKED', reason: 'No reason provided' };
|
||||
}
|
||||
|
||||
return { completed: false };
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract ALL completion markers from output (for phase mode).
|
||||
* Returns an array of all TASK_COMPLETE/TASK_BLOCKED markers found,
|
||||
* plus whether LOOP_COMPLETE or PHASE_COMPLETE was present.
|
||||
*/
|
||||
export function checkForAllCompletions(output: string): {
|
||||
tasks: CompletionCheckResult[];
|
||||
loopComplete: boolean;
|
||||
phaseComplete: boolean;
|
||||
} {
|
||||
const tasks: CompletionCheckResult[] = [];
|
||||
let loopComplete = false;
|
||||
let phaseComplete = false;
|
||||
|
||||
if (output.includes('LOOP_COMPLETE')) {
|
||||
loopComplete = true;
|
||||
}
|
||||
|
||||
if (output.includes('PHASE_COMPLETE')) {
|
||||
phaseComplete = true;
|
||||
}
|
||||
|
||||
// Find all TASK_COMPLETE markers with task info
|
||||
// Format: TASK_COMPLETE: 1.1 - Task description
|
||||
const completeRegex = /TASK_COMPLETE[:\[\s]+(\d+\.\d+)[\]:\-\s]+(.+?)(?:\n|$)/gi;
|
||||
let match: RegExpExecArray | null;
|
||||
while ((match = completeRegex.exec(output)) !== null) {
|
||||
tasks.push({
|
||||
completed: true,
|
||||
marker: 'TASK_COMPLETE',
|
||||
taskId: match[1],
|
||||
taskName: match[2].trim(),
|
||||
});
|
||||
}
|
||||
|
||||
// Find all TASK_BLOCKED markers with task info
|
||||
const blockedRegex = /TASK_BLOCKED[:\[\s]+(\d+\.\d+)[\]:\-\s]+(.+?)(?:\n|$)/gi;
|
||||
while ((match = blockedRegex.exec(output)) !== null) {
|
||||
tasks.push({
|
||||
completed: true,
|
||||
marker: 'TASK_BLOCKED',
|
||||
taskId: match[1],
|
||||
reason: match[2].trim(),
|
||||
});
|
||||
}
|
||||
|
||||
// Find PREREQ_COMPLETE markers
|
||||
const prereqRegex = /PREREQ_COMPLETE[:\[\s]+([^\]:\-\s]+)[\]:\-\s]+(.+?)(?:\n|$)/gi;
|
||||
while ((match = prereqRegex.exec(output)) !== null) {
|
||||
tasks.push({
|
||||
completed: true,
|
||||
marker: 'PREREQ_COMPLETE',
|
||||
taskId: match[1],
|
||||
taskName: match[2].trim(),
|
||||
});
|
||||
}
|
||||
|
||||
// Find PREREQ_ASSUMED markers
|
||||
const assumedRegex = /PREREQ_ASSUMED[:\[\s]+([^\]:\-\s]+)[\]:\-\s]+(.+?)(?:\n|$)/gi;
|
||||
while ((match = assumedRegex.exec(output)) !== null) {
|
||||
tasks.push({
|
||||
completed: true,
|
||||
marker: 'PREREQ_ASSUMED',
|
||||
taskId: match[1],
|
||||
taskName: match[2].trim(),
|
||||
});
|
||||
}
|
||||
|
||||
return { tasks, loopComplete, phaseComplete };
|
||||
}
|
||||
|
||||
@@ -1,136 +1,136 @@
|
||||
import { execa } from 'execa';
|
||||
import { existsSync, readFileSync, writeFileSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { logger } from './logger.js';
|
||||
|
||||
export interface GitCommitOptions {
|
||||
taskName: string;
|
||||
jiraTicketId?: string;
|
||||
cwd?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Check if directory is a git repository
|
||||
*/
|
||||
export async function isGitRepo(cwd: string): Promise<boolean> {
|
||||
const result = await execa('git', ['rev-parse', '--git-dir'], { cwd, reject: false });
|
||||
return result.exitCode === 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* Initialize a git repository if one doesn't exist
|
||||
*/
|
||||
export async function ensureGitRepo(cwd: string): Promise<boolean> {
|
||||
if (await isGitRepo(cwd)) {
|
||||
return true;
|
||||
}
|
||||
|
||||
logger.info('Initializing git repository...');
|
||||
const result = await execa('git', ['init'], { cwd, reject: false });
|
||||
|
||||
if (result.exitCode !== 0) {
|
||||
logger.error(`Failed to initialize git repo: ${result.stderr}`);
|
||||
return false;
|
||||
}
|
||||
|
||||
logger.success('Git repository initialized');
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Required entries for the .gitignore file
|
||||
*/
|
||||
const REQUIRED_GITIGNORE_ENTRIES = ['specs/', 'specs--completed/', '.plan2code-loop', '.plan2code-metrics', 'nul', 'node_modules/'];
|
||||
|
||||
/**
|
||||
* Ensure .gitignore exists with required entries
|
||||
*/
|
||||
export function ensureGitignore(cwd: string): void {
|
||||
const gitignorePath = join(cwd, '.gitignore');
|
||||
let content = '';
|
||||
|
||||
if (existsSync(gitignorePath)) {
|
||||
content = readFileSync(gitignorePath, 'utf-8');
|
||||
}
|
||||
|
||||
const lines = content.split('\n').map(line => line.trim());
|
||||
const missingEntries = REQUIRED_GITIGNORE_ENTRIES.filter(entry => !lines.includes(entry));
|
||||
|
||||
if (missingEntries.length > 0) {
|
||||
const needsNewline = content.length > 0 && !content.endsWith('\n');
|
||||
const newContent = content + (needsNewline ? '\n' : '') + missingEntries.join('\n') + '\n';
|
||||
writeFileSync(gitignorePath, newContent);
|
||||
logger.dim(`Added to .gitignore: ${missingEntries.join(', ')}`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Create a local git commit for a completed task
|
||||
*/
|
||||
export async function createTaskCommit(options: GitCommitOptions): Promise<boolean> {
|
||||
const { taskName, jiraTicketId, cwd = process.cwd() } = options;
|
||||
|
||||
logger.dim(`Git commit check in: ${cwd}`);
|
||||
|
||||
try {
|
||||
// Ensure we have a git repo
|
||||
if (!await ensureGitRepo(cwd)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
// Ensure .gitignore exists with required entries
|
||||
ensureGitignore(cwd);
|
||||
|
||||
// Check if there are any changes to commit
|
||||
const statusResult = await execa('git', ['status', '--porcelain'], { cwd, reject: false });
|
||||
|
||||
// Debug: show what git status returned
|
||||
if (statusResult.stdout?.trim()) {
|
||||
logger.dim(`Git status found changes:\n${statusResult.stdout.slice(0, 500)}`);
|
||||
}
|
||||
|
||||
if (statusResult.exitCode !== 0) {
|
||||
logger.error(`Git status failed: ${statusResult.stderr}`);
|
||||
return false;
|
||||
}
|
||||
|
||||
if (!statusResult.stdout?.trim()) {
|
||||
logger.dim('No changes to commit');
|
||||
return false;
|
||||
}
|
||||
|
||||
// Stage all changes
|
||||
const addResult = await execa('git', ['add', '-A'], { cwd, reject: false });
|
||||
if (addResult.exitCode !== 0) {
|
||||
logger.error(`Git add failed: ${addResult.stderr}`);
|
||||
return false;
|
||||
}
|
||||
|
||||
// Build commit message
|
||||
let commitMessage = taskName;
|
||||
if (jiraTicketId) {
|
||||
commitMessage = `${taskName}\n\n${jiraTicketId}\n\nAI Assisted`;
|
||||
} else {
|
||||
commitMessage = `${taskName}\n\nAI Assisted`;
|
||||
}
|
||||
|
||||
// Create the commit
|
||||
const commitResult = await execa('git', ['commit', '-m', commitMessage], { cwd, reject: false });
|
||||
|
||||
if (commitResult.exitCode !== 0) {
|
||||
// Check if it's just "nothing to commit" vs actual error
|
||||
if (commitResult.stdout?.includes('nothing to commit') || commitResult.stderr?.includes('nothing to commit')) {
|
||||
logger.dim('No changes to commit');
|
||||
return false;
|
||||
}
|
||||
logger.error(`Git commit failed: ${commitResult.stderr || commitResult.stdout}`);
|
||||
return false;
|
||||
}
|
||||
|
||||
logger.success(`Created commit: ${taskName}${jiraTicketId ? ` (${jiraTicketId})` : ''}`);
|
||||
return true;
|
||||
} catch (error) {
|
||||
logger.error(`Failed to create commit: ${error instanceof Error ? error.message : String(error)}`);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
import { execa } from 'execa';
|
||||
import { existsSync, readFileSync, writeFileSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { logger } from './logger.js';
|
||||
|
||||
export interface GitCommitOptions {
|
||||
taskName: string;
|
||||
jiraTicketId?: string;
|
||||
cwd?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Check if directory is a git repository
|
||||
*/
|
||||
export async function isGitRepo(cwd: string): Promise<boolean> {
|
||||
const result = await execa('git', ['rev-parse', '--git-dir'], { cwd, reject: false });
|
||||
return result.exitCode === 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* Initialize a git repository if one doesn't exist
|
||||
*/
|
||||
export async function ensureGitRepo(cwd: string): Promise<boolean> {
|
||||
if (await isGitRepo(cwd)) {
|
||||
return true;
|
||||
}
|
||||
|
||||
logger.info('Initializing git repository...');
|
||||
const result = await execa('git', ['init'], { cwd, reject: false });
|
||||
|
||||
if (result.exitCode !== 0) {
|
||||
logger.error(`Failed to initialize git repo: ${result.stderr}`);
|
||||
return false;
|
||||
}
|
||||
|
||||
logger.success('Git repository initialized');
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Required entries for the .gitignore file
|
||||
*/
|
||||
const REQUIRED_GITIGNORE_ENTRIES = ['specs/', 'specs--completed/', '.plan2code-loop', '.plan2code-metrics', 'nul', 'node_modules/'];
|
||||
|
||||
/**
|
||||
* Ensure .gitignore exists with required entries
|
||||
*/
|
||||
export function ensureGitignore(cwd: string): void {
|
||||
const gitignorePath = join(cwd, '.gitignore');
|
||||
let content = '';
|
||||
|
||||
if (existsSync(gitignorePath)) {
|
||||
content = readFileSync(gitignorePath, 'utf-8');
|
||||
}
|
||||
|
||||
const lines = content.split('\n').map(line => line.trim());
|
||||
const missingEntries = REQUIRED_GITIGNORE_ENTRIES.filter(entry => !lines.includes(entry));
|
||||
|
||||
if (missingEntries.length > 0) {
|
||||
const needsNewline = content.length > 0 && !content.endsWith('\n');
|
||||
const newContent = content + (needsNewline ? '\n' : '') + missingEntries.join('\n') + '\n';
|
||||
writeFileSync(gitignorePath, newContent);
|
||||
logger.dim(`Added to .gitignore: ${missingEntries.join(', ')}`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Create a local git commit for a completed task
|
||||
*/
|
||||
export async function createTaskCommit(options: GitCommitOptions): Promise<boolean> {
|
||||
const { taskName, jiraTicketId, cwd = process.cwd() } = options;
|
||||
|
||||
logger.dim(`Git commit check in: ${cwd}`);
|
||||
|
||||
try {
|
||||
// Ensure we have a git repo
|
||||
if (!await ensureGitRepo(cwd)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
// Ensure .gitignore exists with required entries
|
||||
ensureGitignore(cwd);
|
||||
|
||||
// Check if there are any changes to commit
|
||||
const statusResult = await execa('git', ['status', '--porcelain'], { cwd, reject: false });
|
||||
|
||||
// Debug: show what git status returned
|
||||
if (statusResult.stdout?.trim()) {
|
||||
logger.dim(`Git status found changes:\n${statusResult.stdout.slice(0, 500)}`);
|
||||
}
|
||||
|
||||
if (statusResult.exitCode !== 0) {
|
||||
logger.error(`Git status failed: ${statusResult.stderr}`);
|
||||
return false;
|
||||
}
|
||||
|
||||
if (!statusResult.stdout?.trim()) {
|
||||
logger.dim('No changes to commit');
|
||||
return false;
|
||||
}
|
||||
|
||||
// Stage all changes
|
||||
const addResult = await execa('git', ['add', '-A'], { cwd, reject: false });
|
||||
if (addResult.exitCode !== 0) {
|
||||
logger.error(`Git add failed: ${addResult.stderr}`);
|
||||
return false;
|
||||
}
|
||||
|
||||
// Build commit message
|
||||
let commitMessage = taskName;
|
||||
if (jiraTicketId) {
|
||||
commitMessage = `${taskName}\n\n${jiraTicketId}\nAI Assisted`;
|
||||
} else {
|
||||
commitMessage = `${taskName}\n\nAI Assisted`;
|
||||
}
|
||||
|
||||
// Create the commit
|
||||
const commitResult = await execa('git', ['commit', '-m', commitMessage], { cwd, reject: false });
|
||||
|
||||
if (commitResult.exitCode !== 0) {
|
||||
// Check if it's just "nothing to commit" vs actual error
|
||||
if (commitResult.stdout?.includes('nothing to commit') || commitResult.stderr?.includes('nothing to commit')) {
|
||||
logger.dim('No changes to commit');
|
||||
return false;
|
||||
}
|
||||
logger.error(`Git commit failed: ${commitResult.stderr || commitResult.stdout}`);
|
||||
return false;
|
||||
}
|
||||
|
||||
logger.success(`Created commit: ${taskName}${jiraTicketId ? ` (${jiraTicketId})` : ''}`);
|
||||
return true;
|
||||
} catch (error) {
|
||||
logger.error(`Failed to create commit: ${error instanceof Error ? error.message : String(error)}`);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,19 +1,19 @@
|
||||
export { logger, MASCOT, type Logger } from './logger.js';
|
||||
export {
|
||||
checkForCompletion,
|
||||
checkForAllCompletions,
|
||||
type CompletionMarker,
|
||||
type CompletionCheckResult
|
||||
} from './completion.js';
|
||||
export {
|
||||
executeCommand,
|
||||
type ExecuteOptions,
|
||||
type ExecuteResult
|
||||
} from './process.js';
|
||||
export {
|
||||
createTaskCommit,
|
||||
isGitRepo,
|
||||
ensureGitRepo,
|
||||
ensureGitignore,
|
||||
type GitCommitOptions
|
||||
} from './git.js';
|
||||
export { logger, MASCOT, type Logger } from './logger.js';
|
||||
export {
|
||||
checkForCompletion,
|
||||
checkForAllCompletions,
|
||||
type CompletionMarker,
|
||||
type CompletionCheckResult
|
||||
} from './completion.js';
|
||||
export {
|
||||
executeCommand,
|
||||
type ExecuteOptions,
|
||||
type ExecuteResult
|
||||
} from './process.js';
|
||||
export {
|
||||
createTaskCommit,
|
||||
isGitRepo,
|
||||
ensureGitRepo,
|
||||
ensureGitignore,
|
||||
type GitCommitOptions
|
||||
} from './git.js';
|
||||
|
||||
@@ -1,126 +1,126 @@
|
||||
import chalk from 'chalk';
|
||||
import ora, { type Ora } from 'ora';
|
||||
|
||||
// mascot - our friendly robot assistant
|
||||
export const MASCOT = {
|
||||
// Full mascot for headers
|
||||
full: [
|
||||
' ╭───╮ ',
|
||||
' │ ● │ ',
|
||||
' │ ◡ │ ',
|
||||
' ╰───╯ ',
|
||||
],
|
||||
// Mini mascot for inline use
|
||||
mini: '(◉‿◉)',
|
||||
// Waving mascot for greetings
|
||||
wave: [
|
||||
' ╭───╮ ',
|
||||
' │ ● │ ',
|
||||
' │ ◡ │ ',
|
||||
' ╰───╯ ',
|
||||
],
|
||||
// Celebration mascot for completion
|
||||
celebrate: [
|
||||
' ╭───╮ ',
|
||||
' │ ★ │ ',
|
||||
' │ ◡ │ ',
|
||||
' ╰───╯ ',
|
||||
],
|
||||
};
|
||||
|
||||
export const logger = {
|
||||
info: (message: string) => console.log(chalk.blue('i'), message),
|
||||
success: (message: string) => console.log(chalk.green('+'), message),
|
||||
warning: (message: string) => console.log(chalk.yellow('!'), message),
|
||||
error: (message: string) => console.log(chalk.red('x'), message),
|
||||
|
||||
dim: (message: string) => console.log(chalk.dim(message)),
|
||||
bold: (message: string) => console.log(chalk.bold(message)),
|
||||
|
||||
header: (message: string) => {
|
||||
console.log();
|
||||
console.log(chalk.cyan.bold(message));
|
||||
console.log(chalk.cyan('-'.repeat(message.length)));
|
||||
},
|
||||
|
||||
task: (phase: number, taskId: string, description: string) => {
|
||||
console.log(
|
||||
chalk.cyan(`[Phase ${phase}]`),
|
||||
chalk.yellow(`Task ${taskId}:`),
|
||||
chalk.white(description)
|
||||
);
|
||||
},
|
||||
|
||||
iteration: (num: number, max: number, status: string) => {
|
||||
console.log(
|
||||
chalk.cyan(`[${num}/${max}]`),
|
||||
chalk.white(status)
|
||||
);
|
||||
},
|
||||
|
||||
specContent: (content: string) => {
|
||||
console.log(chalk.dim('-'.repeat(50)));
|
||||
console.log(chalk.white(content));
|
||||
console.log(chalk.dim('-'.repeat(50)));
|
||||
},
|
||||
|
||||
spinner: (text: string): Ora => ora({ text, color: 'cyan' }).start(),
|
||||
|
||||
// Display mascot with optional message
|
||||
mascot: (variant: keyof typeof MASCOT = 'full', message?: string) => {
|
||||
console.log();
|
||||
const mascot = MASCOT[variant];
|
||||
if (Array.isArray(mascot)) {
|
||||
const mascotLines = [...mascot];
|
||||
if (message) {
|
||||
// Add message next to mascot (at the "mouth" line)
|
||||
mascotLines[4] = mascotLines[4] + ' ' + chalk.cyan(message);
|
||||
}
|
||||
mascotLines.forEach((line) => console.log(chalk.yellow(line)));
|
||||
} else {
|
||||
// Mini variant
|
||||
console.log(chalk.yellow(mascot), message ? chalk.cyan(message) : '');
|
||||
}
|
||||
console.log();
|
||||
},
|
||||
|
||||
// Welcome banner with mascot
|
||||
welcome: () => {
|
||||
console.log();
|
||||
console.log(chalk.cyan.bold('═'.repeat(50)));
|
||||
MASCOT.wave.forEach((line, i) => {
|
||||
if (i === 4) {
|
||||
console.log(chalk.yellow(line) + ' ' + chalk.cyan.bold("Hi!"));
|
||||
} else if (i === 5) {
|
||||
console.log(chalk.yellow(line) + ' ' + chalk.dim('I\'m Your Plan2Code assistant'));
|
||||
} else {
|
||||
console.log(chalk.yellow(line));
|
||||
}
|
||||
});
|
||||
console.log(chalk.cyan.bold('═'.repeat(50)));
|
||||
console.log();
|
||||
},
|
||||
|
||||
// Completion celebration with reminder
|
||||
allPhasesComplete: () => {
|
||||
console.log();
|
||||
console.log(chalk.green.bold('═'.repeat(50)));
|
||||
MASCOT.celebrate.forEach((line, i) => {
|
||||
if (i === 4) {
|
||||
console.log(chalk.yellow(line) + ' ' + chalk.green.bold('All phases complete!'));
|
||||
} else if (i === 5) {
|
||||
console.log(chalk.yellow(line) + ' ' + chalk.cyan('Great work!'));
|
||||
} else {
|
||||
console.log(chalk.yellow(line));
|
||||
}
|
||||
});
|
||||
console.log(chalk.green.bold('═'.repeat(50)));
|
||||
console.log();
|
||||
console.log(chalk.cyan.bold('Next Step:'));
|
||||
console.log(chalk.white(' Return to your AI Agent and run the'), chalk.yellow.bold('/plan2code-4--finalize'), chalk.white('step.'));
|
||||
console.log(chalk.dim(' This will ensure quality, completeness, and proper documentation.'));
|
||||
console.log();
|
||||
},
|
||||
};
|
||||
|
||||
export type Logger = typeof logger;
|
||||
import chalk from 'chalk';
|
||||
import ora, { type Ora } from 'ora';
|
||||
|
||||
// mascot - our friendly robot assistant
|
||||
export const MASCOT = {
|
||||
// Full mascot for headers
|
||||
full: [
|
||||
' ╭───╮ ',
|
||||
' │ ● │ ',
|
||||
' │ ◡ │ ',
|
||||
' ╰───╯ ',
|
||||
],
|
||||
// Mini mascot for inline use
|
||||
mini: '(◉‿◉)',
|
||||
// Waving mascot for greetings
|
||||
wave: [
|
||||
' ╭───╮ ',
|
||||
' │ ● │ ',
|
||||
' │ ◡ │ ',
|
||||
' ╰───╯ ',
|
||||
],
|
||||
// Celebration mascot for completion
|
||||
celebrate: [
|
||||
' ╭───╮ ',
|
||||
' │ ★ │ ',
|
||||
' │ ◡ │ ',
|
||||
' ╰───╯ ',
|
||||
],
|
||||
};
|
||||
|
||||
export const logger = {
|
||||
info: (message: string) => console.log(chalk.blue('i'), message),
|
||||
success: (message: string) => console.log(chalk.green('+'), message),
|
||||
warning: (message: string) => console.log(chalk.yellow('!'), message),
|
||||
error: (message: string) => console.log(chalk.red('x'), message),
|
||||
|
||||
dim: (message: string) => console.log(chalk.dim(message)),
|
||||
bold: (message: string) => console.log(chalk.bold(message)),
|
||||
|
||||
header: (message: string) => {
|
||||
console.log();
|
||||
console.log(chalk.cyan.bold(message));
|
||||
console.log(chalk.cyan('-'.repeat(message.length)));
|
||||
},
|
||||
|
||||
task: (phase: number, taskId: string, description: string) => {
|
||||
console.log(
|
||||
chalk.cyan(`[Phase ${phase}]`),
|
||||
chalk.yellow(`Task ${taskId}:`),
|
||||
chalk.white(description)
|
||||
);
|
||||
},
|
||||
|
||||
iteration: (num: number, max: number, status: string) => {
|
||||
console.log(
|
||||
chalk.cyan(`[${num}/${max}]`),
|
||||
chalk.white(status)
|
||||
);
|
||||
},
|
||||
|
||||
specContent: (content: string) => {
|
||||
console.log(chalk.dim('-'.repeat(50)));
|
||||
console.log(chalk.white(content));
|
||||
console.log(chalk.dim('-'.repeat(50)));
|
||||
},
|
||||
|
||||
spinner: (text: string): Ora => ora({ text, color: 'cyan' }).start(),
|
||||
|
||||
// Display mascot with optional message
|
||||
mascot: (variant: keyof typeof MASCOT = 'full', message?: string) => {
|
||||
console.log();
|
||||
const mascot = MASCOT[variant];
|
||||
if (Array.isArray(mascot)) {
|
||||
const mascotLines = [...mascot];
|
||||
if (message) {
|
||||
// Add message next to mascot (at the "mouth" line)
|
||||
mascotLines[4] = mascotLines[4] + ' ' + chalk.cyan(message);
|
||||
}
|
||||
mascotLines.forEach((line) => console.log(chalk.yellow(line)));
|
||||
} else {
|
||||
// Mini variant
|
||||
console.log(chalk.yellow(mascot), message ? chalk.cyan(message) : '');
|
||||
}
|
||||
console.log();
|
||||
},
|
||||
|
||||
// Welcome banner with mascot
|
||||
welcome: () => {
|
||||
console.log();
|
||||
console.log(chalk.cyan.bold('═'.repeat(50)));
|
||||
MASCOT.wave.forEach((line, i) => {
|
||||
if (i === 4) {
|
||||
console.log(chalk.yellow(line) + ' ' + chalk.cyan.bold("Hi!"));
|
||||
} else if (i === 5) {
|
||||
console.log(chalk.yellow(line) + ' ' + chalk.dim('I\'m Your Plan2Code assistant'));
|
||||
} else {
|
||||
console.log(chalk.yellow(line));
|
||||
}
|
||||
});
|
||||
console.log(chalk.cyan.bold('═'.repeat(50)));
|
||||
console.log();
|
||||
},
|
||||
|
||||
// Completion celebration with reminder
|
||||
allPhasesComplete: () => {
|
||||
console.log();
|
||||
console.log(chalk.green.bold('═'.repeat(50)));
|
||||
MASCOT.celebrate.forEach((line, i) => {
|
||||
if (i === 4) {
|
||||
console.log(chalk.yellow(line) + ' ' + chalk.green.bold('All phases complete!'));
|
||||
} else if (i === 5) {
|
||||
console.log(chalk.yellow(line) + ' ' + chalk.cyan('Great work!'));
|
||||
} else {
|
||||
console.log(chalk.yellow(line));
|
||||
}
|
||||
});
|
||||
console.log(chalk.green.bold('═'.repeat(50)));
|
||||
console.log();
|
||||
console.log(chalk.cyan.bold('Next Step:'));
|
||||
console.log(chalk.white(' Return to your AI Agent and run the'), chalk.yellow.bold('/plan2code-4-finalize'), chalk.white('step.'));
|
||||
console.log(chalk.dim(' This will ensure quality, completeness, and proper documentation.'));
|
||||
console.log();
|
||||
},
|
||||
};
|
||||
|
||||
export type Logger = typeof logger;
|
||||
|
||||
@@ -1,81 +1,81 @@
|
||||
import { execa, type Options as ExecaOptions } from 'execa';
|
||||
|
||||
const MAX_OUTPUT_SIZE = 10 * 1024 * 1024; // 10MB
|
||||
|
||||
function truncateOutput(output: string, maxSize: number): string {
|
||||
if (output.length > maxSize) {
|
||||
return output.slice(0, maxSize) + '\n...[truncated]';
|
||||
}
|
||||
return output;
|
||||
}
|
||||
|
||||
export interface ExecuteOptions {
|
||||
command: string;
|
||||
args: string[];
|
||||
cwd: string;
|
||||
timeout: number; // milliseconds
|
||||
env?: Record<string, string>;
|
||||
signal?: AbortSignal; // For cancellation
|
||||
stdin?: string; // Input to pass via stdin as string
|
||||
stdinFile?: string; // Path to file to pipe as stdin
|
||||
}
|
||||
|
||||
export interface ExecuteResult {
|
||||
stdout: string;
|
||||
stderr: string;
|
||||
exitCode: number;
|
||||
timedOut: boolean;
|
||||
cancelled: boolean;
|
||||
duration: number; // milliseconds
|
||||
}
|
||||
|
||||
export async function executeCommand(
|
||||
options: ExecuteOptions
|
||||
): Promise<ExecuteResult> {
|
||||
const startTime = Date.now();
|
||||
|
||||
try {
|
||||
const execaOptions: ExecaOptions = {
|
||||
cwd: options.cwd,
|
||||
timeout: options.timeout,
|
||||
env: { ...process.env, ...options.env },
|
||||
reject: false,
|
||||
all: true,
|
||||
};
|
||||
|
||||
// Add cancellation signal if provided
|
||||
if (options.signal) {
|
||||
(execaOptions as any).cancelSignal = options.signal;
|
||||
}
|
||||
|
||||
// Add stdin input if provided (string or file)
|
||||
if (options.stdin) {
|
||||
(execaOptions as any).input = options.stdin;
|
||||
} else if (options.stdinFile) {
|
||||
(execaOptions as any).inputFile = options.stdinFile;
|
||||
}
|
||||
|
||||
const result = await execa(options.command, options.args, execaOptions);
|
||||
|
||||
return {
|
||||
stdout: truncateOutput(result.stdout || '', MAX_OUTPUT_SIZE),
|
||||
stderr: truncateOutput(result.stderr || '', MAX_OUTPUT_SIZE),
|
||||
exitCode: result.exitCode ?? 1,
|
||||
timedOut: result.timedOut ?? false,
|
||||
cancelled: result.isCanceled ?? false,
|
||||
duration: Date.now() - startTime,
|
||||
};
|
||||
} catch (error: any) {
|
||||
// Check if this was a cancellation
|
||||
const isCancelled = error?.isCanceled || options.signal?.aborted;
|
||||
|
||||
return {
|
||||
stdout: error?.stdout || '',
|
||||
stderr: error?.stderr || (error instanceof Error ? error.message : String(error)),
|
||||
exitCode: isCancelled ? -1 : 1,
|
||||
timedOut: false,
|
||||
cancelled: isCancelled,
|
||||
duration: Date.now() - startTime,
|
||||
};
|
||||
}
|
||||
}
|
||||
import { execa, type Options as ExecaOptions } from 'execa';
|
||||
|
||||
const MAX_OUTPUT_SIZE = 10 * 1024 * 1024; // 10MB
|
||||
|
||||
function truncateOutput(output: string, maxSize: number): string {
|
||||
if (output.length > maxSize) {
|
||||
return output.slice(0, maxSize) + '\n...[truncated]';
|
||||
}
|
||||
return output;
|
||||
}
|
||||
|
||||
export interface ExecuteOptions {
|
||||
command: string;
|
||||
args: string[];
|
||||
cwd: string;
|
||||
timeout: number; // milliseconds
|
||||
env?: Record<string, string>;
|
||||
signal?: AbortSignal; // For cancellation
|
||||
stdin?: string; // Input to pass via stdin as string
|
||||
stdinFile?: string; // Path to file to pipe as stdin
|
||||
}
|
||||
|
||||
export interface ExecuteResult {
|
||||
stdout: string;
|
||||
stderr: string;
|
||||
exitCode: number;
|
||||
timedOut: boolean;
|
||||
cancelled: boolean;
|
||||
duration: number; // milliseconds
|
||||
}
|
||||
|
||||
export async function executeCommand(
|
||||
options: ExecuteOptions
|
||||
): Promise<ExecuteResult> {
|
||||
const startTime = Date.now();
|
||||
|
||||
try {
|
||||
const execaOptions: ExecaOptions = {
|
||||
cwd: options.cwd,
|
||||
timeout: options.timeout,
|
||||
env: { ...process.env, ...options.env },
|
||||
reject: false,
|
||||
all: true,
|
||||
};
|
||||
|
||||
// Add cancellation signal if provided
|
||||
if (options.signal) {
|
||||
(execaOptions as any).cancelSignal = options.signal;
|
||||
}
|
||||
|
||||
// Add stdin input if provided (string or file)
|
||||
if (options.stdin) {
|
||||
(execaOptions as any).input = options.stdin;
|
||||
} else if (options.stdinFile) {
|
||||
(execaOptions as any).inputFile = options.stdinFile;
|
||||
}
|
||||
|
||||
const result = await execa(options.command, options.args, execaOptions);
|
||||
|
||||
return {
|
||||
stdout: truncateOutput(result.stdout || '', MAX_OUTPUT_SIZE),
|
||||
stderr: truncateOutput(result.stderr || '', MAX_OUTPUT_SIZE),
|
||||
exitCode: result.exitCode ?? 1,
|
||||
timedOut: result.timedOut ?? false,
|
||||
cancelled: result.isCanceled ?? false,
|
||||
duration: Date.now() - startTime,
|
||||
};
|
||||
} catch (error: any) {
|
||||
// Check if this was a cancellation
|
||||
const isCancelled = error?.isCanceled || options.signal?.aborted;
|
||||
|
||||
return {
|
||||
stdout: error?.stdout || '',
|
||||
stderr: error?.stderr || (error instanceof Error ? error.message : String(error)),
|
||||
exitCode: isCancelled ? -1 : 1,
|
||||
timedOut: false,
|
||||
cancelled: isCancelled,
|
||||
duration: Date.now() - startTime,
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,20 +1,20 @@
|
||||
{
|
||||
"compilerOptions": {
|
||||
"target": "ES2022",
|
||||
"module": "ESNext",
|
||||
"moduleResolution": "bundler",
|
||||
"lib": ["ES2022"],
|
||||
"outDir": "dist",
|
||||
"rootDir": ".",
|
||||
"strict": true,
|
||||
"esModuleInterop": true,
|
||||
"skipLibCheck": true,
|
||||
"forceConsistentCasingInFileNames": true,
|
||||
"resolveJsonModule": true,
|
||||
"declaration": true,
|
||||
"declarationMap": true,
|
||||
"sourceMap": true
|
||||
},
|
||||
"include": ["src/**/*"],
|
||||
"exclude": ["node_modules", "dist"]
|
||||
}
|
||||
{
|
||||
"compilerOptions": {
|
||||
"target": "ES2022",
|
||||
"module": "ESNext",
|
||||
"moduleResolution": "bundler",
|
||||
"lib": ["ES2022"],
|
||||
"outDir": "dist",
|
||||
"rootDir": ".",
|
||||
"strict": true,
|
||||
"esModuleInterop": true,
|
||||
"skipLibCheck": true,
|
||||
"forceConsistentCasingInFileNames": true,
|
||||
"resolveJsonModule": true,
|
||||
"declaration": true,
|
||||
"declarationMap": true,
|
||||
"sourceMap": true
|
||||
},
|
||||
"include": ["src/**/*"],
|
||||
"exclude": ["node_modules", "dist"]
|
||||
}
|
||||
|
||||
@@ -1,15 +1,15 @@
|
||||
import { defineConfig } from 'tsup';
|
||||
|
||||
export default defineConfig({
|
||||
entry: {
|
||||
'bin/plan2code-loop': 'src/bin/plan2code-loop.ts',
|
||||
index: 'src/index.ts',
|
||||
},
|
||||
format: ['esm'],
|
||||
dts: true,
|
||||
clean: true,
|
||||
sourcemap: true,
|
||||
banner: {
|
||||
js: '#!/usr/bin/env node',
|
||||
},
|
||||
});
|
||||
import { defineConfig } from 'tsup';
|
||||
|
||||
export default defineConfig({
|
||||
entry: {
|
||||
'bin/plan2code-loop': 'src/bin/plan2code-loop.ts',
|
||||
index: 'src/index.ts',
|
||||
},
|
||||
format: ['esm'],
|
||||
dts: false,
|
||||
clean: true,
|
||||
sourcemap: true,
|
||||
banner: {
|
||||
js: '#!/usr/bin/env node',
|
||||
},
|
||||
});
|
||||
|
||||
@@ -6,7 +6,7 @@ Recursive self-improvement toolchain for plan2code contributors. Collects metric
|
||||
|
||||
```bash
|
||||
# Install (one-time, from the plan2code root)
|
||||
node install.js # → "I" (Install All) includes metrics
|
||||
node install.js # → "A" (Install All + dev tools) includes metrics
|
||||
# → or "M" (Metrics only) under the CUSTOM sub-menu
|
||||
|
||||
# After finishing any project spec (steps 1-4):
|
||||
@@ -78,7 +78,7 @@ From the plan2code repo root:
|
||||
node install.js
|
||||
```
|
||||
|
||||
Choose **I** (Install All) to install prompts, loop CLI, and metrics together. Or choose **C** (Custom) then **M** (Metrics) to install just the metrics CLI.
|
||||
Choose **A** (Install All + dev tools) to install prompts, loop CLI, bot, metrics, and status line together. Or choose **C** (Custom) then **M** (Metrics) to install just the metrics CLI.
|
||||
|
||||
The installer handles `npm install`, `npm run build`, and global linking automatically. After install, `plan2code-metrics` is available from any directory.
|
||||
|
||||
@@ -116,9 +116,10 @@ plan2code-metrics
|
||||
| **Collect metrics** | Parse a completed project spec and extract step-by-step metrics into a run JSON |
|
||||
| **Import run data** | Copy a run JSON from another project into the local metrics store |
|
||||
| **View metrics status** | Show aggregated metrics with health indicators and generation deltas |
|
||||
| **Run analysis** | AI-powered diagnosis of weak steps (requires Claude Code or Copilot CLI) |
|
||||
| **Run analysis** | AI-powered diagnosis of weak steps (requires Claude Code, GitHub Copilot CLI, or Devin CLI) |
|
||||
| **Generate improvement proposal** | AI generates surgical prompt edits based on a diagnosis |
|
||||
| **Review and apply** | Interactive diff review to accept/reject individual edits |
|
||||
| **Fetch community submissions** | List open community-feedback GitHub issues, parse + validate their METRICS_JSON payload, import into the local run store, and close them |
|
||||
|
||||
### Each Time You Finish a Spec
|
||||
|
||||
@@ -256,6 +257,7 @@ The analysis and improvement steps require an AI agent. Two backends are support
|
||||
|---------|---------|-------|
|
||||
| **Claude Code** | `claude` | Uses `--print` mode. Recommended. |
|
||||
| **GitHub Copilot CLI** | `copilot` | Uses stdin piping with `--allow-all-tools -s`. |
|
||||
| **Devin CLI** | `devin` | Uses `--print --prompt-file <file> --permission-mode dangerous`. |
|
||||
|
||||
Model selection is interactive — choose from available models when prompted.
|
||||
|
||||
@@ -271,7 +273,8 @@ src/
|
||||
├── analyzer.ts # AI diagnosis via LLM invocation
|
||||
├── improver.ts # AI improvement proposal + validation
|
||||
├── applier.ts # Interactive diff review + file patching
|
||||
├── invoke-llm.ts # Unified LLM invocation (Claude Code / Copilot CLI)
|
||||
├── community.ts # Community-feedback issue parsing + ingestion
|
||||
├── invoke-llm.ts # Unified LLM invocation (Claude / Copilot / Devin)
|
||||
├── index.ts # Public API exports
|
||||
└── prompts/
|
||||
├── analyze.md # AI prompt template for diagnosis
|
||||
|
||||
@@ -1,6 +1,9 @@
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import { avg, rate, buildCohortKey, backfillPromptVersions } from './aggregator.js';
|
||||
import type { PromptVersions } from './types.js';
|
||||
import fs from 'fs';
|
||||
import os from 'os';
|
||||
import path from 'path';
|
||||
import { describe, it, expect, afterEach } from 'vitest';
|
||||
import { avg, rate, buildCohortKey, backfillPromptVersions, cohortKeyForRun, aggregate } from './aggregator.js';
|
||||
import type { PromptVersions, RunMetrics } from './types.js';
|
||||
|
||||
// ── avg() ─────────────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -152,3 +155,99 @@ describe('buildCohortKey', () => {
|
||||
expect(buildCohortKey(ordered)).toBe(buildCohortKey(reversed));
|
||||
});
|
||||
});
|
||||
|
||||
// ── Run fixtures for cohort keying / aggregation ──────────────────────────────
|
||||
|
||||
function makeRun(overrides: Partial<RunMetrics> = {}): RunMetrics {
|
||||
return {
|
||||
schema_version: '1.0',
|
||||
run_id: 'run-20260101-000000-0000',
|
||||
plan2code_version: '1.17.0',
|
||||
prompt_versions: { ...FULL_VERSIONS },
|
||||
project: { name: 'proj', started_at: null, completed_at: null },
|
||||
step1_plan: { present: false, final_confidence: null, confidence_breakdown: null, clarification_rounds: null, tech_stack_revision_rounds: null, verification_gaps_found: null, functional_requirements_count: null, non_functional_requirements_count: null, risk_count: null, phase_count: null },
|
||||
step2_document: { present: false, total_tasks: null, tasks_per_phase: null, phase_count: null, parallel_groups_identified: null, requirement_coverage_percent: null, verification_items_added: null },
|
||||
step3_implement: { present: false, task_completion_rate: null, tasks_completed: null, tasks_total: null, blocker_count: null },
|
||||
step4_finalize: { present: false, completion_rate_at_audit: null, verification_failures_found: null, documentation_updates_needed: null, archival_succeeded: null },
|
||||
user_feedback: null,
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
// ── cohortKeyForRun() ─────────────────────────────────────────────────────────
|
||||
|
||||
describe('cohortKeyForRun', () => {
|
||||
it('keys local runs by the prompt-version hash (unchanged from buildCohortKey)', () => {
|
||||
const run = makeRun({ source: 'local' });
|
||||
expect(cohortKeyForRun(run)).toBe(buildCohortKey(run.prompt_versions));
|
||||
});
|
||||
|
||||
it('treats a run with no source as local', () => {
|
||||
const run = makeRun();
|
||||
delete run.source;
|
||||
expect(cohortKeyForRun(run)).toBe(buildCohortKey(run.prompt_versions));
|
||||
});
|
||||
|
||||
it('keys community runs by plan2code_version, ignoring prompt fingerprints', () => {
|
||||
const run = makeRun({ source: 'community', plan2code_version: '1.17.0' });
|
||||
expect(cohortKeyForRun(run)).toBe('community:v1.17.0');
|
||||
});
|
||||
|
||||
it('groups two community runs of the same version together regardless of prompt fingerprint', () => {
|
||||
const a = makeRun({ source: 'community', plan2code_version: '1.17.0', prompt_versions: { ...FULL_VERSIONS } });
|
||||
const b = makeRun({ source: 'community', plan2code_version: '1.17.0', prompt_versions: backfillPromptVersions({} as PromptVersions) });
|
||||
expect(cohortKeyForRun(a)).toBe(cohortKeyForRun(b));
|
||||
});
|
||||
|
||||
it('separates community runs from different versions', () => {
|
||||
const a = makeRun({ source: 'community', plan2code_version: '1.17.0' });
|
||||
const b = makeRun({ source: 'community', plan2code_version: '1.18.0' });
|
||||
expect(cohortKeyForRun(a)).not.toBe(cohortKeyForRun(b));
|
||||
});
|
||||
});
|
||||
|
||||
// ── aggregate() cohort separation ─────────────────────────────────────────────
|
||||
|
||||
describe('aggregate', () => {
|
||||
let tmpDir: string;
|
||||
|
||||
afterEach(() => {
|
||||
if (tmpDir) fs.rmSync(tmpDir, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
function writeRuns(runs: RunMetrics[]): { runsDir: string; outPath: string } {
|
||||
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'plan2code-agg-test-'));
|
||||
const runsDir = path.join(tmpDir, 'runs');
|
||||
fs.mkdirSync(runsDir, { recursive: true });
|
||||
for (const run of runs) {
|
||||
fs.writeFileSync(path.join(runsDir, `${run.run_id}.json`), JSON.stringify(run), 'utf8');
|
||||
}
|
||||
return { runsDir, outPath: path.join(tmpDir, 'aggregated.json') };
|
||||
}
|
||||
|
||||
it('places local and community runs of the same version in separate cohorts', () => {
|
||||
const local = makeRun({ run_id: 'run-20260101-000001-0001', source: 'local' });
|
||||
const community = makeRun({ run_id: 'run-20260101-000002-0002', source: 'community' });
|
||||
const { runsDir, outPath } = writeRuns([local, community]);
|
||||
|
||||
const result = aggregate(runsDir, outPath);
|
||||
|
||||
expect(result.total_runs).toBe(2);
|
||||
expect(result.cohorts).toHaveLength(2);
|
||||
const communityCohort = result.cohorts.find(c => c.source === 'community');
|
||||
const localCohort = result.cohorts.find(c => c.source === 'local');
|
||||
expect(communityCohort?.cohort_key).toBe('community:v1.17.0');
|
||||
expect(localCohort?.cohort_key).toBe(buildCohortKey(local.prompt_versions));
|
||||
});
|
||||
|
||||
it('never selects a community cohort as current when a local cohort exists', () => {
|
||||
// Community run sorts last by run_id, but current must stay on the local cohort.
|
||||
const local = makeRun({ run_id: 'run-20260101-000001-0001', source: 'local' });
|
||||
const community = makeRun({ run_id: 'run-29991231-235959-9999', source: 'community' });
|
||||
const { runsDir, outPath } = writeRuns([local, community]);
|
||||
|
||||
const result = aggregate(runsDir, outPath);
|
||||
|
||||
expect(result.current_cohort_key).toBe(buildCohortKey(local.prompt_versions));
|
||||
});
|
||||
});
|
||||
|
||||
@@ -35,6 +35,26 @@ export function buildCohortKey(promptVersions: PromptVersions): string {
|
||||
return crypto.createHash('sha256').update(str).digest('hex').slice(0, 12);
|
||||
}
|
||||
|
||||
/**
|
||||
* Cohort key for a single run, chosen by the run's origin.
|
||||
*
|
||||
* Local runs are keyed by their prompt-file SHA-256 fingerprint, which captures
|
||||
* in-development prompt edits that share one unreleased version.
|
||||
*
|
||||
* Community submissions can't reproduce that byte-exact hash: they carry the
|
||||
* installed, platform-transformed prompts (not the raw src/*.md the local
|
||||
* collector hashes), and the payload is LLM-generated. They are keyed instead
|
||||
* by `plan2code_version` -- a reliable, byte-comparison-free identifier, since
|
||||
* every released version ships a fixed set of prompts. This keeps community
|
||||
* cohorts free of the CRLF/whitespace fragility a cross-machine hash would have.
|
||||
*/
|
||||
export function cohortKeyForRun(run: RunMetrics): string {
|
||||
if (run.source === 'community') {
|
||||
return `community:v${run.plan2code_version}`;
|
||||
}
|
||||
return buildCohortKey(run.prompt_versions);
|
||||
}
|
||||
|
||||
/**
|
||||
* Backfill missing PromptVersions fields for old run files (pre-v1.1).
|
||||
*/
|
||||
@@ -99,13 +119,9 @@ function buildCohort(runs: RunMetrics[], cohortKey: string): CohortMetrics {
|
||||
const avgReqCoverage = avg(runs.map(r => r.step2_document.requirement_coverage_percent));
|
||||
const avgVerifItems = avg(runs.map(r => r.step2_document.verification_items_added));
|
||||
|
||||
// Step 3 (loop only)
|
||||
const loopRuns = runs.filter(r => r.step3_implement.used_loop_mode);
|
||||
const avgCompletionRate = avg(loopRuns.map(r => r.step3_implement.task_completion_rate));
|
||||
const avgBlockerCount = avg(loopRuns.map(r => r.step3_implement.blocker_count));
|
||||
const avgTotalIter = avg(loopRuns.map(r => r.step3_implement.total_iterations));
|
||||
const avgIterDuration = avg(loopRuns.map(r => r.step3_implement.avg_iteration_duration_ms));
|
||||
const avgMarkerSuccess = avg(loopRuns.map(r => r.step3_implement.completion_marker_success_rate));
|
||||
// Step 3 averages
|
||||
const avgCompletionRate = avg(runs.map(r => r.step3_implement.task_completion_rate));
|
||||
const avgBlockerCount = avg(runs.map(r => r.step3_implement.blocker_count));
|
||||
|
||||
// Step 4 averages
|
||||
const avgCompletionAtAudit = avg(runs.map(r => r.step4_finalize.completion_rate_at_audit));
|
||||
@@ -120,6 +136,7 @@ function buildCohort(runs: RunMetrics[], cohortKey: string): CohortMetrics {
|
||||
|
||||
return {
|
||||
cohort_key: cohortKey,
|
||||
source: runs[0].source ?? 'local',
|
||||
prompt_versions: runs[0].prompt_versions,
|
||||
run_count: runs.length,
|
||||
run_ids: runIds,
|
||||
@@ -140,12 +157,8 @@ function buildCohort(runs: RunMetrics[], cohortKey: string): CohortMetrics {
|
||||
avg_requirement_coverage_percent: avgReqCoverage,
|
||||
avg_verification_items_added: avgVerifItems,
|
||||
|
||||
loop_run_count: loopRuns.length,
|
||||
avg_task_completion_rate: avgCompletionRate,
|
||||
avg_blocker_count: avgBlockerCount,
|
||||
avg_total_iterations: avgTotalIter,
|
||||
avg_iteration_duration_ms: avgIterDuration,
|
||||
avg_completion_marker_success_rate: avgMarkerSuccess,
|
||||
|
||||
avg_completion_rate_at_audit: avgCompletionAtAudit,
|
||||
avg_verification_failures_found: avgVerifFailures,
|
||||
@@ -165,7 +178,7 @@ export function aggregate(runsDir: string, outputPath: string): AggregatedMetric
|
||||
// Group by cohort key
|
||||
const cohortMap = new Map<string, RunMetrics[]>();
|
||||
for (const run of runs) {
|
||||
const key = buildCohortKey(run.prompt_versions);
|
||||
const key = cohortKeyForRun(run);
|
||||
if (!cohortMap.has(key)) cohortMap.set(key, []);
|
||||
cohortMap.get(key)!.push(run);
|
||||
}
|
||||
@@ -177,9 +190,14 @@ export function aggregate(runsDir: string, outputPath: string): AggregatedMetric
|
||||
}
|
||||
cohorts.sort((a, b) => a.first_seen.localeCompare(b.first_seen));
|
||||
|
||||
// Determine current cohort (most recent)
|
||||
const currentCohortKey = cohorts.length > 0
|
||||
? cohorts[cohorts.length - 1].cohort_key
|
||||
// Determine current cohort (most recent). Prefer local cohorts so an
|
||||
// ingested community submission never becomes the maintainer's "current
|
||||
// generation" for self-improvement; fall back to all cohorts if there are
|
||||
// no local runs yet.
|
||||
const localCohorts = cohorts.filter(c => c.source !== 'community');
|
||||
const currentPool = localCohorts.length > 0 ? localCohorts : cohorts;
|
||||
const currentCohortKey = currentPool.length > 0
|
||||
? currentPool[currentPool.length - 1].cohort_key
|
||||
: null;
|
||||
|
||||
const aggregated: AggregatedMetrics = {
|
||||
@@ -208,13 +226,10 @@ export function loadAggregated(outputPath: string): AggregatedMetrics | null {
|
||||
}
|
||||
|
||||
/**
|
||||
* Import a single run JSON from another project into the local runs dir.
|
||||
* Returns true if imported, false if already present.
|
||||
* Write a run to the local runs dir, deduped by run_id filename.
|
||||
* Returns true if written, false if a file for that run_id already existed.
|
||||
*/
|
||||
export function importRun(runJsonPath: string, runsDir: string): boolean {
|
||||
const content = fs.readFileSync(runJsonPath, 'utf8');
|
||||
const run = JSON.parse(content) as RunMetrics;
|
||||
run.prompt_versions = backfillPromptVersions(run.prompt_versions);
|
||||
export function writeRunFile(run: RunMetrics, runsDir: string): boolean {
|
||||
const destPath = path.join(runsDir, `${run.run_id}.json`);
|
||||
|
||||
if (fs.existsSync(destPath)) {
|
||||
@@ -225,3 +240,14 @@ export function importRun(runJsonPath: string, runsDir: string): boolean {
|
||||
fs.writeFileSync(destPath, JSON.stringify(run, null, 2), 'utf8');
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Import a single run JSON from another project into the local runs dir.
|
||||
* Returns true if imported, false if already present.
|
||||
*/
|
||||
export function importRun(runJsonPath: string, runsDir: string): boolean {
|
||||
const content = fs.readFileSync(runJsonPath, 'utf8');
|
||||
const run = JSON.parse(content) as RunMetrics;
|
||||
run.prompt_versions = backfillPromptVersions(run.prompt_versions);
|
||||
return writeRunFile(run, runsDir);
|
||||
}
|
||||
|
||||
@@ -16,14 +16,14 @@ const ANALYZE_PROMPT_PATH = new URL('../src/prompts/analyze.md', import.meta.url
|
||||
function readPromptFiles(plan2codeRoot: string): Record<string, string> {
|
||||
const srcDir = path.join(plan2codeRoot, 'src');
|
||||
const promptFiles = [
|
||||
'plan2code-1--plan.md',
|
||||
'plan2code-1b--revise-plan.md',
|
||||
'plan2code-2--document.md',
|
||||
'plan2code-3--implement.md',
|
||||
'plan2code-4--finalize.md',
|
||||
'plan2code---init.md',
|
||||
'plan2code---init-update.md',
|
||||
'plan2code---quick-task.md',
|
||||
'plan2code-1-plan.md',
|
||||
'plan2code-1b-revise-plan.md',
|
||||
'plan2code-2-document.md',
|
||||
'plan2code-3-implement.md',
|
||||
'plan2code-4-finalize.md',
|
||||
'plan2code-init.md',
|
||||
'plan2code-init-update.md',
|
||||
'plan2code-quick-task.md',
|
||||
];
|
||||
|
||||
const contents: Record<string, string> = {};
|
||||
@@ -57,12 +57,12 @@ export interface AnalyzerOptions {
|
||||
aggregatedPath: string; // Path to aggregated.json
|
||||
plan2codeRoot: string; // Path to plan2code repo root
|
||||
proposalsDir: string; // Where to save diagnosis output
|
||||
model?: string; // AI model to use (default: claude-opus-4-6)
|
||||
model?: string; // AI model to use (default: agent's default)
|
||||
agent?: AgentType; // Agent to use (default: claude-code)
|
||||
}
|
||||
|
||||
export async function runAnalysis(opts: AnalyzerOptions): Promise<string> {
|
||||
const { aggregatedPath, plan2codeRoot, proposalsDir, model = 'claude-opus-4-6', agent = 'claude-code' } = opts;
|
||||
const { aggregatedPath, plan2codeRoot, proposalsDir, model = 'default', agent = 'claude-code' } = opts;
|
||||
|
||||
// Load aggregated metrics
|
||||
let aggregated: AggregatedMetrics | null = null;
|
||||
@@ -102,7 +102,7 @@ export async function runAnalysis(opts: AnalyzerOptions): Promise<string> {
|
||||
});
|
||||
|
||||
// Invoke Claude
|
||||
console.log(`\nInvoking AI analysis (model: ${model})...`);
|
||||
console.log(`\nInvoking AI analysis (agent: ${agent}, model: ${model === 'default' ? 'user default' : model})...`);
|
||||
console.log('This may take a minute...\n');
|
||||
|
||||
let diagnosisContent: string;
|
||||
|
||||
@@ -5,18 +5,48 @@
|
||||
*/
|
||||
|
||||
import fs from 'fs';
|
||||
import os from 'os';
|
||||
import path from 'path';
|
||||
import { select, input, confirm } from '@inquirer/prompts';
|
||||
import chalk from 'chalk';
|
||||
import ora from 'ora';
|
||||
import { execa } from 'execa';
|
||||
import { collectRun } from './collector.js';
|
||||
import { aggregate, loadAggregated, importRun, loadRunFiles } from './aggregator.js';
|
||||
import { aggregate, loadAggregated, importRun, loadRunFiles, writeRunFile } from './aggregator.js';
|
||||
import { runAnalysis } from './analyzer.js';
|
||||
import { generateImprovement } from './improver.js';
|
||||
import { reviewAndApply } from './applier.js';
|
||||
import type { AggregatedMetrics, CohortMetrics } from './types.js';
|
||||
import { METRIC_TARGETS } from './types.js';
|
||||
import { AGENTS, type AgentType } from './invoke-llm.js';
|
||||
import { listCommunityIssues, closeIssue, ingestCommunityIssues } from './community.js';
|
||||
|
||||
// ── Session state (set at startup via interactive prompts) ───────────────────
|
||||
|
||||
let resolvedPlan2CodeRoot = '';
|
||||
let resolvedProjectRoot = '';
|
||||
|
||||
// ── Persisted user config (~/.plan2code-metrics.json) ──────────────────────
|
||||
|
||||
const USER_CONFIG_PATH = path.join(os.homedir(), '.plan2code-metrics.json');
|
||||
|
||||
interface UserConfig {
|
||||
plan2codeRepoPath?: string;
|
||||
}
|
||||
|
||||
function loadUserConfig(): UserConfig {
|
||||
try {
|
||||
return JSON.parse(fs.readFileSync(USER_CONFIG_PATH, 'utf8'));
|
||||
} catch {
|
||||
return {};
|
||||
}
|
||||
}
|
||||
|
||||
function saveUserConfig(config: UserConfig): void {
|
||||
try {
|
||||
fs.writeFileSync(USER_CONFIG_PATH, JSON.stringify(config, null, 2), 'utf8');
|
||||
} catch { /* best-effort */ }
|
||||
}
|
||||
|
||||
// ── Defaults ──────────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -25,7 +55,7 @@ const DEFAULT_RUNS_SUBDIR = 'runs';
|
||||
const DEFAULT_AGGREGATED_FILE = 'aggregated.json';
|
||||
const DEFAULT_PROPOSALS_SUBDIR = 'proposals';
|
||||
|
||||
function getMetricsDirs(baseDir = process.cwd()) {
|
||||
function getMetricsDirs(baseDir = resolvedProjectRoot) {
|
||||
const metricsDir = path.join(baseDir, DEFAULT_METRICS_DIR);
|
||||
return {
|
||||
metricsDir,
|
||||
@@ -41,8 +71,8 @@ function detectPlan2CodeRoot(): string | null {
|
||||
for (let i = 0; i < 5; i++) {
|
||||
const srcDir = path.join(dir, 'src');
|
||||
if (
|
||||
fs.existsSync(path.join(srcDir, 'plan2code-1--plan.md')) &&
|
||||
fs.existsSync(path.join(srcDir, 'plan2code-2--document.md'))
|
||||
fs.existsSync(path.join(srcDir, 'plan2code-1-plan.md')) &&
|
||||
fs.existsSync(path.join(srcDir, 'plan2code-2-document.md'))
|
||||
) {
|
||||
return dir;
|
||||
}
|
||||
@@ -89,13 +119,7 @@ async function selectAgentAndModel(): Promise<{ agent: AgentType; model: string
|
||||
choices: Object.values(AGENTS).map(a => ({ name: a.displayName, value: a.name })),
|
||||
});
|
||||
|
||||
const agentDef = AGENTS[agentChoice];
|
||||
const model = await select({
|
||||
message: 'AI model:',
|
||||
choices: agentDef.models.map(m => ({ name: m.label, value: m.value })),
|
||||
});
|
||||
|
||||
return { agent: agentChoice, model };
|
||||
return { agent: agentChoice, model: AGENTS[agentChoice].defaultModel };
|
||||
}
|
||||
|
||||
// ── Flow: Collect metrics ─────────────────────────────────────────────────────
|
||||
@@ -107,8 +131,8 @@ async function flowCollect(): Promise<void> {
|
||||
const sourceChoice = await select({
|
||||
message: 'Where is the completed project spec?',
|
||||
choices: [
|
||||
{ name: 'Active spec directory (specs/<feature-name>/)', value: 'active' },
|
||||
{ name: 'Archived spec (specs--completed/<feature-name>/)', value: 'archived' },
|
||||
{ name: 'Active spec directory (specs/<feature-name>/)', value: 'active' },
|
||||
{ name: 'Custom path', value: 'custom' },
|
||||
],
|
||||
});
|
||||
@@ -118,10 +142,10 @@ async function flowCollect(): Promise<void> {
|
||||
|
||||
if (sourceChoice === 'active' || sourceChoice === 'archived') {
|
||||
const baseSubdir = sourceChoice === 'active' ? 'specs' : 'specs--completed';
|
||||
const baseDir = path.join(process.cwd(), baseSubdir);
|
||||
const baseDir = path.join(resolvedProjectRoot, baseSubdir);
|
||||
|
||||
if (!fs.existsSync(baseDir)) {
|
||||
console.log(chalk.red(`No ${baseSubdir}/ directory found in ${process.cwd()}`));
|
||||
console.log(chalk.red(`No ${baseSubdir}/ directory found in ${resolvedProjectRoot}`));
|
||||
return;
|
||||
}
|
||||
|
||||
@@ -152,7 +176,7 @@ async function flowCollect(): Promise<void> {
|
||||
}
|
||||
|
||||
// Detect plan2code root
|
||||
const plan2codeRoot = detectPlan2CodeRoot() ?? process.cwd();
|
||||
const plan2codeRoot = resolvedPlan2CodeRoot;
|
||||
const plan2codeVersion = detectPlan2CodeVersion(plan2codeRoot);
|
||||
|
||||
const { runsDir, aggregatedPath } = getMetricsDirs();
|
||||
@@ -179,7 +203,7 @@ async function flowCollect(): Promise<void> {
|
||||
console.log(chalk.bold('Collected:'));
|
||||
console.log(` Step 1 (Plan): ${metrics.step1_plan.present ? chalk.green('✓') : chalk.gray('—')}`);
|
||||
console.log(` Step 2 (Document): ${metrics.step2_document.present ? chalk.green('✓') : chalk.gray('—')}`);
|
||||
console.log(` Step 3 (Implement): ${metrics.step3_implement.present ? (metrics.step3_implement.used_loop_mode ? chalk.green('✓ (loop)') : chalk.yellow('✓ (manual)')) : chalk.gray('—')}`);
|
||||
console.log(` Step 3 (Implement): ${metrics.step3_implement.present ? chalk.green('✓') : chalk.gray('—')}`);
|
||||
console.log(` Step 4 (Finalize): ${metrics.step4_finalize.present ? chalk.green('✓') : chalk.gray('—')}`);
|
||||
console.log(` User Feedback: ${metrics.user_feedback ? chalk.green(`✓ (${metrics.user_feedback.overall_rating}/10)`) : chalk.gray('—')}`);
|
||||
|
||||
@@ -299,18 +323,37 @@ async function flowViewStatus(): Promise<void> {
|
||||
return;
|
||||
}
|
||||
|
||||
const localCount = aggregated.cohorts.filter(c => c.source !== 'community').length;
|
||||
const communityCount = aggregated.cohorts.length - localCount;
|
||||
|
||||
console.log();
|
||||
console.log(chalk.bold(`Total runs: ${aggregated.total_runs} | Generations: ${aggregated.cohorts.length}`));
|
||||
console.log(chalk.bold(
|
||||
`Total runs: ${aggregated.total_runs} | Generations: ${localCount}` +
|
||||
(communityCount > 0 ? ` | Community cohorts: ${communityCount}` : '')
|
||||
));
|
||||
console.log(chalk.gray(`Last updated: ${aggregated.last_updated}`));
|
||||
|
||||
// Community runs carry no wall-clock timestamps, so their first_seen/last_seen
|
||||
// fall back to the run_id string; only render a Period line for real ISO dates.
|
||||
const isIsoDate = (s?: string): boolean => !!s && /^\d{4}-\d{2}-\d{2}/.test(s);
|
||||
|
||||
for (let i = 0; i < aggregated.cohorts.length; i++) {
|
||||
const cohort = aggregated.cohorts[i];
|
||||
const isCurrent = cohort.cohort_key === aggregated.current_cohort_key;
|
||||
const label = isCurrent ? chalk.bold.green('[CURRENT]') : '';
|
||||
const isCommunity = cohort.source === 'community';
|
||||
|
||||
console.log();
|
||||
console.log(chalk.bold(`Generation ${i + 1} (sha:${cohort.cohort_key}) — ${cohort.run_count} run(s) ${label}`));
|
||||
console.log(chalk.gray(` Period: ${cohort.first_seen?.slice(0, 10)} → ${cohort.last_seen?.slice(0, 10)}`));
|
||||
if (isCommunity) {
|
||||
// cohort_key is already `community:v<version>` — no sha: prefix.
|
||||
console.log(chalk.bold(`Community feedback (${cohort.cohort_key}) — ${cohort.run_count} run(s) ${label}`));
|
||||
} else {
|
||||
console.log(chalk.bold(`Generation ${i + 1} (sha:${cohort.cohort_key}) — ${cohort.run_count} run(s) ${label}`));
|
||||
}
|
||||
if (isIsoDate(cohort.first_seen)) {
|
||||
const end = isIsoDate(cohort.last_seen) ? cohort.last_seen.slice(0, 10) : cohort.first_seen.slice(0, 10);
|
||||
console.log(chalk.gray(` Period: ${cohort.first_seen.slice(0, 10)} → ${end}`));
|
||||
}
|
||||
|
||||
// Step 1 metrics
|
||||
if (cohort.avg_confidence != null || cohort.avg_clarification_rounds != null) {
|
||||
@@ -334,15 +377,13 @@ async function flowViewStatus(): Promise<void> {
|
||||
console.log(` avg_verification_items: ${metricStatus('avg_verification_items_added', cohort.avg_verification_items_added)}`);
|
||||
}
|
||||
|
||||
// Step 3 metrics (loop only)
|
||||
if (cohort.loop_run_count > 0) {
|
||||
console.log(chalk.bold(` Step 3 (Implement) — ${cohort.loop_run_count} loop run(s):`));
|
||||
// Step 3 metrics
|
||||
if (cohort.avg_task_completion_rate != null || cohort.avg_blocker_count != null) {
|
||||
console.log(chalk.bold(' Step 3 (Implement):'));
|
||||
if (cohort.avg_task_completion_rate != null)
|
||||
console.log(` avg_task_completion_rate: ${metricStatus('avg_task_completion_rate', cohort.avg_task_completion_rate)}`);
|
||||
if (cohort.avg_blocker_count != null)
|
||||
console.log(` avg_blocker_count: ${metricStatus('avg_blocker_count', cohort.avg_blocker_count)}`);
|
||||
if (cohort.avg_completion_marker_success_rate != null)
|
||||
console.log(` avg_marker_success_rate: ${metricStatus('avg_completion_marker_success_rate', cohort.avg_completion_marker_success_rate)}`);
|
||||
}
|
||||
|
||||
// Step 4 metrics
|
||||
@@ -363,8 +404,10 @@ async function flowViewStatus(): Promise<void> {
|
||||
console.log(` feedback_count: ${chalk.white(String(cohort.feedback_count))}`);
|
||||
}
|
||||
|
||||
// Compare with previous generation
|
||||
if (i > 0) {
|
||||
// Compare with the previous cohort -- but only within the same population.
|
||||
// Community and local cohorts measure different things; a cross-source
|
||||
// delta (e.g. a community cohort vs the last local generation) is noise.
|
||||
if (i > 0 && aggregated.cohorts[i - 1].source === cohort.source) {
|
||||
const prev = aggregated.cohorts[i - 1];
|
||||
const deltas: string[] = [];
|
||||
if (cohort.avg_confidence != null && prev.avg_confidence != null) {
|
||||
@@ -376,7 +419,7 @@ async function flowViewStatus(): Promise<void> {
|
||||
deltas.push(`completion ${d >= 0 ? chalk.green(`▲${(d * 100).toFixed(1)}%`) : chalk.red(`▼${(Math.abs(d) * 100).toFixed(1)}%`)}`);
|
||||
}
|
||||
if (deltas.length > 0) {
|
||||
console.log(chalk.gray(` vs Gen ${i}: ${deltas.join(' ')}`));
|
||||
console.log(chalk.gray(` ${isCommunity ? 'vs prior version' : `vs Gen ${i}`}: ${deltas.join(' ')}`));
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -402,7 +445,7 @@ async function flowRunAnalysis(): Promise<string | null> {
|
||||
console.log(chalk.bold.cyan('── Run Analysis (Diagnose Weak Steps) ──'));
|
||||
|
||||
const { aggregatedPath, proposalsDir } = getMetricsDirs();
|
||||
const plan2codeRoot = detectPlan2CodeRoot() ?? process.cwd();
|
||||
const plan2codeRoot = resolvedPlan2CodeRoot;
|
||||
|
||||
const aggregated = loadAggregated(aggregatedPath);
|
||||
if (!aggregated || aggregated.total_runs === 0) {
|
||||
@@ -460,7 +503,7 @@ async function flowGenerateProposal(): Promise<void> {
|
||||
console.log(chalk.bold.cyan('── Generate Improvement Proposal ──'));
|
||||
|
||||
const { aggregatedPath, proposalsDir, runsDir } = getMetricsDirs();
|
||||
const plan2codeRoot = detectPlan2CodeRoot() ?? process.cwd();
|
||||
const plan2codeRoot = resolvedPlan2CodeRoot;
|
||||
|
||||
// Find diagnosis files
|
||||
let diagFiles: string[] = [];
|
||||
@@ -540,7 +583,7 @@ async function flowReviewAndApply(): Promise<void> {
|
||||
console.log(chalk.bold.cyan('── Review and Apply a Proposal ──'));
|
||||
|
||||
const { proposalsDir } = getMetricsDirs();
|
||||
const plan2codeRoot = detectPlan2CodeRoot() ?? process.cwd();
|
||||
const plan2codeRoot = resolvedPlan2CodeRoot;
|
||||
|
||||
if (!fs.existsSync(proposalsDir)) {
|
||||
console.log(chalk.yellow('No proposals directory found. Generate a proposal first.'));
|
||||
@@ -577,6 +620,187 @@ async function flowReviewAndApply(): Promise<void> {
|
||||
});
|
||||
}
|
||||
|
||||
// ── Flow: Delete metrics data ─────────────────────────────────────────────────
|
||||
|
||||
async function flowDelete(): Promise<void> {
|
||||
console.log();
|
||||
console.log(chalk.bold.cyan('── Delete Metrics Data ──'));
|
||||
|
||||
const { metricsDir, runsDir, aggregatedPath, proposalsDir } = getMetricsDirs();
|
||||
|
||||
if (!fs.existsSync(metricsDir)) {
|
||||
console.log(chalk.yellow('No metrics data found (.plan2code-metrics/ does not exist).'));
|
||||
return;
|
||||
}
|
||||
|
||||
const scope = await select({
|
||||
message: 'What do you want to delete?',
|
||||
choices: [
|
||||
{ name: 'Delete specific run(s)', value: 'select' },
|
||||
{ name: 'Delete all runs and aggregated data', value: 'runs' },
|
||||
{ name: 'Delete everything (runs, aggregated data, proposals)', value: 'all' },
|
||||
{ name: 'Cancel', value: 'cancel' },
|
||||
],
|
||||
});
|
||||
|
||||
if (scope === 'cancel') return;
|
||||
|
||||
if (scope === 'select') {
|
||||
const runs = loadRunFiles(runsDir);
|
||||
if (runs.length === 0) {
|
||||
console.log(chalk.yellow('No run files found.'));
|
||||
return;
|
||||
}
|
||||
|
||||
const choices = runs.map(r => ({
|
||||
name: `${r.run_id} ${r.project.name} v${r.plan2code_version}`,
|
||||
value: r.run_id,
|
||||
}));
|
||||
|
||||
// Select runs one at a time since @inquirer/prompts select is single-choice
|
||||
const toDelete: string[] = [];
|
||||
let selecting = true;
|
||||
while (selecting) {
|
||||
const remaining = choices.filter(c => !toDelete.includes(c.value));
|
||||
if (remaining.length === 0) break;
|
||||
|
||||
const chosen = await select({
|
||||
message: `Select a run to delete (${toDelete.length} selected so far):`,
|
||||
choices: [
|
||||
...remaining,
|
||||
{ name: toDelete.length > 0 ? `Done selecting (delete ${toDelete.length})` : 'Cancel', value: '__done__' },
|
||||
],
|
||||
});
|
||||
|
||||
if (chosen === '__done__') {
|
||||
selecting = false;
|
||||
} else {
|
||||
toDelete.push(chosen);
|
||||
console.log(chalk.gray(` + ${chosen}`));
|
||||
}
|
||||
}
|
||||
|
||||
if (toDelete.length === 0) return;
|
||||
|
||||
const confirmed = await confirm({
|
||||
message: `Delete ${toDelete.length} run(s)? This cannot be undone.`,
|
||||
default: false,
|
||||
});
|
||||
if (!confirmed) return;
|
||||
|
||||
let deleted = 0;
|
||||
for (const runId of toDelete) {
|
||||
const filePath = path.join(runsDir, `${runId}.json`);
|
||||
try {
|
||||
fs.unlinkSync(filePath);
|
||||
deleted++;
|
||||
} catch {
|
||||
console.log(chalk.yellow(` Could not delete ${runId}.json`));
|
||||
}
|
||||
}
|
||||
|
||||
console.log(chalk.green(`✓ Deleted ${deleted} run(s).`));
|
||||
|
||||
// Re-aggregate with remaining runs
|
||||
const remainingRuns = loadRunFiles(runsDir);
|
||||
if (remainingRuns.length > 0) {
|
||||
aggregate(runsDir, aggregatedPath);
|
||||
console.log(chalk.gray(` Aggregated metrics updated (${remainingRuns.length} runs remaining).`));
|
||||
} else {
|
||||
try { fs.unlinkSync(aggregatedPath); } catch { /* ignore */ }
|
||||
console.log(chalk.gray(' No runs remaining — aggregated data removed.'));
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
// scope === 'runs' or 'all'
|
||||
const label = scope === 'all'
|
||||
? 'ALL metrics data (runs, aggregated data, and proposals)'
|
||||
: 'all runs and aggregated data';
|
||||
|
||||
const confirmed = await confirm({
|
||||
message: `Delete ${label}? This cannot be undone.`,
|
||||
default: false,
|
||||
});
|
||||
if (!confirmed) return;
|
||||
|
||||
// Delete run files
|
||||
let runCount = 0;
|
||||
if (fs.existsSync(runsDir)) {
|
||||
const files = fs.readdirSync(runsDir).filter(f => f.endsWith('.json'));
|
||||
for (const f of files) {
|
||||
try { fs.unlinkSync(path.join(runsDir, f)); runCount++; } catch { /* ignore */ }
|
||||
}
|
||||
}
|
||||
|
||||
// Delete aggregated file
|
||||
try { fs.unlinkSync(aggregatedPath); } catch { /* ignore */ }
|
||||
|
||||
console.log(chalk.green(`✓ Deleted ${runCount} run(s) and aggregated data.`));
|
||||
|
||||
if (scope === 'all') {
|
||||
// Delete proposals
|
||||
let proposalCount = 0;
|
||||
if (fs.existsSync(proposalsDir)) {
|
||||
const files = fs.readdirSync(proposalsDir);
|
||||
for (const f of files) {
|
||||
try { fs.unlinkSync(path.join(proposalsDir, f)); proposalCount++; } catch { /* ignore */ }
|
||||
}
|
||||
try { fs.rmdirSync(proposalsDir); } catch { /* ignore */ }
|
||||
}
|
||||
console.log(chalk.green(`✓ Deleted ${proposalCount} proposal file(s).`));
|
||||
}
|
||||
}
|
||||
|
||||
// ── Flow: Fetch community submissions ────────────────────────────────────────
|
||||
|
||||
const COMMUNITY_REPO = 'jparkerweb/plan2code';
|
||||
|
||||
async function flowFetchCommunitySubmissions(): Promise<void> {
|
||||
console.log();
|
||||
console.log(chalk.bold.cyan('── Fetch Community Submissions ──'));
|
||||
|
||||
try {
|
||||
await execa('gh', ['auth', 'status']);
|
||||
} catch {
|
||||
console.log(chalk.red('`gh` CLI not found or not authenticated — install/auth `gh` to use this feature.'));
|
||||
return;
|
||||
}
|
||||
|
||||
let issues;
|
||||
try {
|
||||
issues = await listCommunityIssues(COMMUNITY_REPO);
|
||||
} catch (err) {
|
||||
console.log(chalk.red(`Failed to list community-feedback issues: ${err instanceof Error ? err.message : String(err)}`));
|
||||
return;
|
||||
}
|
||||
|
||||
if (issues.length === 0) {
|
||||
console.log(chalk.yellow('No open community-feedback issues found.'));
|
||||
return;
|
||||
}
|
||||
|
||||
const { runsDir, aggregatedPath } = getMetricsDirs();
|
||||
|
||||
const tally = await ingestCommunityIssues(issues, COMMUNITY_REPO, runsDir, { writeRunFile, closeIssue });
|
||||
|
||||
for (const num of tally.malformedIssues) {
|
||||
console.log(chalk.yellow(` Skipping issue #${num}: malformed or missing METRICS_JSON payload.`));
|
||||
}
|
||||
for (const num of tally.closeFailedIssues) {
|
||||
console.log(chalk.yellow(` Imported issue #${num} but failed to close it (still open; will retry next fetch).`));
|
||||
}
|
||||
|
||||
if (tally.imported > 0) {
|
||||
aggregate(runsDir, aggregatedPath);
|
||||
}
|
||||
|
||||
console.log();
|
||||
console.log(chalk.green(
|
||||
`✓ Imported: ${tally.imported} Skipped (duplicate): ${tally.skippedDuplicate} Skipped (malformed): ${tally.skippedMalformed} Closed: ${tally.closed} Close failed: ${tally.closeFailed}`
|
||||
));
|
||||
}
|
||||
|
||||
// ── Main menu ─────────────────────────────────────────────────────────────────
|
||||
|
||||
export async function runCLI(): Promise<void> {
|
||||
@@ -585,18 +809,58 @@ export async function runCLI(): Promise<void> {
|
||||
console.log(chalk.gray('Recursive self-improvement toolchain for plan2code contributors'));
|
||||
console.log();
|
||||
|
||||
// Check if we're in (or near) a plan2code repo
|
||||
const plan2codeRoot = detectPlan2CodeRoot();
|
||||
if (!plan2codeRoot) {
|
||||
console.log(chalk.yellow('⚠ Could not detect plan2code repository (src/plan2code-*.md not found nearby).'));
|
||||
console.log(chalk.gray(' Prompt version hashing and analysis will be limited.'));
|
||||
console.log();
|
||||
} else {
|
||||
const version = detectPlan2CodeVersion(plan2codeRoot);
|
||||
console.log(chalk.gray(`plan2code root: ${plan2codeRoot} (v${version})`));
|
||||
console.log();
|
||||
// ── Prompt for paths ────────────────────────────────────────────────────────
|
||||
|
||||
const userConfig = loadUserConfig();
|
||||
const savedRoot = userConfig.plan2codeRepoPath && fs.existsSync(userConfig.plan2codeRepoPath)
|
||||
? userConfig.plan2codeRepoPath
|
||||
: null;
|
||||
const detectedRoot = savedRoot ?? detectPlan2CodeRoot();
|
||||
|
||||
const s2cPath = await input({
|
||||
message: 'Path to plan2code repo:',
|
||||
default: detectedRoot ?? undefined,
|
||||
validate: (v) => {
|
||||
if (!v.trim()) return 'Path is required';
|
||||
const resolved = path.resolve(v.trim());
|
||||
if (!fs.existsSync(resolved)) return 'Directory not found';
|
||||
return true;
|
||||
},
|
||||
});
|
||||
resolvedPlan2CodeRoot = path.resolve(s2cPath.trim());
|
||||
|
||||
// Persist the path for next run
|
||||
if (resolvedPlan2CodeRoot !== userConfig.plan2codeRepoPath) {
|
||||
saveUserConfig({ ...userConfig, plan2codeRepoPath: resolvedPlan2CodeRoot });
|
||||
}
|
||||
|
||||
// Warn if the path doesn't look like a plan2code repo
|
||||
const hasSrcPrompts =
|
||||
fs.existsSync(path.join(resolvedPlan2CodeRoot, 'src', 'plan2code-1-plan.md')) &&
|
||||
fs.existsSync(path.join(resolvedPlan2CodeRoot, 'src', 'plan2code-2-document.md'));
|
||||
if (!hasSrcPrompts) {
|
||||
console.log(chalk.yellow('⚠ No src/plan2code-*.md prompts found at that path. Hashing and analysis may be limited.'));
|
||||
}
|
||||
|
||||
const projPath = await input({
|
||||
message: 'Path to project repo (metrics source):',
|
||||
default: process.cwd(),
|
||||
validate: (v) => {
|
||||
if (!v.trim()) return 'Path is required';
|
||||
const resolved = path.resolve(v.trim());
|
||||
if (!fs.existsSync(resolved)) return 'Directory not found';
|
||||
return true;
|
||||
},
|
||||
});
|
||||
resolvedProjectRoot = path.resolve(projPath.trim());
|
||||
|
||||
// Display resolved paths
|
||||
const version = detectPlan2CodeVersion(resolvedPlan2CodeRoot);
|
||||
console.log();
|
||||
console.log(chalk.gray(`plan2code repo: ${resolvedPlan2CodeRoot} (v${version})`));
|
||||
console.log(chalk.gray(`project repo: ${resolvedProjectRoot}`));
|
||||
console.log();
|
||||
|
||||
let continueLoop = true;
|
||||
while (continueLoop) {
|
||||
const action = await select({
|
||||
@@ -608,6 +872,8 @@ export async function runCLI(): Promise<void> {
|
||||
{ name: 'Run analysis (diagnose weak steps)', value: 'analyze' },
|
||||
{ name: 'Generate improvement proposal', value: 'propose' },
|
||||
{ name: 'Review and apply a proposal', value: 'apply' },
|
||||
{ name: 'Fetch community submissions', value: 'fetch-community' },
|
||||
{ name: chalk.red('Delete metrics data'), value: 'delete' },
|
||||
{ name: 'Exit', value: 'exit' },
|
||||
],
|
||||
});
|
||||
@@ -631,6 +897,12 @@ export async function runCLI(): Promise<void> {
|
||||
case 'apply':
|
||||
await flowReviewAndApply();
|
||||
break;
|
||||
case 'fetch-community':
|
||||
await flowFetchCommunitySubmissions();
|
||||
break;
|
||||
case 'delete':
|
||||
await flowDelete();
|
||||
break;
|
||||
case 'exit':
|
||||
continueLoop = false;
|
||||
break;
|
||||
|
||||
@@ -36,6 +36,57 @@ function readFileSafe(filePath: string): string | null {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract METRICS_JSON HTML comment blocks from a file.
|
||||
* Format: <!-- METRICS_JSON {"key": value, ...} -->
|
||||
* A file may contain multiple blocks (e.g., overview.md has document + finalize).
|
||||
* If `stepFilter` is provided, returns only the block with matching "step" field.
|
||||
* Returns parsed object or null if not found / invalid.
|
||||
*/
|
||||
export function extractMetricsJson(content: string, stepFilter?: string): Record<string, unknown> | null {
|
||||
const re = /<!--\s*METRICS_JSON\s+(\{[\s\S]*?\})\s*-->/g;
|
||||
let match: RegExpExecArray | null;
|
||||
while ((match = re.exec(content)) !== null) {
|
||||
try {
|
||||
const parsed = JSON.parse(match[1]) as Record<string, unknown>;
|
||||
if (!stepFilter || parsed['step'] === stepFilter) {
|
||||
return parsed;
|
||||
}
|
||||
} catch {
|
||||
// try next match
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Strip the "## Success Criteria" section so its checkboxes aren't counted as tasks. */
|
||||
function stripSuccessCriteriaSection(overview: string): string {
|
||||
return overview.replace(/^## Success Criteria\s*\n[\s\S]*?(?=\n## |\n*$)/m, '');
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract canonical task counts from the Completion Summary section.
|
||||
* Supports two formats found in real specs:
|
||||
* - Bold text: "**Completion Rate:** 100% (5/5 tasks)"
|
||||
* - Table row: "| Total Tasks | 28 (28/28 complete — 100%) |"
|
||||
* Returns null when no parseable counts are found.
|
||||
*/
|
||||
function parseCompletionSummaryTaskCount(overview: string): { completed: number; total: number } | null {
|
||||
const summaryMatch = overview.match(/## Completion Summary\s*\n([\s\S]*?)(?=\n## |\n*$)/);
|
||||
if (!summaryMatch) return null;
|
||||
const summary = summaryMatch[1];
|
||||
|
||||
// "Completion Rate: X% (Y/Z tasks)" — handles bold markdown and various separators
|
||||
const rateMatch = summary.match(/\**Completion(?:\s+Rate)?\**[:\s*]+\d{1,3}%\s*\((\d+)\/(\d+)\s*tasks?\)/i);
|
||||
if (rateMatch) return { completed: parseInt(rateMatch[1], 10), total: parseInt(rateMatch[2], 10) };
|
||||
|
||||
// "Total Tasks | 28 (28/28 complete"
|
||||
const tableMatch = summary.match(/Total Tasks\s*\|\s*(\d+)\s*\((\d+)\/(\d+)\s*complete/i);
|
||||
if (tableMatch) return { completed: parseInt(tableMatch[2], 10), total: parseInt(tableMatch[3], 10) };
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
function generateRunId(): string {
|
||||
const now = new Date();
|
||||
const ts = now.toISOString().replace(/[-:T]/g, '').slice(0, 14);
|
||||
@@ -48,14 +99,14 @@ function generateRunId(): string {
|
||||
export function collectPromptVersions(plan2codeRoot: string): PromptVersions {
|
||||
const srcDir = path.join(plan2codeRoot, 'src');
|
||||
return {
|
||||
plan: sha256File(path.join(srcDir, 'plan2code-1--plan.md')),
|
||||
revise_plan: sha256File(path.join(srcDir, 'plan2code-1b--revise-plan.md')),
|
||||
document: sha256File(path.join(srcDir, 'plan2code-2--document.md')),
|
||||
implement: sha256File(path.join(srcDir, 'plan2code-3--implement.md')),
|
||||
finalize: sha256File(path.join(srcDir, 'plan2code-4--finalize.md')),
|
||||
init: sha256File(path.join(srcDir, 'plan2code---init.md')),
|
||||
init_update: sha256File(path.join(srcDir, 'plan2code---init-update.md')),
|
||||
quick_task: sha256File(path.join(srcDir, 'plan2code---quick-task.md')),
|
||||
plan: sha256File(path.join(srcDir, 'plan2code-1-plan.md')),
|
||||
revise_plan: sha256File(path.join(srcDir, 'plan2code-1b-revise-plan.md')),
|
||||
document: sha256File(path.join(srcDir, 'plan2code-2-document.md')),
|
||||
implement: sha256File(path.join(srcDir, 'plan2code-3-implement.md')),
|
||||
finalize: sha256File(path.join(srcDir, 'plan2code-4-finalize.md')),
|
||||
init: sha256File(path.join(srcDir, 'plan2code-init.md')),
|
||||
init_update: sha256File(path.join(srcDir, 'plan2code-init-update.md')),
|
||||
quick_task: sha256File(path.join(srcDir, 'plan2code-quick-task.md')),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -91,15 +142,58 @@ export function collectStep1(specDir: string): Step1PlanMetrics {
|
||||
draftFiles.sort();
|
||||
const latestDraft = readFileSafe(draftFiles[draftFiles.length - 1]) ?? '';
|
||||
|
||||
// Parse "Confidence: XX%" — look for overall or total confidence
|
||||
let finalConfidence: number | null = null;
|
||||
const confMatch = latestDraft.match(/(?:Overall|Final|Total)?\s*[Cc]onfidence[:\s]+(\d{1,3})%/);
|
||||
if (confMatch) {
|
||||
finalConfidence = parseInt(confMatch[1], 10);
|
||||
} else {
|
||||
// Try table format: | Confidence | 92 |
|
||||
const tableMatch = latestDraft.match(/[|]\s*[Cc]onfidence\s*[|]\s*(\d{1,3})/);
|
||||
if (tableMatch) finalConfidence = parseInt(tableMatch[1], 10);
|
||||
// ── Primary source: METRICS_JSON HTML comment ──
|
||||
// Format: <!-- METRICS_JSON {"confidence": 95, "clarification_rounds": 0, ...} -->
|
||||
const metricsJson = extractMetricsJson(latestDraft);
|
||||
if (metricsJson) {
|
||||
const num = (key: string): number | null => {
|
||||
const v = metricsJson[key];
|
||||
return typeof v === 'number' ? v : null;
|
||||
};
|
||||
const bd = metricsJson['confidence_breakdown'] as Record<string, number> | undefined;
|
||||
return {
|
||||
present: true,
|
||||
final_confidence: num('confidence'),
|
||||
confidence_breakdown: bd ? {
|
||||
requirements: bd.requirements ?? null,
|
||||
feasibility: bd.feasibility ?? null,
|
||||
integration: bd.integration ?? null,
|
||||
risk: bd.risk ?? null,
|
||||
} : null,
|
||||
clarification_rounds: num('clarification_rounds'),
|
||||
tech_stack_revision_rounds: num('tech_stack_revision_rounds'),
|
||||
verification_gaps_found: num('verification_gaps_found'),
|
||||
functional_requirements_count: num('functional_requirements_count'),
|
||||
non_functional_requirements_count: num('non_functional_requirements_count'),
|
||||
risk_count: num('risk_count'),
|
||||
phase_count: num('phase_count'),
|
||||
};
|
||||
}
|
||||
|
||||
// ── Fallback: structured ## Planning Metrics appendix ──
|
||||
// Handles both plain (confidence: 95) and bold (**confidence:** 95) field formats
|
||||
// Planning Metrics is typically the last section, so we capture greedily to end,
|
||||
// but stop at the next ## heading or --- if present.
|
||||
const metricsSection = latestDraft.match(/^##\s+Planning\s+Metrics\s*\n([\s\S]+?)(?:\n##\s|\n---)/m)
|
||||
?? latestDraft.match(/^##\s+Planning\s+Metrics\s*\n([\s\S]+)/m);
|
||||
const metricsBlock = metricsSection?.[1] ?? '';
|
||||
const parseMetricField = (field: string): number | null => {
|
||||
// Match plain "field: N" or bold "**field:** N"
|
||||
const m = metricsBlock.match(new RegExp(`^\\**${field}\\**[:\\s*]+?(\\d+)`, 'm'));
|
||||
return m ? parseInt(m[1], 10) : null;
|
||||
};
|
||||
|
||||
// ── Confidence ──
|
||||
let finalConfidence = parseMetricField('confidence');
|
||||
if (finalConfidence === null) {
|
||||
// Fallback: prose scraping — handle bold markdown like **Confidence:** 95%
|
||||
const confMatch = latestDraft.match(/(?:Overall|Final|Total)?\s*\**[Cc]onfidence\**[:\s*]+(\d{1,3})%/);
|
||||
if (confMatch) {
|
||||
finalConfidence = parseInt(confMatch[1], 10);
|
||||
} else {
|
||||
const tableMatch = latestDraft.match(/[|]\s*[Cc]onfidence\s*[|]\s*(\d{1,3})/);
|
||||
if (tableMatch) finalConfidence = parseInt(tableMatch[1], 10);
|
||||
}
|
||||
}
|
||||
|
||||
// Confidence breakdown (requirements, feasibility, integration, risk)
|
||||
@@ -117,34 +211,52 @@ export function collectStep1(specDir: string): Step1PlanMetrics {
|
||||
};
|
||||
}
|
||||
|
||||
// Count ### FR- / ### NFR- headings
|
||||
const frCount = (latestDraft.match(/###\s+FR-/g) ?? []).length;
|
||||
const nfrCount = (latestDraft.match(/###\s+NFR-/g) ?? []).length;
|
||||
// ── Clarification rounds ──
|
||||
let clarificationRounds = parseMetricField('clarification_rounds');
|
||||
|
||||
// Count risk table rows (lines starting with | that contain risk-level keywords)
|
||||
const riskRows = latestDraft.match(/^\s*[|][^|]*(?:High|Medium|Low|Critical)[^|]*[|]/gm) ?? [];
|
||||
const riskCount = riskRows.length || null;
|
||||
// ── Verification gaps ──
|
||||
let verificationGaps = parseMetricField('verification_gaps_found');
|
||||
|
||||
// Count ## Phase headings
|
||||
const phaseCount = (latestDraft.match(/^##\s+Phase\s+\d/gm) ?? []).length || null;
|
||||
// ── Functional / non-functional requirements ──
|
||||
let frCount = parseMetricField('functional_requirements_count');
|
||||
if (frCount === null) {
|
||||
frCount = (latestDraft.match(/###\s+FR-/g) ?? []).length || null;
|
||||
}
|
||||
let nfrCount = parseMetricField('non_functional_requirements_count');
|
||||
if (nfrCount === null) {
|
||||
nfrCount = (latestDraft.match(/###\s+NFR-/g) ?? []).length || null;
|
||||
}
|
||||
|
||||
// Clarification rounds from conversation file
|
||||
let clarificationRounds: number | null = null;
|
||||
// ── Risk count ──
|
||||
let riskCount = parseMetricField('risk_count');
|
||||
if (riskCount === null) {
|
||||
const riskRows = latestDraft.match(/^\s*[|][^|]*(?:High|Medium|Low|Critical)[^|]*[|]/gm) ?? [];
|
||||
riskCount = riskRows.length || null;
|
||||
}
|
||||
|
||||
// ── Phase count ──
|
||||
let phaseCount = parseMetricField('phase_count');
|
||||
if (phaseCount === null) {
|
||||
phaseCount = (latestDraft.match(/^##\s+Phase\s+\d/gm) ?? []).length || null;
|
||||
}
|
||||
|
||||
// ── Tech stack revision rounds (conversation file only) ──
|
||||
let techStackRevisions: number | null = null;
|
||||
let verificationGaps: number | null = null;
|
||||
|
||||
// Fall back to conversation file scraping for fields not found in appendix
|
||||
if (convFiles.length > 0) {
|
||||
convFiles.sort();
|
||||
const latestConv = readFileSafe(convFiles[convFiles.length - 1]) ?? '';
|
||||
// Count heading repetitions as clarification rounds (## Clarification or ## Round)
|
||||
const clarRounds = (latestConv.match(/^##\s+(?:Clarification|Round)\s+\d/gm) ?? []).length;
|
||||
clarificationRounds = clarRounds || null;
|
||||
// Tech stack revision rounds
|
||||
if (clarificationRounds === null) {
|
||||
const clarRounds = (latestConv.match(/^##\s+(?:Clarification|Round)\s+\d/gm) ?? []).length;
|
||||
clarificationRounds = clarRounds || null;
|
||||
}
|
||||
const techRounds = (latestConv.match(/^##\s+(?:Tech\s+Stack|Technology)\s+Revision/gmi) ?? []).length;
|
||||
techStackRevisions = techRounds || null;
|
||||
// Verification gaps found
|
||||
const gapMatches = latestConv.match(/(?:verification\s+gap|gap\s+found|missing\s+requirement)/gi) ?? [];
|
||||
verificationGaps = gapMatches.length || null;
|
||||
if (verificationGaps === null) {
|
||||
const gapMatches = latestConv.match(/(?:verification\s+gap|gap\s+found|missing\s+requirement)/gi) ?? [];
|
||||
verificationGaps = gapMatches.length || null;
|
||||
}
|
||||
}
|
||||
|
||||
return {
|
||||
@@ -154,8 +266,8 @@ export function collectStep1(specDir: string): Step1PlanMetrics {
|
||||
clarification_rounds: clarificationRounds,
|
||||
tech_stack_revision_rounds: techStackRevisions,
|
||||
verification_gaps_found: verificationGaps,
|
||||
functional_requirements_count: frCount || null,
|
||||
non_functional_requirements_count: nfrCount || null,
|
||||
functional_requirements_count: frCount,
|
||||
non_functional_requirements_count: nfrCount,
|
||||
risk_count: riskCount,
|
||||
phase_count: phaseCount,
|
||||
};
|
||||
@@ -173,13 +285,31 @@ export function collectStep2(specDir: string): Step2DocumentMetrics {
|
||||
requirement_coverage_percent: null, verification_items_added: null };
|
||||
}
|
||||
|
||||
// Count all checkbox tasks: - [ ], - [x], - [!], - [/]
|
||||
const allTasks = (overview.match(/^\s*-\s+\[[ x!/?]\]/gm) ?? []).length;
|
||||
// ── Primary source: METRICS_JSON in overview.md ──
|
||||
const step2Json = extractMetricsJson(overview, 'document');
|
||||
if (step2Json) {
|
||||
const num = (key: string): number | null => {
|
||||
const v = step2Json[key];
|
||||
return typeof v === 'number' ? v : null;
|
||||
};
|
||||
const tpp = step2Json['tasks_per_phase'];
|
||||
return {
|
||||
present: true,
|
||||
total_tasks: num('total_tasks'),
|
||||
tasks_per_phase: Array.isArray(tpp) ? tpp : null,
|
||||
phase_count: num('phase_count'),
|
||||
parallel_groups_identified: num('parallel_groups_identified') ?? 0,
|
||||
requirement_coverage_percent: num('requirement_coverage_percent'),
|
||||
verification_items_added: num('verification_items_added'),
|
||||
};
|
||||
}
|
||||
|
||||
// ── Fallback: regex scraping ──
|
||||
|
||||
// Detect Parallel Execution Groups table
|
||||
const parallelGroups = (overview.match(/Parallel\s+Execution\s+Group/gi) ?? []).length;
|
||||
|
||||
// Count phase files
|
||||
// Count phase files (scanned first so we can use sum as a total_tasks fallback)
|
||||
let phaseCount = 0;
|
||||
let tasksPerPhase: number[] = [];
|
||||
try {
|
||||
@@ -190,13 +320,31 @@ export function collectStep2(specDir: string): Step2DocumentMetrics {
|
||||
phaseCount = phaseFiles.length;
|
||||
for (const pf of phaseFiles) {
|
||||
const content = readFileSafe(path.join(specDir, pf)) ?? '';
|
||||
const count = (content.match(/^\s*-\s+\[[ x!/?]\]/gm) ?? []).length;
|
||||
const count = (content.match(/^\s*-\s+\[[ x!/?]\]\s+\*\*Task\s+\d/gm) ?? []).length;
|
||||
tasksPerPhase.push(count);
|
||||
}
|
||||
} catch {
|
||||
// leave empty
|
||||
}
|
||||
|
||||
// total_tasks priority:
|
||||
// 1. Completion Summary canonical count
|
||||
// 2. Sum of phase file task counts
|
||||
// 3. Stripped overview checkboxes (excluding Success Criteria)
|
||||
const summaryCount = parseCompletionSummaryTaskCount(overview);
|
||||
const phaseSum = tasksPerPhase.length > 0 ? tasksPerPhase.reduce((a, b) => a + b, 0) : 0;
|
||||
const strippedOverview = stripSuccessCriteriaSection(overview);
|
||||
const strippedCheckboxCount = (strippedOverview.match(/^\s*-\s+\[[ x!/?]\]/gm) ?? []).length;
|
||||
|
||||
let totalTasks: number;
|
||||
if (summaryCount) {
|
||||
totalTasks = summaryCount.total;
|
||||
} else if (phaseSum > 0) {
|
||||
totalTasks = phaseSum;
|
||||
} else {
|
||||
totalTasks = strippedCheckboxCount;
|
||||
}
|
||||
|
||||
// Requirement coverage: look for coverage percentage in overview
|
||||
let reqCoverage: number | null = null;
|
||||
const covMatch = overview.match(/[Cc]overage[:\s]+(\d{1,3})%/);
|
||||
@@ -207,7 +355,7 @@ export function collectStep2(specDir: string): Step2DocumentMetrics {
|
||||
|
||||
return {
|
||||
present: true,
|
||||
total_tasks: allTasks || null,
|
||||
total_tasks: totalTasks || null,
|
||||
tasks_per_phase: tasksPerPhase.length > 0 ? tasksPerPhase : null,
|
||||
phase_count: phaseCount || null,
|
||||
parallel_groups_identified: parallelGroups || 0,
|
||||
@@ -216,124 +364,89 @@ export function collectStep2(specDir: string): Step2DocumentMetrics {
|
||||
};
|
||||
}
|
||||
|
||||
// ── Step 3: Implement (loop data) ─────────────────────────────────────────────
|
||||
// ── Step 3: Implement ──────────────────────────────────────────────────────────
|
||||
|
||||
export function collectStep3(specDir: string): Step3ImplementMetrics {
|
||||
const loopDir = path.join(specDir, '.plan2code-loop');
|
||||
const iterLogPath = path.join(loopDir, 'iteration.log');
|
||||
const configPath = path.join(loopDir, 'config.json');
|
||||
|
||||
if (!fs.existsSync(loopDir)) {
|
||||
return { present: true, used_loop_mode: false, loop_mode: null,
|
||||
task_completion_rate: null, tasks_completed: null, tasks_total: null,
|
||||
blocker_count: null, blocker_categories: null, total_iterations: null,
|
||||
avg_iteration_duration_ms: null, exit_code_distribution: null,
|
||||
completion_marker_success_rate: null };
|
||||
// Discover phase-*.md files
|
||||
let phaseFiles: string[] = [];
|
||||
try {
|
||||
const files = fs.readdirSync(specDir);
|
||||
phaseFiles = files
|
||||
.filter(f => /^phase-\d+\.md$/i.test(f))
|
||||
.sort();
|
||||
} catch {
|
||||
// leave empty
|
||||
}
|
||||
|
||||
// Parse config.json for loop mode
|
||||
let loopMode: 'task' | 'phase' | null = null;
|
||||
const configContent = readFileSafe(configPath);
|
||||
if (configContent) {
|
||||
try {
|
||||
const config = JSON.parse(configContent);
|
||||
loopMode = config.loopMode ?? null;
|
||||
} catch {
|
||||
// ignore
|
||||
}
|
||||
}
|
||||
|
||||
// Parse NDJSON iteration.log
|
||||
const iterLogContent = readFileSafe(iterLogPath);
|
||||
if (!iterLogContent) {
|
||||
return { present: true, used_loop_mode: true, loop_mode: loopMode,
|
||||
task_completion_rate: null, tasks_completed: null, tasks_total: null,
|
||||
blocker_count: null, blocker_categories: null, total_iterations: null,
|
||||
avg_iteration_duration_ms: null, exit_code_distribution: null,
|
||||
completion_marker_success_rate: null };
|
||||
}
|
||||
|
||||
interface IterEntry {
|
||||
iteration: number;
|
||||
timestamp: string;
|
||||
duration: number;
|
||||
exitCode: number;
|
||||
status: string;
|
||||
completionMarker?: string;
|
||||
}
|
||||
|
||||
const entries: IterEntry[] = [];
|
||||
for (const line of iterLogContent.split('\n')) {
|
||||
const trimmed = line.trim();
|
||||
if (!trimmed) continue;
|
||||
try {
|
||||
entries.push(JSON.parse(trimmed));
|
||||
} catch {
|
||||
// skip malformed lines
|
||||
}
|
||||
}
|
||||
|
||||
const totalIterations = entries.length;
|
||||
if (totalIterations === 0) {
|
||||
return { present: true, used_loop_mode: true, loop_mode: loopMode,
|
||||
task_completion_rate: null, tasks_completed: null, tasks_total: null,
|
||||
blocker_count: null, blocker_categories: null, total_iterations: 0,
|
||||
avg_iteration_duration_ms: null, exit_code_distribution: null,
|
||||
completion_marker_success_rate: null };
|
||||
}
|
||||
|
||||
// Exit code distribution
|
||||
const exitCodes: Record<string, number> = {};
|
||||
for (const e of entries) {
|
||||
const key = String(e.exitCode ?? 'other');
|
||||
exitCodes[key] = (exitCodes[key] ?? 0) + 1;
|
||||
}
|
||||
|
||||
// Average duration
|
||||
const durations = entries.map(e => e.duration).filter(d => d != null && d > 0);
|
||||
const avgDuration = durations.length > 0
|
||||
? Math.round(durations.reduce((a, b) => a + b, 0) / durations.length)
|
||||
: null;
|
||||
|
||||
// Completion markers
|
||||
const markerEntries = entries.filter(e => e.completionMarker && e.completionMarker.trim() !== '');
|
||||
const taskCompleteEntries = markerEntries.filter(e =>
|
||||
e.completionMarker?.startsWith('TASK_COMPLETE')
|
||||
);
|
||||
const blockedEntries = markerEntries.filter(e =>
|
||||
e.completionMarker?.startsWith('TASK_BLOCKED')
|
||||
);
|
||||
|
||||
// Completion marker success rate: entries where a valid marker was detected vs total
|
||||
const validMarkerCount = markerEntries.length;
|
||||
const markerSuccessRate = totalIterations > 0
|
||||
? validMarkerCount / totalIterations
|
||||
: null;
|
||||
|
||||
// Task counts from markers (used for completion_marker_success_rate)
|
||||
const blockerCount = blockedEntries.length;
|
||||
|
||||
// Blocker categories (parse from "TASK_BLOCKED: X.Y - reason")
|
||||
const blockerCategories: string[] = [];
|
||||
for (const e of blockedEntries) {
|
||||
const match = e.completionMarker?.match(/TASK_BLOCKED:\s*[\d.]+\s*-\s*(.+)/);
|
||||
if (match) {
|
||||
// Normalize to snake_case category
|
||||
const reason = match[1].toLowerCase().replace(/\s+/g, '_').replace(/[^a-z0-9_]/g, '');
|
||||
if (!blockerCategories.includes(reason)) {
|
||||
blockerCategories.push(reason);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Count tasks from overview.md checkboxes (consistent numerator and denominator)
|
||||
// Read overview.md for Completion Summary
|
||||
const overviewPath = path.join(specDir, 'overview.md');
|
||||
const overview = readFileSafe(overviewPath);
|
||||
const summaryCount = overview ? parseCompletionSummaryTaskCount(overview) : null;
|
||||
|
||||
// ── Primary source: METRICS_JSON in overview.md (from finalize step) ──
|
||||
if (overview) {
|
||||
const step3Json = extractMetricsJson(overview, 'finalize');
|
||||
if (step3Json) {
|
||||
const num = (key: string): number | null => {
|
||||
const v = step3Json[key];
|
||||
return typeof v === 'number' ? v : null;
|
||||
};
|
||||
const total = num('tasks_total');
|
||||
const completed = num('tasks_completed');
|
||||
return {
|
||||
present: true,
|
||||
task_completion_rate: total && total > 0 && completed != null ? completed / total : null,
|
||||
tasks_completed: completed,
|
||||
tasks_total: total,
|
||||
blocker_count: num('blocker_count') ?? 0,
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
// ── Fallback: regex scraping ──
|
||||
|
||||
// present = true if phase files exist OR overview has completion data
|
||||
const present = phaseFiles.length > 0 || summaryCount != null;
|
||||
if (!present) {
|
||||
return { present: false, task_completion_rate: null, tasks_completed: null,
|
||||
tasks_total: null, blocker_count: null };
|
||||
}
|
||||
|
||||
// Count checkboxes in phase files
|
||||
let phaseTotal = 0;
|
||||
let phaseCompleted = 0;
|
||||
let blockerCount = 0;
|
||||
for (const pf of phaseFiles) {
|
||||
const content = readFileSafe(path.join(specDir, pf)) ?? '';
|
||||
phaseTotal += (content.match(/^\s*-\s+\[[ x!/?]\]\s+\*\*Task\s+\d/gm) ?? []).length;
|
||||
phaseCompleted += (content.match(/^\s*-\s+\[x\]\s+\*\*Task\s+\d/gm) ?? []).length;
|
||||
blockerCount += (content.match(/^\s*-\s+\[!\]\s+\*\*Task\s+\d/gm) ?? []).length;
|
||||
}
|
||||
|
||||
// Count checkboxes in overview (stripped of Success Criteria)
|
||||
let overviewTotal = 0;
|
||||
let overviewCompleted = 0;
|
||||
if (overview) {
|
||||
const stripped = stripSuccessCriteriaSection(overview);
|
||||
overviewTotal = (stripped.match(/^\s*-\s+\[[ x!/?]\]/gm) ?? []).length;
|
||||
overviewCompleted = (stripped.match(/^\s*-\s+\[x\]/gm) ?? []).length;
|
||||
}
|
||||
|
||||
// Task counts priority:
|
||||
// 1. Completion Summary in overview.md
|
||||
// 2. Checkbox counting in phase files
|
||||
// 3. Checkbox counting in overview.md (stripped of Success Criteria)
|
||||
let tasksTotal: number | null = null;
|
||||
let tasksCompleted: number | null = null;
|
||||
if (overview) {
|
||||
tasksTotal = (overview.match(/^\s*-\s+\[[ x!/?]\]/gm) ?? []).length || null;
|
||||
tasksCompleted = (overview.match(/^\s*-\s+\[x\]/gm) ?? []).length || null;
|
||||
if (summaryCount) {
|
||||
tasksTotal = summaryCount.total;
|
||||
tasksCompleted = summaryCount.completed;
|
||||
} else if (phaseTotal > 0) {
|
||||
tasksTotal = phaseTotal;
|
||||
tasksCompleted = phaseCompleted;
|
||||
} else if (overviewTotal > 0) {
|
||||
tasksTotal = overviewTotal;
|
||||
tasksCompleted = overviewCompleted;
|
||||
}
|
||||
|
||||
const taskCompletionRate = tasksTotal && tasksTotal > 0 && tasksCompleted != null
|
||||
@@ -342,17 +455,10 @@ export function collectStep3(specDir: string): Step3ImplementMetrics {
|
||||
|
||||
return {
|
||||
present: true,
|
||||
used_loop_mode: true,
|
||||
loop_mode: loopMode,
|
||||
task_completion_rate: taskCompletionRate,
|
||||
tasks_completed: tasksCompleted || null,
|
||||
tasks_completed: tasksCompleted,
|
||||
tasks_total: tasksTotal,
|
||||
blocker_count: blockerCount,
|
||||
blocker_categories: blockerCategories.length > 0 ? blockerCategories : null,
|
||||
total_iterations: totalIterations,
|
||||
avg_iteration_duration_ms: avgDuration,
|
||||
exit_code_distribution: exitCodes,
|
||||
completion_marker_success_rate: markerSuccessRate,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -380,10 +486,35 @@ export function collectStep4(specDir: string): Step4FinalizeMetrics {
|
||||
archival_succeeded: archivalSucceeded };
|
||||
}
|
||||
|
||||
// Completion rate: completed tasks / total tasks
|
||||
const totalTasks = (overview.match(/^\s*-\s+\[[ x!/?]\]/gm) ?? []).length;
|
||||
const completedTasks = (overview.match(/^\s*-\s+\[x\]/gm) ?? []).length;
|
||||
const completionRate = totalTasks > 0 ? completedTasks / totalTasks : null;
|
||||
// ── Primary source: METRICS_JSON in overview.md ──
|
||||
const step4Json = extractMetricsJson(overview, 'finalize');
|
||||
if (step4Json) {
|
||||
const num = (key: string): number | null => {
|
||||
const v = step4Json[key];
|
||||
return typeof v === 'number' ? v : null;
|
||||
};
|
||||
return {
|
||||
present: true,
|
||||
completion_rate_at_audit: num('completion_rate_at_audit'),
|
||||
verification_failures_found: num('verification_failures_found') ?? 0,
|
||||
documentation_updates_needed: num('documentation_updates_needed'),
|
||||
archival_succeeded: archivalSucceeded,
|
||||
};
|
||||
}
|
||||
|
||||
// ── Fallback: regex scraping ──
|
||||
|
||||
// Completion rate: prefer Completion Summary, fall back to stripped overview checkboxes
|
||||
let completionRate: number | null = null;
|
||||
const summaryCount = parseCompletionSummaryTaskCount(overview);
|
||||
if (summaryCount) {
|
||||
completionRate = summaryCount.total > 0 ? summaryCount.completed / summaryCount.total : null;
|
||||
} else {
|
||||
const stripped = stripSuccessCriteriaSection(overview);
|
||||
const totalTasks = (stripped.match(/^\s*-\s+\[[ x!/?]\]/gm) ?? []).length;
|
||||
const completedTasks = (stripped.match(/^\s*-\s+\[x\]/gm) ?? []).length;
|
||||
completionRate = totalTasks > 0 ? completedTasks / totalTasks : null;
|
||||
}
|
||||
|
||||
// Verification failures: look for [!] tasks or "verification failed" text
|
||||
const blockedTasks = (overview.match(/^\s*-\s+\[!\]/gm) ?? []).length;
|
||||
@@ -476,28 +607,11 @@ export async function collectRun(opts: CollectorOptions): Promise<RunMetrics> {
|
||||
// leave null
|
||||
}
|
||||
|
||||
// Try to get started_at from earliest iteration log entry
|
||||
const loopDir = path.join(specDir, '.plan2code-loop');
|
||||
const iterLogPath = path.join(loopDir, 'iteration.log');
|
||||
const iterLogContent = readFileSafe(iterLogPath);
|
||||
if (iterLogContent) {
|
||||
const lines = iterLogContent.split('\n').filter(l => l.trim());
|
||||
if (lines.length > 0) {
|
||||
try {
|
||||
const first = JSON.parse(lines[0]);
|
||||
if (first.timestamp) startedAt = first.timestamp;
|
||||
} catch { /* ignore */ }
|
||||
try {
|
||||
const last = JSON.parse(lines[lines.length - 1]);
|
||||
if (last.timestamp) completedAt = last.timestamp;
|
||||
} catch { /* ignore */ }
|
||||
}
|
||||
}
|
||||
|
||||
const metrics: RunMetrics = {
|
||||
schema_version: '1.0',
|
||||
run_id: runId,
|
||||
plan2code_version: plan2codeVersion,
|
||||
source: 'local',
|
||||
prompt_versions: collectPromptVersions(plan2codeRoot),
|
||||
project: {
|
||||
name: projectName,
|
||||
|
||||
@@ -0,0 +1,243 @@
|
||||
import fs from 'fs';
|
||||
import os from 'os';
|
||||
import path from 'path';
|
||||
import { describe, it, expect, afterEach } from 'vitest';
|
||||
import { parseSubmissionPayload, ingestCommunityIssues } from './community.js';
|
||||
import { writeRunFile } from './aggregator.js';
|
||||
import type { RunMetrics } from './types.js';
|
||||
|
||||
function issueBody(payload: Record<string, unknown>): string {
|
||||
return `Some issue text.\n\n<!-- METRICS_JSON ${JSON.stringify(payload)} -->\n`;
|
||||
}
|
||||
|
||||
const VALID_PAYLOAD = {
|
||||
schema_version: '1.0',
|
||||
run_id: 'run-20260715-143000-a1b2',
|
||||
plan2code_version: '1.15.3',
|
||||
prompt_versions_short: {
|
||||
plan: 'abc123def456', revise_plan: 'a', document: 'b', implement: 'c',
|
||||
finalize: 'd', init: 'e', init_update: 'f', quick_task: 'g',
|
||||
},
|
||||
step1: {
|
||||
final_confidence: 95,
|
||||
confidence_breakdown: { requirements: 24, feasibility: 23, integration: 24, risk: 22 },
|
||||
clarification_rounds: 0,
|
||||
tech_stack_revision_rounds: 0,
|
||||
verification_gaps_found: 0,
|
||||
functional_requirements_count: 8,
|
||||
non_functional_requirements_count: 6,
|
||||
risk_count: 7,
|
||||
phase_count: 4,
|
||||
},
|
||||
step2: {
|
||||
total_tasks: 28,
|
||||
phase_count: 4,
|
||||
parallel_groups_identified: 1,
|
||||
requirement_coverage_percent: 100,
|
||||
verification_items_added: 3,
|
||||
},
|
||||
step3: {
|
||||
task_completion_rate: 0.96,
|
||||
tasks_completed: 27,
|
||||
tasks_total: 28,
|
||||
blocker_count: 1,
|
||||
},
|
||||
step4: {
|
||||
completion_rate_at_audit: 0.96,
|
||||
verification_failures_found: 1,
|
||||
documentation_updates_needed: 2,
|
||||
},
|
||||
user_feedback: {
|
||||
overall_rating: 8,
|
||||
rating_reason: 'good stuff',
|
||||
what_went_well: 'well',
|
||||
what_went_poorly: 'poorly',
|
||||
},
|
||||
};
|
||||
|
||||
// ── parseSubmissionPayload() ─────────────────────────────────────────────────
|
||||
|
||||
describe('parseSubmissionPayload', () => {
|
||||
it('parses a fully valid payload into a correctly-shaped RunMetrics', () => {
|
||||
const result = parseSubmissionPayload(issueBody(VALID_PAYLOAD));
|
||||
expect(result).not.toBeNull();
|
||||
expect(result!.run_id).toBe('run-20260715-143000-a1b2');
|
||||
expect(result!.schema_version).toBe('1.0');
|
||||
expect(result!.plan2code_version).toBe('1.15.3');
|
||||
expect(result!.source).toBe('community');
|
||||
expect(result!.step1_plan.present).toBe(true);
|
||||
expect(result!.step1_plan.final_confidence).toBe(95);
|
||||
expect(result!.step2_document.present).toBe(true);
|
||||
expect(result!.step2_document.total_tasks).toBe(28);
|
||||
expect(result!.step3_implement.present).toBe(true);
|
||||
expect(result!.step3_implement.tasks_completed).toBe(27);
|
||||
expect(result!.step4_finalize.present).toBe(true);
|
||||
expect(result!.step4_finalize.completion_rate_at_audit).toBe(0.96);
|
||||
expect(result!.user_feedback).toEqual({
|
||||
overall_rating: 8,
|
||||
rating_reason: 'good stuff',
|
||||
what_went_well: 'well',
|
||||
what_went_poorly: 'poorly',
|
||||
});
|
||||
});
|
||||
|
||||
it('returns null when run_id is missing', () => {
|
||||
const { run_id, ...withoutRunId } = VALID_PAYLOAD;
|
||||
const result = parseSubmissionPayload(issueBody(withoutRunId));
|
||||
expect(result).toBeNull();
|
||||
});
|
||||
|
||||
it('returns null when user_feedback.overall_rating is a string instead of a number', () => {
|
||||
const badPayload = {
|
||||
...VALID_PAYLOAD,
|
||||
user_feedback: { ...VALID_PAYLOAD.user_feedback, overall_rating: 'eight' },
|
||||
};
|
||||
const result = parseSubmissionPayload(issueBody(badPayload));
|
||||
expect(result).toBeNull();
|
||||
});
|
||||
|
||||
it('returns null when schema_version is not exactly "1.0"', () => {
|
||||
const badPayload = { ...VALID_PAYLOAD, schema_version: '2.0' };
|
||||
const result = parseSubmissionPayload(issueBody(badPayload));
|
||||
expect(result).toBeNull();
|
||||
});
|
||||
|
||||
it('sets all four steps present:false when only user_feedback is included', () => {
|
||||
const minimalPayload = {
|
||||
schema_version: '1.0',
|
||||
run_id: 'run-20260715-150000-c3d4',
|
||||
plan2code_version: '1.15.3',
|
||||
user_feedback: VALID_PAYLOAD.user_feedback,
|
||||
};
|
||||
const result = parseSubmissionPayload(issueBody(minimalPayload));
|
||||
expect(result).not.toBeNull();
|
||||
expect(result!.step1_plan.present).toBe(false);
|
||||
expect(result!.step2_document.present).toBe(false);
|
||||
expect(result!.step3_implement.present).toBe(false);
|
||||
expect(result!.step4_finalize.present).toBe(false);
|
||||
});
|
||||
|
||||
it('backfills missing prompt_versions_short keys with the sha256:missing sentinel', () => {
|
||||
const partialPayload = {
|
||||
...VALID_PAYLOAD,
|
||||
prompt_versions_short: { plan: 'abc123def456' },
|
||||
};
|
||||
const result = parseSubmissionPayload(issueBody(partialPayload));
|
||||
expect(result).not.toBeNull();
|
||||
expect(result!.prompt_versions.plan).toBe('abc123def456');
|
||||
expect(result!.prompt_versions.revise_plan).toBe('sha256:missing');
|
||||
expect(result!.prompt_versions.document).toBe('sha256:missing');
|
||||
expect(result!.prompt_versions.implement).toBe('sha256:missing');
|
||||
expect(result!.prompt_versions.finalize).toBe('sha256:missing');
|
||||
expect(result!.prompt_versions.init).toBe('sha256:missing');
|
||||
expect(result!.prompt_versions.init_update).toBe('sha256:missing');
|
||||
expect(result!.prompt_versions.quick_task).toBe('sha256:missing');
|
||||
});
|
||||
});
|
||||
|
||||
// ── writeRunFile() ────────────────────────────────────────────────────────────
|
||||
|
||||
describe('writeRunFile', () => {
|
||||
let tmpDir: string;
|
||||
|
||||
afterEach(() => {
|
||||
if (tmpDir) fs.rmSync(tmpDir, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
const RUN: RunMetrics = {
|
||||
schema_version: '1.0',
|
||||
run_id: 'run-20260715-160000-e5f6',
|
||||
plan2code_version: '1.15.3',
|
||||
prompt_versions: {
|
||||
plan: 'sha256:missing', revise_plan: 'sha256:missing', document: 'sha256:missing',
|
||||
implement: 'sha256:missing', finalize: 'sha256:missing', init: 'sha256:missing',
|
||||
init_update: 'sha256:missing', quick_task: 'sha256:missing',
|
||||
},
|
||||
project: { name: '', started_at: null, completed_at: null },
|
||||
step1_plan: { present: false, final_confidence: null, confidence_breakdown: null, clarification_rounds: null, tech_stack_revision_rounds: null, verification_gaps_found: null, functional_requirements_count: null, non_functional_requirements_count: null, risk_count: null, phase_count: null },
|
||||
step2_document: { present: false, total_tasks: null, tasks_per_phase: null, phase_count: null, parallel_groups_identified: null, requirement_coverage_percent: null, verification_items_added: null },
|
||||
step3_implement: { present: false, task_completion_rate: null, tasks_completed: null, tasks_total: null, blocker_count: null },
|
||||
step4_finalize: { present: false, completion_rate_at_audit: null, verification_failures_found: null, documentation_updates_needed: null, archival_succeeded: null },
|
||||
user_feedback: null,
|
||||
};
|
||||
|
||||
it('writes a new run_id to an empty runsDir and returns true', () => {
|
||||
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'plan2code-metrics-test-'));
|
||||
const result = writeRunFile(RUN, tmpDir);
|
||||
expect(result).toBe(true);
|
||||
expect(fs.existsSync(path.join(tmpDir, `${RUN.run_id}.json`))).toBe(true);
|
||||
});
|
||||
|
||||
it('returns false and does not overwrite when the same run_id already exists', () => {
|
||||
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'plan2code-metrics-test-'));
|
||||
writeRunFile(RUN, tmpDir);
|
||||
const modified = { ...RUN, plan2code_version: '9.9.9' };
|
||||
const result = writeRunFile(modified, tmpDir);
|
||||
expect(result).toBe(false);
|
||||
const onDisk = JSON.parse(fs.readFileSync(path.join(tmpDir, `${RUN.run_id}.json`), 'utf8'));
|
||||
expect(onDisk.plan2code_version).toBe('1.15.3');
|
||||
});
|
||||
});
|
||||
|
||||
// ── ingestCommunityIssues() ───────────────────────────────────────────────────
|
||||
|
||||
describe('ingestCommunityIssues', () => {
|
||||
const goodBody = issueBody(VALID_PAYLOAD);
|
||||
const badBody = 'an issue with no METRICS_JSON payload';
|
||||
|
||||
it('imports a new run and closes its issue', async () => {
|
||||
const closed: number[] = [];
|
||||
const tally = await ingestCommunityIssues(
|
||||
[{ number: 1, body: goodBody }],
|
||||
'owner/repo',
|
||||
'/runs',
|
||||
{ writeRunFile: () => true, closeIssue: async (_r, n) => { closed.push(n); } },
|
||||
);
|
||||
expect(tally.imported).toBe(1);
|
||||
expect(tally.skippedDuplicate).toBe(0);
|
||||
expect(tally.closed).toBe(1);
|
||||
expect(closed).toEqual([1]);
|
||||
});
|
||||
|
||||
it('closes an already-imported (duplicate) issue instead of skipping the close', async () => {
|
||||
const closed: number[] = [];
|
||||
const tally = await ingestCommunityIssues(
|
||||
[{ number: 7, body: goodBody }],
|
||||
'owner/repo',
|
||||
'/runs',
|
||||
{ writeRunFile: () => false, closeIssue: async (_r, n) => { closed.push(n); } },
|
||||
);
|
||||
expect(tally.imported).toBe(0);
|
||||
expect(tally.skippedDuplicate).toBe(1);
|
||||
expect(tally.closed).toBe(1); // the close-retry: duplicates are still closed
|
||||
expect(closed).toEqual([7]);
|
||||
});
|
||||
|
||||
it('records a close failure without throwing and leaves the issue for a later retry', async () => {
|
||||
const tally = await ingestCommunityIssues(
|
||||
[{ number: 9, body: goodBody }],
|
||||
'owner/repo',
|
||||
'/runs',
|
||||
{ writeRunFile: () => true, closeIssue: async () => { throw new Error('network'); } },
|
||||
);
|
||||
expect(tally.imported).toBe(1);
|
||||
expect(tally.closed).toBe(0);
|
||||
expect(tally.closeFailed).toBe(1);
|
||||
expect(tally.closeFailedIssues).toEqual([9]);
|
||||
});
|
||||
|
||||
it('skips and reports a malformed issue without writing or closing it', async () => {
|
||||
let wrote = false;
|
||||
let closeCalled = false;
|
||||
const tally = await ingestCommunityIssues(
|
||||
[{ number: 3, body: badBody }],
|
||||
'owner/repo',
|
||||
'/runs',
|
||||
{ writeRunFile: () => { wrote = true; return true; }, closeIssue: async () => { closeCalled = true; } },
|
||||
);
|
||||
expect(tally.skippedMalformed).toBe(1);
|
||||
expect(tally.malformedIssues).toEqual([3]);
|
||||
expect(wrote).toBe(false);
|
||||
expect(closeCalled).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,276 @@
|
||||
/**
|
||||
* community.ts
|
||||
* Ingestion side of the community feedback flow: list/parse/close
|
||||
* `community-feedback`-labeled GitHub issues on jparkerweb/plan2code.
|
||||
*/
|
||||
|
||||
import { execa } from 'execa';
|
||||
import type {
|
||||
RunMetrics,
|
||||
PromptVersions,
|
||||
UserFeedback,
|
||||
Step1PlanMetrics,
|
||||
Step2DocumentMetrics,
|
||||
Step3ImplementMetrics,
|
||||
Step4FinalizeMetrics,
|
||||
} from './types.js';
|
||||
import { extractMetricsJson } from './collector.js';
|
||||
import { backfillPromptVersions } from './aggregator.js';
|
||||
|
||||
// ── GitHub interaction (via gh CLI) ──────────────────────────────────────────
|
||||
|
||||
export interface CommunityIssue {
|
||||
number: number;
|
||||
body: string;
|
||||
}
|
||||
|
||||
interface RawIssue {
|
||||
number: number;
|
||||
title: string;
|
||||
body: string;
|
||||
labels: { name: string }[];
|
||||
}
|
||||
|
||||
const COMMUNITY_LABEL = 'community-feedback';
|
||||
const FEEDBACK_TITLE_PREFIX = '[Feedback]';
|
||||
const METRICS_JSON_MARKER = /<!--\s*METRICS_JSON\s+\{/;
|
||||
|
||||
export async function listCommunityIssues(repo: string): Promise<CommunityIssue[]> {
|
||||
// Fetch open issues broadly rather than by label alone. Browser/print-tier
|
||||
// submissions from outside contributors can lose the `community-feedback`
|
||||
// label: GitHub only honors the `labels=` query param on issues/new for
|
||||
// users with triage/push access, so the label is silently dropped for
|
||||
// community members without `gh`. We therefore also match by the `[Feedback]`
|
||||
// title prefix and the METRICS_JSON marker. Only OPEN issues are considered
|
||||
// (closed/done submissions are already processed); parseSubmissionPayload is
|
||||
// the final gate that rejects anything without a valid payload.
|
||||
const result = await execa('gh', [
|
||||
'issue', 'list',
|
||||
'--repo', repo,
|
||||
'--state', 'open',
|
||||
'--limit', '1000',
|
||||
'--json', 'number,title,body,labels',
|
||||
]);
|
||||
const raw = JSON.parse(result.stdout) as RawIssue[];
|
||||
return raw
|
||||
.filter((issue) =>
|
||||
issue.labels.some((l) => l.name === COMMUNITY_LABEL) ||
|
||||
issue.title.startsWith(FEEDBACK_TITLE_PREFIX) ||
|
||||
METRICS_JSON_MARKER.test(issue.body)
|
||||
)
|
||||
.map((issue) => ({ number: issue.number, body: issue.body }));
|
||||
}
|
||||
|
||||
export async function closeIssue(repo: string, issueNumber: number): Promise<void> {
|
||||
await execa('gh', ['issue', 'close', String(issueNumber), '--repo', repo]);
|
||||
}
|
||||
|
||||
// ── Ingestion control flow (I/O injected so it is unit-testable) ──────────────
|
||||
|
||||
export interface IngestionDeps {
|
||||
writeRunFile: (run: RunMetrics, runsDir: string) => boolean;
|
||||
closeIssue: (repo: string, issueNumber: number) => Promise<void>;
|
||||
}
|
||||
|
||||
export interface IngestionTally {
|
||||
imported: number;
|
||||
skippedDuplicate: number;
|
||||
skippedMalformed: number;
|
||||
closed: number;
|
||||
closeFailed: number;
|
||||
malformedIssues: number[];
|
||||
closeFailedIssues: number[];
|
||||
}
|
||||
|
||||
/**
|
||||
* Process a batch of community issues: parse each payload, write new runs
|
||||
* (deduped by run_id), and close every open issue idempotently.
|
||||
*
|
||||
* The close is attempted on the duplicate path too: a submission that imported
|
||||
* on an earlier run but failed to close would otherwise be seen as a duplicate
|
||||
* forever and never closed again, leaving the issue open and reprocessed on
|
||||
* every fetch. Malformed issues are reported (not fixed up) and left open.
|
||||
*
|
||||
* I/O (writeRunFile/closeIssue) is injected so the control flow can be unit
|
||||
* tested without a live `gh`. Returns a tally; the caller owns all logging.
|
||||
*/
|
||||
export async function ingestCommunityIssues(
|
||||
issues: CommunityIssue[],
|
||||
repo: string,
|
||||
runsDir: string,
|
||||
deps: IngestionDeps,
|
||||
): Promise<IngestionTally> {
|
||||
const tally: IngestionTally = {
|
||||
imported: 0, skippedDuplicate: 0, skippedMalformed: 0,
|
||||
closed: 0, closeFailed: 0, malformedIssues: [], closeFailedIssues: [],
|
||||
};
|
||||
|
||||
for (const issue of issues) {
|
||||
const run = parseSubmissionPayload(issue.body);
|
||||
if (!run) {
|
||||
tally.skippedMalformed++;
|
||||
tally.malformedIssues.push(issue.number);
|
||||
continue;
|
||||
}
|
||||
|
||||
if (deps.writeRunFile(run, runsDir)) {
|
||||
tally.imported++;
|
||||
} else {
|
||||
tally.skippedDuplicate++;
|
||||
}
|
||||
|
||||
try {
|
||||
await deps.closeIssue(repo, issue.number);
|
||||
tally.closed++;
|
||||
} catch {
|
||||
tally.closeFailed++;
|
||||
tally.closeFailedIssues.push(issue.number);
|
||||
}
|
||||
}
|
||||
|
||||
return tally;
|
||||
}
|
||||
|
||||
// ── Payload parsing (type-only validation, per NFR-5) ────────────────────────
|
||||
|
||||
function isString(v: unknown): v is string {
|
||||
return typeof v === 'string';
|
||||
}
|
||||
|
||||
function isNumber(v: unknown): v is number {
|
||||
return typeof v === 'number';
|
||||
}
|
||||
|
||||
function numOrNull(v: unknown): number | null {
|
||||
return isNumber(v) ? v : null;
|
||||
}
|
||||
|
||||
/** Type-check `keys` off `raw` (object or not) into a { [key]: number | null } map. */
|
||||
function pickNumbers<K extends string>(raw: Record<string, unknown>, keys: readonly K[]): Record<K, number | null> {
|
||||
const result = {} as Record<K, number | null>;
|
||||
for (const key of keys) result[key] = numOrNull(raw[key]);
|
||||
return result;
|
||||
}
|
||||
|
||||
const STEP1_ABSENT: Step1PlanMetrics = {
|
||||
present: false, final_confidence: null, confidence_breakdown: null,
|
||||
clarification_rounds: null, tech_stack_revision_rounds: null,
|
||||
verification_gaps_found: null, functional_requirements_count: null,
|
||||
non_functional_requirements_count: null, risk_count: null, phase_count: null,
|
||||
};
|
||||
|
||||
function parseStep1(raw: unknown): Step1PlanMetrics {
|
||||
if (raw == null || typeof raw !== 'object') return STEP1_ABSENT;
|
||||
const step1 = raw as Record<string, unknown>;
|
||||
const bdRaw = step1['confidence_breakdown'];
|
||||
const breakdown = bdRaw != null && typeof bdRaw === 'object'
|
||||
? pickNumbers(bdRaw as Record<string, unknown>, ['requirements', 'feasibility', 'integration', 'risk'])
|
||||
: null;
|
||||
return {
|
||||
present: true,
|
||||
confidence_breakdown: breakdown,
|
||||
...pickNumbers(step1, [
|
||||
'final_confidence', 'clarification_rounds', 'tech_stack_revision_rounds',
|
||||
'verification_gaps_found', 'functional_requirements_count',
|
||||
'non_functional_requirements_count', 'risk_count', 'phase_count',
|
||||
]),
|
||||
};
|
||||
}
|
||||
|
||||
const STEP2_ABSENT: Step2DocumentMetrics = {
|
||||
present: false, total_tasks: null, tasks_per_phase: null,
|
||||
phase_count: null, parallel_groups_identified: null,
|
||||
requirement_coverage_percent: null, verification_items_added: null,
|
||||
};
|
||||
|
||||
function parseStep2(raw: unknown): Step2DocumentMetrics {
|
||||
if (raw == null || typeof raw !== 'object') return STEP2_ABSENT;
|
||||
const step2 = raw as Record<string, unknown>;
|
||||
return {
|
||||
present: true,
|
||||
tasks_per_phase: null,
|
||||
...pickNumbers(step2, [
|
||||
'total_tasks', 'phase_count', 'parallel_groups_identified',
|
||||
'requirement_coverage_percent', 'verification_items_added',
|
||||
]),
|
||||
};
|
||||
}
|
||||
|
||||
const STEP3_ABSENT: Step3ImplementMetrics = {
|
||||
present: false, task_completion_rate: null, tasks_completed: null, tasks_total: null, blocker_count: null,
|
||||
};
|
||||
|
||||
function parseStep3(raw: unknown): Step3ImplementMetrics {
|
||||
if (raw == null || typeof raw !== 'object') return STEP3_ABSENT;
|
||||
const step3 = raw as Record<string, unknown>;
|
||||
return {
|
||||
present: true,
|
||||
...pickNumbers(step3, ['task_completion_rate', 'tasks_completed', 'tasks_total', 'blocker_count']),
|
||||
};
|
||||
}
|
||||
|
||||
const STEP4_ABSENT: Step4FinalizeMetrics = {
|
||||
present: false, completion_rate_at_audit: null, verification_failures_found: null,
|
||||
documentation_updates_needed: null, archival_succeeded: null,
|
||||
};
|
||||
|
||||
function parseStep4(raw: unknown): Step4FinalizeMetrics {
|
||||
if (raw == null || typeof raw !== 'object') return STEP4_ABSENT;
|
||||
const step4 = raw as Record<string, unknown>;
|
||||
const archivalRaw = step4['archival_succeeded'];
|
||||
return {
|
||||
present: true,
|
||||
archival_succeeded: typeof archivalRaw === 'boolean' ? archivalRaw : null,
|
||||
...pickNumbers(step4, ['completion_rate_at_audit', 'verification_failures_found', 'documentation_updates_needed']),
|
||||
};
|
||||
}
|
||||
|
||||
function parsePromptVersionsShort(raw: unknown): PromptVersions {
|
||||
const partial: Partial<PromptVersions> = {};
|
||||
if (raw != null && typeof raw === 'object') {
|
||||
const pv = raw as Record<string, unknown>;
|
||||
for (const key of ['plan', 'revise_plan', 'document', 'implement', 'finalize', 'init', 'init_update', 'quick_task'] as const) {
|
||||
const v = pv[key];
|
||||
if (isString(v)) partial[key] = v;
|
||||
}
|
||||
}
|
||||
return backfillPromptVersions(partial as PromptVersions);
|
||||
}
|
||||
|
||||
export function parseSubmissionPayload(body: string): RunMetrics | null {
|
||||
const parsed = extractMetricsJson(body);
|
||||
if (!parsed) return null;
|
||||
|
||||
if (parsed['schema_version'] !== '1.0') return null;
|
||||
if (!isString(parsed['run_id'])) return null;
|
||||
if (!isString(parsed['plan2code_version'])) return null;
|
||||
|
||||
const feedbackRaw = parsed['user_feedback'];
|
||||
if (feedbackRaw == null || typeof feedbackRaw !== 'object') return null;
|
||||
const feedback = feedbackRaw as Record<string, unknown>;
|
||||
if (!isNumber(feedback['overall_rating'])) return null;
|
||||
if (!isString(feedback['rating_reason'])) return null;
|
||||
if (!isString(feedback['what_went_well'])) return null;
|
||||
if (!isString(feedback['what_went_poorly'])) return null;
|
||||
|
||||
const userFeedback: UserFeedback = {
|
||||
overall_rating: feedback['overall_rating'],
|
||||
rating_reason: feedback['rating_reason'],
|
||||
what_went_well: feedback['what_went_well'],
|
||||
what_went_poorly: feedback['what_went_poorly'],
|
||||
};
|
||||
|
||||
return {
|
||||
schema_version: '1.0',
|
||||
run_id: parsed['run_id'],
|
||||
plan2code_version: parsed['plan2code_version'],
|
||||
source: 'community',
|
||||
prompt_versions: parsePromptVersionsShort(parsed['prompt_versions_short']),
|
||||
project: { name: '', started_at: null, completed_at: null },
|
||||
step1_plan: parseStep1(parsed['step1']),
|
||||
step2_document: parseStep2(parsed['step2']),
|
||||
step3_implement: parseStep3(parsed['step3']),
|
||||
step4_finalize: parseStep4(parsed['step4']),
|
||||
user_feedback: userFeedback,
|
||||
};
|
||||
}
|
||||
@@ -5,18 +5,18 @@ import type { PromptEdit } from './types.js';
|
||||
// ── Shared fixtures ───────────────────────────────────────────────────────────
|
||||
|
||||
const PROMPT_CONTENTS: Record<string, string> = {
|
||||
'plan2code-1--plan.md': 'This is the plan prompt content. It has some text here.',
|
||||
'plan2code-2--document.md': 'Document prompt with repeated text. repeated text. Done.',
|
||||
'plan2code-3--implement.md': 'Implement prompt content.',
|
||||
'plan2code-1-plan.md': 'This is the plan prompt content. It has some text here.',
|
||||
'plan2code-2-document.md': 'Document prompt with repeated text. repeated text. Done.',
|
||||
'plan2code-3-implement.md': 'Implement prompt content.',
|
||||
};
|
||||
|
||||
function makeEdit(overrides: Partial<PromptEdit> = {}): PromptEdit {
|
||||
return {
|
||||
file: 'plan2code-1--plan.md',
|
||||
file: 'plan2code-1-plan.md',
|
||||
rationale: 'test rationale',
|
||||
expected_metric_impact: 'test impact',
|
||||
char_count_before: PROMPT_CONTENTS['plan2code-1--plan.md'].length,
|
||||
char_count_after: PROMPT_CONTENTS['plan2code-1--plan.md'].length,
|
||||
char_count_before: PROMPT_CONTENTS['plan2code-1-plan.md'].length,
|
||||
char_count_after: PROMPT_CONTENTS['plan2code-1-plan.md'].length,
|
||||
char_count_delta: 0,
|
||||
old_text: 'some text',
|
||||
new_text: 'better text',
|
||||
@@ -59,10 +59,10 @@ describe('validateEdit', () => {
|
||||
|
||||
it('warns when old_text appears multiple times', () => {
|
||||
const edit = makeEdit({
|
||||
file: 'plan2code-2--document.md',
|
||||
file: 'plan2code-2-document.md',
|
||||
old_text: 'repeated text',
|
||||
char_count_before: PROMPT_CONTENTS['plan2code-2--document.md'].length,
|
||||
char_count_after: PROMPT_CONTENTS['plan2code-2--document.md'].length,
|
||||
char_count_before: PROMPT_CONTENTS['plan2code-2-document.md'].length,
|
||||
char_count_after: PROMPT_CONTENTS['plan2code-2-document.md'].length,
|
||||
});
|
||||
const result = validateEdit(edit, PROMPT_CONTENTS);
|
||||
expect(result.valid).toBe(true);
|
||||
|
||||
@@ -32,14 +32,14 @@ function generateProposalId(): string {
|
||||
function readPromptFiles(plan2codeRoot: string): Record<string, string> {
|
||||
const srcDir = path.join(plan2codeRoot, 'src');
|
||||
const promptFiles = [
|
||||
'plan2code-1--plan.md',
|
||||
'plan2code-1b--revise-plan.md',
|
||||
'plan2code-2--document.md',
|
||||
'plan2code-3--implement.md',
|
||||
'plan2code-4--finalize.md',
|
||||
'plan2code---init.md',
|
||||
'plan2code---init-update.md',
|
||||
'plan2code---quick-task.md',
|
||||
'plan2code-1-plan.md',
|
||||
'plan2code-1b-revise-plan.md',
|
||||
'plan2code-2-document.md',
|
||||
'plan2code-3-implement.md',
|
||||
'plan2code-4-finalize.md',
|
||||
'plan2code-init.md',
|
||||
'plan2code-init-update.md',
|
||||
'plan2code-quick-task.md',
|
||||
];
|
||||
|
||||
const contents: Record<string, string> = {};
|
||||
@@ -154,7 +154,7 @@ export interface ImproverResult {
|
||||
}
|
||||
|
||||
export async function generateImprovement(opts: ImproverOptions): Promise<ImproverResult> {
|
||||
const { diagnosisPath, plan2codeRoot, proposalsDir, runsDir, model = 'claude-opus-4-6', agent = 'claude-code' } = opts;
|
||||
const { diagnosisPath, plan2codeRoot, proposalsDir, runsDir, model = 'default', agent = 'claude-code' } = opts;
|
||||
|
||||
// Load diagnosis
|
||||
let diagnosisContent: string;
|
||||
@@ -193,7 +193,7 @@ export async function generateImprovement(opts: ImproverOptions): Promise<Improv
|
||||
});
|
||||
|
||||
// Invoke Claude
|
||||
console.log(`\nInvoking AI improvement proposal (model: ${model})...`);
|
||||
console.log(`\nInvoking AI improvement proposal (agent: ${agent}, model: ${model === 'default' ? 'user default' : model})...`);
|
||||
console.log('This may take a minute...\n');
|
||||
|
||||
let aiResponse: string;
|
||||
|
||||
@@ -1,7 +1,8 @@
|
||||
/**
|
||||
* invoke-llm.ts
|
||||
* Unified LLM invocation for plan2code-metrics.
|
||||
* Supports Claude Code (temp file → stdin) and Copilot CLI (stdin string).
|
||||
* Supports Claude Code (temp file → stdin), Copilot CLI (stdin string), and
|
||||
* Devin CLI (temp prompt file).
|
||||
* Mirrors the agent pattern from plan2code-loop.
|
||||
*/
|
||||
|
||||
@@ -12,13 +13,13 @@ import { tmpdir } from 'os';
|
||||
|
||||
// ── Agent definitions ────────────────────────────────────────────────────────
|
||||
|
||||
export type AgentType = 'claude-code' | 'copilot-cli';
|
||||
export type AgentType = 'claude-code' | 'copilot-cli' | 'devin-cli';
|
||||
|
||||
export interface AgentDef {
|
||||
name: AgentType;
|
||||
displayName: string;
|
||||
command: string;
|
||||
models: Array<{ value: string; label: string }>;
|
||||
defaultModel: string;
|
||||
}
|
||||
|
||||
export const AGENTS: Record<AgentType, AgentDef> = {
|
||||
@@ -26,23 +27,19 @@ export const AGENTS: Record<AgentType, AgentDef> = {
|
||||
name: 'claude-code',
|
||||
displayName: 'Claude Code',
|
||||
command: 'claude',
|
||||
models: [
|
||||
{ value: 'claude-opus-4-6', label: 'Claude Opus 4.6 (Recommended)' },
|
||||
{ value: 'claude-sonnet-4-6', label: 'Claude Sonnet 4.6 (faster)' },
|
||||
],
|
||||
defaultModel: 'default',
|
||||
},
|
||||
'copilot-cli': {
|
||||
name: 'copilot-cli',
|
||||
displayName: 'GitHub Copilot CLI',
|
||||
command: 'copilot',
|
||||
models: [
|
||||
{ value: 'claude-sonnet-4', label: 'Claude Sonnet 4 (Default)' },
|
||||
{ value: 'claude-sonnet-4.5', label: 'Claude Sonnet 4.5' },
|
||||
{ value: 'claude-opus-4.5', label: 'Claude Opus 4.5' },
|
||||
{ value: 'gpt-5', label: 'GPT-5' },
|
||||
{ value: 'gpt-5-mini', label: 'GPT-5 Mini' },
|
||||
{ value: 'gemini-3-pro-preview', label: 'Gemini 3 Pro' },
|
||||
],
|
||||
defaultModel: 'claude-sonnet-4',
|
||||
},
|
||||
'devin-cli': {
|
||||
name: 'devin-cli',
|
||||
displayName: 'Devin CLI',
|
||||
command: 'devin',
|
||||
defaultModel: 'default',
|
||||
},
|
||||
};
|
||||
|
||||
@@ -69,7 +66,8 @@ export async function invokeLLM(opts: InvokeLLMOptions): Promise<string> {
|
||||
'--print',
|
||||
'--dangerously-skip-permissions',
|
||||
];
|
||||
if (model) {
|
||||
// Only add --model if not using default
|
||||
if (model && model !== 'default') {
|
||||
args.push('--model', model);
|
||||
}
|
||||
|
||||
@@ -81,10 +79,28 @@ export async function invokeLLM(opts: InvokeLLMOptions): Promise<string> {
|
||||
} finally {
|
||||
try { unlinkSync(tempFile); } catch { /* ignore cleanup errors */ }
|
||||
}
|
||||
} else if (agent === 'devin-cli') {
|
||||
// Devin CLI: load prompt from a temp file, run single-turn, auto-approve tool calls
|
||||
const tempFile = join(tmpdir(), `plan2code-metrics-prompt-${Date.now()}.txt`);
|
||||
writeFileSync(tempFile, prompt, 'utf-8');
|
||||
|
||||
try {
|
||||
const args: string[] = ['--print', '--prompt-file', tempFile, '--permission-mode', 'dangerous'];
|
||||
// Only add --model if not using default
|
||||
if (model && model !== 'default') {
|
||||
args.push('--model', model);
|
||||
}
|
||||
|
||||
const result = await execa(def.command, args, { timeout });
|
||||
return result.stdout;
|
||||
} finally {
|
||||
try { unlinkSync(tempFile); } catch { /* ignore cleanup errors */ }
|
||||
}
|
||||
} else {
|
||||
// Copilot CLI: pipe prompt via stdin string
|
||||
const args: string[] = [];
|
||||
if (model) {
|
||||
// Only add --model if not using default
|
||||
if (model && model !== 'default') {
|
||||
args.push('--model', model);
|
||||
}
|
||||
args.push('--allow-all-tools', '-s');
|
||||
|
||||
@@ -35,7 +35,6 @@ The following are the current contents of the plan2code workflow prompt files be
|
||||
| avg_verification_items_added (Step 2) | ≤ 1.5 | lower is better |
|
||||
| avg_task_completion_rate (Step 3) | ≥ 0.95 | higher is better |
|
||||
| avg_blocker_count (Step 3) | ≤ 1.5 | lower is better |
|
||||
| avg_completion_marker_success_rate (Step 3) | ≥ 0.95 | higher is better |
|
||||
| avg_verification_failures_found (Step 4) | ≤ 1.0 | lower is better |
|
||||
| archival_success_rate (Step 4) | ≥ 0.99 | higher is better |
|
||||
| avg_user_rating (Feedback) | ≥ 7.0 | higher is better (1-10 scale, null if no feedback) |
|
||||
@@ -77,7 +76,7 @@ For each step (1–4), provide:
|
||||
|
||||
Numbered list. For each underperforming metric:
|
||||
1. **Metric:** [metric name] | **Value:** [actual] | **Target:** [target]
|
||||
- **File:** [plan2code-X--name.md]
|
||||
- **File:** [plan2code-X-name.md]
|
||||
- **Section:** [specific heading or section name]
|
||||
- **Hypothesis:** [specific gap in the prompt that would explain the metric miss]
|
||||
- **Confidence:** [High/Medium/Low] — [reason for confidence level]
|
||||
|
||||
@@ -28,7 +28,7 @@ The following diagnosis was produced by the analysis step:
|
||||
|
||||
1. **Char limit:** Each target file MUST stay under 11,000 characters after your edit is applied. This is enforced by code — edits that violate it will be automatically rejected.
|
||||
2. **Maximum 5 edits per cycle.** Focus on the highest-impact changes only.
|
||||
3. **Each edit must cite a specific metric** in its `expected_metric_impact` field (e.g., "avg_completion_marker_success_rate", "avg_blocker_count").
|
||||
3. **Each edit must cite a specific metric** in its `expected_metric_impact` field (e.g., "avg_task_completion_rate", "avg_blocker_count").
|
||||
4. **`old_text` must be verbatim** from the file. Copy-paste exactly — include surrounding whitespace/newlines as they appear. Edits with mismatched old_text will be automatically rejected.
|
||||
5. **Do NOT modify Role sections** (lines starting with "You are" at the top of each file) or step headings (lines starting with `#`).
|
||||
6. **Each edit must be independently applicable** — no edit should depend on another edit being applied first.
|
||||
@@ -41,7 +41,6 @@ The following diagnosis was produced by the analysis step:
|
||||
|
||||
- Target the specific sections identified in "Recommended Improvement Targets" from the diagnosis.
|
||||
- For high `avg_clarification_rounds`: Add more upfront specification examples or decision criteria to Step 1.
|
||||
- For low `avg_completion_marker_success_rate`: Clarify or simplify the completion marker format in Step 3.
|
||||
- For high `avg_blocker_count`: Add blocker-recovery guidance or prerequisite check instructions.
|
||||
- For low `avg_confidence`: Strengthen the confidence calculation instructions with clearer rubrics.
|
||||
- For low `avg_parallel_groups`: Add explicit guidance for identifying parallel tasks in Step 2.
|
||||
@@ -68,7 +67,7 @@ Produce EXACTLY the following sections. The JSON block is parsed by machine —
|
||||
```json
|
||||
[
|
||||
{
|
||||
"file": "plan2code-X--name.md",
|
||||
"file": "plan2code-X-name.md",
|
||||
"rationale": "One sentence explaining what this edit fixes",
|
||||
"expected_metric_impact": "avg_metric_name: expected direction and magnitude",
|
||||
"char_count_before": 0,
|
||||
|
||||