22 Commits

Author SHA1 Message Date
jparkerweb 1907281b40 Add a GitHub Issues backend to Pathfinder
Chart Step 1 now asks where the map should live: local files under
gitignored specs/ (the default) or a pathfinder:map issue whose decision
questions are its sub-issues, driven by gh. The pick is recorded as the
first Ground rules bullet and never re-asked; Auto-Discovery resolves
either backend, and an issue URL routes straight to github.

On github the local model maps onto the tracker's own primitives rather
than being simulated in issue bodies: sub-issues, native issue
dependencies for blocking, the assignee as the claim, one
pathfinder:<type>-<mode> label so type and mode cannot drift, and an
Answer comment plus a close reason that distinguishes a decision from a
question ruled out of scope. Both wiring calls key on the database id,
not the number. The map body therefore carries no checklist at all.

Sketches, secrets and the PLAN-DRAFT stay on local disk regardless --
plan2code-1-plan discovers its input with ls specs/ and has no notion of
a tracker. github is only offered after a five-check preflight, and the
offer names the repo's visibility, because a map on a public tracker
publishes the destination and the codebase recon. The Step 3 recon is
held until Step 6 so the no-fog off-ramp leaves no litter.

Depth lives in a seventh reference file; the orchestrator was
recompressed to absorb the new section within the 11,000 char limit by
removing duplication with its references. Adapted from the GitHub
tracker doc behind Matt Pocock's wayfinder skill (MIT).

AI Assisted
2026-08-08 19:00:26 -07:00
jparkerweb cf9fe6d80f Point plan2code-bot at the GitHub repo URL
AI Assisted
2026-08-08 12:39:52 -07:00
jparkerweb 0c17ef1454 Release v2.0.0
Ports upstream v1.17.0 and v2.0.0. Breaking: Gemini CLI is no longer an
install target. Adds Pathfinder Step 0, community feedback submission and
ingestion, and Devin CLI as an agent backend.

AI Assisted
2026-08-08 12:39:37 -07:00
jparkerweb 0767c6b4c7 Update the Claude Code status line
AI Assisted
2026-08-08 12:39:29 -07:00
jparkerweb b78b3cd6a3 Document Pathfinder and community feedback across the contributor docs
Moves the README's deep material into .readme/ (walkthrough, autonomous
loop, status line, metrics, test bot) and brings AGENTS.md, the .agents-docs
set and QUICK-REFERENCE.md in line with Step 0 and the metrics ingestion
flow.

AI Assisted
2026-08-08 12:39:29 -07:00
jparkerweb 474c76e564 Ingest community feedback submissions in plan2code-metrics
Adds a CLI flow that lists open community-feedback issues via gh, validates
and parses each METRICS_JSON payload, imports them deduped by run_id,
re-aggregates and closes the issue. Community runs cohort by Plan2Code
version rather than prompt fingerprint, and local cohorts stay preferred so
ingested feedback never displaces the maintainer's current generation. Adds
Devin CLI as an AI backend.

AI Assisted
2026-08-08 12:39:28 -07:00
jparkerweb 161e424632 Add Devin CLI as a selectable agent backend in plan2code-loop
Upstream replaced GitHub Copilot CLI with Devin; Plan2Code keeps both, so
Claude Code, Copilot CLI and Devin are all selectable and existing Copilot
selections keep working.

AI Assisted
2026-08-08 12:39:28 -07:00
jparkerweb 09aa565559 Drop Gemini CLI as an install target and register Pathfinder
Removes the .gemini/commands/*.toml surface and the build-time inlining it
required; every remaining target resolves reference directives at runtime.
The uninstall entry is retained behind an uninstallOnly flag so .toml files
from earlier versions can still be cleaned up.

AI Assisted
2026-08-08 12:39:13 -07:00
jparkerweb 0501260098 Add Pathfinder as an optional Step 0 and offload prompt depth to references
Pathfinder charts an idea that is too big to plan as a map of decision
questions under specs/<idea>/pathfinder/, clearing one per session until it
can hand a PLAN-DRAFT to Step 1. Finalize now archives pathfinder/ with the
spec and revise-plan no longer deletes it — both previously described the
cleanup target in wording that pointed at questions/. Init-update Step 7 and
review Session End move to reference files with inline fallbacks.

AI Assisted
2026-08-08 12:39:13 -07:00
jparkerweb 48a7cf68bd Pin LF line endings and renormalize the working tree
scripts/validate-char-count.js measures characters on disk, so a CRLF
checkout added ~1 char per line and pushed the 11k-budget prompt files over
the limit depending on how the repo was cloned. Pins eol=lf, marks binaries,
and renormalizes the files that had CRLF.

AI Assisted
2026-08-08 12:38:58 -07:00
jparkerweb 30860a7654 Redesign docs site and README around an airmail postcard theme
Adds a postage-stamp favicon set and banner, removes three orphaned images,
and records the site and README as protected divergences so an upstream sync
can't revert them.

AI Assisted
2026-08-08 12:32:17 -07:00
jparkerweb fda0146969 URL updates 2026-08-01 13:04:39 -07:00
jparkerweb 75605e5e41 update URLs 2026-08-01 12:56:52 -07:00
jparkerweb 3b18b42e30 Merge pull request #3 from jparkerweb/statusline-worktree
Add git worktree awareness to status line (v1.16.1)
2026-07-31 22:24:02 -07:00
jparkerweb 8e457fa6e4 Bump to v1.16.1 for status line git worktree awareness 2026-07-31 22:20:38 -07:00
jparkerweb 6740e26261 Add git worktree awareness to status line
Sessions in a linked worktree previously showed only the worktree's
directory name, so the underlying repository was unidentifiable and the
session looked like an unrelated project.

The project segment now renders as "repo | worktree". Repo identity comes
from git rev-parse --git-common-dir, so it is correct regardless of how the
worktree directory was named. A leading repo prefix is stripped from the
worktree name, and the name collapses to a bare marker when it merely
restates the branch already on screen -- matched across / _ . - separators
and type prefixes like feature/, and only when the branch is displayed.

Gated behind the new items.worktree config flag (default on). Costs one
extra timeout-bounded git call, skipped entirely outside git repos.
2026-07-31 22:20:38 -07:00
jparkerweb aedfe944b2 Merge pull request #2 from jparkerweb/handoff
Release v1.16.0: install handoff skill, add publish skill, sync upstream
2026-07-20 22:39:25 -07:00
jparkerweb 9b3c06b498 Address PR review feedback: update step label + release-list limit 2026-07-20 22:36:35 -07:00
jparkerweb 8fe1bfa805 Set v1.16.0 releaseDate to 2026-07-20 2026-07-20 22:33:24 -07:00
jparkerweb 68542fd778 Release v1.16.0: install handoff skill, add publish skill, sync upstream
- Make plan2code-handoff an installed workflow prompt: move to src/,
  register in install.js SOURCE_PROMPTS, remove repo-local-only copy
- Add repo-local maintainer /plan2code-publish skill (cuts GitHub Releases;
  excluded from install.js)
- Review workflow next-step suggestion made context-aware (upstream v1.15.4)
- Doc updates: list handoff in README, QUICK-REFERENCE, architecture table;
  add "Adding a New Workflow Prompt / Skill" checklist to dev-commands
- Fix broken AGENTS.md index links: rename .agents-docs loop/metrics files
  to plan2code-* to match references
- Bump version.json to 1.16.0 to align with package.json and CHANGELOG
2026-07-20 22:31:37 -07:00
jparkerweb 4656a6ebb7 Merge pull request #1 from jparkerweb/handoff
Add plan2code-handoff skill for conversation handoff documents
2026-07-20 22:09:45 -07:00
jparkerweb 7e9de755f1 Add plan2code-handoff skill for conversation handoff documents
AI Assisted
2026-07-20 22:07:21 -07:00
86 changed files with 8358 additions and 4774 deletions
+34 -9
View File
@@ -5,11 +5,23 @@
```
plan2code/
├── src/ # Source workflow prompts (9 markdown files)
── plan2code-review-references/ # Reference files for review skill
├── verification-protocol.md # Deep verification + confidence calibration
├── dimensions.md # 11 dimensions with detailed checklists
── false-positives.md # Known false-positive patterns
├── src/ # Source workflow prompts (11 markdown files)
── plan2code-0-pathfinder-references/ # Reference files for pathfinder skill
├── chart.md # MODE A: destination + frontier grills, templates
├── grilling.md # Folded-in grilling + domain-modeling
── questions.md # On-disk question-file format + markers
│ │ ├── resolve.md # Per-type resolution + graduating the fog
│ │ ├── handoff.md # Clearing gate + PLAN-DRAFT handoff
│ │ └── trail.md # Every-response map visual + pathed resume command
│ ├── plan2code-review-references/ # Reference files for review skill
│ │ ├── verification-protocol.md # Deep verification + confidence calibration
│ │ ├── dimensions.md # 11 dimensions with detailed checklists
│ │ ├── false-positives.md # Known false-positive patterns
│ │ └── session-end.md # Next-step routing at review session end
│ ├── plan2code-init-update-references/ # Reference files for init-update skill
│ │ └── ai-agent-file-sync.md # Step 7: replace AI configs with AGENTS.md refs
│ └── plan2code-4-finalize-references/ # Reference files for finalize skill
│ └── community-feedback-submission.md # STEP 6.5 payload schema + submission tiers
├── plan2code-loop/ # Autonomous loop CLI tool (Node.js/TypeScript)
│ ├── src/ # TypeScript source
│ └── dist/ # Built output (tsup)
@@ -27,6 +39,8 @@ plan2code/
│ └── local-commands/ # For per-project installation (.claude/, etc.)
├── .husky/ # Git hooks (husky)
│ └── pre-commit # Runs character count validation
├── .claude/ # Repo-local Claude Code config (NOT installed by install.js)
│ └── skills/ # Maintainer-only dev skills, e.g. plan2code-publish/
├── docs/ # Documentation and assets
├── specs/ # Feature specs (if any in-progress)
├── install.js # Interactive installer (Node.js)
@@ -52,13 +66,15 @@ plan2code/
|------|------|---------|
| `plan2code-init.md` | Init | Generate AGENTS.md as index + `.agents-docs/` section files (progressive discovery) |
| `plan2code-init-update.md` | Update | Update AGENTS.md with learnings; detects and routes edits to `.agents-docs/` files |
| `plan2code-quick-task.md` | 0 | Lightweight planning for small tasks |
| `plan2code-0-pathfinder.md` | 0 | Chart a foggy idea as a local map of decision questions under `specs/<idea>/pathfinder/`, resolve one per session, hand a seeded PLAN-DRAFT to Step 1 |
| `plan2code-quick-task.md` | quick | Lightweight planning for small tasks (standalone — not a pipeline step) |
| `plan2code-1-plan.md` | 1 | Requirements analysis & architecture |
| `plan2code-1b-revise-plan.md` | 1b | Mid-implementation revisions |
| `plan2code-2-document.md` | 2 | Create implementation specs |
| `plan2code-3-implement.md` | 3 | Execute implementation (phase by phase) |
| `plan2code-review.md` | review | Post-implementation comprehensive review |
| `plan2code-4-finalize.md` | 4 | Validate, summarize, feedback, archive (7 steps) |
| `plan2code-4-finalize.md` | 4 | Validate, summarize, feedback, archive (steps 17, +optional 6.5) |
| `plan2code-handoff.md` | handoff | Compact the conversation into a self-contained handoff document for a fresh session |
## Naming Convention
@@ -80,9 +96,18 @@ Some workflows use companion reference files for depth that exceeds the 11k char
**How the installer handles them:**
- **Skill-directory platforms** (Claude Code, Agents, Crush, Devin): reference files are nested as `<skill-name>/references/`. Read paths use canonical `references/<file>.md`.
- **Flat-file platforms** (Windsurf, Cursor, Copilot, Continue): reference files are placed as a sibling directory. The installer rewrites Read paths to the sibling directory name (e.g., `plan2code-review-references/<file>.md`).
- **TOML platforms** (Gemini CLI): reference files are skipped — TOML embeds content inline, so Read directives won't resolve. The orchestrator's inline fallback text covers this.
Reference files are NOT subject to the 11,000 character limit. Currently only the review workflow uses this pattern — it serves as the POC for potential adoption by other workflows.
Reference files are NOT subject to the 11,000 character limit. The review workflow pioneered this pattern (`verification-protocol`, `dimensions`, `false-positives`, `session-end`); the init-update workflow also uses it (`ai-agent-file-sync` for its Step 7), `plan2code-4-finalize.md` uses it for STEP 6.5 (`community-feedback-submission`), and `plan2code-0-pathfinder.md` leans on it hardest (`chart`, `grilling`, `questions`, `resolve`, `handoff`, `trail` — the orchestrator is a dispatcher, the depth lives in the references). Other workflows can adopt it when a source file's detail exceeds the 11k limit.
**`Read` directives must sit at column 0.** `install.js` matches `/^(Read\s+)references\//gm` for the flat-file path rewrite — anchored, with no leading-whitespace tolerance. An indented or bulleted `Read` line is silently skipped, so flat-file platforms ship a `references/` path that does not exist there (they receive the reference dir as a *sibling*, named `plan2code-<name>-references/`).
## Repo-Local Skills (`.claude/skills/`)
Maintainer-only Claude Code skills committed to the repo but **deliberately excluded** from `install.js` — they are dev tooling, not shipped product, so they never install to `~/.claude/skills/` and carry no version bump of their own (a product-version bump would wrongly imply a user-facing release); changelog mentions fold into the current version's entry.
- `plan2code-publish/` — cuts a GitHub Release from the top `CHANGELOG.md` entry once `CHANGELOG.md` / `version.json` / `package.json` agree and the version is ahead of the latest published release. Delegates tag creation to `gh release create --target main`.
**Warning:** anything named `plan2code-*` placed under `~/.claude/skills/` is deleted by the installer's uninstall (`uninstallFiles()` in `install.js`) and by every re-install's pre-copy cleanup in `install()` (both match `/^plan2code-/` for the Claude Code skills target). Keep these skills repo-local only.
## Status Line
+2
View File
@@ -17,4 +17,6 @@
- **Workflow file character limit:** All `src/plan2code-*.md` files must be ≤ 11,000 characters. A husky pre-commit hook enforces this. The 11,000 limit leaves buffer for platform-specific YAML headers (106-142 chars) to stay under Windsurf's 12,000 char limit.
- **Metrics internal prompts have no char limit:** Files in `plan2code-metrics/src/prompts/` are NOT subject to the 11,000 char limit — only `src/plan2code-*.md` consumer-facing prompts are.
- **User Feedback table format:** The `## User Feedback` markdown table in `overview.md` has a strict format the collector regex depends on. Field names must be exactly `Rating`, `Reason`, `Went Well`, `Went Poorly`. Pipe characters in values must be escaped as `\|`.
- **PLAN-DRAFT confidence numbers are scraped by regex:** when a `specs/<feature>/PLAN-DRAFT-*.md` contains no `<!-- METRICS_JSON ... -->` comment, `collector.ts` falls back to prose scraping (`collector.ts:186-241`). The overall-confidence pattern requires a literal `%`, but the four *breakdown* patterns (`collector.ts:201-204`) do **not**`/[Rr]equirements?[:\s|]+(\d{1,2})/` and its siblings match a bare dimension word followed by whitespace, a colon, or a pipe and then digits. So a PLAN-DRAFT written by anything other than `/plan2code-1-plan` Phase 7 must keep both the `%` sign **and** bare `Requirements` / `Feasibility` / `Integration` / `Risk` followed by a number off the page — including innocent table rows like `| Requirements | 11 |`. Otherwise the metrics pipeline records a planning-step confidence that no planning step produced. `/plan2code-0-pathfinder` works around this by hyphenating the labels (`Requirements-clarity 22/25`), which breaks the character class.
- **Reference file sizing guideline:** Files in `src/plan2code-*-references/` directories target ~100-200 lines each (soft guideline; evaluate splitting above 300). They are NOT subject to the 11,000 character limit. The pre-commit hook (`validate-char-count.js`) only checks `src/plan2code-*.md` flat files — subdirectory contents are automatically excluded.
- **The splitting guideline has a hard ceiling — reference files cannot always be split:** two constraints bound it. (1) Each new reference costs the orchestrator a column-0 `Read references/<file>.md` line plus its fallback blockquote (~150-200 chars), and orchestrators near the 11,000 limit have no room to spend. (2) **The path rewrite is not recursive**`syncPrompts()` rewrites `Read references/…` paths on the *orchestrator's* content only, so a `Read references/…` directive placed *inside* a reference file is never rewritten for flat-file targets; it ships as a dangling instruction pointing at a path that does not exist there. When a reference legitimately exceeds 300 lines (e.g. `plan2code-0-pathfinder-references/chart.md`), that is an accepted trade-off, not an oversight.
+27 -2
View File
@@ -61,9 +61,8 @@ cd plan2code-metrics && npm run build # Build the CLI
| VS Code Copilot | `.prompt.md` | — | — | YAML frontmatter |
| Codeium | `.md` | — | — | YAML frontmatter |
| Claude Code (Skills) | `SKILL.md` in subdir | `.claude/skills/<skill-name>/` | `~/.claude/skills/<skill-name>/` | YAML frontmatter + `disable-model-invocation: true` |
| Agent Skills (Amp · Devin · Gemini CLI · OpenCode · Zed) | `SKILL.md` in subdir | `.agents/skills/<skill-name>/` | `~/.agents/skills/<skill-name>/` | YAML frontmatter (no disable flag) |
| Agent Skills (Amp · Devin · OpenCode · Zed) | `SKILL.md` in subdir | `.agents/skills/<skill-name>/` | `~/.agents/skills/<skill-name>/` | YAML frontmatter (no disable flag) |
| Crush | `SKILL.md` in subdir | — (global only) | `~/.config/crush/skills/<skill-name>/` (Unix) / `%LOCALAPPDATA%\crush\skills\<skill-name>\` (Windows) | YAML frontmatter |
| Gemini CLI (TOML) | `.toml` | `.gemini/commands/` | `~/.gemini/commands/` | None (TOML fields: `description`, `prompt`) |
| Pi (pi.dev) | `.md` | `.pi/prompts/` | `~/.pi/agent/prompts/` | YAML frontmatter (`description`) |
## Editing Workflow Prompts
@@ -74,3 +73,29 @@ When modifying workflow prompts in `src/`:
2. Run `node install.js` to regenerate distribution files
3. Test the workflow in your AI tool of choice
4. The `dist/` folder is regenerated automatically — don't edit files there directly
## Adding a New Workflow Prompt / Skill
Adding a new prompt to `src/` is more than dropping in a file — the installer is
driven by an explicit registry and several docs enumerate the command set. When
you add a `src/plan2code-<name>.md`, do **all** of the following so nothing drifts:
1. **Create the source file** `src/plan2code-<name>.md` — body content **only**,
no YAML frontmatter (the installer generates frontmatter per platform). Keep
it **under 11,000 characters** (`npm test` enforces this).
2. **Register it in the installer.** Add an entry to the `SOURCE_PROMPTS` array in
`install.js` (`source`, `stepNumber`, `name`, `displayName`, `description`, and
`isUtility: true` for non-numbered utilities). If `stepNumber` is a non-numeric
label (e.g. `'handoff'`), add a matching case to `generateStepLabel()` so the
generated description reads correctly.
3. **Update every doc that lists the command set** — keep these in sync, they are
the canonical inventories:
- `README.md` — command table ("When to Use")
- `QUICK-REFERENCE.md` — Commands table
- `.agents-docs/AGENTS-architecture.md` — "Workflow Prompts (in `src/`)" table
- `CHANGELOG.md` — add an entry under the current version
- `docs/index.html`**only** if the new prompt belongs to the core pipeline
shown there; utilities (like `init`, `quick-task`, `handoff`) are deliberately
omitted from that curated marketing list.
4. **Regenerate and validate:** run `node install.js` (regenerates `dist/`) and
`npm test` (character-count validator now covers the new file).
@@ -32,6 +32,9 @@ plan2code-metrics # Run (fully interactive, no flags)
| Analyze | AI diagnosis of weak metrics |
| Propose | AI improvement proposals with validation |
| Apply | Interactive diff review + file patching |
| Fetch community submissions | List/parse/import open community-feedback GitHub issues from jparkerweb/plan2code, close on success |
Community submissions arrive as GitHub issues labeled `community-feedback` on `jparkerweb/plan2code`, created by the finalize prompt's post-Step-6 submission flow; the "Fetch community submissions" option requires an authenticated `gh` CLI to list/close them.
## Key Source Files
@@ -40,11 +43,12 @@ plan2code-metrics # Run (fully interactive, no flags)
| `types.ts` | All interfaces (`RunMetrics`, `UserFeedback`, `CohortMetrics`, etc.) + `METRIC_TARGETS` |
| `collector.ts` | Reads project artifacts → run JSON (parses plan drafts, overview.md, loop logs) |
| `aggregator.ts` | Merges runs by prompt generation (SHA cohort) → `aggregated.json` |
| `community.ts` | Lists/parses/closes `community-feedback`-labeled GitHub issues via `gh` CLI |
| `analyzer.ts` | AI diagnosis via `prompts/analyze.md` template |
| `improver.ts` | AI proposals via `prompts/improve.md` + validation (char count, old_text match) |
| `applier.ts` | Interactive diff review + file patching |
| `cli.ts` | Menu-driven interactive CLI (100% prompts, no flags) |
| `invoke-llm.ts` | Unified LLM interface (Claude Code or Copilot CLI) |
| `invoke-llm.ts` | Unified LLM interface (Claude Code, GitHub Copilot CLI, or Devin CLI) |
## User Feedback
@@ -66,6 +70,7 @@ Feedback is collected during finalize (Step 5) or retroactively via the CLI. Pip
- **Claude Code** (recommended): `claude` CLI with `--inputFile` for prompt delivery
- **GitHub Copilot CLI**: `copilot` CLI with stdin prompt delivery
- **Devin CLI**: `devin` CLI with `--print --prompt-file <file> --permission-mode dangerous`
## Metric Targets
+159
View File
@@ -0,0 +1,159 @@
---
name: plan2code-publish
description: "Publish a GitHub Release for jparkerweb/plan2code whenever CHANGELOG.md's top version is ahead of the latest published release on GitHub — after verifying CHANGELOG.md, version.json, and package.json all agree on the version. Use this skill when the user says 'publish a release', 'create a GitHub release', 'cut a release', 'tag a release', 'tag and release', 'is the changelog published', or otherwise mentions publishing/releasing/tagging this repo."
---
# Plan2Code Release Publisher
Publish a GitHub Release for `jparkerweb/plan2code` whenever `CHANGELOG.md`'s top version is ahead of the latest published release on GitHub. This turns the merged CHANGELOG entry on `main` into an actual GitHub Release (which also creates the `vX.Y.Z` git tag).
**Repo-local by design.** This skill lives in the repo's `.claude/skills/` and is intentionally NOT wired into `install.js` — it is a maintainer dev tool, not part of the shipped product, so it is never installed to `~/.claude/skills/`. Do **not** copy it there: the uninstaller (`uninstallFiles()` in `install.js`) and every re-install's pre-copy cleanup in `install()` both delete every entry matching `/^plan2code-/` under `~/.claude/skills/`, so a copy placed there would be silently removed.
## Workflow
### Step 1 — Preflight: clean working tree
Run `git status --porcelain`. If the output is non-empty, stop immediately — do not switch branches or take any other action. Tell the user:
> Working tree has uncommitted changes. Commit or stash them, then re-run this skill.
### Step 2 — Switch to main and pull
If the working tree is clean, switch to `main` and pull latest. Run these as two separate, non-chained commands (never `&&`/`;`-chain `git`/`gh` commands — Windows PowerShell 5.1 rejects `&&`):
```bash
git checkout main
git pull origin main
```
### Step 3 — Read the three version sources
Read all three files at the repo root and extract each version:
- `CHANGELOG.md` — parse the first `## vX.Y.Z` heading. Strip the leading `v``$CHANGELOG_VERSION`. Note: plan2code CHANGELOG headings are `## vX.Y.Z` (v-prefixed, **no** date and **no** brackets) — different from a `## [x.y.z] - YYYY-MM-DD` format.
- `version.json` — the `version` field → `$VERSION_JSON`. Also read its `releaseDate` field → `$RELEASE_DATE` (informational only — shown at the confirm step, never part of the sync gate; use `(none)` if absent).
- `package.json` (root) → the `version` field → `$PACKAGE_JSON`.
Also capture from `CHANGELOG.md` the **section body** for the top version: everything from the `## vX.Y.Z` heading line itself (heading **included**) up to — but not including — the next `## v` heading, or end-of-file if there is none. Call this `$SECTION`. Keep the `## vX.Y.Z` heading in `$SECTION`; plan2code release bodies include it.
### Step 4 — Version-sync preflight (STOP on mismatch)
All three versions must be identical. This enforces the repo's documented invariant (`.agents-docs/AGENTS-code-style.md` → "Version sync"): `CHANGELOG.md`, `version.json`, and `package.json` must always show the same version number.
If `$CHANGELOG_VERSION`, `$VERSION_JSON`, and `$PACKAGE_JSON` are **not** all equal, **stop** — do not read the release, do not publish. Report exactly which files disagree:
> ⚠️ Version files are out of sync — refusing to publish. The repo requires `CHANGELOG.md`, `version.json`, and `package.json` to match.
>
> - **CHANGELOG.md:** $CHANGELOG_VERSION
> - **version.json:** $VERSION_JSON
> - **package.json:** $PACKAGE_JSON
>
> Fix the mismatch first, then re-run this skill. To realign: pick the intended version (normally the highest / newest CHANGELOG entry) and update the other two files to match — see the "Version sync" gotcha in `.agents-docs/AGENTS-code-style.md`.
This skill never modifies these files — it only reads and compares them.
### Step 5 — Determine the highest published release
List every published (non-draft) release and take the numerically-highest semver tag. Do **not** rely on `gh release view` / GitHub's "Latest" flag: that flag returns whatever release is *marked* latest — normally the newest semver, but a maintainer can manually pin it to an older release, which would make the Step 6 ahead-comparison misfire.
```bash
gh release list --repo jparkerweb/plan2code --limit 100 --json tagName,isDraft -q '.[] | select(.isDraft==false) | .tagName'
```
Strip any leading `v` from each returned tag, compare them as numeric `(major, minor, patch)` tuples (the same rule as Step 6), and take the maximum → `$RELEASE_VERSION`. If the command returns no tags or exits non-zero for **any** reason (no releases yet, transient error, etc.), set `$RELEASE_VERSION = "0.0.0"` — no error-text matching is needed.
### Step 6 — Compare versions
Compare `$CHANGELOG_VERSION` vs `$RELEASE_VERSION` as a numeric `(major, minor, patch)` tuple. Never do a plain string/lexicographic compare — e.g. `"1.9.0" > "1.10.0"` is true as strings but wrong numerically.
### Step 7 — Not ahead: no-op
If `$CHANGELOG_VERSION``$RELEASE_VERSION`, print a simple status message showing both versions and stop:
> CHANGELOG top version ($CHANGELOG_VERSION) is not ahead of the latest published release ($RELEASE_VERSION). Nothing to publish.
No error is raised and no release is created.
### Step 8 — Ahead: compute and confirm
If `$CHANGELOG_VERSION` > `$RELEASE_VERSION`, compute:
- `tag = "v$CHANGELOG_VERSION"`
- `title = "v$CHANGELOG_VERSION"` (plan2code keeps the `v` prefix in release titles)
- `notes` = the header line `# What's New 🎉`, then one blank line, then `$SECTION` verbatim (`$SECTION` already starts with the `## vX.Y.Z` heading). Build this as a real multi-line string with **actual newlines** — the `\n\n` shorthand shown elsewhere means "a blank line," never the literal two-character sequence `\` + `n`. Getting this wrong would run the header and the first CHANGELOG heading together with a stray `\n\n` in the published body.
Present all three to the user and wait for an explicit answer before any write. Offer the optional decorative title suffix — a plain `vX.Y.Z` title is the default, but a release may append one (e.g. `v1.14.0 - 🔍 Review workflow`):
> 🚀 [Publish Plan2Code Release]
>
> CHANGELOG is ahead of the latest published release:
> - **Current release:** $RELEASE_VERSION
> - **CHANGELOG top version:** $CHANGELOG_VERSION
> - **version.json releaseDate:** $RELEASE_DATE (informational — read from `version.json`)
>
> Proposed release:
> - **Tag:** $tag
> - **Title:** $title
> - **Notes:**
> ```
> $notes
> ```
>
> Publish this release? Reply **yes** to publish as-is, **no** to cancel, or provide a decorative suffix to append to the title (e.g. `⇢ 🎆 Feature Name` or `- 🔍 Feature Name`).
If the user supplies a suffix, set `title = "v$CHANGELOG_VERSION " + <suffix>` (single space join) and proceed to publish. The tag and notes are unaffected by the suffix.
### Step 9 — Publish (on approval)
On approval, write `$notes` to a temp file — never pass multiline text inline via `--notes`, that regresses into a quoting bug — then create the release targeting `main`. Two requirements for the temp file: `$notes` must already hold **real newlines** (per Step 8) because `printf '%s'` / `WriteAllText` write it byte-for-byte — a literal `\n` in the string lands literally in the release body; and it MUST be **UTF-8** because the notes contain emoji (`🎉`, `🐛`, `✨`, `🔧`).
**bash (preferred in this environment):**
```bash
NOTES_FILE=$(mktemp)
printf '%s' "$notes" > "$NOTES_FILE"
gh release create "$tag" --repo jparkerweb/plan2code --title "$title" --notes-file "$NOTES_FILE" --target main
rm -f "$NOTES_FILE"
```
**PowerShell:** do NOT use `Set-Content` — under Windows PowerShell 5.1 it writes ANSI/UTF-16 by default and mangles the emoji into `??`. Write UTF-8 **without BOM** (a BOM would leak into the release body):
```powershell
$NotesFile = [System.IO.Path]::GetTempFileName()
[System.IO.File]::WriteAllText($NotesFile, $notes, [System.Text.UTF8Encoding]::new($false))
gh release create "$tag" --repo jparkerweb/plan2code --title "$title" --notes-file "$NotesFile" --target main
Remove-Item -Path $NotesFile
```
Always delete the temp file afterward, regardless of whether `gh release create` succeeded or failed.
**Failure handling — already-exists classification:** if `gh release create` exits non-zero, inspect the error text.
- If and only if it contains the substring `already exists` (real output: `HTTP 422: Validation Failed` / `Release.tag_name already exists`), report this to the user as already published, not as a raw CLI error:
> This version ($CHANGELOG_VERSION) was already published as a release — nothing more to do.
- Every other failure (auth, network, permissions, etc.) must be surfaced to the user verbatim. Never silently reclassify a genuine failure as "already published."
### Step 10 — Verify
Confirm the release now exists and report its URL:
```bash
gh release view "$tag" --repo jparkerweb/plan2code
```
Report the release URL to the user.
## Rules
- **Repo-local only** — this skill is not part of the installed product; never add it to `install.js`, and never copy it to `~/.claude/skills/` (the uninstaller deletes `plan2code-*` entries there).
- **Read-only on version files** — never modify `CHANGELOG.md`, `version.json`, or `package.json`; this skill only reads and compares them.
- **Version-sync gate is hard** (Step 4) — if the three version sources disagree, stop and report; do not publish a release from an inconsistent repo.
- **Never run a local `git tag` or `git push`** — tag creation is delegated entirely to `gh release create --target main`.
- **Always use `--notes-file`**, never inline multiline `--notes`, and always write the notes file as UTF-8 (no BOM) so emoji survive.
- **Always clean up the temp notes file**, even on a mid-run error.
- **Always get explicit approval before any write** (Step 8) — no release is created without a yes (or a yes-with-suffix).
- **Derive the current version from the highest published semver tag** (Step 5), never from GitHub's manually-pinnable "Latest" flag. On an empty or failed release list, treat it as "no prior release" (baseline `0.0.0`) — no error-text matching needed.
- **On a `gh release create` failure** (Step 9), only reclassify as already-published when the error contains `already exists` — every other failure must be shown verbatim, never swallowed.
- **Idempotent and safe to re-run** at any time — re-running after a successful publish hits the Step 7 no-op; re-running after a race-lost publish hits the Step 9 already-exists handling.
File diff suppressed because one or more lines are too long
+26
View File
@@ -1,3 +1,29 @@
# ─────────────────────────────────────────────────────────────────────────────
# Line endings
#
# Git stores LF, and every text file is checked out as LF on all platforms.
# This is not cosmetic: scripts/validate-char-count.js measures the characters
# actually on disk, so a CRLF checkout adds ~1 character per line. The workflow
# prompts in src/plan2code-*.md run close to their 11,000 character budget
# (several sit above 10,800), and a CRLF working tree pushes them over — turning
# `npm test` into a check that passes or fails depending on how the repo was
# cloned. Pinning eol=lf makes the count reproducible everywhere.
# ─────────────────────────────────────────────────────────────────────────────
* text=auto eol=lf
# ─────────────────────────────────────────────────────────────────────────────
# Binary — never line-ending-converted, never diffed as text
# ─────────────────────────────────────────────────────────────────────────────
*.png binary
*.jpg binary
*.jpeg binary
*.gif binary
*.ico binary
*.mp4 binary
*.webm binary
*.woff binary
*.woff2 binary
# The encrypted sync-repo blob is base64 text but must never have its bytes
# altered by line-ending normalization. Treat it as binary so autocrlf/eol
# settings can never corrupt the ciphertext.
+1
View File
@@ -11,6 +11,7 @@ plan2code-metrics/package-lock.json
.plan2code-metrics
nul
.cognition/
handoffs/
node_modules/
package-lock.json
SYNC.md
+64
View File
@@ -0,0 +1,64 @@
# Autonomous Loop
`plan2code-loop` is a CLI that works through your spec's tasks on its own, one agent call at a
time. It is an **alternative to Step 3**, not a replacement — the four-step workflow and the specs it
produces are unchanged.
← [Back to README](../README.md)
---
## When to use it instead of Step 3
| Approach | Best for |
|----------|----------|
| `/plan2code-3-implement` | Interactive control, reviewing each phase, logic that needs your judgment |
| `plan2code-loop` | Straightforward implementations, batch work, overnight runs |
The loop reads the same `overview.md` and phase files. You can start with the loop and finish by
hand, or the reverse — the checkboxes are the only handoff.
---
## Install
```bash
# From the plan2code root directory
node install.js # A (everything + dev tools) — or C → O (loop only)
```
## Run
```bash
plan2code-loop # fully interactive
```
It will:
1. Find specs in `./specs/`
2. Let you pick one if there are several
3. Offer to resume an existing session
4. Ask for a JIRA ticket ID, which agent to drive, the loop mode, and a max iteration count
Then, per iteration: read `overview.md` and the phase files, find the first unchecked task (or
phase), implement it, mark the checkbox, repeat — until everything is done or it hits the iteration
cap.
---
## Loop modes
| Mode | Each agent call | Git commits | Best for |
|------|-----------------|-------------|----------|
| **One task per loop** (default) | Implements a single task | The Node controller commits after each task | Smaller models, cautious execution |
| **One phase per loop** | Implements every task in a phase | The agent commits after each task, with the JIRA ID | Larger context windows, tightly related tasks |
Session state lives per-spec in `specs/<feature>/.plan2code-loop/`, so each feature's progress
stays isolated.
---
## Full documentation
Architecture, completion markers, agent adapters, and configuration:
[`plan2code-loop/README.md`](../plan2code-loop/README.md)
+59
View File
@@ -0,0 +1,59 @@
# Metrics &amp; Self-Improvement
`plan2code-metrics` closes the loop on the workflow itself: it collects data from your finished
specs, aggregates it across runs and prompt generations, then uses AI to diagnose which step is
underperforming and propose edits to the workflow prompts.
Aimed at **contributors and heavy users** — you don't need it to use Plan2Code.
← [Back to README](../README.md)
---
## The habit
One thing to remember: **collect after every finished spec.** Everything else is on demand.
```bash
# Install once, from the plan2code root
node install.js # A (everything + dev tools) — or C → M (metrics only)
# After finishing a spec (steps 14)
cd your-project
plan2code-metrics # → "Collect metrics" → pick the spec dir → done, ~5 seconds
```
Then, when you're curious or have a few runs banked:
```bash
plan2code-metrics # → "View metrics status" the dashboard
# → "Run analysis" AI diagnosis of weak steps
# → "Generate improvement proposal" concrete prompt edits
# → "Review and apply" patch src/plan2code-*.md
```
## How much data you need
| Runs | What you get |
|------|--------------|
| **1** | Raw data and a basic dashboard. Start here. |
| **3+** | AI analysis unlocks. Pattern detection starts working. |
| **510+** | Averages stabilise; generation-over-generation comparisons become meaningful. |
You're looking for trends, not individual scores.
---
## Sending feedback upstream
`/plan2code-4-finalize` can submit an anonymised metrics payload to the maintainers as a
`community-feedback` issue on the repo. Community runs are cohorted by the Plan2Code version that
produced them, so your data improves the prompts everyone installs — without displacing the
maintainer's own measurements.
---
## Full documentation
Data model, aggregation, cohorts, analysis prompts, and the ingestion flow:
[`plan2code-metrics/README.md`](../plan2code-metrics/README.md)
+54
View File
@@ -0,0 +1,54 @@
# Claude Code Status Line
A persistent three-line status bar for Claude Code: model, project, git branch, uncommitted diff
stats, session duration and cost, context-window usage, and plan quota.
It reads everything from the JSON Claude Code already sends on stdin — **no API calls, no auth, no
background processes.** Optional, and unrelated to the workflow itself.
← [Back to README](../README.md)
---
## What it looks like
On Pro / Max / Teams accounts, where rate limits are available:
```
◦ ◦ Opus 5 / high │ plan2code │ feature/PCWEB-11702-pathfinder │ +12 -3
╭●╮ ┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄
├■┤ 3h 5m ($4.62) │ ▰▰▰▰▰▱▱▱▱▱▱▱ 42% (84K) │ 5h: 28% · 7d: 61%
```
On Enterprise / Bedrock / Vertex / pay-as-you-go, where they aren't, the last segment becomes session
token counts instead:
```
◦ ◦ Sonnet 5 │ plan2code │ main │ +12 -3
╭●╮ ┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄
├■┤ 2m ($0.18) │ ▰▰▱▱▱▱▱▱▱▱▱▱ 18% │ 88k in · 3k out
```
The icons down the left are Planny, the project mascot.
---
## Install
```bash
node install.js # A (everything + dev tools) — or C → S (status line only)
```
That copies the script to `~/.claude/plan2code-statusline.js`, writes a default config to
`~/.claude/statusline-config.json` (an existing config is preserved), and registers it in
`~/.claude/settings.json`.
If you already have a custom `statusLine` entry, the installer asks before replacing it and backs the
old one up. Uninstalling removes the script and the settings entry but leaves your config file alone.
---
## Full documentation
Config options, compact mode, thresholds, and troubleshooting:
[`src/statusline-claude/README.md`](../src/statusline-claude/README.md)
+46
View File
@@ -0,0 +1,46 @@
# Workflow Test Bot
`plan2code-bot` drives the whole workflow end to end with no human in the loop — init, plan,
document, implement, finalize — to test that the prompts still hold together.
Built for **maintainers**. If you're using Plan2Code to ship features, you don't need this.
← [Back to README](../README.md)
---
## Two modes, auto-detected
| Condition | Mode | What it does |
|-----------|------|--------------|
| No `AGENTS.md` in the working directory | **New project** | Invents an app idea, creates a subdirectory, writes `IDEA.md`, runs init, then all four steps |
| `AGENTS.md` present | **Enhancement** | Reads the existing codebase, proposes a realistic enhancement, writes `IDEA.md`, then runs plan → finalize |
## Install
```bash
node install.js # I (everything) — or C → B (bot only)
```
## Run
```bash
# New project mode — start from an empty directory
mkdir /tmp/test-bot && cd /tmp/test-bot
plan2code-bot
# Enhancement mode — start from a project that has an AGENTS.md
cd my-project
plan2code-bot
# Or seed the idea generator yourself
plan2code-bot --idea "web app that renders the current weather as vector images"
```
---
## Full documentation
Agent SDK wiring, run evaluation, and scoring:
[`plan2code-bot/README.md`](../plan2code-bot/README.md) ·
[`plan2code-bot/EVALUATION-SYSTEM.md`](../plan2code-bot/EVALUATION-SYSTEM.md)
+103
View File
@@ -0,0 +1,103 @@
# Walkthrough: one feature, start to finish
A REST API for task management, from a sentence to archived specs. Every block below is a **separate
conversation** — that's the whole trick.
← [Back to README](../README.md)
---
## Session 1 — Plan
```
You: /plan2code-1-plan
I want to build a REST API for a task management application.
AI: 🤔 [REQUIREMENTS ANALYSIS]
Six questions. Four of them I need answered before I can size this.
...
AI: 🤔 [TECH STACK]
Proposing Fastify + Postgres + Zod. I don't pick this alone — sign off?
You: approved
AI: 🤔 [ARCHITECTURE DESIGN]
...
AI: 🤔 [TRANSITION DECISION]
Confidence: 92% (gate is 90%). Assumptions written to the draft: 3.
→ specs/task-api/PLAN-DRAFT-20260804.md
```
Two files land: the draft, and a `PLAN-CONVERSATION-*.md` log of how you got there.
> **Started foggy instead?** Run `/plan2code-0-pathfinder` first. When its map clears it writes a
> `PLAN-DRAFT` that Step 1 picks up at Phase 4, with requirements and scope already answered.
---
## Session 2 — Document
```
You: /plan2code-2-document
specs/task-api/PLAN-DRAFT-20260804.md
AI: 📝 [DOCUMENTATION]
Creating specs/task-api/overview.md
Creating specs/task-api/Phase 1.md Project setup (6 tasks)
Creating specs/task-api/Phase 2.md Data model (8 tasks)
Creating specs/task-api/Phase 3.md API endpoints (11 tasks)
Creating specs/task-api/Phase 4.md Authentication (7 tasks)
Parallel execution groups: Phase 3 and Phase 4 don't share files.
```
---
## Sessions 3…N — Implement, one phase each
```
You: /plan2code-3-implement
specs/task-api/overview.md
AI: ⚡ [PHASE 1: Project setup]
(Phase 1 is the next unchecked phase.)
...
✓ 6/6 tasks. Checkboxes updated in Phase 1.md and overview.md.
Next: Phase 2 — Data model (8 tasks). New conversation.
```
Repeat. New conversation each time. It always finds the next open phase itself.
When you reach a parallel group, it offers you the choice — open a second agent, take the other
phase, and the `[/]` marks keep them from colliding.
**Optional, any time:** `/plan2code-review` for an independent second opinion on what just landed.
---
## Final session — Finalize
```
You: /plan2code-4-finalize
specs/task-api/overview.md
AI: 🧹 [VALIDATION]
All 32 tasks verified against the phase specs. 2 gaps found and fixed.
AI: 🧹 [DOCUMENTATION REVIEW]
README needs the new /tasks endpoints. AGENTS.md is current.
AI: 🧹 [SPEC CLEANUP]
Moved specs/task-api/ → specs--completed/task-api/
Implementation complete.
```
---
## If requirements move mid-build
Don't patch the code and hope the specs catch up. Run `/plan2code-1b-revise-plan` — it edits the
specs (and only the specs), so the drawing and the build stay in agreement.
+2 -2
View File
@@ -1,6 +1,6 @@
# AGENTS.md
This file provides guidance to AI coding agents like Claude Code (claude.ai/code), Cursor AI, Codex, Gemini CLI, GitHub Copilot, Devin, Zed, and other AI coding assistants when working with code in this repository.
This file provides guidance to AI coding agents like Claude Code (claude.ai/code), Cursor AI, Codex, GitHub Copilot, Devin, Zed, and other AI coding assistants when working with code in this repository.
## Project Overview
@@ -52,7 +52,7 @@ Details: [Code Style & Gotchas](./.agents-docs/AGENTS-code-style.md)
## Mascot
The project has a mascot called "Planny" - an ASCII art robot that appears in installer output and workflow prompts. Mascot variants are defined in `MASCOT` constant in `install.js` and appear in workflow markdown files.
The project has a mascot called "Planny" an ASCII art robot that appears in installer output and workflow prompts. Mascot variants are defined in the `MASCOT` constant in `install.js` and appear in workflow markdown files.
```
╭───╮
+93
View File
@@ -2,6 +2,99 @@
All notable changes to Plan2Code will be documented in this file.
## v2.1.0
### ✨ Added
- **GitHub Issues backend for Pathfinder** — `/plan2code-0-pathfinder` no longer assumes local files. Chart Step 1 now asks, HITL and never self-picked, where the map should live: **local** (the default — files under gitignored `specs/<idea>/pathfinder/`, private and solo) or **github** (a `pathfinder:map` issue whose decision questions are its sub-issues, driven by the `gh` CLI). The pick is recorded as the first `## Ground rules` bullet, never re-asked and never switched mid-map, and Auto-Discovery resolves either backend — an issue URL or number as the argument routes straight to `github`.
On `github`, the local model maps onto the tracker's own primitives rather than being simulated in issue bodies: a question is a **sub-issue** (`sub_issues` endpoint), blocking is GitHub's **native issue dependencies** (`dependencies/blocked_by`, so the frontier renders in GitHub's UI without opening the map), the claim is the **assignee**, `Type:` becomes a single `pathfinder:<type>-<mode>` label so type and mode cannot drift, `Locked: yes` becomes `pathfinder:locked`, and resolution is an `## Answer` comment followed by a close — `completed` for a decision, `not planned` for a question ruled out of scope. Both wiring calls key on the issue's **database id**, not its `#number`. The map body therefore carries no question checklist at all: the frontier is a live query, which removes the single largest source of drift in the local backend.
Three things stay on local disk whatever the backend: runnable sketches (`specs/<idea>/pathfinder/sketch-NN/`), anything secret, and the `PLAN-DRAFT-<date>.md``/plan2code-1-plan` discovers its input with `ls specs/` and has no notion of a tracker, so a draft that existed only as an issue would be invisible to the rest of the pipeline.
Guardrails carried over from the local backend's assumptions: `github` is only offered after a five-check preflight (`gh` present, authenticated, GitHub remote, issues enabled, push access), and the offer must name the repo's **visibility** in the same breath, because a map on a public tracker publishes the destination, the rejected alternatives, and the codebase recon. The Step 3 recon is held in-session and published at Step 6, so the Step 4 no-fog off-ramp leaves no litter on a shared tracker. A `gh` failure mid-session stops the session rather than falling back to local files, which would fork the map.
Depth lives in a seventh reference file, `src/plan2code-0-pathfinder-references/github-issues.md` — preflight, label set, the local↔GitHub equivalence table, create-then-wire charting, the frontier query, resolve and out-of-scope flows, reconcile, the trail footer, handoff, and a failure-mode table. Adapted from the GitHub tracker doc behind Matt Pocock's [`wayfinder`](https://github.com/mattpocock/skills/tree/main/skills/engineering/wayfinder) skill (MIT).
### 🔧 Changed
- **`src/plan2code-0-pathfinder.md` recompressed** to absorb the new `## Backend` section within the 11,000-character workflow-file limit — duplication between the orchestrator and its reference files was removed (the local layout and marker legend now live only in `questions.md`; Form A/B footer detail only in `trail.md`), and step text tightened. No behaviour was dropped.
## v2.0.0
Ports upstream v1.17.0 and v2.0.0 into Plan2Code.
### 💥 Breaking
- **Gemini CLI is no longer an install target.** The `.gemini/commands/*.toml` surface is removed, along with the build-time machinery it required: `generateTomlContent()`, `writeTomlToDestination()`, `inlineReferenceContent()`, and the `dest.type === 'toml'` branch in `syncPrompts()`. Every remaining target resolves `Read references/*.md` directives at runtime, so reference content no longer needs inlining at build time.
**Uninstall still cleans up Gemini files.** The `.gemini/commands` entry is deliberately retained in the uninstall target list (labelled `Gemini CLI (legacy — uninstall only)`) behind a new `uninstallOnly` flag, so `.toml` files written by v1.x installs can still be removed. Do not add it back to `LOCAL_DESTINATIONS` / `GLOBAL_DESTINATIONS`.
All other platforms are unaffected — Claude Code (skills *and* commands), Cursor, Windsurf, Continue, Codeium, GitHub Copilot, VS Code Copilot, Pi, Crush, Amp, Devin, OpenCode, and Zed all install exactly as before.
### ✨ Added
- **Pathfinder workflow** — new `/plan2code-0-pathfinder` command, an optional Step 0 for an idea too big and unclear to plan yet: where you can feel the shape of the work but can't write it down as requirements. Adapted from Matt Pocock's [`wayfinder`](https://github.com/mattpocock/skills/tree/main/skills/engineering/wayfinder) skill (MIT), reworked for Plan2Code's local `specs/` workflow.
Pathfinder names a **destination**, then charts the way to it as a map of decision **questions** under `specs/<idea>/pathfinder/``map.md` as the index plus one `questions/NN-<slug>.md` file per decision. It resolves **one question per session** (research excepted), and each resolution clears the fog ahead of it, graduating whatever became specifiable into fresh questions. When nothing is left to decide, it writes `specs/<idea>/PLAN-DRAFT-<date>.md` carrying the status string `/plan2code-1-plan` already recognizes, so planning resumes at Phase 4 in the same folder with Requirements, System Context, and Scope pre-answered.
**Grilling is batched** — up to three *independent* probes per turn instead of one probe per round trip, delivered either through the environment's structured question tool or as numbered prose Q blocks, chosen per batch by a detail test. Probes are written in plain English, and any probe you skip is re-asked rather than quietly dropped. It **plans, it never builds**: four question types — `grill` (HITL, the default), `research` (AFK, resolved by background subagents in parallel), `sketch` (HITL), and `legwork`.
Map state uses the house checkbox vocabulary — `[ ]` open (the frontier), `[/]` claimed, `[x]` resolved, `[!]` blocked, `[-]` out of scope. `questions/` is ground truth and `map.md` is a rebuildable index: every Work session reconciles the two before choosing, which self-heals drift and recovers claims left by a crashed session. Every response ends with a **Trail Footer** — a one-line path from `START` to the `⚑` destination, a numbered legend, a plain-English confidence line, and exactly one closer chosen by turn type (a turn that asks you something never emits a resume command).
Depth lives in six new reference files under `src/plan2code-0-pathfinder-references/` (`chart`, `grilling`, `questions`, `resolve`, `handoff`, `trail`), so the orchestrator stays a dispatcher and the skill has no external skill dependencies.
- **Community feedback submission** — `/plan2code-4-finalize` gains STEP 6.5: after archival, assembles a METRICS_JSON payload from the completed run and submits it as a `community-feedback`-labeled GitHub issue on `jparkerweb/plan2code`, with a tiered fallback (`gh` CLI issue create → browser-opened prefilled issue → printed URL) for environments without `gh`. Payload schema and submission tiers live in the new `plan2code-4-finalize-references/community-feedback-submission.md`. Step 5 now asks for explicit submission consent and skips straight to Step 6 when declined.
- **Community submission ingestion in plan2code-metrics** — new "Fetch community submissions" CLI flow (`community.ts`) lists open feedback issues via `gh`, validates and parses each `METRICS_JSON` payload (type-only validation; malformed submissions are skipped and logged, not fixed up), imports them into the local run store deduped by `run_id`, re-aggregates, and closes each imported issue. Matches an open issue by the `community-feedback` label OR the `[Feedback]` title prefix OR the `METRICS_JSON` marker, so browser/print-tier submissions from outside contributors are still picked up; paginates fully (`--limit 1000`); and closes issues idempotently even on the duplicate path.
- **Community runs cohort by Plan2Code version** — ingested community runs are keyed into cohorts by their `plan2code_version` rather than by a prompt-file fingerprint, since community submissions carry the installed, platform-transformed prompts and an LLM-generated payload. Runs now carry a `source` (`local`/`community`) tag, and `current_cohort_key` prefers local cohorts so ingested feedback never displaces the maintainer's current prompt generation.
- **Devin CLI as an AI backend** for both `plan2code-metrics` (`invoke-llm.ts`) and `plan2code-loop` (`agents/devin-cli.ts`) — `devin --print --prompt-file <file> --permission-mode dangerous`. Unlike upstream, which replaced GitHub Copilot CLI with Devin, Plan2Code keeps **both**: Claude Code, GitHub Copilot CLI, and Devin CLI are all selectable. Existing Copilot CLI selections keep working.
### 🔧 Changed
- **README rebuilt** around a shorter, task-first structure, with the deep material split into a new `.readme/` folder: `walkthrough.md`, `autonomous-loop.md`, `status-line.md`, `metrics.md`, `test-bot.md`.
- **Docs site and README redesigned** around an "airmail" postcard theme — a fixed four-sided airmail-chevron page frame, sticky header, and the workflow presented as six posted letters, with Pathfinder and the optional Review step both surfaced. Adds a postage-stamp favicon set (`favicon.svg` / `.ico` / `.png` / `apple-touch-icon.png`) and a new README banner; removes three orphaned images (`desk.jpg`, `install-script.jpg`, `plan2code.jpg`).
- **`/plan2code-4-finalize` archives `pathfinder/` with the spec** — STEP 6 now names `pathfinder/` in the move list and no longer describes the cleanup target as "research or scratch files," wording that pointed an agent straight at `pathfinder/questions/`. The map is the rationale record behind the plan, in the same class as `PLAN-CONVERSATION-*.md`.
- **`/plan2code-1b-revise-plan` no longer deletes `pathfinder/`** — its Step 6 cleanup had the same "research or scratch files" wording.
- **`/plan2code-quick-task` is no longer labelled "Step 0"** — pathfinder now owns step 0, and quick-task was never a pipeline step. It registers as a utility (like `init`, `review`, and `handoff`), so its generated description reads `Plan2Code Quick Task: Quick Task Mode`. Filename, skill name, and command path are unchanged.
- **`/plan2code-handoff` asks where to save** — the OS temp directory is now the default, with `./handoffs/` or any other path available on request. Adds a spec-awareness section: when the session worked inside `specs/<feature>/`, the handoff cites the in-progress `phase-X.md` and its actual checkbox state rather than relying on conversation memory.
- **`/plan2code-init-update` Step 7 offloaded to a reference file** — the AI Agent File Sync detail moves to `plan2code-init-update-references/ai-agent-file-sync.md` with an inline fallback. The `CLAUDE.md` MANDATORY-FIRST-STEP template is unchanged.
- **`/plan2code-review` Session End offloaded to a reference file** — next-step routing moves to `plan2code-review-references/session-end.md` with an inline fallback.
- **Status line: context-bar token count suppressed on token-usage accounts** — the bar's `(84k)` reads the same `context_window.total_input_tokens` the `in`/`out` usage segment already shows on Enterprise/Bedrock/Vertex/PAYG accounts. It now renders only on Pro/Max/Teams (rate-limit) accounts, where no other segment carries an absolute token count. The `items.contextTokens` flag still turns it off entirely.
- **`aggregator.ts` refactor** — extracted `writeRunFile()` (dedup-by-`run_id` write) out of `importRun()` so the community ingestion path can reuse it without a source file path; `collector.ts` now exports `extractMetricsJson()` for the same reason.
- **`.agents-docs/AGENTS-code-style.md`** documents a metrics gotcha: when a `PLAN-DRAFT-*.md` carries no `METRICS_JSON` comment, `collector.ts` scrapes it by regex, and the four confidence-*breakdown* patterns match a bare dimension word plus a number **without** requiring a `%` — so even a table row like `| Requirements | 11 |` gets ingested as a planning confidence score.
- **`.agents-docs/AGENTS-architecture.md`** documents the column-0 requirement for `Read references/*.md` directives — `install.js` anchors its flat-file path-rewrite regex at `^`, so an indented `Read` line is silently skipped.
### 🐛 Fixed
- Broken review-command row and column alignment in `QUICK-REFERENCE.md`.
- `plan2code-loop` banner misspelled the mascot as "Plany".
## v1.16.1
### ✨ Added
- **Status line: git worktree awareness** — a session running in a linked git worktree now renders the project segment as `repo ⑂ worktree` (e.g. `plan2code ⑂ spike`) instead of only the worktree's directory name, which previously made the session look like an unrelated project
- Repo identity resolved from `git rev-parse --git-common-dir`, so it is correct regardless of how the worktree directory was named (bare `<name>.git` main repos included)
- A leading repo prefix is stripped from the worktree name (`plan2code-user-auth``user-auth`), and the name collapses to a bare `⑂` when it merely restates the branch already on screen — matched across `/ _ . -` separators and type prefixes like `feature/`, and only when the branch is actually displayed
- Toggleable via the new `items.worktree` config flag (default on); one extra timeout-bounded `git` call, skipped outside git repos
## v1.16.0
### ✨ Added
- **`/plan2code-handoff` skill** — compacts the current conversation into a self-contained handoff document (written to gitignored `./handoffs/<timestamp>-handoff.md`) so a fresh session or another agent can resume the work
- Always captures a confirmed **Next task**: infers a candidate from context and requires the user to confirm or fill it in before the file is written
- References plan specs, logs, and files by path rather than copying them; strips secrets; suggests follow-on skills and verification steps
- Repo-safe: checks `git check-ignore` and warns (without silently editing `.gitignore`) when `handoffs/` isn't ignored in an arbitrary repo
- **Repo-local release publisher skill** — new `/plan2code-publish` maintainer skill in `.claude/skills/` cuts a GitHub Release from the top `CHANGELOG.md` entry once `CHANGELOG.md`, `version.json`, and `package.json` agree and the version is ahead of the latest published release. Dev tooling only — deliberately excluded from `install.js`, never installed to `~/.claude/skills/`.
### 🐛 Fixed
- **Review workflow next-step suggestion made context-aware** — `/plan2code-review` Session End now reconciles three signals: session context (what preceded the review in the conversation), the user's review intent, and on-disk spec state gathered shell-agnostically — a file-search tool's empty result is never treated as proof that no specs exist. Suggestions render only at actual session end, cite their evidence and its source, and conflicting signals ask one targeted question instead of guessing.
## v1.15.4
### ✨ Added
+21 -11
View File
@@ -2,23 +2,29 @@
## Commands
| Step | Command | Input | Output |
| ------ | ------------------------------- | --------------- | ----------------------------------- |
| Init | /plan2code-init | None | AGENTS.md file |
| Update | /plan2code-init-update | AGENTS.md | Updated AGENTS.md |
| 0 | /plan2code-quick-task | Requirements | Conversational plan |
| review | /plan2code-review | Scope guidance | Review findings + fixes |
| 1 | /plan2code-1-plan | Requirements | PLAN-CONVERSATION-<date>.md + PLAN-DRAFT-<date>.md |
| 1b | /plan2code-1b-revise-plan | Specs + changes | Updated specs |
| 2 | /plan2code-2-document | PLAN-DRAFT.md | overview.md + Phase files |
| 3 | /plan2code-3-implement | overview.md | Implemented code |
| 4 | /plan2code-4-finalize | overview.md | Archived specs |
| Step | Command | Input | Output |
| ------- | ----------------------------- | --------------- | -------------------------------------------------- |
| Init | /plan2code-init | None | AGENTS.md file |
| Update | /plan2code-init-update | AGENTS.md | Updated AGENTS.md |
| 0 | /plan2code-0-pathfinder | A foggy idea | pathfinder/map.md *or* GitHub Issues + PLAN-DRAFT-<date>.md |
| quick | /plan2code-quick-task | Requirements | Conversational plan (standalone — not a pipeline step) |
| review | /plan2code-review | Scope guidance | Review findings + fixes |
| 1 | /plan2code-1-plan | Requirements | PLAN-CONVERSATION-<date>.md + PLAN-DRAFT-<date>.md |
| 1b | /plan2code-1b-revise-plan | Specs + changes | Updated specs |
| 2 | /plan2code-2-document | PLAN-DRAFT.md | overview.md + Phase files |
| 3 | /plan2code-3-implement | overview.md | Implemented code |
| 4 | /plan2code-4-finalize | overview.md | Archived specs |
| handoff | /plan2code-handoff | Conversation | Self-contained handoff doc in handoffs/ |
## File Structure
```
specs/
└── <feature-name>/
├── pathfinder/ # From Step 0 (optional, if charted locally)
│ ├── map.md # the map: destination, decisions, fog
│ └── questions/NN-<slug>.md # one decision question per file
│ # (GitHub Issues backend: map issue + sub-issues instead)
├── PLAN-DRAFT-<date>.md # From Step 1 (verified plan)
├── PLAN-CONVERSATION-<date>.md # From Step 1 (conversation log)
├── overview.md # From Step 2
@@ -76,6 +82,10 @@ New to a project?
Learned something during a session?
└── /plan2code-init-update → Add learnings to AGENTS.md
Too unclaer to plan? (big idea, don't yet know what the questions are)
└── /plan2code-0-pathfinder → chart it, clear one decision per session
└── then → /plan2code-1-plan (resumes at Phase 4)
Is it a quick, small task?
├── Yes → /plan2code-quick-task (standalone)
└── No → /plan2code-1-plan (full workflow)
+235 -489
View File
@@ -1,553 +1,299 @@
# Plan2Code: AI-Assisted Software Development Workflow
A structured 4-step workflow for developing features and projects with AI assistance. This methodology emphasizes thorough planning before implementation, ensuring well-documented, maintainable code.
# Plan2Code
<img src="docs/desk.jpg" alt="Plan2Code Workflow" height="275">
<img src="docs/banner.png" alt="Plan2Code — send the plan, the build follows" style="max-width:1024px;">
## Overview
**A spec-driven workflow for AI coding agents. Send the plan — the build follows.**
```
🤔 📝 ⚡ 🧹
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ Step 1 │ │ Step 2 │ │ Step 3 │ │ Step 4 │
│ PLAN │ --> │ DOCUMENT │ --> │ IMPLEMENT │ --> │ FINALIZE │
└─────────────────┘ └─────────────────┘ └─────────────────┘ └─────────────────┘
New Chat New Chat New Chat (per phase) New Chat
```
An AI agent is a fine builder and a terrible client. Plan2Code stops making it both: you approve a
plan, the plan becomes a set of phase documents in your repo, and the agent builds to those documents
one phase at a time. Progress lives in files instead of chat history — so the next session, the next
agent, and the next engineer all start from the same specs.
| Command | When to Use |
|----------------------------------|----------------------------------------------------------|
| `/plan2code-init` | Generate AGENTS.md file for new/existing projects |
| `/plan2code-init-update` | Update AGENTS.md with new learnings from coding sessions |
| `/plan2code-quick-task` | Small, quick tasks that don't need full workflow |
| `/plan2code-review` | Post-phase code review (after key features or milestones)|
| `/plan2code-1-plan` | Starting a new feature (full planning) |
| `/plan2code-1b-revise-plan` | Requirements change mid-implementation |
| `/plan2code-2-document` | After planning, create implementation specs |
| `/plan2code-3-implement` | Execute implementation (one phase per conversation) |
| `/plan2code-4-finalize` | All phases complete, ready to archive |
Six commands, each posted separately. Two of them are optional.
**Key Rules:**
- Start NEW conversation for each step (and each implementation phase)
- ONE phase per conversation
- Reply "approved" to complete phases
- 90% confidence required before planning completes
See [QUICK-REFERENCE.md](QUICK-REFERENCE.md) for full reference card.
Version 2.0.0 · MIT · 📖 [plan2code.jparkerweb.com](https://plan2code.jparkerweb.com)
---
## Installation
## Install
Plan2Code includes an interactive installer that generates and installs workflow files for all major AI coding assistants.
Requires [Node.js](https://nodejs.org/) 14 or later. Re-run any time to update.
<img src="docs/install-script.jpg" width="600">
### Prerequisites
The install script requires **Node.js** (v14 or later). If you don't have Node.js installed:
1. Download from [nodejs.org](https://nodejs.org/) (recommended)
2. Or use a package manager:
- **macOS:** `brew install node`
- **Windows:** `winget install OpenJS.NodeJS` or `choco install nodejs`
- **Linux:** `sudo apt install nodejs` (Debian/Ubuntu) or `sudo dnf install nodejs` (Fedora)
### Supported Platforms
- Claude Code
- Cursor
- Windsurf
- Continue
- Codeium (IntelliJ)
- GitHub Copilot CLI
- VS Code GitHub Copilot
- Gemini CLI
- Crush
- Pi (pi.dev)
- Amp
- OpenCode
- Devin
- Zed
### Install via npx (Recommended — No Clone Required)
Run the interactive installer directly using `npx` with your preferred GitHub authentication method:
**If you use SSH keys:**
```bash
npx git+ssh://git@github.com/jparkerweb/plan2code.git
npx --allow-git=all git+https://github.com/jparkerweb/plan2code.git
```
**If you use HTTPS authentication:**
```bash
npx git+https://github.com/jparkerweb/plan2code.git
This fetches the installer to a temp directory, runs it, writes the slash commands for whichever
tools you pick, and cleans up after itself. The installed commands work independently from then on.
Either route lands you on the same menu:
```
╔═════════════════════════════════════════════════════════╗
║ INSTALL PLAN2CODE ║
╠═════════════════════════════════════════════════════════╣
║ I. INSTALL Install Plan2Code for all platforms ║
║ A. ALL Install Plan2Code + dev tools ║
║ U. UNINSTALL Remove Plan2Code files ║
║ C. CUSTOM Advanced options ║
║ Q. QUIT Exit ║
╚═════════════════════════════════════════════════════════╝
```
This downloads the installer to a temporary location, runs it, installs the workflow files to your machine, and cleans up automatically. The installed workflows remain on your system and work independently. To update or reinstall, simply run the command again.
**Supported tools:** Claude Code · Cursor · Windsurf · Continue · Codeium (IntelliJ) ·
GitHub Copilot CLI · VS Code Copilot · Crush · Pi · Amp · OpenCode · Devin · Zed
### Standard Installation (Clone Method)
#### No Clone Required
Run the interactive installer directly using `npx` with your preferred GitHub authentication method:
**If you use SSH keys:**
```bash
npx git+ssh://git@github.com/jparkerweb/plan2code.git
```
**If you use HTTPS authentication:**
```bash
npx git+https://github.com/jparkerweb/plan2code.git
```
This downloads the installer to a temporary location, runs it, installs the workflow files to your machine, and cleans up automatically. The installed workflows remain on your system and work independently. To update or reinstall, simply run the command again.
#### Standard Installation (Clone Method)
<details>
<summary>Prefer to clone?</summary>
```bash
# Clone the repository
git clone https://github.com/jparkerweb/plan2code.git
cd plan2code
# Run the interactive installer
node install.js
# OPTIONAL: Install dev dependencies (ONLY if you plan to modify/contribute to Plan2Code)
npm install
```
The installer displays an interactive menu:
```
╔════════════════════════════════════════════════════════════════╗
║ INSTALL PLAN2CODE ║
╠════════════════════════════════════════════════════════════════╣
║ I. INSTALL Install Plan2Code for all platforms ║
║ U. UNINSTALL Remove Plan2Code files ║
║ C. CUSTOM Advanced options ║
║ Q. QUIT Exit ║
╚════════════════════════════════════════════════════════════════╝
SELECT OPTION (I, U, C, Q) [I]:
# Only if you plan to modify or contribute to Plan2Code itself
npm install && npx husky
```
</details>
---
## Status Line — Claude CLI (Optional)
## The workflow
A persistent three-line status bar for Claude Code showing model, project, git branch, uncommitted diff stats, context usage, and plan quota. Included in `node install.js``A` (Install All + dev tools), or available individually via Custom → Status Line. See [src/statusline-claude/README.md](src/statusline-claude/README.md) for details.
```
┌╴╴╴╴╴╴╴╴╴╴╴╴┐
╎0 PATHFINDER╎ optional · new in 2.0 · for an idea too big or unclear to plan
└╴╴╴╴╴╴┬╴╴╴╴╴┘
┌────────────┐ ┌────────────┐ ┌────────────┐ ┌╴╴╴╴╴╴╴╴╴╴╴╴┐ ┌────────────┐
│ 1 PLAN │─>│ 2 DOCUMENT │─>│3 IMPLEMENT │─>╎ REVIEW ╎─>│ 4 FINALIZE │
│decide what │ │ draw it as │ │build to the│ ╎ optional ╎ │verify, sum,│
│ to build │ │phase specs │ │ drawing │ ╎ any time ╎ │ archive │
└────────────┘ └────────────┘ └────────────┘ └╴╴╴╴╴╴╴╴╴╴╴╴┘ └────────────┘
new chat new chat new chat/phase new chat new chat
├◀─────────────── one feature, start to archive ──────────────────▶┤
```
Every box is its own conversation. That is not a style preference — planning context leaking into
implementation is where most agent drift starts.
| Command | Use it when |
|---------|-------------|
| `/plan2code-0-pathfinder` | The idea is too big and unclear to plan. Charts it as decisions, clears one per session, hands a hot plan draft to Step 1 |
| `/plan2code-1-plan` | Starting a feature. Full requirements → architecture pass |
| `/plan2code-2-document` | Planning is done. Turn the plan into phase specs |
| `/plan2code-3-implement` | Build the next phase (one per conversation) |
| `/plan2code-review` | Independent second opinion on local changes, then optional fixes |
| `/plan2code-4-finalize` | All phases done. Validate, summarize, archive |
| `/plan2code-init` | Generate this repo's `AGENTS.md` so every agent starts informed |
| `/plan2code-init-update` | Fold what you learned this session back into `AGENTS.md` |
| `/plan2code-quick-task` | A small change that doesn't warrant the full sequence |
| `/plan2code-1b-revise-plan` | Requirements moved mid-build. Revise the specs, not the code |
| `/plan2code-handoff` | Compact this conversation into a doc the next one resumes from |
---
## Important: Start Fresh Conversations
## The four rules that do most of the work
**Start a new conversation/chat session before each step.** This includes:
**1 · A fresh conversation for each step, and each implementation phase.**
Step 3 gets a new chat per phase, not one chat for all of them.
- Step 1: New conversation
- Step 2: New conversation
- Step 3: New conversation **for each phase** (Phase 1, Phase 2, etc.)
- Step 4: New conversation
**2 · No code until the plan hits 90% confidence.**
Step 1 will not finalize below the threshold. Under it, the agent keeps asking and keeps reading your
code — and writes every assumption down where you can argue with it.
Fresh conversations prevent context pollution and ensure the AI focuses on the current task with the relevant specifications.
**3 · Checkboxes are the state, not the chat.**
Progress lives in the spec files. Any agent, any session, resumes cold from them.
**4 · Reply `approved` to close a phase.**
Nothing advances on a guess about what you meant.
---
## The Workflow Steps
### Step 1: Planning Mode 🤔
**Purpose:** Thoroughly analyze requirements and design the solution architecture before writing any code.
**AI Role:** Senior software architect and technical product manager
**Phases (completed one at a time):**
1. **Requirements Analysis** - Extract functional/non-functional requirements, identify ambiguities
2. **System Context Examination** - Review existing codebase, identify integration points
3. **Tech Stack** - Recommend and confirm all technologies (requires user sign-off)
4. **Architecture Design** - Propose patterns, define components, design interfaces/schemas
5. **Technical Specification** - Break down implementation phases, identify risks
6. **Transition Decision** - Finalize plan when confidence reaches 90%+
**Output:** `specs/<feature-name>/PLAN-DRAFT-<date>.md` and `specs/<feature-name>/PLAN-CONVERSATION-<date>.md` (date format: YYYYMMDD)
**Key Behaviors:**
- AI stops after each phase for clarification
- Must reach 90% confidence before finalizing
- All assumptions are documented
- User must approve tech stack decisions
---
### Step 2: Documentation Mode 📝
**Purpose:** Transform the planning output into structured, actionable implementation documents.
**Required Context:** Attach or reference the `specs/<feature-name>/PLAN-DRAFT-<date>.md` from Step 1 (or provide the planning conversation).
**Output Structure:**
```
specs/
└── <feature-name>/
├── overview.md # High-level overview with phase checkboxes and parallel groups
├── Phase 1.md # Detailed tasks for Phase 1
├── Phase 2.md # Detailed tasks for Phase 2
└── Phase N.md # ...additional phases
```
The `overview.md` includes a "Parallel Execution Groups" section that identifies which phases can be run simultaneously in separate agent instances.
**Document Format:**
- Each phase file contains detailed one-story-point tasks
- All tasks have checkboxes `[ ]` for progress tracking
- Each phase is self-contained (developer needs no prior context)
- Unit/E2E testing excluded unless explicitly requested
---
### Step 3: Implementation Mode ⚡
**Purpose:** Execute the implementation following the documented specifications.
**AI Role:** Senior software engineer
**Required Context:** Provide the path to `specs/<feature-name>/overview.md`. The command will auto-detect the next uncompleted phase and read the corresponding `Phase X.md` file automatically.
**Workflow:**
1. Identify the next uncompleted phase (unchecked in `overview.md`)
2. Check for parallel execution options (if phases can run simultaneously)
3. Implement ALL tasks in that phase exactly as specified
4. Update `Phase X.md` checkboxes as tasks complete `[x]`
5. Update `overview.md` phase checkbox when phase completes
6. Perform code review to ensure nothing was missed
7. Add completion summary to the phase document
**Parallel Execution:** If the next phase is part of a parallel-eligible group, you'll be prompted to choose which phase to implement. This allows running multiple agent instances simultaneously on different phases that don't conflict with each other.
**Key Rules:**
- **Start a new conversation for EACH phase**
- Work on ONE phase per conversation (unless told otherwise)
- Follow specifications EXACTLY as documented
- Keep checkboxes updated (enables progress tracking across sessions)
- Do NOT run tests unless specified in phase tasks
---
### Review Mode 🔬 (Optional)
**Purpose:** Comprehensive post-implementation code review with adaptive scope and spec compliance checking.
**AI Role:** Critical review specialist -- independent second opinion
**When to use:** After completing implementation phases, especially key features or milestones. Can be run after any phase, not just before finalization.
**What it does:**
1. Determines review scope — detects conversation context, user-specified scope, or gathers git changes as fallback
2. Understands context and determines review strategy
3. Analyzes across 11 dimensions (correctness, security, performance, spec compliance, etc.)
4. Generates findings ranked by severity (Critical, Warning, Suggestion)
5. Offers to fix issues, then provides commit guidance
**Key Behaviors:**
- Uses companion reference files for deep verification, detailed dimension checklists, and false-positive detection (loaded automatically during review)
- Reviews all changed files across 11 dimensions; prioritizes by risk when batches exceed 50 files
- Spec-aware when `specs/` exists; works standalone without specs
- Review phase is read-only; fixes only on user request with verification
---
### Step 4: Finalization Mode 🧹
**Purpose:** Validate implementation, create summaries, and archive documentation.
**Required Context:** Attach or reference the `specs/<feature-name>/` directory contents.
**Steps:**
1. **Validation** - Verify all tasks implemented correctly, check for issues
2. **Summary** - Document what was built and list all modified/created files
3. **Documentation Review** - Identify any needed README/CHANGELOG updates
4. **Spec Cleanup** - Move completed specs to `specs--completed/<implementation-name>/`
5. **Final Confirmation** - Confirm completion
---
## How to Use
After running `node install.js`, use the slash commands directly in your AI tool:
```
/plan2code-1-plan # Start planning a new feature
/plan2code-2-document # Create implementation docs from plan
/plan2code-3-implement # Begin/continue implementation
/plan2code-review # Post-phase code review (optional)
/plan2code-4-finalize # Wrap up after all phases complete
```
---
## Complete Workflow Example
### Starting a New Project
**Session 1 - Planning (New Chat):**
```
User: [Paste or invoke Step 1 prompt]
I want to build a REST API for a task management application.
AI: 🤔 [REQUIREMENTS ANALYSIS]
... asks clarifying questions, works through phases ...
AI: 🤔 [TRANSITION DECISION]
Confidence: 92%. Creating specs/task-api/PLAN-DRAFT-20250204.md...
```
**Session 2 - Documentation (New Chat):**
```
User: [Paste or invoke Step 2 prompt]
[Attach: specs/task-api/PLAN-DRAFT-20250204.md]
AI: 📝 [DOCUMENTATION]
Creating specs/task-api/overview.md...
Creating specs/task-api/Phase 1.md...
Creating specs/task-api/Phase 2.md...
...
```
**Session 3 - Implementation Phase 1 (New Chat):**
```
User: [Paste or invoke Step 3 prompt]
[Provide: specs/task-api/overview.md]
AI: ⚡ [PHASE 1: Project Setup]
(Auto-detected Phase 1 as next uncompleted phase)
Implementing tasks...
✓ Phase 1 complete. Updated checkboxes in Phase 1.md and overview.md.
```
**Session 4 - Implementation Phase 2 (New Chat):**
```
User: [Paste or invoke Step 3 prompt]
[Provide: specs/task-api/overview.md]
AI: ⚡ [PHASE 2: Database Models]
(Auto-detected Phase 2 as next uncompleted phase)
Implementing tasks...
✓ Phase 2 complete. Updated checkboxes in Phase 2.md and overview.md.
```
**Sessions 5-N - Continue Implementation (New Chat for each phase):**
```
... repeat for each remaining phase ...
```
**Final Session - Finalization (New Chat):**
```
User: [Paste or invoke Step 4 prompt]
[Provide: specs/task-api/overview.md]
AI: 🧹 [VALIDATION]
Verifying implementation...
AI: 🧹 [SPEC CLEANUP]
Moving to specs--completed/task-api/
Implementation complete!
```
---
## Progress Tracking
The checkbox system enables seamless progress tracking across multiple sessions:
**Phase Status (overview.md):**
| Checkbox | Status | Meaning |
|----------|--------|---------|
| `[ ]` | Pending | Not yet started |
| `[/]` | In Progress | Agent actively working (or paused/aborted) |
| `[x]` | Complete | Finished and approved |
```markdown
## Phases
- [x] Phase 1: Project Setup
- [x] Phase 2: Database Models
- [/] Phase 3: API Endpoints <- In progress (agent working)
- [ ] Phase 4: Authentication <- Next available
```
**Task Status (phase-X.md):**
```markdown
## Tasks
- [x] Create routes file
- [x] Implement GET /tasks
- [ ] Implement POST /tasks <- Current task
- [ ] Implement PUT /tasks/:id
- [ ] Implement DELETE /tasks/:id
```
The `[/]` status enables parallel execution - multiple agents can work on different phases simultaneously, and you can see which phases are actively being worked on.
---
## Best Practices
1. **Start fresh conversations** - New chat for each step and each implementation phase
2. **Always attach specs** - The AI needs the spec files to understand the current state
3. **Don't skip planning** - The upfront investment prevents costly rework later
4. **Confirm tech stack** - Ensure AI gets explicit approval before architecture design
5. **One phase at a time** - Keeps conversations focused and manageable
6. **Update checkboxes immediately** - Maintains accurate progress state
7. **Review phase output** - Verify each phase before moving to the next
8. **Keep spec files** - The completed folder serves as project documentation
---
## What to Attach at Each Step
| Step | Required Input |
| ------------------ | ---------------------------------------------------- |
| Step 1 (Plan) | None (describe your feature/project) |
| Step 2 (Document) | `specs/<feature>/PLAN-DRAFT-<date>.md` or planning conversation |
| Step 3 (Implement) | `specs/<feature>/overview.md` (auto-detects phase) |
| Review (Optional) | Scope guidance (e.g., "review last 2 phases", "just the auth module", "whole PR"). Auto-detects changes and specs if no scope given. |
| Step 4 (Finalize) | `specs/<feature>/overview.md` |
---
## Autonomous Loop (Alternative to Step 3)
For hands-off implementation, Plan2Code includes an optional autonomous loop CLI that iterates through your spec tasks automatically.
> **Note:** The loop is an **alternative** to `/plan2code-3-implement`, not a replacement. Use the manual Step 3 workflow when you want direct control over each phase, or use the loop when you prefer autonomous execution.
### When to Use Each
| Approach | Best For |
|----------|----------|
| `/plan2code-3-implement` | Interactive control, reviewing each phase, complex logic requiring human judgment |
| `plan2code-loop` | Straightforward implementations, batch processing, overnight runs |
### Installing the Loop
```bash
# From the plan2code root directory:
# Option 1: Install everything (recommended)
node install.js # Select I at the menu
# Option 2: Install loop only
node install.js # Select C, then O at the menu
```
### Using the Loop
```bash
# Run the loop - fully interactive
plan2code-loop
```
The CLI will:
1. Auto-detect specs in `./specs/` directory
2. Let you select a spec if multiple are found
3. Prompt to continue if an existing session is found
4. Ask for JIRA ticket ID, agent selection, loop mode, and max iterations
### Loop Modes
| Mode | Behavior | Git Commits | Best For |
|------|----------|-------------|----------|
| **One task per loop** (default) | Each agent call implements one task | Node controller commits after each task | Smaller models, cautious execution |
| **One phase per loop** | Each agent call implements all tasks in a phase | LLM commits after each task (with JIRA ID) | Smart models with larger context windows, related tasks |
Session state is stored per-spec in `specs/<feature>/.plan2code-loop/`, keeping each feature's progress isolated.
The loop will:
1. Read your `overview.md` and phase files
2. Find the first unchecked task (or phase, in phase mode)
3. Implement it and mark the checkbox complete
4. Repeat until all tasks are done or max iterations reached
See [plan2code-loop/](plan2code-loop/) for full documentation.
---
## File Structure After Complete Implementation
## What lands in your repo
```
your-project/
├── specs/
│ └── another-feature/ # In-progress feature
│ ├── overview.md
└── Phase 1.md
│ └── task-api/ ← in progress
│ ├── pathfinder/ ← only if you charted it in Step 0
│ ├── map.md the destination, the decisions, the fog
│ │ └── questions/NN-<slug>.md one decision per file
│ ├── PLAN-DRAFT-20260804.md ← Step 1: the verified plan
│ ├── PLAN-CONVERSATION-*.md ← Step 1: how you got there
│ ├── overview.md ← Step 2: phase list + parallel groups
│ └── Phase 1.md … Phase N.md ← Step 2: one-point tasks, self-contained
├── specs--completed/
│ └── feature-name/
│ ├── overview.md # Archived with completion summary
│ ├── Phase 1.md # All checkboxes marked [x]
│ ├── Phase 2.md
│ └── ...
├── your project files...
└── README.md
│ └── auth-refresh/ ← Step 4 files finished work here
└── ...your code
```
`specs/` is gitignored by default — it's your working drawing, not a deliverable. Share a folder
deliberately with `git add -f` when you want to.
### Progress marks
| Mark | Status | Meaning |
|------|--------|---------|
| `[ ]` | Open | Unclaimed. Any agent picks it up cold. |
| `[/]` | In progress | Claimed right now — which is how two agents run parallel phases without colliding. |
| `[x]` | Done | Built, self-reviewed against the spec, approved by you. |
```markdown
## Phases
- [x] Phase 1: Project setup
- [x] Phase 2: Data model
- [/] Phase 3: API endpoints ← an agent is on this now
- [ ] Phase 4: Authentication ← next available
```
Step 2 marks which phases don't share files. Open a second agent on one of those, and the `[/]` marks
keep the two out of each other's way.
---
## Customization
## The six steps in detail
Feel free to modify these prompts to fit your workflow:
Each one travels on its own — a fresh conversation, opened and closed, with the specs on disk as the
only thing carried between them.
- **Add testing phases** - Uncomment/add testing requirements in Step 2
- **Adjust confidence threshold** - Change the 90% threshold in Step 1
- **Modify output structure** - Customize the specs folder organization
- **Add code review steps** - Enhance Step 3 with additional review gates
### 0 · Pathfinder 🧭 — optional, new in 2.0
Some ideas are too big and unclear to plan: you can feel the shape of the work but you can't write
it as requirements, so planning would just invent the answers. Pathfinder finds the *way* to the
destination; Step 1 then walks it.
1. **Name the destination** — one or two lines fixing what this effort is finding its way to. Settled
first, because it fixes scope. It also asks where the map should live: **local files** under
gitignored `specs/` (private, solo — the default), or **GitHub Issues** (a map issue with one
sub-issue per decision, native blocking, so your team can see and work the frontier in the tracker).
2. **Chart the map** — a breadth-first grilling surfaces the open decisions. Anything you can phrase
*sharply* becomes a question file; anything you can only sense stays listed as fog.
3. **Clear one question per session** — resolving a question burns off the fog behind it, graduating
whatever just became sharp into new questions.
4. **Hand off** — when nothing is left to decide, it writes a `PLAN-DRAFT` that
`/plan2code-1-plan` resumes from at Phase 4, with requirements, context, and scope already
answered.
**Question types:** `grill` (a decision only you can make — the default) · `research` (a fact gates
it; background agents resolve these, several in parallel) · `sketch` (you need something concrete to
react to) · `legwork` (manual work that has to happen before a decision is possible).
It never answers its own questions, and it **plans, it never builds.** When the urge to just build it
arrives, the map is done. Skip Step 0 entirely when you already know what you're building.
**Out:** `specs/<feature>/pathfinder/map.md` + `questions/` (or a `pathfinder:map` issue and its
sub-issues) → `PLAN-DRAFT-<date>.md`. The draft is always a local file — that is what Step 1 reads.
### 1 · Plan 🤔
The agent works as a senior architect through six phases, stopping for you after each: requirements
analysis · system context (reading your actual codebase) · tech stack (needs your explicit sign-off) ·
architecture design · technical specification · transition decision.
It won't finalize below **90% confidence**, and every assumption it makes is written into the draft.
**In:** a description of the feature. **Out:** `PLAN-DRAFT-<date>.md` + `PLAN-CONVERSATION-<date>.md`
### 2 · Document 📝
The plan becomes the drawing. One `overview.md` with the phase checklist, plus one file per phase of
one-story-point tasks. Each phase is **self-contained** — an agent opening `Phase 3.md` cold needs
nothing else to build it. Unit and E2E tests are excluded unless you ask for them.
The overview also identifies the **parallel execution groups**: phases with no shared files or
dependencies, safe to run in separate agents at once.
**In:** the `PLAN-DRAFT`. **Out:** `overview.md` + `Phase 1…N.md`
### 3 · Implement ⚡
Point it at `overview.md` and it does the rest: finds the next unchecked phase, implements every task
exactly as specified, ticks tasks off as they land, then reviews its own work against the spec and
writes a completion summary.
One phase per conversation. It won't run tests unless the phase says to.
**In:** `specs/<feature>/overview.md`. **Out:** working code, and updated checkboxes.
### Review 🔬 — optional, any time
An independent second opinion, not a rubber stamp. It figures out its own scope (conversation
context, your instruction, or the git diff as a fallback), analyses across 11 dimensions, and ranks
findings Critical / Warning / Suggestion. Every finding cites a file and a line, or it gets dropped —
and the review pass is read-only. It fixes things only if you ask, and verifies each fix afterwards.
Spec-aware when `specs/` exists, and works fine without it. Most useful right after a planning or
implementation step, but there's no wrong time to run it.
### 4 · Finalize 🧹
Validates every task against its phase spec, writes the summary and the list of files touched, flags
the docs that drifted (`README`, `CHANGELOG`, `AGENTS.md`), then archives the whole spec folder —
`pathfinder/` included — to `specs--completed/`. That folder is the record of *why* the code looks
like this.
**In:** `specs/<feature>/overview.md`. **Out:** archived specs.
---
## What to bring to each step
| Step | Required input |
|------|----------------|
| 0 · Pathfinder | Nothing to start — just describe the idea. To continue: the feature name; it finds its own map |
| 1 · Plan | Nothing — describe the feature |
| 2 · Document | `specs/<feature>/PLAN-DRAFT-<date>.md`, or the planning conversation |
| 3 · Implement | `specs/<feature>/overview.md` — it detects the phase itself |
| Review | Scope guidance, e.g. "the last two phases", "just the auth module", "the whole PR". Auto-detects if you give none |
| 4 · Finalize | `specs/<feature>/overview.md` |
---
## Troubleshooting
**Slash commands/workflows not recognized:**
**Slash commands aren't recognised.** Re-run `node install.js` for that platform and restart your AI
tool. For a per-project install, check the directory isn't gitignored.
- Ensure you ran `node install.js` and selected the appropriate platform
- Restart your AI tool after installation
- For per-project installation, ensure the directory isn't in `.gitignore`
**The agent starts coding during planning.** The prompts forbid it, but models drift. Say: "Stay in
planning mode. Do not write code yet."
**AI jumps ahead to implementation during planning:**
**The agent doesn't know what to implement.** Give it the path to `overview.md` — it reads the phase
file itself from there.
- The prompts explicitly forbid this, but if it happens, remind the AI: "Stay in planning mode. Do not write code yet."
**You lost track between sessions.** `overview.md` has the phase status; the phase files have the
task status. That's the whole state.
**AI doesn't know what to implement:**
**The agent isn't following the spec.** Point at the specific phase document and tell it to re-read
the requirements.
- Make sure you provided the path to `overview.md`
- The AI will auto-detect the next phase and read the corresponding `Phase X.md` file
**Too many or too few phases.** Fix it in Step 2 — a phase should be a logical grouping of work, not
a fixed size.
**Lost progress between sessions:**
---
- Check `overview.md` for phase status
- Review individual phase files for task completion status
## Customizing
**AI not following spec exactly:**
The prompts are yours to edit. Common changes: add testing requirements in Step 2, move the 90%
confidence threshold in Step 1, restructure the `specs/` layout, or add review gates to Step 3.
Source files live in `src/`; re-run `node install.js` to push your edits out to every platform.
- Reference the specific phase document and ask it to re-read the requirements
---
**Too many/few phases:**
## Dive deeper
- Adjust during Step 2 (Documentation) - phases should represent logical groupings of work
Core reference:
- **[QUICK-REFERENCE.md](QUICK-REFERENCE.md)** — the one-page card: commands, inputs, outputs, decision tree
- **[.readme/walkthrough.md](.readme/walkthrough.md)** — one feature from a sentence to archived specs, session by session
- **[AGENTS.md](AGENTS.md)** — architecture and contributor guide for this repo
- **[CHANGELOG.md](CHANGELOG.md)** — what changed, and why
Optional tooling — none of it is required to use the workflow:
- **[.readme/autonomous-loop.md](.readme/autonomous-loop.md)** — `plan2code-loop`, a hands-off alternative to Step 3
- **[.readme/status-line.md](.readme/status-line.md)** — three-line Claude Code status bar: model, context, quota, diff
- **[.readme/metrics.md](.readme/metrics.md)** — `plan2code-metrics`, measuring and improving the prompts themselves
- **[.readme/test-bot.md](.readme/test-bot.md)** — `plan2code-bot`, maintainer harness that runs the whole workflow unattended
Binary file not shown.

After

Width:  |  Height:  |  Size: 6.1 KiB

BIN
View File
Binary file not shown.

After

Width:  |  Height:  |  Size: 127 KiB

BIN
View File
Binary file not shown.

Before

Width:  |  Height:  |  Size: 185 KiB

BIN
View File
Binary file not shown.

Before

Width:  |  Height:  |  Size: 24 KiB

After

Width:  |  Height:  |  Size: 1.9 KiB

BIN
View File
Binary file not shown.

Before

Width:  |  Height:  |  Size: 2.4 KiB

After

Width:  |  Height:  |  Size: 584 B

+33
View File
@@ -0,0 +1,33 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 64 64" width="64" height="64" role="img" aria-label="Plan2Code postage stamp">
<defs>
<!-- Perforated stamp silhouette: white keeps, black bites out.
Few, deep notches so the scalloped edge still reads at 16px. -->
<mask id="perf">
<rect width="64" height="64" fill="#000"/>
<rect x="2" y="2" width="60" height="60" fill="#fff"/>
<g fill="#000">
<circle cx="2" cy="2" r="5"/><circle cx="14" cy="2" r="5"/><circle cx="26" cy="2" r="5"/>
<circle cx="38" cy="2" r="5"/><circle cx="50" cy="2" r="5"/><circle cx="62" cy="2" r="5"/>
<circle cx="2" cy="62" r="5"/><circle cx="14" cy="62" r="5"/><circle cx="26" cy="62" r="5"/>
<circle cx="38" cy="62" r="5"/><circle cx="50" cy="62" r="5"/><circle cx="62" cy="62" r="5"/>
<circle cx="2" cy="14" r="5"/><circle cx="2" cy="26" r="5"/>
<circle cx="2" cy="38" r="5"/><circle cx="2" cy="50" r="5"/>
<circle cx="62" cy="14" r="5"/><circle cx="62" cy="26" r="5"/>
<circle cx="62" cy="38" r="5"/><circle cx="62" cy="50" r="5"/>
</g>
</mask>
</defs>
<g mask="url(#perf)">
<!-- printed stamp -->
<rect x="2" y="2" width="60" height="60" fill="#D33A38"/>
<!-- the 2, set as a franking-machine numeral -->
<g fill="#FBFAF6">
<rect x="16" y="13" width="32" height="8"/>
<rect x="40" y="13" width="8" height="22"/>
<rect x="16" y="27" width="32" height="8"/>
<rect x="16" y="27" width="8" height="24"/>
<rect x="16" y="43" width="32" height="8"/>
</g>
</g>
</svg>

After

Width:  |  Height:  |  Size: 1.6 KiB

+365 -1203
View File
File diff suppressed because it is too large Load Diff
Binary file not shown.

Before

Width:  |  Height:  |  Size: 147 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 95 KiB

+42 -98
View File
@@ -131,12 +131,20 @@ const SOURCE_PROMPTS = [
description: 'Update existing AGENTS.md with new learnings',
isUtility: true
},
{
source: 'plan2code-0-pathfinder.md',
stepNumber: '0',
name: 'pathfinder',
displayName: 'Pathfinder Mode',
description: 'charting of a foggy idea as a map of decision questions, cleared one at a time'
},
{
source: 'plan2code-quick-task.md',
stepNumber: 0,
stepNumber: 'quick',
name: 'quick-task',
displayName: 'Quick Task Mode',
description: 'Lightweight planning for small tasks'
description: 'Lightweight planning for small tasks',
isUtility: true
},
{
source: 'plan2code-1-plan.md',
@@ -180,6 +188,14 @@ const SOURCE_PROMPTS = [
name: 'finalize',
displayName: 'Finalization Mode',
description: 'Validate, summarize, and archive completed work'
},
{
source: 'plan2code-handoff.md',
stepNumber: 'handoff',
name: 'handoff',
displayName: 'Handoff Mode',
description: 'Compact the conversation into a self-contained handoff document for a fresh session',
isUtility: true
}
];
@@ -199,7 +215,10 @@ function generateSkillName(prompt) {
// Helper function to generate step label for descriptions
function generateStepLabel(prompt) {
if (prompt.stepNumber === 'init') return 'Init';
if (prompt.stepNumber === 'update') return 'Update';
if (prompt.stepNumber === 'review') return 'Review';
if (prompt.stepNumber === 'handoff') return 'Handoff';
if (prompt.stepNumber === 'quick') return 'Quick Task';
return `Step ${prompt.stepNumber}`;
}
@@ -215,12 +234,6 @@ function generateSkillHeader(prompt, disableModelInvocation = false) {
return lines.join('\n');
}
// Helper function to generate complete .toml file content for Gemini CLI
function generateTomlContent(prompt, sourceContent) {
const stepLabel = generateStepLabel(prompt);
const desc = `Plan2Code ${stepLabel}: ${prompt.displayName} - user-initiated ${prompt.description || 'workflow step'}`;
return `description = "${desc}"\nprompt = '''\n${sourceContent}\n'''\n`;
}
// Destination configurations for project-level installation (local)
const LOCAL_DESTINATIONS = [
@@ -325,16 +338,10 @@ const LOCAL_DESTINATIONS = [
header: (prompt) => generateSkillHeader(prompt, true),
},
{
name: 'Agent Skills (Amp · Devin · Gemini CLI · OpenCode)',
name: 'Agent Skills (Amp · Devin · OpenCode)',
dir: '.agents/skills',
type: 'skill',
header: (prompt) => generateSkillHeader(prompt, false),
},
{
name: 'Gemini CLI',
dir: '.gemini/commands',
type: 'toml',
filePattern: (prompt) => `${generateFilename(prompt, '')}.toml`,
}
];
@@ -421,7 +428,7 @@ const GLOBAL_DESTINATIONS = [
header: (prompt) => generateSkillHeader(prompt, true),
},
{
name: 'Agent Skills (Amp · Devin · Gemini CLI · OpenCode · Zed)',
name: 'Agent Skills (Amp · Devin · OpenCode · Zed)',
dir: '.agents/skills',
type: 'skill',
header: (prompt) => generateSkillHeader(prompt, true),
@@ -432,12 +439,6 @@ const GLOBAL_DESTINATIONS = [
type: 'skill',
header: (prompt) => generateSkillHeader(prompt, false),
},
{
name: 'Gemini CLI',
dir: '.gemini/commands',
type: 'toml',
filePattern: (prompt) => `${generateFilename(prompt, '')}.toml`,
}
];
// Helper function to get VS Code Copilot prompts directory based on platform
@@ -566,7 +567,7 @@ const INSTALL_TARGETS = [
icon: '◉ '
},
{
name: 'Agent Skills (Amp · Devin · Gemini CLI · OpenCode · Zed)',
name: 'Agent Skills (Amp · Devin · OpenCode · Zed)',
id: 'agent-skills',
dir: getAgentSkillsDir,
sourceDir: '.agents/skills',
@@ -587,12 +588,17 @@ const INSTALL_TARGETS = [
: '~/.config/crush/skills',
icon: '◉ '
},
// Gemini CLI is no longer an install target. This entry is RETAINED FOR UNINSTALL
// ONLY so `.gemini/commands/plan2code-*.toml` files written by earlier versions can
// still be cleaned up. Do not add a matching entry back to LOCAL_DESTINATIONS /
// GLOBAL_DESTINATIONS.
{
name: 'Gemini CLI',
name: 'Gemini CLI (legacy — uninstall only)',
id: 'gemini',
dir: '.gemini/commands',
type: 'toml',
filePattern: /^plan2code-.*\.toml$/,
uninstallOnly: true,
icon: '◉ '
},
{
@@ -636,7 +642,7 @@ function displayHeader() {
console.log('║ ╰───╯ Welcome to Plan2Code! ║');
console.log('║ ║');
console.log('║ G L O B A L I N S T A L L A T I O N S Y S T E M ║');
console.log('║ https://jparkerweb.github.io/plan2code ║');
console.log('║ https://github.com/jparkerweb/plan2code ║');
console.log('║ ║');
console.log('╚═════════════════════════════════════════════════════════╝');
console.log(COLORS.RESET);
@@ -837,25 +843,6 @@ function copyDirRecursive(src, dest, stats) {
}
}
/**
* Inline reference file content into orchestrator content.
* Replaces each `Read references/X.md` directive with the file's actual content,
* wrapped in markers. Used for TOML targets (Gemini CLI), which cannot resolve
* relative file reads at runtime.
*/
function inlineReferenceContent(sourceContent, srcRefDir) {
if (!fs.existsSync(srcRefDir)) return sourceContent;
return sourceContent.replace(
/^Read\s+references\/([^\s]+\.md)(.*)$/gm,
(match, filename, suffix) => {
const refPath = path.join(srcRefDir, filename);
if (!fs.existsSync(refPath)) return match;
const refContent = fs.readFileSync(refPath, 'utf8');
return `<!-- BEGIN inlined references/${filename}${suffix} -->\n${refContent}\n<!-- END inlined references/${filename} -->`;
}
);
}
/**
* Copy a reference directory (e.g., plan2code-review-references/) to a destination.
* Used by syncPrompts() to distribute reference files alongside orchestrator files.
@@ -879,48 +866,6 @@ function copyReferenceDirectory(srcRefDir, destRefDir, stats, quiet = false) {
}
}
/**
* Write a TOML file to a destination for Gemini CLI
*/
function writeTomlToDestination(rootDir, baseDir, dest, prompt, sourceContent, stats, quiet = false) {
const filename = dest.filePattern(prompt);
const destDir = path.join(rootDir, baseDir, dest.dir);
const filePath = path.join(destDir, filename);
// Build TOML content
const tomlContent = generateTomlContent(prompt, sourceContent);
// Ensure destination directory exists
fs.mkdirSync(destDir, { recursive: true });
// Check if file needs updating
let needsUpdate = true;
if (fs.existsSync(filePath)) {
try {
const existingContent = fs.readFileSync(filePath, 'utf8');
needsUpdate = existingContent !== tomlContent;
} catch (err) {
// File exists but can't be read, will try to write
}
}
if (needsUpdate) {
try {
fs.writeFileSync(filePath, tomlContent, 'utf8');
if (!quiet) {
console.log(` ${COLORS.GREEN}▰▰▰${COLORS.RESET} ${filename}`);
}
} catch (err) {
console.error(` ${COLORS.RED}✖✖✖${COLORS.RESET} ${filename}: ${err.message}`);
stats.errors++;
return;
}
stats.updated++;
} else {
stats.skipped++;
}
}
/**
* Sync prompts from src/ to dist/ directories
* This generates the distribution files before installation
@@ -981,13 +926,6 @@ function syncPrompts(quiet = false) {
const destRefDir = path.join(projectRoot, base, dest.dir, skillName, 'references');
copyReferenceDirectory(srcRefDir, destRefDir, stats, quiet);
}
} else if (dest.type === 'toml') {
// TOML targets (Gemini CLI) cannot resolve relative file reads at runtime,
// so inline reference content directly into the prompt body.
const contentForToml = hasRefDir
? inlineReferenceContent(sourceContent, srcRefDir)
: sourceContent;
writeTomlToDestination(projectRoot, base, dest, prompt, contentForToml, stats, quiet);
} else {
// For flat-file destinations, rewrite Read directive paths to sibling directory name
const contentForDest = hasRefDir
@@ -1022,11 +960,16 @@ function syncPrompts(quiet = false) {
/**
* Get targets to process based on selected platforms
*/
function getTargets(platformIds = null) {
function getTargets(platformIds = null, { includeLegacy = false } = {}) {
// Entries flagged `uninstallOnly` are platforms we no longer install to, kept so their
// files from earlier versions can still be cleaned up. Only uninstall may see them.
const pool = includeLegacy
? INSTALL_TARGETS
: INSTALL_TARGETS.filter(t => !t.uninstallOnly);
if (platformIds && platformIds.length > 0) {
return INSTALL_TARGETS.filter(t => platformIds.includes(t.id));
return pool.filter(t => platformIds.includes(t.id));
}
return INSTALL_TARGETS;
return pool;
}
// ============================================================================
@@ -1307,7 +1250,8 @@ async function install(targets = null) {
*/
function uninstallFiles(targets = null) {
const homeDir = os.homedir();
const targetList = targets || getTargets();
// includeLegacy: uninstall must also clear platforms we no longer install to
const targetList = targets || getTargets(null, { includeLegacy: true });
let totalRemoved = 0;
let totalErrors = 0;
@@ -1520,7 +1464,7 @@ function runInteractive() {
const question = (prompt) => new Promise(resolve => rl.question(prompt, resolve));
async function main() {
// Task 2.1: Main menu display
// Main menu display
displayHeader();
console.log(`${COLORS.BLUE}${COLORS.BRIGHT}Version Info:${COLORS.RESET} ${projectVersion.name} ${projectVersion.version}`);
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "plan2code",
"version": "1.15.4",
"version": "2.1.0",
"private": true,
"bin": {
"plan2code": "./install.js"
+1 -1
View File
@@ -48,7 +48,7 @@ Examples:
plan2code-bot --resume
Documentation:
https://jparkerweb.github.io/plan2code
https://github.com/jparkerweb/plan2code
`);
}
+4 -2
View File
@@ -109,6 +109,7 @@ The scratchpad is managed by the LLM itself - after each task, the AI appends no
|-------|--------|
| Claude Code | Supported |
| GitHub Copilot CLI | Supported |
| Devin CLI | Supported |
The loop uses your configured default model for each agent.
@@ -119,7 +120,7 @@ $ plan2code-loop
╭──────────────────────────────────────╮
│ │
│ 🔮 Plany's Loop │
│ 🔮 Planny's Loop
│ Autonomous Implementation │
│ │
╰──────────────────────────────────────╯
@@ -196,7 +197,8 @@ src/
├── cli.ts # Interactive prompts
├── agents/ # Agent implementations
│ ├── claude-code.ts
── copilot-cli.ts
── copilot-cli.ts
│ └── devin-cli.ts
├── prompt/ # Prompt building
│ ├── templates.ts
│ └── builder.ts
+1
View File
@@ -20,6 +20,7 @@
"automation",
"claude",
"copilot",
"devin",
"spec-driven"
],
"license": "MIT",
+81
View File
@@ -0,0 +1,81 @@
import type { Agent, AgentConfig, AgentExecutionOptions, AgentExecutionResult } from './types.js';
import { executeCommand } from '../utils/process.js';
import { writeFileSync, unlinkSync } from 'fs';
import { join } from 'path';
import { tmpdir } from 'os';
const devinCliConfig: AgentConfig = {
name: 'devin-cli',
displayName: 'Devin CLI',
command: 'devin',
models: [
{ value: 'default', label: 'Default (use Devin config)' },
],
defaultModel: 'default',
flags: {
prompt: '--print',
promptFile: '--prompt-file',
model: '--model',
skipPermissions: '--permission-mode',
},
};
class DevinCliAgent implements Agent {
readonly config = devinCliConfig;
async execute(options: AgentExecutionOptions): Promise<AgentExecutionResult> {
// Devin CLI takes the prompt via --prompt-file rather than stdin
const tempFile = join(tmpdir(), `plan2code-prompt-${Date.now()}.txt`);
writeFileSync(tempFile, options.prompt, 'utf-8');
try {
const args: string[] = [
this.config.flags.prompt, // --print for non-interactive mode
this.config.flags.promptFile!, tempFile, // --prompt-file <path>
this.config.flags.skipPermissions, 'dangerous', // --permission-mode dangerous (auto-approve all tools)
];
// Only add --model if not using default
if (options.model && options.model !== 'default') {
args.push(this.config.flags.model, options.model);
}
const result = await executeCommand({
command: this.config.command,
args,
cwd: options.cwd,
timeout: options.timeout,
signal: options.signal,
});
return {
stdout: result.stdout,
stderr: result.stderr,
exitCode: result.exitCode,
timedOut: result.timedOut,
cancelled: result.cancelled,
duration: result.duration,
};
} finally {
// Clean up temp file
try {
unlinkSync(tempFile);
} catch {
// Ignore cleanup errors
}
}
}
async isAvailable(): Promise<boolean> {
// Run devin --version to verify it's actually installed and working
const result = await executeCommand({
command: this.config.command,
args: ['--version'],
cwd: process.cwd(),
timeout: 5000,
});
return result.exitCode === 0;
}
}
export const devinCliAgent = new DevinCliAgent();
+3
View File
@@ -9,11 +9,14 @@ export type {
export { agentRegistry } from './registry.js';
export { claudeCodeAgent } from './claude-code.js';
export { copilotCliAgent } from './copilot-cli.js';
export { devinCliAgent } from './devin-cli.js';
// Register all agents
import { agentRegistry } from './registry.js';
import { claudeCodeAgent } from './claude-code.js';
import { copilotCliAgent } from './copilot-cli.js';
import { devinCliAgent } from './devin-cli.js';
agentRegistry.register(claudeCodeAgent);
agentRegistry.register(copilotCliAgent);
agentRegistry.register(devinCliAgent);
+1
View File
@@ -14,6 +14,7 @@ export interface AgentConfig {
model: string;
skipPermissions: string;
silent?: string;
promptFile?: string;
};
}
+1 -1
View File
@@ -1,7 +1,7 @@
export type LoopMode = 'task' | 'phase';
export interface SessionConfig {
agent: string; // "claude-code" | "copilot-cli"
agent: string; // "claude-code" | "copilot-cli" | "devin-cli"
model: string; // Selected model
maxIterations: number; // 5-50
timeout: number; // Base timeout in minutes per iteration attempt
+5 -2
View File
@@ -116,9 +116,10 @@ plan2code-metrics
| **Collect metrics** | Parse a completed project spec and extract step-by-step metrics into a run JSON |
| **Import run data** | Copy a run JSON from another project into the local metrics store |
| **View metrics status** | Show aggregated metrics with health indicators and generation deltas |
| **Run analysis** | AI-powered diagnosis of weak steps (requires Claude Code or Copilot CLI) |
| **Run analysis** | AI-powered diagnosis of weak steps (requires Claude Code, GitHub Copilot CLI, or Devin CLI) |
| **Generate improvement proposal** | AI generates surgical prompt edits based on a diagnosis |
| **Review and apply** | Interactive diff review to accept/reject individual edits |
| **Fetch community submissions** | List open community-feedback GitHub issues, parse + validate their METRICS_JSON payload, import into the local run store, and close them |
### Each Time You Finish a Spec
@@ -256,6 +257,7 @@ The analysis and improvement steps require an AI agent. Two backends are support
|---------|---------|-------|
| **Claude Code** | `claude` | Uses `--print` mode. Recommended. |
| **GitHub Copilot CLI** | `copilot` | Uses stdin piping with `--allow-all-tools -s`. |
| **Devin CLI** | `devin` | Uses `--print --prompt-file <file> --permission-mode dangerous`. |
Model selection is interactive — choose from available models when prompted.
@@ -271,7 +273,8 @@ src/
├── analyzer.ts # AI diagnosis via LLM invocation
├── improver.ts # AI improvement proposal + validation
├── applier.ts # Interactive diff review + file patching
├── invoke-llm.ts # Unified LLM invocation (Claude Code / Copilot CLI)
├── community.ts # Community-feedback issue parsing + ingestion
├── invoke-llm.ts # Unified LLM invocation (Claude / Copilot / Devin)
├── index.ts # Public API exports
└── prompts/
├── analyze.md # AI prompt template for diagnosis
+102 -3
View File
@@ -1,6 +1,9 @@
import { describe, it, expect } from 'vitest';
import { avg, rate, buildCohortKey, backfillPromptVersions } from './aggregator.js';
import type { PromptVersions } from './types.js';
import fs from 'fs';
import os from 'os';
import path from 'path';
import { describe, it, expect, afterEach } from 'vitest';
import { avg, rate, buildCohortKey, backfillPromptVersions, cohortKeyForRun, aggregate } from './aggregator.js';
import type { PromptVersions, RunMetrics } from './types.js';
// ── avg() ─────────────────────────────────────────────────────────────────────
@@ -152,3 +155,99 @@ describe('buildCohortKey', () => {
expect(buildCohortKey(ordered)).toBe(buildCohortKey(reversed));
});
});
// ── Run fixtures for cohort keying / aggregation ──────────────────────────────
function makeRun(overrides: Partial<RunMetrics> = {}): RunMetrics {
return {
schema_version: '1.0',
run_id: 'run-20260101-000000-0000',
plan2code_version: '1.17.0',
prompt_versions: { ...FULL_VERSIONS },
project: { name: 'proj', started_at: null, completed_at: null },
step1_plan: { present: false, final_confidence: null, confidence_breakdown: null, clarification_rounds: null, tech_stack_revision_rounds: null, verification_gaps_found: null, functional_requirements_count: null, non_functional_requirements_count: null, risk_count: null, phase_count: null },
step2_document: { present: false, total_tasks: null, tasks_per_phase: null, phase_count: null, parallel_groups_identified: null, requirement_coverage_percent: null, verification_items_added: null },
step3_implement: { present: false, task_completion_rate: null, tasks_completed: null, tasks_total: null, blocker_count: null },
step4_finalize: { present: false, completion_rate_at_audit: null, verification_failures_found: null, documentation_updates_needed: null, archival_succeeded: null },
user_feedback: null,
...overrides,
};
}
// ── cohortKeyForRun() ─────────────────────────────────────────────────────────
describe('cohortKeyForRun', () => {
it('keys local runs by the prompt-version hash (unchanged from buildCohortKey)', () => {
const run = makeRun({ source: 'local' });
expect(cohortKeyForRun(run)).toBe(buildCohortKey(run.prompt_versions));
});
it('treats a run with no source as local', () => {
const run = makeRun();
delete run.source;
expect(cohortKeyForRun(run)).toBe(buildCohortKey(run.prompt_versions));
});
it('keys community runs by plan2code_version, ignoring prompt fingerprints', () => {
const run = makeRun({ source: 'community', plan2code_version: '1.17.0' });
expect(cohortKeyForRun(run)).toBe('community:v1.17.0');
});
it('groups two community runs of the same version together regardless of prompt fingerprint', () => {
const a = makeRun({ source: 'community', plan2code_version: '1.17.0', prompt_versions: { ...FULL_VERSIONS } });
const b = makeRun({ source: 'community', plan2code_version: '1.17.0', prompt_versions: backfillPromptVersions({} as PromptVersions) });
expect(cohortKeyForRun(a)).toBe(cohortKeyForRun(b));
});
it('separates community runs from different versions', () => {
const a = makeRun({ source: 'community', plan2code_version: '1.17.0' });
const b = makeRun({ source: 'community', plan2code_version: '1.18.0' });
expect(cohortKeyForRun(a)).not.toBe(cohortKeyForRun(b));
});
});
// ── aggregate() cohort separation ─────────────────────────────────────────────
describe('aggregate', () => {
let tmpDir: string;
afterEach(() => {
if (tmpDir) fs.rmSync(tmpDir, { recursive: true, force: true });
});
function writeRuns(runs: RunMetrics[]): { runsDir: string; outPath: string } {
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'plan2code-agg-test-'));
const runsDir = path.join(tmpDir, 'runs');
fs.mkdirSync(runsDir, { recursive: true });
for (const run of runs) {
fs.writeFileSync(path.join(runsDir, `${run.run_id}.json`), JSON.stringify(run), 'utf8');
}
return { runsDir, outPath: path.join(tmpDir, 'aggregated.json') };
}
it('places local and community runs of the same version in separate cohorts', () => {
const local = makeRun({ run_id: 'run-20260101-000001-0001', source: 'local' });
const community = makeRun({ run_id: 'run-20260101-000002-0002', source: 'community' });
const { runsDir, outPath } = writeRuns([local, community]);
const result = aggregate(runsDir, outPath);
expect(result.total_runs).toBe(2);
expect(result.cohorts).toHaveLength(2);
const communityCohort = result.cohorts.find(c => c.source === 'community');
const localCohort = result.cohorts.find(c => c.source === 'local');
expect(communityCohort?.cohort_key).toBe('community:v1.17.0');
expect(localCohort?.cohort_key).toBe(buildCohortKey(local.prompt_versions));
});
it('never selects a community cohort as current when a local cohort exists', () => {
// Community run sorts last by run_id, but current must stay on the local cohort.
const local = makeRun({ run_id: 'run-20260101-000001-0001', source: 'local' });
const community = makeRun({ run_id: 'run-29991231-235959-9999', source: 'community' });
const { runsDir, outPath } = writeRuns([local, community]);
const result = aggregate(runsDir, outPath);
expect(result.current_cohort_key).toBe(buildCohortKey(local.prompt_versions));
});
});
+44 -10
View File
@@ -35,6 +35,26 @@ export function buildCohortKey(promptVersions: PromptVersions): string {
return crypto.createHash('sha256').update(str).digest('hex').slice(0, 12);
}
/**
* Cohort key for a single run, chosen by the run's origin.
*
* Local runs are keyed by their prompt-file SHA-256 fingerprint, which captures
* in-development prompt edits that share one unreleased version.
*
* Community submissions can't reproduce that byte-exact hash: they carry the
* installed, platform-transformed prompts (not the raw src/*.md the local
* collector hashes), and the payload is LLM-generated. They are keyed instead
* by `plan2code_version` -- a reliable, byte-comparison-free identifier, since
* every released version ships a fixed set of prompts. This keeps community
* cohorts free of the CRLF/whitespace fragility a cross-machine hash would have.
*/
export function cohortKeyForRun(run: RunMetrics): string {
if (run.source === 'community') {
return `community:v${run.plan2code_version}`;
}
return buildCohortKey(run.prompt_versions);
}
/**
* Backfill missing PromptVersions fields for old run files (pre-v1.1).
*/
@@ -116,6 +136,7 @@ function buildCohort(runs: RunMetrics[], cohortKey: string): CohortMetrics {
return {
cohort_key: cohortKey,
source: runs[0].source ?? 'local',
prompt_versions: runs[0].prompt_versions,
run_count: runs.length,
run_ids: runIds,
@@ -157,7 +178,7 @@ export function aggregate(runsDir: string, outputPath: string): AggregatedMetric
// Group by cohort key
const cohortMap = new Map<string, RunMetrics[]>();
for (const run of runs) {
const key = buildCohortKey(run.prompt_versions);
const key = cohortKeyForRun(run);
if (!cohortMap.has(key)) cohortMap.set(key, []);
cohortMap.get(key)!.push(run);
}
@@ -169,9 +190,14 @@ export function aggregate(runsDir: string, outputPath: string): AggregatedMetric
}
cohorts.sort((a, b) => a.first_seen.localeCompare(b.first_seen));
// Determine current cohort (most recent)
const currentCohortKey = cohorts.length > 0
? cohorts[cohorts.length - 1].cohort_key
// Determine current cohort (most recent). Prefer local cohorts so an
// ingested community submission never becomes the maintainer's "current
// generation" for self-improvement; fall back to all cohorts if there are
// no local runs yet.
const localCohorts = cohorts.filter(c => c.source !== 'community');
const currentPool = localCohorts.length > 0 ? localCohorts : cohorts;
const currentCohortKey = currentPool.length > 0
? currentPool[currentPool.length - 1].cohort_key
: null;
const aggregated: AggregatedMetrics = {
@@ -200,13 +226,10 @@ export function loadAggregated(outputPath: string): AggregatedMetrics | null {
}
/**
* Import a single run JSON from another project into the local runs dir.
* Returns true if imported, false if already present.
* Write a run to the local runs dir, deduped by run_id filename.
* Returns true if written, false if a file for that run_id already existed.
*/
export function importRun(runJsonPath: string, runsDir: string): boolean {
const content = fs.readFileSync(runJsonPath, 'utf8');
const run = JSON.parse(content) as RunMetrics;
run.prompt_versions = backfillPromptVersions(run.prompt_versions);
export function writeRunFile(run: RunMetrics, runsDir: string): boolean {
const destPath = path.join(runsDir, `${run.run_id}.json`);
if (fs.existsSync(destPath)) {
@@ -217,3 +240,14 @@ export function importRun(runJsonPath: string, runsDir: string): boolean {
fs.writeFileSync(destPath, JSON.stringify(run, null, 2), 'utf8');
return true;
}
/**
* Import a single run JSON from another project into the local runs dir.
* Returns true if imported, false if already present.
*/
export function importRun(runJsonPath: string, runsDir: string): boolean {
const content = fs.readFileSync(runJsonPath, 'utf8');
const run = JSON.parse(content) as RunMetrics;
run.prompt_versions = backfillPromptVersions(run.prompt_versions);
return writeRunFile(run, runsDir);
}
+83 -7
View File
@@ -10,14 +10,16 @@ import path from 'path';
import { select, input, confirm } from '@inquirer/prompts';
import chalk from 'chalk';
import ora from 'ora';
import { execa } from 'execa';
import { collectRun } from './collector.js';
import { aggregate, loadAggregated, importRun, loadRunFiles } from './aggregator.js';
import { aggregate, loadAggregated, importRun, loadRunFiles, writeRunFile } from './aggregator.js';
import { runAnalysis } from './analyzer.js';
import { generateImprovement } from './improver.js';
import { reviewAndApply } from './applier.js';
import type { AggregatedMetrics, CohortMetrics } from './types.js';
import { METRIC_TARGETS } from './types.js';
import { AGENTS, type AgentType } from './invoke-llm.js';
import { listCommunityIssues, closeIssue, ingestCommunityIssues } from './community.js';
// ── Session state (set at startup via interactive prompts) ───────────────────
@@ -321,18 +323,37 @@ async function flowViewStatus(): Promise<void> {
return;
}
const localCount = aggregated.cohorts.filter(c => c.source !== 'community').length;
const communityCount = aggregated.cohorts.length - localCount;
console.log();
console.log(chalk.bold(`Total runs: ${aggregated.total_runs} | Generations: ${aggregated.cohorts.length}`));
console.log(chalk.bold(
`Total runs: ${aggregated.total_runs} | Generations: ${localCount}` +
(communityCount > 0 ? ` | Community cohorts: ${communityCount}` : '')
));
console.log(chalk.gray(`Last updated: ${aggregated.last_updated}`));
// Community runs carry no wall-clock timestamps, so their first_seen/last_seen
// fall back to the run_id string; only render a Period line for real ISO dates.
const isIsoDate = (s?: string): boolean => !!s && /^\d{4}-\d{2}-\d{2}/.test(s);
for (let i = 0; i < aggregated.cohorts.length; i++) {
const cohort = aggregated.cohorts[i];
const isCurrent = cohort.cohort_key === aggregated.current_cohort_key;
const label = isCurrent ? chalk.bold.green('[CURRENT]') : '';
const isCommunity = cohort.source === 'community';
console.log();
console.log(chalk.bold(`Generation ${i + 1} (sha:${cohort.cohort_key}) — ${cohort.run_count} run(s) ${label}`));
console.log(chalk.gray(` Period: ${cohort.first_seen?.slice(0, 10)}${cohort.last_seen?.slice(0, 10)}`));
if (isCommunity) {
// cohort_key is already `community:v<version>` — no sha: prefix.
console.log(chalk.bold(`Community feedback (${cohort.cohort_key}) — ${cohort.run_count} run(s) ${label}`));
} else {
console.log(chalk.bold(`Generation ${i + 1} (sha:${cohort.cohort_key}) — ${cohort.run_count} run(s) ${label}`));
}
if (isIsoDate(cohort.first_seen)) {
const end = isIsoDate(cohort.last_seen) ? cohort.last_seen.slice(0, 10) : cohort.first_seen.slice(0, 10);
console.log(chalk.gray(` Period: ${cohort.first_seen.slice(0, 10)}${end}`));
}
// Step 1 metrics
if (cohort.avg_confidence != null || cohort.avg_clarification_rounds != null) {
@@ -383,8 +404,10 @@ async function flowViewStatus(): Promise<void> {
console.log(` feedback_count: ${chalk.white(String(cohort.feedback_count))}`);
}
// Compare with previous generation
if (i > 0) {
// Compare with the previous cohort -- but only within the same population.
// Community and local cohorts measure different things; a cross-source
// delta (e.g. a community cohort vs the last local generation) is noise.
if (i > 0 && aggregated.cohorts[i - 1].source === cohort.source) {
const prev = aggregated.cohorts[i - 1];
const deltas: string[] = [];
if (cohort.avg_confidence != null && prev.avg_confidence != null) {
@@ -396,7 +419,7 @@ async function flowViewStatus(): Promise<void> {
deltas.push(`completion ${d >= 0 ? chalk.green(`${(d * 100).toFixed(1)}%`) : chalk.red(`${(Math.abs(d) * 100).toFixed(1)}%`)}`);
}
if (deltas.length > 0) {
console.log(chalk.gray(` vs Gen ${i}: ${deltas.join(' ')}`));
console.log(chalk.gray(` ${isCommunity ? 'vs prior version' : `vs Gen ${i}`}: ${deltas.join(' ')}`));
}
}
}
@@ -729,6 +752,55 @@ async function flowDelete(): Promise<void> {
}
}
// ── Flow: Fetch community submissions ────────────────────────────────────────
const COMMUNITY_REPO = 'jparkerweb/plan2code';
async function flowFetchCommunitySubmissions(): Promise<void> {
console.log();
console.log(chalk.bold.cyan('── Fetch Community Submissions ──'));
try {
await execa('gh', ['auth', 'status']);
} catch {
console.log(chalk.red('`gh` CLI not found or not authenticated — install/auth `gh` to use this feature.'));
return;
}
let issues;
try {
issues = await listCommunityIssues(COMMUNITY_REPO);
} catch (err) {
console.log(chalk.red(`Failed to list community-feedback issues: ${err instanceof Error ? err.message : String(err)}`));
return;
}
if (issues.length === 0) {
console.log(chalk.yellow('No open community-feedback issues found.'));
return;
}
const { runsDir, aggregatedPath } = getMetricsDirs();
const tally = await ingestCommunityIssues(issues, COMMUNITY_REPO, runsDir, { writeRunFile, closeIssue });
for (const num of tally.malformedIssues) {
console.log(chalk.yellow(` Skipping issue #${num}: malformed or missing METRICS_JSON payload.`));
}
for (const num of tally.closeFailedIssues) {
console.log(chalk.yellow(` Imported issue #${num} but failed to close it (still open; will retry next fetch).`));
}
if (tally.imported > 0) {
aggregate(runsDir, aggregatedPath);
}
console.log();
console.log(chalk.green(
`✓ Imported: ${tally.imported} Skipped (duplicate): ${tally.skippedDuplicate} Skipped (malformed): ${tally.skippedMalformed} Closed: ${tally.closed} Close failed: ${tally.closeFailed}`
));
}
// ── Main menu ─────────────────────────────────────────────────────────────────
export async function runCLI(): Promise<void> {
@@ -800,6 +872,7 @@ export async function runCLI(): Promise<void> {
{ name: 'Run analysis (diagnose weak steps)', value: 'analyze' },
{ name: 'Generate improvement proposal', value: 'propose' },
{ name: 'Review and apply a proposal', value: 'apply' },
{ name: 'Fetch community submissions', value: 'fetch-community' },
{ name: chalk.red('Delete metrics data'), value: 'delete' },
{ name: 'Exit', value: 'exit' },
],
@@ -824,6 +897,9 @@ export async function runCLI(): Promise<void> {
case 'apply':
await flowReviewAndApply();
break;
case 'fetch-community':
await flowFetchCommunitySubmissions();
break;
case 'delete':
await flowDelete();
break;
+2 -1
View File
@@ -43,7 +43,7 @@ function readFileSafe(filePath: string): string | null {
* If `stepFilter` is provided, returns only the block with matching "step" field.
* Returns parsed object or null if not found / invalid.
*/
function extractMetricsJson(content: string, stepFilter?: string): Record<string, unknown> | null {
export function extractMetricsJson(content: string, stepFilter?: string): Record<string, unknown> | null {
const re = /<!--\s*METRICS_JSON\s+(\{[\s\S]*?\})\s*-->/g;
let match: RegExpExecArray | null;
while ((match = re.exec(content)) !== null) {
@@ -611,6 +611,7 @@ export async function collectRun(opts: CollectorOptions): Promise<RunMetrics> {
schema_version: '1.0',
run_id: runId,
plan2code_version: plan2codeVersion,
source: 'local',
prompt_versions: collectPromptVersions(plan2codeRoot),
project: {
name: projectName,
+243
View File
@@ -0,0 +1,243 @@
import fs from 'fs';
import os from 'os';
import path from 'path';
import { describe, it, expect, afterEach } from 'vitest';
import { parseSubmissionPayload, ingestCommunityIssues } from './community.js';
import { writeRunFile } from './aggregator.js';
import type { RunMetrics } from './types.js';
function issueBody(payload: Record<string, unknown>): string {
return `Some issue text.\n\n<!-- METRICS_JSON ${JSON.stringify(payload)} -->\n`;
}
const VALID_PAYLOAD = {
schema_version: '1.0',
run_id: 'run-20260715-143000-a1b2',
plan2code_version: '1.15.3',
prompt_versions_short: {
plan: 'abc123def456', revise_plan: 'a', document: 'b', implement: 'c',
finalize: 'd', init: 'e', init_update: 'f', quick_task: 'g',
},
step1: {
final_confidence: 95,
confidence_breakdown: { requirements: 24, feasibility: 23, integration: 24, risk: 22 },
clarification_rounds: 0,
tech_stack_revision_rounds: 0,
verification_gaps_found: 0,
functional_requirements_count: 8,
non_functional_requirements_count: 6,
risk_count: 7,
phase_count: 4,
},
step2: {
total_tasks: 28,
phase_count: 4,
parallel_groups_identified: 1,
requirement_coverage_percent: 100,
verification_items_added: 3,
},
step3: {
task_completion_rate: 0.96,
tasks_completed: 27,
tasks_total: 28,
blocker_count: 1,
},
step4: {
completion_rate_at_audit: 0.96,
verification_failures_found: 1,
documentation_updates_needed: 2,
},
user_feedback: {
overall_rating: 8,
rating_reason: 'good stuff',
what_went_well: 'well',
what_went_poorly: 'poorly',
},
};
// ── parseSubmissionPayload() ─────────────────────────────────────────────────
describe('parseSubmissionPayload', () => {
it('parses a fully valid payload into a correctly-shaped RunMetrics', () => {
const result = parseSubmissionPayload(issueBody(VALID_PAYLOAD));
expect(result).not.toBeNull();
expect(result!.run_id).toBe('run-20260715-143000-a1b2');
expect(result!.schema_version).toBe('1.0');
expect(result!.plan2code_version).toBe('1.15.3');
expect(result!.source).toBe('community');
expect(result!.step1_plan.present).toBe(true);
expect(result!.step1_plan.final_confidence).toBe(95);
expect(result!.step2_document.present).toBe(true);
expect(result!.step2_document.total_tasks).toBe(28);
expect(result!.step3_implement.present).toBe(true);
expect(result!.step3_implement.tasks_completed).toBe(27);
expect(result!.step4_finalize.present).toBe(true);
expect(result!.step4_finalize.completion_rate_at_audit).toBe(0.96);
expect(result!.user_feedback).toEqual({
overall_rating: 8,
rating_reason: 'good stuff',
what_went_well: 'well',
what_went_poorly: 'poorly',
});
});
it('returns null when run_id is missing', () => {
const { run_id, ...withoutRunId } = VALID_PAYLOAD;
const result = parseSubmissionPayload(issueBody(withoutRunId));
expect(result).toBeNull();
});
it('returns null when user_feedback.overall_rating is a string instead of a number', () => {
const badPayload = {
...VALID_PAYLOAD,
user_feedback: { ...VALID_PAYLOAD.user_feedback, overall_rating: 'eight' },
};
const result = parseSubmissionPayload(issueBody(badPayload));
expect(result).toBeNull();
});
it('returns null when schema_version is not exactly "1.0"', () => {
const badPayload = { ...VALID_PAYLOAD, schema_version: '2.0' };
const result = parseSubmissionPayload(issueBody(badPayload));
expect(result).toBeNull();
});
it('sets all four steps present:false when only user_feedback is included', () => {
const minimalPayload = {
schema_version: '1.0',
run_id: 'run-20260715-150000-c3d4',
plan2code_version: '1.15.3',
user_feedback: VALID_PAYLOAD.user_feedback,
};
const result = parseSubmissionPayload(issueBody(minimalPayload));
expect(result).not.toBeNull();
expect(result!.step1_plan.present).toBe(false);
expect(result!.step2_document.present).toBe(false);
expect(result!.step3_implement.present).toBe(false);
expect(result!.step4_finalize.present).toBe(false);
});
it('backfills missing prompt_versions_short keys with the sha256:missing sentinel', () => {
const partialPayload = {
...VALID_PAYLOAD,
prompt_versions_short: { plan: 'abc123def456' },
};
const result = parseSubmissionPayload(issueBody(partialPayload));
expect(result).not.toBeNull();
expect(result!.prompt_versions.plan).toBe('abc123def456');
expect(result!.prompt_versions.revise_plan).toBe('sha256:missing');
expect(result!.prompt_versions.document).toBe('sha256:missing');
expect(result!.prompt_versions.implement).toBe('sha256:missing');
expect(result!.prompt_versions.finalize).toBe('sha256:missing');
expect(result!.prompt_versions.init).toBe('sha256:missing');
expect(result!.prompt_versions.init_update).toBe('sha256:missing');
expect(result!.prompt_versions.quick_task).toBe('sha256:missing');
});
});
// ── writeRunFile() ────────────────────────────────────────────────────────────
describe('writeRunFile', () => {
let tmpDir: string;
afterEach(() => {
if (tmpDir) fs.rmSync(tmpDir, { recursive: true, force: true });
});
const RUN: RunMetrics = {
schema_version: '1.0',
run_id: 'run-20260715-160000-e5f6',
plan2code_version: '1.15.3',
prompt_versions: {
plan: 'sha256:missing', revise_plan: 'sha256:missing', document: 'sha256:missing',
implement: 'sha256:missing', finalize: 'sha256:missing', init: 'sha256:missing',
init_update: 'sha256:missing', quick_task: 'sha256:missing',
},
project: { name: '', started_at: null, completed_at: null },
step1_plan: { present: false, final_confidence: null, confidence_breakdown: null, clarification_rounds: null, tech_stack_revision_rounds: null, verification_gaps_found: null, functional_requirements_count: null, non_functional_requirements_count: null, risk_count: null, phase_count: null },
step2_document: { present: false, total_tasks: null, tasks_per_phase: null, phase_count: null, parallel_groups_identified: null, requirement_coverage_percent: null, verification_items_added: null },
step3_implement: { present: false, task_completion_rate: null, tasks_completed: null, tasks_total: null, blocker_count: null },
step4_finalize: { present: false, completion_rate_at_audit: null, verification_failures_found: null, documentation_updates_needed: null, archival_succeeded: null },
user_feedback: null,
};
it('writes a new run_id to an empty runsDir and returns true', () => {
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'plan2code-metrics-test-'));
const result = writeRunFile(RUN, tmpDir);
expect(result).toBe(true);
expect(fs.existsSync(path.join(tmpDir, `${RUN.run_id}.json`))).toBe(true);
});
it('returns false and does not overwrite when the same run_id already exists', () => {
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'plan2code-metrics-test-'));
writeRunFile(RUN, tmpDir);
const modified = { ...RUN, plan2code_version: '9.9.9' };
const result = writeRunFile(modified, tmpDir);
expect(result).toBe(false);
const onDisk = JSON.parse(fs.readFileSync(path.join(tmpDir, `${RUN.run_id}.json`), 'utf8'));
expect(onDisk.plan2code_version).toBe('1.15.3');
});
});
// ── ingestCommunityIssues() ───────────────────────────────────────────────────
describe('ingestCommunityIssues', () => {
const goodBody = issueBody(VALID_PAYLOAD);
const badBody = 'an issue with no METRICS_JSON payload';
it('imports a new run and closes its issue', async () => {
const closed: number[] = [];
const tally = await ingestCommunityIssues(
[{ number: 1, body: goodBody }],
'owner/repo',
'/runs',
{ writeRunFile: () => true, closeIssue: async (_r, n) => { closed.push(n); } },
);
expect(tally.imported).toBe(1);
expect(tally.skippedDuplicate).toBe(0);
expect(tally.closed).toBe(1);
expect(closed).toEqual([1]);
});
it('closes an already-imported (duplicate) issue instead of skipping the close', async () => {
const closed: number[] = [];
const tally = await ingestCommunityIssues(
[{ number: 7, body: goodBody }],
'owner/repo',
'/runs',
{ writeRunFile: () => false, closeIssue: async (_r, n) => { closed.push(n); } },
);
expect(tally.imported).toBe(0);
expect(tally.skippedDuplicate).toBe(1);
expect(tally.closed).toBe(1); // the close-retry: duplicates are still closed
expect(closed).toEqual([7]);
});
it('records a close failure without throwing and leaves the issue for a later retry', async () => {
const tally = await ingestCommunityIssues(
[{ number: 9, body: goodBody }],
'owner/repo',
'/runs',
{ writeRunFile: () => true, closeIssue: async () => { throw new Error('network'); } },
);
expect(tally.imported).toBe(1);
expect(tally.closed).toBe(0);
expect(tally.closeFailed).toBe(1);
expect(tally.closeFailedIssues).toEqual([9]);
});
it('skips and reports a malformed issue without writing or closing it', async () => {
let wrote = false;
let closeCalled = false;
const tally = await ingestCommunityIssues(
[{ number: 3, body: badBody }],
'owner/repo',
'/runs',
{ writeRunFile: () => { wrote = true; return true; }, closeIssue: async () => { closeCalled = true; } },
);
expect(tally.skippedMalformed).toBe(1);
expect(tally.malformedIssues).toEqual([3]);
expect(wrote).toBe(false);
expect(closeCalled).toBe(false);
});
});
+276
View File
@@ -0,0 +1,276 @@
/**
* community.ts
* Ingestion side of the community feedback flow: list/parse/close
* `community-feedback`-labeled GitHub issues on jparkerweb/plan2code.
*/
import { execa } from 'execa';
import type {
RunMetrics,
PromptVersions,
UserFeedback,
Step1PlanMetrics,
Step2DocumentMetrics,
Step3ImplementMetrics,
Step4FinalizeMetrics,
} from './types.js';
import { extractMetricsJson } from './collector.js';
import { backfillPromptVersions } from './aggregator.js';
// ── GitHub interaction (via gh CLI) ──────────────────────────────────────────
export interface CommunityIssue {
number: number;
body: string;
}
interface RawIssue {
number: number;
title: string;
body: string;
labels: { name: string }[];
}
const COMMUNITY_LABEL = 'community-feedback';
const FEEDBACK_TITLE_PREFIX = '[Feedback]';
const METRICS_JSON_MARKER = /<!--\s*METRICS_JSON\s+\{/;
export async function listCommunityIssues(repo: string): Promise<CommunityIssue[]> {
// Fetch open issues broadly rather than by label alone. Browser/print-tier
// submissions from outside contributors can lose the `community-feedback`
// label: GitHub only honors the `labels=` query param on issues/new for
// users with triage/push access, so the label is silently dropped for
// community members without `gh`. We therefore also match by the `[Feedback]`
// title prefix and the METRICS_JSON marker. Only OPEN issues are considered
// (closed/done submissions are already processed); parseSubmissionPayload is
// the final gate that rejects anything without a valid payload.
const result = await execa('gh', [
'issue', 'list',
'--repo', repo,
'--state', 'open',
'--limit', '1000',
'--json', 'number,title,body,labels',
]);
const raw = JSON.parse(result.stdout) as RawIssue[];
return raw
.filter((issue) =>
issue.labels.some((l) => l.name === COMMUNITY_LABEL) ||
issue.title.startsWith(FEEDBACK_TITLE_PREFIX) ||
METRICS_JSON_MARKER.test(issue.body)
)
.map((issue) => ({ number: issue.number, body: issue.body }));
}
export async function closeIssue(repo: string, issueNumber: number): Promise<void> {
await execa('gh', ['issue', 'close', String(issueNumber), '--repo', repo]);
}
// ── Ingestion control flow (I/O injected so it is unit-testable) ──────────────
export interface IngestionDeps {
writeRunFile: (run: RunMetrics, runsDir: string) => boolean;
closeIssue: (repo: string, issueNumber: number) => Promise<void>;
}
export interface IngestionTally {
imported: number;
skippedDuplicate: number;
skippedMalformed: number;
closed: number;
closeFailed: number;
malformedIssues: number[];
closeFailedIssues: number[];
}
/**
* Process a batch of community issues: parse each payload, write new runs
* (deduped by run_id), and close every open issue idempotently.
*
* The close is attempted on the duplicate path too: a submission that imported
* on an earlier run but failed to close would otherwise be seen as a duplicate
* forever and never closed again, leaving the issue open and reprocessed on
* every fetch. Malformed issues are reported (not fixed up) and left open.
*
* I/O (writeRunFile/closeIssue) is injected so the control flow can be unit
* tested without a live `gh`. Returns a tally; the caller owns all logging.
*/
export async function ingestCommunityIssues(
issues: CommunityIssue[],
repo: string,
runsDir: string,
deps: IngestionDeps,
): Promise<IngestionTally> {
const tally: IngestionTally = {
imported: 0, skippedDuplicate: 0, skippedMalformed: 0,
closed: 0, closeFailed: 0, malformedIssues: [], closeFailedIssues: [],
};
for (const issue of issues) {
const run = parseSubmissionPayload(issue.body);
if (!run) {
tally.skippedMalformed++;
tally.malformedIssues.push(issue.number);
continue;
}
if (deps.writeRunFile(run, runsDir)) {
tally.imported++;
} else {
tally.skippedDuplicate++;
}
try {
await deps.closeIssue(repo, issue.number);
tally.closed++;
} catch {
tally.closeFailed++;
tally.closeFailedIssues.push(issue.number);
}
}
return tally;
}
// ── Payload parsing (type-only validation, per NFR-5) ────────────────────────
function isString(v: unknown): v is string {
return typeof v === 'string';
}
function isNumber(v: unknown): v is number {
return typeof v === 'number';
}
function numOrNull(v: unknown): number | null {
return isNumber(v) ? v : null;
}
/** Type-check `keys` off `raw` (object or not) into a { [key]: number | null } map. */
function pickNumbers<K extends string>(raw: Record<string, unknown>, keys: readonly K[]): Record<K, number | null> {
const result = {} as Record<K, number | null>;
for (const key of keys) result[key] = numOrNull(raw[key]);
return result;
}
const STEP1_ABSENT: Step1PlanMetrics = {
present: false, final_confidence: null, confidence_breakdown: null,
clarification_rounds: null, tech_stack_revision_rounds: null,
verification_gaps_found: null, functional_requirements_count: null,
non_functional_requirements_count: null, risk_count: null, phase_count: null,
};
function parseStep1(raw: unknown): Step1PlanMetrics {
if (raw == null || typeof raw !== 'object') return STEP1_ABSENT;
const step1 = raw as Record<string, unknown>;
const bdRaw = step1['confidence_breakdown'];
const breakdown = bdRaw != null && typeof bdRaw === 'object'
? pickNumbers(bdRaw as Record<string, unknown>, ['requirements', 'feasibility', 'integration', 'risk'])
: null;
return {
present: true,
confidence_breakdown: breakdown,
...pickNumbers(step1, [
'final_confidence', 'clarification_rounds', 'tech_stack_revision_rounds',
'verification_gaps_found', 'functional_requirements_count',
'non_functional_requirements_count', 'risk_count', 'phase_count',
]),
};
}
const STEP2_ABSENT: Step2DocumentMetrics = {
present: false, total_tasks: null, tasks_per_phase: null,
phase_count: null, parallel_groups_identified: null,
requirement_coverage_percent: null, verification_items_added: null,
};
function parseStep2(raw: unknown): Step2DocumentMetrics {
if (raw == null || typeof raw !== 'object') return STEP2_ABSENT;
const step2 = raw as Record<string, unknown>;
return {
present: true,
tasks_per_phase: null,
...pickNumbers(step2, [
'total_tasks', 'phase_count', 'parallel_groups_identified',
'requirement_coverage_percent', 'verification_items_added',
]),
};
}
const STEP3_ABSENT: Step3ImplementMetrics = {
present: false, task_completion_rate: null, tasks_completed: null, tasks_total: null, blocker_count: null,
};
function parseStep3(raw: unknown): Step3ImplementMetrics {
if (raw == null || typeof raw !== 'object') return STEP3_ABSENT;
const step3 = raw as Record<string, unknown>;
return {
present: true,
...pickNumbers(step3, ['task_completion_rate', 'tasks_completed', 'tasks_total', 'blocker_count']),
};
}
const STEP4_ABSENT: Step4FinalizeMetrics = {
present: false, completion_rate_at_audit: null, verification_failures_found: null,
documentation_updates_needed: null, archival_succeeded: null,
};
function parseStep4(raw: unknown): Step4FinalizeMetrics {
if (raw == null || typeof raw !== 'object') return STEP4_ABSENT;
const step4 = raw as Record<string, unknown>;
const archivalRaw = step4['archival_succeeded'];
return {
present: true,
archival_succeeded: typeof archivalRaw === 'boolean' ? archivalRaw : null,
...pickNumbers(step4, ['completion_rate_at_audit', 'verification_failures_found', 'documentation_updates_needed']),
};
}
function parsePromptVersionsShort(raw: unknown): PromptVersions {
const partial: Partial<PromptVersions> = {};
if (raw != null && typeof raw === 'object') {
const pv = raw as Record<string, unknown>;
for (const key of ['plan', 'revise_plan', 'document', 'implement', 'finalize', 'init', 'init_update', 'quick_task'] as const) {
const v = pv[key];
if (isString(v)) partial[key] = v;
}
}
return backfillPromptVersions(partial as PromptVersions);
}
export function parseSubmissionPayload(body: string): RunMetrics | null {
const parsed = extractMetricsJson(body);
if (!parsed) return null;
if (parsed['schema_version'] !== '1.0') return null;
if (!isString(parsed['run_id'])) return null;
if (!isString(parsed['plan2code_version'])) return null;
const feedbackRaw = parsed['user_feedback'];
if (feedbackRaw == null || typeof feedbackRaw !== 'object') return null;
const feedback = feedbackRaw as Record<string, unknown>;
if (!isNumber(feedback['overall_rating'])) return null;
if (!isString(feedback['rating_reason'])) return null;
if (!isString(feedback['what_went_well'])) return null;
if (!isString(feedback['what_went_poorly'])) return null;
const userFeedback: UserFeedback = {
overall_rating: feedback['overall_rating'],
rating_reason: feedback['rating_reason'],
what_went_well: feedback['what_went_well'],
what_went_poorly: feedback['what_went_poorly'],
};
return {
schema_version: '1.0',
run_id: parsed['run_id'],
plan2code_version: parsed['plan2code_version'],
source: 'community',
prompt_versions: parsePromptVersionsShort(parsed['prompt_versions_short']),
project: { name: '', started_at: null, completed_at: null },
step1_plan: parseStep1(parsed['step1']),
step2_document: parseStep2(parsed['step2']),
step3_implement: parseStep3(parsed['step3']),
step4_finalize: parseStep4(parsed['step4']),
user_feedback: userFeedback,
};
}
+26 -2
View File
@@ -1,7 +1,8 @@
/**
* invoke-llm.ts
* Unified LLM invocation for plan2code-metrics.
* Supports Claude Code (temp file → stdin) and Copilot CLI (stdin string).
* Supports Claude Code (temp file → stdin), Copilot CLI (stdin string), and
* Devin CLI (temp prompt file).
* Mirrors the agent pattern from plan2code-loop.
*/
@@ -12,7 +13,7 @@ import { tmpdir } from 'os';
// ── Agent definitions ────────────────────────────────────────────────────────
export type AgentType = 'claude-code' | 'copilot-cli';
export type AgentType = 'claude-code' | 'copilot-cli' | 'devin-cli';
export interface AgentDef {
name: AgentType;
@@ -34,6 +35,12 @@ export const AGENTS: Record<AgentType, AgentDef> = {
command: 'copilot',
defaultModel: 'claude-sonnet-4',
},
'devin-cli': {
name: 'devin-cli',
displayName: 'Devin CLI',
command: 'devin',
defaultModel: 'default',
},
};
// ── Invocation ───────────────────────────────────────────────────────────────
@@ -72,6 +79,23 @@ export async function invokeLLM(opts: InvokeLLMOptions): Promise<string> {
} finally {
try { unlinkSync(tempFile); } catch { /* ignore cleanup errors */ }
}
} else if (agent === 'devin-cli') {
// Devin CLI: load prompt from a temp file, run single-turn, auto-approve tool calls
const tempFile = join(tmpdir(), `plan2code-metrics-prompt-${Date.now()}.txt`);
writeFileSync(tempFile, prompt, 'utf-8');
try {
const args: string[] = ['--print', '--prompt-file', tempFile, '--permission-mode', 'dangerous'];
// Only add --model if not using default
if (model && model !== 'default') {
args.push('--model', model);
}
const result = await execa(def.command, args, { timeout });
return result.stdout;
} finally {
try { unlinkSync(tempFile); } catch { /* ignore cleanup errors */ }
}
} else {
// Copilot CLI: pipe prompt via stdin string
const args: string[] = [];
+7 -1
View File
@@ -66,6 +66,11 @@ export interface RunMetrics {
schema_version: '1.0';
run_id: string;
plan2code_version: string;
// Origin of the run. Local runs are collected on the maintainer's machine;
// community runs are ingested from GitHub feedback issues. Absent on pre-v1.17
// run files, which are treated as 'local'. Drives cohort keying (see
// cohortKeyForRun in aggregator.ts).
source?: 'local' | 'community';
prompt_versions: PromptVersions;
project: {
name: string;
@@ -102,7 +107,8 @@ export interface PromptProposal {
// Aggregated metrics schema
export interface CohortMetrics {
cohort_key: string; // hash of sorted prompt_versions
cohort_key: string; // local: hash of sorted prompt_versions; community: `community:v<version>`
source?: 'local' | 'community';
prompt_versions: PromptVersions;
run_count: number;
run_ids: string[];
@@ -0,0 +1,663 @@
# Chart Playbook
> Part of plan2code-0-pathfinder — loaded at the top of MODE A (Chart). Expands the numbered Chart steps.
>
> **Backend note.** Steps 0, 2, and 4 — the gates and the grills — are identical either way, and so is every judgement call below (fog vs question, in scope vs out, the destination test). What differs is where Steps 3 and 5-7 put the bytes: on `**Backend:** github`, `github-issues.md` overrides the `map.md` and question-file templates here, the single-pass rule under *Step 6: Numbering and dependency order*, and the timing of the recon. Read it alongside this file, not instead of it.
Charting produces a map and a set of question files. It resolves nothing by hand. Every judgement below serves one goal: put a sharp question on the map for everything you can phrase now, and leave everything else honestly in the fog.
---
## Step 0: The intent gate
Pathfinder builds an apparatus — a directory, a map, a file per decision. That apparatus earns its keep only when the way to the destination is genuinely foggy. For a small or already-clear ask it is pure overhead, and creating it before the human has agreed to it is the fastest way to make Pathfinder feel heavy and get in the way. So the gate runs **before the first byte hits disk**.
You already have the idea name from Auto-Discovery. Do NOT create the directory yet. Say, in substance:
> "This is Pathfinder. Nothing exists for `<idea>` yet. Pathfinder charts a map of the open decisions when an idea is big or unclear to plan — but that is overhead if this is small or already clear. Three ways to go:
> - **Chart it** — I map the open decisions, one per session, then hand a draft to `/plan2code-1-plan`.
> - **Straight to `/plan2code-1-plan`** — the way looks clear enough to plan now.
> - **`/plan2code-quick-task`** — small enough to just do.
>
> My read: `<recommendation with a one-line reason>`. Which?"
Rules for the gate:
- **It is HITL.** You recommend; the human chooses. Never self-select "chart" and start creating files because it is the default path — that is exactly the failure this gate exists to stop.
- **Read the request honestly.** A one-line bugfix, or a change with no open decisions, is not a charting job — recommend an off-ramp and mean it. Reserve "chart" for real fog: several unsettled decisions, unclear scope, or competing designs.
- **No disk writes.** Naming the idea and talking is free. Creating `specs/<idea>/pathfinder/` is not — it waits for an explicit "chart."
- **On an off-ramp, route and STOP.** Point at `/plan2code-1-plan` or `/plan2code-quick-task`, create nothing, end the session. If `AGENTS.md` is absent, mention `/plan2code-init` first, as with any handoff.
This gate and the Step 4 no-fog off-ramp are the same instinct at two moments: the gate is the human's call before any work; the off-ramp is your call once the breadth-first grill has proven there was no fog after all. Either one ending the session without a map is a success, not a failure.
---
## Step 2: The destination grill
The destination is settled **first** because it fixes scope. Every later judgement — is this a question or fog, is this in scope or past the edge, is this map cleared — is measured against it. A vague destination makes all three unanswerable, and you will spend the rest of the map arguing about boundaries instead of decisions.
A destination is **one or two lines** describing what exists when the map clears. It is not a feature description. It is not a value proposition. It names the artifact and draws the edge.
### The script
Recommend an answer with each probe — the human corrects faster than they compose. Follow the grilling playbook for tone, cadence, and the batching rules; this is the content.
**Six probes, two batches.** They do not all pass the independence test, so they split:
| Batch | Probes | Why they go together |
|---|---|---|
| First | 1 (artifact), 2 (person), 6 (forcing function) | Each stands alone. None reads differently under the others' answers. |
| Second | 3 (sacrificial boundary), 4 (smallest arrival), 5 (arrival signal) | All three presuppose an artifact and an actor. Sending them before batch 1 lands asks the human to draw an edge around something unnamed. |
Two round trips, not six. Three probes each — exactly the cap, so neither batch needs splitting.
Batch 2 bends the independence test on purpose. The arrival signal (5) can shift under the smallest arrival (4), so strictly it should be held back — but holding it costs a third round trip to catch a conflict that is rare and cheap to spot. The trade is to send them together and reconcile at the recap: if the smallest arrival comes back materially smaller than the artifact you were told about, re-check the arrival signal against it before writing the destination. A knowing trade here, not a licence to batch dependent probes elsewhere.
**Both batches go out as numbered Q blocks — this grill is the other standing exception to the tool-first rule.** Probes 2, 3, 4, and 5 need the human's own phrasing — the destination is written into `map.md` verbatim as agreed, so a clicked option label is not something you can write down. That is the detail test's first row, four times over. Probe 6 names categories but the category is the worthless half of the answer: "deadline" changes nothing, "Q3 close, and the SEC audit lands Nov 1" changes the delivery question, the testing posture, and the out-of-scope line at once. Only probe 1 would survive a picker on its own, and it rides in a Q block anyway, because one tripping probe downgrades the whole batch. Do not reach for the structured question tool here.
**Probe 1 — the artifact**
> "When this map is cleared, what exists that does not exist now: a plan you hand to `/plan2code-1-plan`, a decision locked before anyone plans, or a change already made in the codebase? My guess: a plan."
*Fishing for:* the shape of the destination. Push back if the answer is "the feature working" — that is past the edge of every pathfinder map. Say so plainly: "That is the build. The map ends at the plan for the build."
**Probe 2 — the person on the other end**
> "Who uses the result, and what do they do with it the day it lands?"
*Fishing for:* the actor and the moment of use. Vague actors ("users", "the business") produce vague scope. Push until you get a role someone could name in an approval — compliance officer, on-call SRE, tenant admin.
**Probe 3 — the sacrificial boundary**
> "Name one thing a reasonable person would assume is part of this that you are willing to say is NOT part of it."
*Fishing for:* the first `## Out of scope` bullet. This probe does more work than any other. A destination nobody has excluded anything from has not been thought about. If the human cannot name one, offer two candidates and make them reject one.
**Probe 4 — the smallest arrival**
> "What is the smallest version that would still count as arriving? If only that existed, would you call it done or would you feel cheated?"
*Fishing for:* the difference between the destination and the wish list. Everything above the smallest arrival is a candidate for out of scope or for a later effort.
**Probe 5 — the arrival signal**
> "How do you know you have arrived — what do you look at?"
*Fishing for:* a checkable condition. "It feels right" is not one. "Every open decision has an answer and I can hand the draft to planning without re-litigating format" is one.
**Probe 6 — the forcing function**
> "What made this surface now? A deadline, an incident, an audit, a customer?"
*Fishing for:* constraints that will shape half the questions and that nobody volunteers unprompted. A regulatory deadline changes the delivery question, the testing posture question, and the out-of-scope line all at once.
**The probe names above are internal labels, not headings the human reads.** Head each Q block plainly — *What you end up with*, *Who uses it*, *What's not included*, *Smallest version that counts*, *How you know it's done*, *Why now* — and keep the probe text itself as plain as the quotes above. "The sacrificial boundary" and "the arrival signal" mean something to this playbook and nothing to the person answering. Full rule in the grilling playbook, *Say it in plain English*.
### Worked example — same idea, two destinations
**Idea:** "we should let people export audit logs"
**BAD destination**
> Let users export audit logs so they have their data.
Why it fails, concretely:
| Failure | Consequence downstream |
|---|---|
| No artifact named | Nobody knows whether the map clears at a plan or at shipped code |
| "Users" is not a role | The authorization question cannot even be phrased |
| No edge | Live streaming, SIEM push, and a schema redesign all argue their way in |
| No arrival signal | The Clearing Gate has nothing to check against |
| "their data" is a rationale, not a boundary | Every fog bullet reads as in scope |
**GOOD destination**
> A locked implementation plan for a compliance officer to export a filtered range of audit events from the admin UI and receive them as a single downloadable file. The map ends at the plan, not at shipped code. Continuous streaming to external systems is not on the route.
Three sentences, two lines of substance: artifact (a plan), actor (compliance officer), trigger surface (admin UI), shape of the result (one downloadable file), and an explicit edge (no streaming). Every one of those clauses will be cited later when you decide whether something is a question, fog, or out of scope.
**Write the destination into `map.md` verbatim as agreed.** Do not improve it afterwards. If it needs changing, change it with the human present — a silently redrawn destination invalidates every scope call already made.
---
## Step 3: Codebase recon
Recon is `legwork · AFK` — you do it alone, and you write it down **already resolved**. It exists so that no later session re-reads the same directories, and so that the handoff carries the ground truth `/plan2code-1-plan` Phase 2 (System Context Examination) would otherwise have to rediscover.
`questions/00-codebase-context.md` is always `00`. It always exists. It is created with `State: resolved` and a filled `## Answer` in the same write.
### What to explore
| Area | What to establish | Where to look |
|---|---|---|
| **Directory structure** | The map of the repo at the depth that matters for this destination — not every folder, the ones the work will touch | Top-level listing, then two levels into the relevant subtrees |
| **Key components** | The modules that would be read, changed, or called. Name, path, responsibility | Entry points, route/controller registries, service layers |
| **Patterns and conventions** | How this codebase does the thing you are about to plan: error handling, validation, config, module layout, naming, async style | Two or three recent files in the target area, plus `AGENTS.md` |
| **Integration points** | External systems, queues, storage, auth providers, feature-flag services the work will cross | Config files, environment variable references, client wrappers |
| **Technical debt in the blast radius** | Only debt the destination would collide with. Not a repo-wide audit | Long files in the target area, duplicated helpers, stale TODO markers with no owner |
| **System boundaries** | What this effort owns versus what it merely calls. Where the change stops | Package boundaries, ownership files, API contracts |
Two disciplines keep this file useful:
- **Verify behaviour against actual code, never against a filename.** A file called `auditLogger.ts` may log nothing.
- **Scope the recon to the destination.** A recon of the whole repo is unreadable and stale in a week. If a subtree cannot plausibly be touched by the destination, say so in one line and move on.
Record what you could **not** determine as an explicit gap. Gaps at recon time are often the first real fog bullets.
### The literal file
```markdown
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
# Codebase context
Type: legwork · AFK
State: resolved
Blocked by: none
Claimed: 2026-08-03 09:12
Locked: no
## Question
What does this codebase already provide, constrain, and forbid for an operator-initiated
audit-log export? Establish structure, components, conventions, integrations, debt in the
blast radius, and boundaries — enough that no later session re-reads the same ground and
enough to hand to `/plan2code-1-plan` as its system context.
## Answer
### Directory structure
- `src/api/` — Express routers, one file per resource. `src/api/admin/` is the admin surface.
- `src/services/` — business logic; the only layer allowed to touch `src/db/`.
- `src/db/` — Knex query builders and migrations. `audit_events` lives here.
- `src/jobs/` — BullMQ workers. Existing precedent for long-running work.
- `src/web/admin/` — React admin UI, TanStack Query, colocated route components.
- `test/` — Vitest, mirroring `src/` one-to-one.
Untouched by this destination: `src/billing/`, `src/web/marketing/`.
### Key components
| Component | Path | Responsibility |
|---|---|---|
| `auditEvents.record()` | `src/services/auditEvents.ts` | Sole writer of `audit_events`; called from 31 sites |
| `adminRouter` | `src/api/admin/index.ts` | Mounts admin routes; applies `requireAdmin` |
| `requireAdmin` | `src/api/middleware/auth.ts` | Session check plus role check; no per-tenant scoping today |
| `reportQueue` | `src/jobs/reportQueue.ts` | BullMQ queue used by the existing billing report export |
| `signedUrl()` | `src/services/storage.ts` | S3 pre-signed URL helper, fixed 15-minute expiry |
### Patterns and conventions
- Errors: typed error classes thrown from services, mapped to HTTP by `errorHandler`. Never raw `res.status(500)`.
- Validation: Zod schema per route, exported next to the handler.
- Config: everything through `src/config.ts`; no direct `process.env` reads outside it.
- Async: `async`/`await` throughout. No callback style remains.
- Long-running work: enqueue to BullMQ, return `202` with a job id. Established by billing reports.
- Tests: Vitest, colocated fixtures, no shared mutable state between cases.
### Integration points
- **Postgres 15** via Knex. `audit_events` is ~180M rows, partitioned monthly.
- **Redis** backing BullMQ.
- **S3** for generated artifacts; the billing export already writes there.
- **SES** for transactional mail; templates in `src/mail/templates/`.
- No SIEM, log-shipping, or streaming integration exists today.
### Technical debt in the blast radius
- `audit_events` has an index on `(tenant_id, created_at)` but none on `actor_id`. Any
actor-filtered export will sequential-scan a partition.
- `requireAdmin` does not scope by tenant — a platform admin currently sees all tenants.
Any authorization decision here inherits that gap.
- The billing export writes CSV by hand-rolled string concatenation with no escaping.
Do not copy it.
### System boundaries
Owned by this effort: a read path over `audit_events`, an admin UI surface, an artifact
written to S3, and a delivery notification. Not owned: the write path (`auditEvents.record()`
is untouched), the audit event schema, tenancy semantics in `requireAdmin`.
### Gaps
- Retention policy for `audit_events` is not expressed anywhere in code. Someone outside
engineering owns it.
- No load figures exist for the largest tenant's monthly event count.
**Gist:** Node/Express/Knex/React with an established BullMQ-to-S3 export precedent from
billing; `audit_events` is 180M rows partitioned monthly with no `actor_id` index, and
`requireAdmin` has no per-tenant scoping.
## Evidence
- `src/api/admin/index.ts`, `src/api/middleware/auth.ts`
- `src/services/auditEvents.ts`, `src/services/storage.ts`
- `src/jobs/reportQueue.ts` and the billing export job it drives
- `src/db/migrations/20240914_partition_audit_events.js`
- `AGENTS.md` (conventions section)
```
---
## Step 4: The breadth-first frontier grill
The destination grill went **deep on one thing**. This grill goes **wide on everything**. They are different activities and mixing them is the most common way to produce a bad map.
| | Destination grill (Step 2) | Frontier grill (Step 4) |
|---|---|---|
| Goal | One or two settled lines | An inventory of open decisions |
| Movement | Drill until it is precise | Fan out until you stop finding new areas |
| Follow-ups | Chase every hedge | One clarifier at most, then move on |
| Success | The human commits to a boundary | You can name the areas, sharp and unsharp alike |
| Failure mode | Accepting a slogan | Solving a question instead of finding the next one |
**You are not resolving anything here.** You are taking inventory. The instant an answer feels satisfying, you are probably going deep.
### Recognising that you have gone deep
Watch for these. Any one of them means stop and pull back:
- You have asked three consecutive probes about the same area.
- You are discussing an implementation detail (a column type, a library, a retry count) rather than a decision.
- You are proposing a design instead of asking what has to be decided.
- The human is enjoying it. Depth is more fun than breadth; that is exactly why it steals the session.
- You have written something down that reads like an answer.
### Pulling back
Say it out loud so the human tracks the move, then jump:
> "Good — that is one for the map, not for now. Parking it as *Export format*. Different corner: who is allowed to run an export at all?"
Two mechanics keep the fan-out honest:
1. **Round-robin the areas.** Before you start, list the axes you intend to cross: data, surface, permissions, volume, delivery, failure, operations, testing. Take one probe per axis before any second probe on any axis.
2. **Ask for the axis you have not touched.** Near the end: "What have I not asked about that would embarrass us to discover in week three?"
**Breadth-first is the ideal batch.** One probe per axis means the probes are independent by construction — that is what breadth-first *means* — so this grill should run as batches of three, not as a stream of singles. Seven axes is three turns. If you catch yourself wanting to batch two probes on the same axis, that is depth wearing a batch's clothes; pull back.
### Sample breadth probes
Each opens a different axis. Send 1-3 as one batch and 4-6 as the next, then probe 7 alongside the "what have I not asked about" closer above; note each answer and move.
**They go out as numbered Q blocks — this grill is one of the two standing exceptions to the tool-first rule.** Several of the probes do name alternatives, so they would pass the detail test on its own terms, and that is exactly the trap: the output of this grill is not a decision, it is a *sort* into sharp question or fog, and sorting takes the elaboration around the answer. A clicked label leaves you nothing to sort with. The structured question tool earns its keep in MODE B, where a claimed question already has named alternatives and the sorting is long done.
1. **Data** — "What is the smallest and largest thing an operator could reasonably ask for in one export? Give me both ends."
2. **Surface** — "Where does this start: a button in the admin UI, a scheduled thing, an API call someone scripts?"
3. **Permissions** — "Who is allowed to run one, and can they export events about people other than themselves?"
4. **Volume and time** — "If an export takes four minutes, is that fine, bad, or a redesign?"
5. **Delivery and failure** — "The export succeeds but the download link expires before they click it. What should have happened?"
6. **Operations** — "Six months from now someone asks who exported what. Does this feature audit itself?"
7. **Testing** — "What would you need to see pass before you would let this near a customer's compliance data?" *(This one always runs — see the mandatory testing-posture question below.)*
Record each answer as one line in your working notes with an area label. At the end of the grill you will have two piles: lines you can turn into a sharp question, and lines you cannot. The second pile is the fog.
---
## The fog-vs-question test
This is the single most important judgement in the skill. Get it wrong toward questions and the map fills with unanswerable stubs that block the frontier. Get it wrong toward fog and the map has nothing takeable on it.
> **The test is whether you can STATE the question precisely now — not whether you can ANSWER it now.**
- **Write a question file** when the question is already sharp — *even if it is blocked and nobody can act on it yet*. Blocked questions are real questions; they get `Blocked by:` and a `[!]` row and they wait. Blocked is not the same as unformed.
- **Leave it in `## Not yet specified`** when you cannot yet phrase it that sharply. You can see there is something there; you cannot say what is being asked.
**Do not pre-slice the fog.** A fog patch is deliberately coarser than a question. One patch may graduate into three questions, or one, or none once the frontier reaches it. Splitting fog into question-shaped fragments before it is sharp invents a structure that the answers will contradict, and it costs a later session the work of deleting your guesses. Write the patch as loosely as the view allows.
A useful forcing check: **could a different person, reading only this line, know what a good answer looks like?** If yes, it is a question. If they would have to ask you what you meant, it is fog.
### Worked examples
| Candidate | Verdict | Reasoning |
|---|---|---|
| "CSV or JSONL for the export file?" | **Question** | Two named options, one decision, one sitting. A reader knows what an answer looks like. `grill · HITL`. |
| "Which roles may export events about other users?" | **Question**, blocked | Sharp today, but it depends on the tenant-scoping decision. Write it, set `Blocked by:`, mark the row `[!]`. Blockedness never demotes a sharp question to fog. |
| "Something about how big exports behave" | **Fog** | "Big" has no meaning until the volume ceiling lands. You cannot say whether the question is about pagination, streaming, timeouts, or refusal. One line in `## Not yet specified`. |
| "There is probably something about PII redaction" | **Fog** | The area is visible, the question is not. Once Legal answers, this may graduate into *which fields are redacted*, *who configures it*, and *does redaction apply to the actor or the subject* — or into nothing, if the answer is "export raw." Slicing it now guesses all three. |
| "How should we architect the export pipeline?" | **Neither — split it** | No single answer closes it; it bundles at least four decisions (sync vs queued, storage target, artifact lifetime, notification). Ask what the parts are. The sharp parts become questions, the rest becomes fog. A candidate no single answer closes is not a question. |
**What never belongs in `## Not yet specified`:** anything already decided (it is a resolved question with a gist on its row), anything that already has a question file, and anything past the destination (that is out of scope).
---
## Out of scope versus fog
Fog gathers **only toward the destination**. The destination fixes the scope, so work beyond it is not dim — it is *excluded*. It is not fog, and it must never sit in `## Not yet specified`, where a later session would try to graduate it.
**The distinction is SCOPE, not sharpness.** This is the part people get wrong. An out-of-scope item can be perfectly sharp — "should the export push to Splunk on a schedule?" is a crisp question with a crisp answer. It is still out of scope, because the destination said the map ends at an operator-initiated export. Sharpness decides *fog versus question*. Position relative to the destination decides *in scope versus out*.
| | Fog (`## Not yet specified`) | Out of scope (`## Out of scope`) |
|---|---|---|
| Position | Before the destination | Past the destination |
| Why it is not a question | Cannot be phrased sharply yet | Could be phrased perfectly — it just is not ours |
| Future | Graduates into questions as the frontier advances | Never graduates |
| Reopening | Automatic, as answers land | Only if the destination is redrawn — and then as a fresh effort, not a resumption |
| The act | An admission of ignorance | A scoping decision |
Ruling something out of scope is a **scoping act, not a step on the route**. When a question you already created turns out to sit past the destination — mis-scoped in during charting, or exposed by a later answer — set `State: out-of-scope`, mark its row `[-]`, and leave one line in `## Out of scope` giving the gist and the reason, linking the question by name. It does not get an `## Answer` and it is not a decision the route walked.
Watch for **stranded** questions: a live question whose `Blocked by:` names something now out of scope will never unblock. Re-frame its `## Question` to drop the dependency, or rule it out too. Never leave it sitting.
---
## Step 6: Numbering and dependency order
Upstream wayfinder creates every unit first and wires the blocking edges in a **second pass**, because a server-side tracker assigns ids and nothing can reference a sibling until it has one. On `**Backend:** local` that constraint does not exist — **you choose `NN` yourself**, so charting is a **single pass**: decide the order, then write each file complete, `Blocked by:` filled at the moment of writing.
(On `**Backend:** github` the constraint comes back, and so does the two-pass shape. Rules 1, 6, 7, and 8 below still hold — they are about dependency reasoning, not about ids. Rules 2-5, which are about `NN`, are replaced by sub-issue order; see `github-issues.md`.)
The rules:
1. **Sort by dependency before you write anything.** Sketch the edges on paper first: which decisions must land before which others can even be discussed.
2. **Blockers get lower numbers.** If *Row-count ceiling* blocks *Delivery channel*, the ceiling is `02` and delivery is `05`. This makes `Blocked by: 02` readable at a glance and makes "lowest `NN` first" on the frontier a sane traversal order.
3. **`00` is always the codebase context.** Never anything else.
4. **`NN` is never reused and never renumbered.** Not when a question is ruled out of scope, not when one is deleted, not to close a gap in the sequence. Links and `Blocked by:` lines would rot silently. Gaps in the numbering are normal and harmless.
5. **The next number is max + 1**, computed from the directory listing, not from the map.
6. **Only depend on what genuinely gates the question.** A `Blocked by:` chain that is really a preference for reading order strangles the frontier. Ask: could this question be answered — badly but honestly — without the blocker? If yes, it is not blocked.
7. **Cycles are a phrasing bug.** If A blocks B and B blocks A, the two are one decision. Merge them or re-frame one to drop the edge.
8. **Refer by name in prose.** Bare numbers appear only on `Blocked by:` lines.
A worked ordering for the audit-log export map:
| `NN` | Name | Type | Blocked by | Why here |
|---|---|---|---|---|
| `00` | Codebase context | `legwork · AFK` | none | Always first, always resolved |
| `01` | Export format | `grill · HITL` | none | Nothing gates it; it gates the artifact shape |
| `02` | Row-count ceiling | `research · AFK` | none | A fact about the data, independent of every preference |
| `03` | Export authorization | `grill · HITL` | none | Independent axis; can be argued today |
| `04` | Testing posture | `grill · HITL` | none | Mandatory; independent of everything else |
| `05` | Delivery channel | `grill · HITL` | `02` | Synchronous download versus queued link turns entirely on volume |
| `06` | SIEM push connector | `grill · HITL` | none | Created, then immediately ruled out of scope during charting — `State: out-of-scope`, no `## Answer` |
---
## The `map.md` template
Below is a complete, realistic map for the audit-log export effort **partway through MODE B**, after two questions have resolved — it shows every marker in use. At Chart Step 5 the same file carries `**Status:** Charting`, an empty `## Question Checklist`, and no `[x]` row except `00`. Copy the structure exactly, including the HTML comments — they are written **for the next session**, which has no memory of this one.
`**Status:**` is `Charting` while Step 5 and Step 6 run, becomes `Working` at Step 7, and becomes `Cleared` only at handoff. It is how the orchestrator routes a fresh session, so it must be accurate before you stop.
Set `**Confidence:**` honestly at Step 5 and re-score it every time a question resolves. Charting scores are low by construction — that is the point. The Clearing Gate needs every dimension at 18/25 or better. Never inflate to make the gate pass.
Write the four dimension labels **hyphenated exactly as shown**`Requirements-clarity`, `Feasibility-technical`, `Integration-points`, `Risk-assessment` — in the `NN/25` form, with no percent symbol anywhere in the file. The metrics collector scrapes planning documents by regex for a bare dimension word followed by whitespace, a colon, or a pipe and then digits; the hyphen breaks that match. A percent sign or a bare `Requirements 18` would be ingested as a completed planning step's confidence score that no planning step ever produced.
Say once, at Step 5: *"This map lives in gitignored `specs/` — local to you, not shared. `git add -f` it to track it."*
```markdown
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
# Map: audit-log-export
**Status:** Working
**Updated:** 2026-08-03
**Confidence:** Requirements-clarity 18/25 · Feasibility-technical 14/25 · Integration-points 16/25 · Risk-assessment 14/25
<!-- Status: Charting while the map is being built (Chart Steps 1-6) -> Working once the
checklist is indexed (Chart Step 7) -> Cleared only when the Clearing Gate passes.
A fresh session routes on this line, so it must be correct before the session ends. -->
<!-- Confidence: four dimensions, each scored out of 25, re-scored at every resolution.
The Clearing Gate requires all four at 18/25 or better. Score against evidence. -->
## Destination
A locked implementation plan for a compliance officer to export a filtered range of audit
events from the admin UI and receive them as a single downloadable file. The map ends at the
plan, not at shipped code. Continuous streaming to external systems is not on the route.
<!-- Settled at Chart Step 2 and quoted as agreed. Every question is measured against it:
in scope or past the edge, still needed or now moot. Change it only with the human
present — a silent redraw invalidates every scope call already made. -->
## Ground rules
<!-- Standing constraints for every session on this map. Read before choosing a question,
obeyed while resolving it. Nothing here is re-asked. -->
- `AGENTS.md` exists and governs. Its conventions are not re-litigated by any question here.
- One question _file_ per session. `research` questions may run as parallel subagents.
- Grill probes are batched per the grilling playbook — at most three per turn, through the structured question tool unless the detail test forces prose Q blocks.
- Questions are put to the human in plain English. Technical terms only where the term is the decision.
- HITL questions are answered by the human in their own words. Never self-answered.
- No new runtime dependency is assumed without a `research` question backing it.
- Compliance language is reviewed by Dana before anything user-facing is finalised.
- Sketches are throwaway and live only under `pathfinder/sketch-NN/`.
## Glossary
<!-- Terms this effort uses precisely. Prevents two sessions meaning different things by
the same word — the cheapest correctness win on the whole map. -->
| Term | Meaning here | Avoid |
|---|---|---|
| Audit event | One row in `audit_events`: actor, tenant, action, target, timestamp, payload | log line, activity record |
| Compliance officer | Tenant-scoped role that reviews activity; not a platform admin | admin, auditor |
| Export | One operator-initiated request producing one artifact for one filtered range | download, dump, extract |
| Retention window | How far back `audit_events` is queryable; owned outside engineering | archive period |
| Signed artifact | The generated file plus a checksum a recipient can verify independently | signed file, bundle |
## Question Checklist
<!-- Rebuilt from questions/ every session — the files are ground truth, this is an index.
[ ] open (the frontier) · [/] claimed · [x] resolved · [!] open but blocked
[-] out of scope. Resolved rows carry the one-line gist from the question's Answer. -->
- [x] [Codebase context](./questions/00-codebase-context.md) — Node/Express/Knex/React with a BullMQ-to-S3 export precedent; `audit_events` is 180M rows partitioned monthly, no `actor_id` index, `requireAdmin` has no tenant scoping.
- [x] [Export format](./questions/01-export-format.md) — CSV with a UTF-8 BOM and RFC 4180 quoting, plus a sidecar SHA-256 manifest; JSONL rejected because recipients open these in Excel.
- [/] [Row-count ceiling](./questions/02-row-count-ceiling.md)
- [ ] [Export authorization](./questions/03-export-authorization.md)
- [ ] [Testing posture](./questions/04-testing-posture.md)
- [!] [Delivery channel](./questions/05-delivery-channel.md) — Blocked by 02
- [-] [SIEM push connector](./questions/06-siem-push-connector.md) — out of scope, see below
## Not yet specified
<!-- The fog: in-scope areas you can see but cannot yet phrase as a question. Graduates into
question files as answers land, and the graduated bullet is deleted from here.
Do NOT pre-slice these into question-sized pieces — one bullet may become three
questions, or none. Nothing already decided, already a question, or out of scope. -->
- How far back an export may reach. There is a retention answer somewhere outside engineering
and nobody has it yet; until then we cannot say whether the question is about a hard limit,
a warning, or a per-tenant setting.
- Redaction of event payloads. Legal may say "export raw", in which case this evaporates —
or it may become several decisions about which fields, who configures them, and whether the
rule follows the actor or the subject. Revisit after Export authorization.
- What happens when an export range straddles a monthly partition that was migrated mid-range.
Cannot phrase this sharply until Export format is applied to a real query plan.
- Whether the export feature audits itself, and if so at what granularity. Suspect this is one
small question but it may turn on the authorization model.
## Out of scope
<!-- Work consciously ruled past the destination. Never graduates; returns only if the
destination is redrawn, and then as a fresh effort. One line each: gist plus why. -->
- [SIEM push connector](./questions/06-siem-push-connector.md) — scheduled push to Splunk or
similar. The destination ends at an operator-initiated export; anything continuous is a
different effort with a different owner.
- Redesign of the `audit_events` schema. The write path is untouched by this destination;
changing it would pull in all 31 call sites of `auditEvents.record()`.
- Adding the missing `actor_id` index. Real, and it will hurt, but it is a database change
with its own review path. Recorded here so the plan can reference it as a dependency
rather than absorb it.
```
---
## The question-file template
Five contiguous `Key: value` lines after the H1. Not YAML. No frontmatter delimiters. No `- [ ]` checkboxes anywhere inside a question file — use plain bullets, including for legwork checklists.
`## Question` is written at charting. `## Answer` is appended only when the question resolves. `## Evidence` holds sources, links, and artifacts, and a `research` subagent writes into it during Chart Step 8 without deciding anything.
### An open question
```markdown
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
# Export authorization
Type: grill · HITL
State: open
Blocked by: none
Claimed: none
Locked: yes
## Question
Who may run an audit-log export, and over whose events?
Three sub-decisions, all of which must land together because any two of them constrain the third:
- Which role gates the export action — the existing `admin` role, a new `compliance` role, or
a per-tenant grant?
- May an exporter include events where they are the actor, or must self-events be excluded to
keep the export usable as evidence?
- `requireAdmin` currently does not scope by tenant, so a platform admin sees every tenant's
events. Does the export inherit that, or does it enforce a tenant scope the rest of the
admin surface does not?
Recommended answer to react to: a new tenant-scoped `compliance` role; self-events included
but flagged in a column; the export enforces tenant scope even though the surrounding admin
surface does not.
Marked `Locked: yes` — the third sub-decision creates a precedent that later admin features
will follow, and reversing it later means re-auditing every export already delivered.
## Evidence
- `src/api/middleware/auth.ts``requireAdmin` checks session and role, no tenant predicate.
- Codebase context records the same gap under technical debt.
```
### A resolved question
```markdown
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
# Export format
Type: grill · HITL
State: resolved
Blocked by: none
Claimed: 2026-08-03 10:41
Locked: yes
## Question
What file format does an export produce, and what does a recipient need in order to trust the
file has not been altered?
## Answer
**Decision.** CSV, UTF-8 with a byte-order mark, RFC 4180 quoting, one header row, timestamps
in ISO 8601 UTC. Alongside it a sidecar `.sha256` manifest listing the artifact filename and
its digest.
**Rejected.**
- *JSONL* — better for nested payloads and trivially streamable, but every named recipient
opens these in Excel and would need a conversion step during an audit. The people who
prefer JSONL are not the people receiving the file.
- *XLSX* — solves the Excel encoding problems outright, but adds a generation library and
makes byte-level verification of the artifact meaningfully harder.
- *Detached signature instead of a checksum* — real integrity guarantees, but requires key
management nobody has scoped, and no recipient has asked to verify a signature.
**Consequences.**
- Nested `payload` is flattened to one JSON string column. Anyone needing structure parses
that column.
- The BOM is required or Excel mangles non-ASCII actor names. This must be an explicit test.
- Do not reuse the billing export's CSV writer — it concatenates strings with no escaping.
A quoting-correct writer is now in scope for the plan.
- The checksum makes the artifact self-verifying, which lets Delivery channel consider a
short-lived link without weakening the integrity story.
**Gist:** CSV with a UTF-8 BOM and RFC 4180 quoting, plus a sidecar SHA-256 manifest; JSONL
rejected because recipients open these in Excel.
## Evidence
- RFC 4180, sections 2.5-2.7 (quoting and embedded delimiters).
- `src/jobs/billingExport.ts` — the hand-rolled writer that must not be copied.
- Dana confirmed on 2026-08-03 that external auditors accept a published checksum.
```
---
## Step 4: The no-fog off-ramp
If the breadth-first grill surfaces **no fog** — every area you opened produced either a settled answer or a question you could phrase immediately, and `## Not yet specified` would be empty — then the way to the destination is already visible. The journey is small enough to plan directly and a map would be pure overhead.
Do this, in order:
1. **Keep `questions/00-codebase-context.md`.** This is the important part. It is exactly what `/plan2code-1-plan` Phase 2 (System Context Examination) has to produce anyway, and you have already produced it. Throwing the recon away to "clean up" wastes the most valuable artifact of the session.
2. **Do not create `map.md`.** The off-ramp fires at Step 4 and the map is not written until Step 5, so in the normal flow there is nothing to delete — just stop before writing it. (If you reached Step 4 with a `map.md` already on disk, delete it: a map with empty fog and no open questions will confuse the next session into resuming something that does not exist.)
3. **Tell the user plainly what happened and what to do:**
> "No fog surfaced — the way to the destination is already visible, so this does not need a map. I kept the codebase recon at `specs/audit-log-export/pathfinder/questions/00-codebase-context.md`; attach that file to a `/plan2code-1-plan` session and it covers Phase 2 outright."
4. **Stop.** Do not chart anyway "just in case", do not create questions, do not start planning in this session.
Be honest about the trigger. If two areas are genuinely unformed, that is fog and the map earns its keep. The off-ramp is for the case where you fanned out across every axis and kept landing on solid ground.
---
## Step 6: The mandatory testing-posture question
**Every map includes a `grill · HITL` question on testing posture.** No exceptions, including maps where testing feels obvious.
The reason is mechanical: `/plan2code-1-plan` Phase 1 asks for three things — testing types, whether tests run after each implementation phase, and a coverage target — and `/plan2code-2-document` **string-matches** that answer, either appending a testing block to every phase, creating a dedicated final testing phase, or omitting testing entirely. A map that clears without this answer hands the human a planning session that stalls on its first question. Ask it while there is still someone in the room.
Record the answer in the literals downstream matches on, not in paraphrase. ("Phase" here names `/plan2code-1-plan`'s implementation phases and is a downstream contract string — it is never a pathfinder unit of work.)
It is almost never blocked. Give it whatever number the dependency ordering leaves free, and expect it to sit on the frontier from day one.
```markdown
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
# Testing posture
Type: grill · HITL
State: open
Blocked by: none
Claimed: none
Locked: no
## Question
What testing does this work carry, so `/plan2code-1-plan` Phase 1 can be answered without
stalling? Three parts, all needed:
- **Types** — unit, integration, end-to-end, some combination, or none.
- **Phase testing** — record one of the two literals `/plan2code-2-document` matches:
`Run after each phase` (a testing block closes every implementation phase) or
`Dedicated phase only` (one final testing phase).
- **Coverage target** — record one of Phase 1's three literals: `Critical paths`,
`Moderate (~60-80%)`, or `Comprehensive (>80%)`.
Recommended answer to react to: unit plus integration; `Run after each phase`;
`Critical paths`. Rationale — this touches compliance data, so the correctness of the
CSV writer and the authorization predicate must be pinned by tests, but the admin UI is thin
enough that end-to-end coverage would cost more than it catches.
Two specific cases worth naming in the answer regardless of the general posture, because
Codebase context shows both are easy to get wrong here:
- The UTF-8 BOM survives Excel round-tripping for non-ASCII actor names.
- The tenant-scope predicate actually excludes other tenants' events, asserted against seeded
cross-tenant data rather than a mock.
Note for whoever resolves this: the coverage target is a number the human owns. Do not infer
it from the codebase's current coverage, and do not soften it to whatever the repo already
achieves.
## Evidence
- `AGENTS.md` records Vitest as the runner with fixtures colocated under `test/`.
- Codebase context: no end-to-end harness exists today; adding one is a real cost, not a flag.
```
@@ -0,0 +1,449 @@
# GitHub Issues Backend
> Part of plan2code-0-pathfinder — loaded at the top of EVERY session whose map lives on GitHub Issues. It re-expresses the local-file model in issue terms: where the map lives, where a question lives, how blocking, claiming, and resolving are done, and what stays on local disk regardless.
>
> **Local-file maps never load this file.** If `## Ground rules` says `**Backend:** local`, close it and use `questions.md`.
Everything the skill says about *judgement* is unchanged by the backend: the fog-vs-question test, the destination grill, one question per session, HITL is never self-answered, the Clearing Gate rubric. This file changes only *where the bytes go*.
---
## Why a second backend exists
Local files are private scratch — `specs/` is gitignored, so the map is yours alone and nobody else can see it, comment on it, or resolve a question in parallel. That is the right default for a solo effort.
A map on GitHub Issues buys three things local files cannot:
| | Local files | GitHub Issues |
|---|---|---|
| Visibility | One machine, one person | Anyone with repo access, in a UI they already have open |
| Blocking | A `Blocked by:` line only an agent reads | Native issue dependencies — GitHub greys out blocked issues in its own UI |
| Concurrency | One session at a time by construction | Several people can work unblocked questions at once; the assignee is a real lock |
It costs three things too, and the human must know all three before they pick it:
1. **Issues on a public repo are public.** The destination, the rejected alternatives, the codebase recon, the technical debt in the blast radius — all of it is world-readable the moment it is written. Never chart to a public repo's tracker anything that would embarrass the project or leak a customer.
2. **It writes to shared state.** A local map costs nothing to abandon. Twelve stale issues labelled `pathfinder:grill-hitl` on a team's tracker is litter someone has to clean.
3. **It needs `gh`, auth, and issues enabled.** More that can break, in a step whose whole job is to remove friction.
---
## Preflight — before the first write
Run these once, at Chart Step 1, **before** offering GitHub as an option. Any failure means GitHub is not offered at all; say why in one line and continue with local files.
| # | Check | Command | On failure |
|---|---|---|---|
| 1 | `gh` is installed | `gh --version` | Not offered — "no `gh` on this machine" |
| 2 | Authenticated | `gh auth status` | Not offered — "`gh` is not logged in" |
| 3 | Inside a repo with a GitHub remote | `gh repo view --json nameWithOwner,visibility,hasIssuesEnabled` | Not offered — "no GitHub remote here" |
| 4 | Issues are enabled | same call, `hasIssuesEnabled` | Not offered — "issues are disabled on this repo" |
| 5 | Write access | `gh api repos/<owner>/<repo> --jq .permissions.push` | Not offered — read-only access cannot chart |
Record `nameWithOwner` and `visibility` from check 3 — **`visibility` is not optional detail.** If it is `PUBLIC`, the offer must say so in the same breath, e.g. *"GitHub Issues — note `jparkerweb/plan2code` is public, so the whole map is world-readable."*
Once a map exists, preflight shrinks to checks 1 and 2. A session that cannot reach `gh` cannot work a GitHub map: say so and stop, rather than silently starting a local one.
### Labels
Create the label set at **Chart Step 5**, with the map issue — never during preflight, which must stay read-only until the human has actually picked `github`. `--force` makes it idempotent, so it is safe to re-run every session:
```bash
gh label create "pathfinder:map" --color 5319E7 --description "Pathfinder map" --force
gh label create "pathfinder:grill-hitl" --color 1D76DB --description "Decision only the human can make" --force
gh label create "pathfinder:research-afk" --color 0E8A16 --description "Fact-finding, agent alone" --force
gh label create "pathfinder:sketch-hitl" --color FBCA04 --description "Human reacts to something concrete" --force
gh label create "pathfinder:legwork-hitl" --color D93F0B --description "Manual work needing a human" --force
gh label create "pathfinder:legwork-afk" --color D93F0B --description "Manual work the agent can do" --force
gh label create "pathfinder:locked" --color B60205 --description "Hard to reverse; consequences recorded" --force
gh label create "pathfinder:out-of-scope" --color CFD3D7 --description "Ruled past the destination" --force
```
**Type and mode share one label**`grill-hitl`, not `grill` plus `hitl` — for the same reason the local `Type:` line is one token: two labels can drift apart, and a `research` question that has quietly become HITL is a question nobody is driving.
---
## The equivalence table
This is the whole mapping. Everything below expands a row.
| Local file model | GitHub Issues model |
|---|---|
| `specs/<idea>/pathfinder/map.md` | One issue, labelled `pathfinder:map`, titled `Map: <idea>` |
| `questions/NN-<slug>.md` | A **sub-issue** of the map, titled with the question name |
| `NN` ordering | The map's sub-issue order — the order they were created, which is dependency order |
| `## Question` in the file | The issue body |
| `## Answer` appended | A comment on the issue, opening `## Answer` |
| `## Evidence` | A comment opening `## Evidence` (a research subagent writes its own) |
| `Type:` line | The `pathfinder:<type>-<mode>` label |
| `State: open` | Issue open, **no assignee** |
| `State: claimed` | Issue open, **assigned** |
| `State: resolved` | Issue **closed as completed**, with an `## Answer` comment |
| `State: out-of-scope` | Issue **closed as not planned**, labelled `pathfinder:out-of-scope`, no `## Answer` |
| `Blocked by: 02, 04` | Native issue dependencies (`dependencies/blocked_by`) |
| `Locked: yes` | The `pathfinder:locked` label |
| `Claimed: <timestamp>` | GitHub's own assignment event in the timeline |
| `## Question Checklist` in `map.md` | **Nothing** — the frontier is a live query, not a written list |
| `## Not yet specified`, `## Out of scope`, `## Ground rules`, `## Destination`, `## Glossary` | The same sections, in the map issue body |
| `sketch-NN/` | Still local disk — see *What stays on local disk* |
| `PLAN-DRAFT-<YYYYMMDD>.md` | Still local disk — see *Handoff* |
### The checklist is deleted, not ported
In local mode `map.md` carries a `## Question Checklist` because a directory of files has no queryable state. GitHub has queryable state, so **the map issue body carries no checklist at all.** Closed questions get one line each under `## Decisions so far`; open questions are not listed anywhere.
This kills the single largest source of drift in the local backend — a checklist that disagrees with the files — and it is why Work Step 2's reconcile pass is much shorter here.
---
## Refer by name
Unchanged, and harder to get right here because GitHub hands you a number for everything. In prose the human reads, write `[Export format](https://github.com/o/r/issues/42)` — never `#42`, never "issue 42", never a bare number. A wall of `#42, #43, #44` is illegible; names read at a glance.
Bare `#<n>` appears in exactly two places: inside a fallback `Blocked by:` body line when native dependencies are unavailable, and inside a `gh` command.
---
## Chart Steps 3-4 — hold the recon, protect the off-ramp
In `local` mode Step 3 writes `questions/00-codebase-context.md` the moment the recon is done, because a file in gitignored scratch costs nothing if the session then takes the Step 4 off-ramp. **On a shared tracker it costs something**: a stray issue nobody asked for, on a repo other people are watching.
So on `github`, Step 3 does the recon and **holds it in the session**. It becomes an issue at Step 6, alongside the other questions.
If the Step 4 breadth-first grill surfaces **no fog**, the off-ramp fires before anything has been created:
1. Write the recon to `specs/<idea>/pathfinder/questions/00-codebase-context.md` — a **local file**, exactly as the local backend would. It is what `/plan2code-1-plan` Phase 2 needs, and it is too valuable to throw away.
2. Create **nothing** on the tracker. No map issue, no question issues, no labels.
3. Tell the user plainly and STOP.
The tracker only ever sees an effort that earned a map.
---
## Creating the map (Chart Step 5)
Title is `Map: <idea>` — the kebab-case idea name, verbatim, so `gh issue list --label pathfinder:map` reads as an index of efforts.
```bash
gh issue create --label "pathfinder:map" --title "Map: audit-log-export" --body-file - <<'EOF'
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
**Status:** Charting
**Updated:** 2026-08-08
**Confidence:** Requirements-clarity 8/25 · Feasibility-technical 6/25 · Integration-points 6/25 · Risk-assessment 5/25
<!-- Status: Charting (Chart Steps 5-6) -> Working (Chart Step 7) -> Cleared at the gate.
A fresh session routes on this line, so it must be correct before the session ends. -->
## Destination
A locked implementation plan for a compliance officer to export a filtered range of audit
events from the admin UI and receive them as a single downloadable file. The map ends at the
plan, not at shipped code. Continuous streaming to external systems is not on the route.
## Ground rules
- **Backend:** github — this issue is the map; questions are its sub-issues.
- `AGENTS.md` exists and governs. Its conventions are not re-litigated by any question here.
- One question _issue_ per session. `research` questions may run as parallel subagents.
- HITL questions are answered by the human in their own words. Never self-answered.
- Sketches are throwaway and live on local disk only, under `specs/audit-log-export/pathfinder/sketch-<issue>/`.
## Glossary
| Term | Meaning here | Avoid |
|---|---|---|
| Audit event | One row in `audit_events`: actor, tenant, action, target, timestamp, payload | log line |
## Decisions so far
<!-- The index — one line per CLOSED question: enough to judge relevance, then open the
issue for the detail. Open questions are NOT listed; they are open sub-issues. -->
## Not yet specified
<!-- The fog: in-scope areas you can see but cannot yet phrase as a question. Graduates into
sub-issues as answers land, and the graduated bullet is deleted from here. -->
## Out of scope
<!-- Work consciously ruled past the destination. Never graduates. One line each: gist plus why. -->
EOF
```
Two things that must be exact:
- **`**Status:**` is still a literal line in the body.** It is how a fresh session routes, exactly as in local mode. `Charting``Working``Cleared`.
- **`**Backend:** github` is the first `## Ground rules` bullet.** It is how a fresh session knows to load this file at all. Without it, a session that opens the map issue has no way to know which playbook it is in.
Confidence keeps the hyphenated `Requirements-clarity 8/25` form for the same reason it does in local mode — the metrics collector scrapes bare dimension words followed by digits, and would ingest a planning confidence nobody scored.
**Say once, at Step 5:** *"The map lives on `<owner>/<repo>`'s issue tracker — `<PUBLIC or PRIVATE>`, so `<world-readable / visible to anyone with repo access>`. Everything charted here is visible there."*
---
## Creating the questions (Chart Step 6) — two passes, not one
The local backend charts in a **single pass** because you choose `NN` yourself and can write `Blocked by: 02` into a file before `02` exists. **On GitHub that is impossible** — an issue has no id until the server assigns one, and a dependency edge needs the blocker's id. So charting here reverts to upstream's shape:
**Pass 1 — create every question issue, in dependency order.** Blockers first. The creation order becomes the sub-issue order, which becomes the reading order for the frontier and the trail, so it is doing the job `NN` does locally. Capture each new issue's number *and* database id as you go.
```bash
# Create, capturing the URL; the number is its last path segment.
gh issue create --label "pathfinder:grill-hitl" --title "Export format" --body-file - <<'EOF'
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
## Question
What file format does an export produce, and what does a recipient need in order to trust
the file has not been altered?
Recommended answer to react to: CSV with a UTF-8 BOM plus a sidecar SHA-256 manifest.
EOF
# The database id — needed for BOTH wiring steps below. Not the #number, not the node_id.
gh api repos/<owner>/<repo>/issues/<number> --jq .id
```
**Pass 2 — wire the structure.** Two edges per question, both keyed on **database ids**:
```bash
# a) Attach as a sub-issue of the map. sub_issue_id is the CHILD's database id.
gh api --method POST repos/<owner>/<repo>/issues/<map-number>/sub_issues \
-F sub_issue_id=<child-db-id>
# b) Add each blocking edge. issue_id is the BLOCKER's database id.
gh api --method POST repos/<owner>/<repo>/issues/<blocked-number>/dependencies/blocked_by \
-F issue_id=<blocker-db-id>
```
**The database id is the single most common failure in this backend.** `gh api repos/o/r/issues/42 --jq .id` returns something like `2716143027`. The `42` is the *number*; `I_kwDO...` is the *node id*. The node id is rejected outright. The *number* is worse: a small integer like `42` is itself a perfectly valid database id — of some unrelated issue created years ago — so the call can succeed and silently attach the wrong thing. Fetch `.id` for every issue you are about to reference, and never hand-assemble one.
Charting still writes `00-codebase-context`'s equivalent — the Step 3 recon you held. Create it **first**, labelled `pathfinder:legwork-afk`, post the recon as an `## Answer` comment, and close it as completed in the same pass. It is resolved on arrival, exactly as in local mode, and it is what the handoff's `## System Context` is built from.
### If the endpoints are unavailable
Sub-issues and dependencies are recent GitHub features. On an instance that rejects either endpoint, fall back in the body — and say plainly, once, that the frontier will not render in GitHub's UI:
| Missing | Fallback |
|---|---|
| Sub-issues | Put `Part of #<map>` on the first line of each question body, and a task list of the questions in the map body |
| Dependencies | Put `Blocked by: #12, #14` on its own line at the top of the question body |
Prefer the native mechanisms every time they work. The whole reason to pay GitHub's costs is that the human sees the frontier in the UI without opening the map.
---
## The frontier query (Work Step 3)
The frontier is every question that is **open, unblocked, and unassigned**. Lowest position in sub-issue order wins — the same traversal `lowest NN first` gives locally.
```bash
# 1. The map's children, in order, with the state you need to filter on.
gh api repos/<owner>/<repo>/issues/<map-number>/sub_issues \
--jq '.[] | {number, title, state, assignee: .assignee.login,
blocked: .issue_dependencies_summary.blocked_by,
labels: [.labels[].name]}'
```
Then, in order:
1. Drop anything `state: closed` — that is resolved or out of scope.
2. Drop anything with an `assignee` — claimed by another session.
3. Drop anything still blocked.
4. The first survivor is the next question.
For step 3, `issue_dependencies_summary.blocked_by` counts **open** blockers, which is exactly the live gate — a blocker that closes drops the count without anyone editing anything. **Treat it as a fast pre-filter, not the authority.** If it comes back `null` or absent from the list response, or you need to *name* the blockers for the trail footer or a fully-blocked report, ask the endpoint that owns the answer:
```bash
gh api repos/<owner>/<repo>/issues/<n>/dependencies/blocked_by --jq '.[] | {number, title, state, reason: .state_reason}'
```
A question is unblocked when every blocker listed there is closed. That call is also the only way to see the next trap:
**A blocker closed as `not planned` is out of scope and will never resolve.** Its dependent is not merely blocked, it is *stranded* — the same trap as locally. Re-frame the dependent's body to drop the dependency, cut the edge, or rule it out too. Never leave it sitting: the summary count cannot tell you the difference, so this check is on you.
```bash
# Cut a dependency edge. The blocker's database id goes in the PATH here, not the body.
gh api --method DELETE \
repos/<owner>/<repo>/issues/<blocked-number>/dependencies/blocked_by/<blocker-db-id>
```
---
## Claim (Work Step 4)
```bash
gh issue edit <n> --add-assignee "@me"
```
**The session's first write, before any work.** The assignee *is* the claim — GitHub timestamps it for you, so there is no `Claimed:` line to maintain. An open, unassigned question is unclaimed; that is the whole protocol.
Unlike the local backend, other people may genuinely be working this map at the same time. Re-read the issue immediately after assigning; if someone else's login is on it, you lost the race — release yours and take the next frontier item.
---
## Resolve (Work Step 7)
Three writes, in this order. The order matters: the answer must exist before the issue closes, or a crash between them leaves a closed question with no decision in it.
```bash
# 1. The answer, as a comment. Same anatomy as a local ## Answer:
# decision, rejected alternatives with reasons, consequences, one-line **Gist:**.
gh issue comment <n> --body-file - <<'EOF'
## Answer
**Decision.** CSV, UTF-8 with a byte-order mark, RFC 4180 quoting, one header row.
Alongside it a sidecar `.sha256` manifest.
**Rejected.**
- *JSONL* — trivially streamable, but every named recipient opens these in Excel.
- *XLSX* — fixes Excel encoding, but adds a library and makes byte-level verification harder.
**Consequences.**
- Do not reuse the billing export's CSV writer; it concatenates strings with no escaping.
- Makes [Delivery channel](https://github.com/o/r/issues/45) sharper — the artifact is
self-verifying, so a short-lived link no longer weakens the integrity story.
**Gist:** CSV with a UTF-8 BOM and RFC 4180 quoting, plus a sidecar SHA-256 manifest;
JSONL rejected because recipients open these in Excel.
EOF
# 2. Close as completed.
gh issue close <n> --reason completed
# 3. Append the gist to the map's Decisions so far (read body, edit, write back).
gh issue view <map-number> --json body --jq .body > /tmp/map.md
# ...append: - [Export format](<issue-url>) — <gist>
gh issue edit <map-number> --body-file /tmp/map.md
```
Then bump `**Updated:**` and re-score `**Confidence:**` in the same map edit.
**Never edit the question body to hold the answer.** The body is the question as asked; the comment is the answer. Editing the body rewrites history and destroys the record of what was actually put to the human — which is half of why the answer is defensible three weeks later.
### Editing the map body safely
Every map mutation is read-modify-write on a body other sessions may be editing concurrently. Read it fresh immediately before the edit, apply your change to *that* text, and write it straight back. Never edit from a copy you read at the top of the session — you will silently revert whatever landed in between.
---
## Ruling a question out of scope (Work Step 8)
```bash
gh issue edit <n> --add-label "pathfinder:out-of-scope"
gh issue close <n> --reason "not planned"
```
Then one line under the map's `## Out of scope`, giving the name as a link plus the reason. **No `## Answer` comment** — there is no decision here, only a scope boundary. Add one comment saying why it is out, so the closed issue explains itself.
`not planned` versus `completed` is the load-bearing distinction: it is how a later session tells a decision that was made from a question that was ruled off the route, and it is what GitHub's UI shows at a glance. Getting it backwards puts a scope boundary into the Provenance table of the plan.
---
## Reconcile (Work Step 2)
Much shorter here — the tracker holds the state, so there is no checklist to rebuild. Three checks:
| Check | Symptom | Repair |
|---|---|---|
| Crashed mid-answer | Open, assigned, and an `## Answer` comment already exists | The comment wins. Close as completed, add the gist to Decisions so far, say so. |
| Stale claim | Open, assigned, no `## Answer`, and the assignee is you from a dead session | Unassign, say so, put it back on the frontier. **If it is someone else's login, leave it** — that is a live session, not a crash. |
| Index drift | A closed, completed question with no line under `## Decisions so far` | Read its `## Answer` comment, append the gist. |
Then re-read `## Not yet specified` in full — that part is identical to local mode, and the bullet left behind after its question exists is just as corrosive here.
---
## What stays on local disk
Three things never move to the tracker, whatever the backend:
| Artifact | Where | Why |
|---|---|---|
| Runnable sketches | `specs/<idea>/pathfinder/sketch-<issue-number>/` | Throwaway code has no business in an issue, and the quarantine rule (never in the project's own source tree) is unchanged. Link the path from the issue and note the reader needs the repo checked out. |
| `PLAN-DRAFT-<YYYYMMDD>.md` | `specs/<idea>/` | `/plan2code-1-plan` reads a **file**. This is a hard downstream contract — see Handoff. |
| Anything secret | Nowhere | Credentials, tokens, customer data. A `legwork` checklist says *where* a credential lives, never what it is — and on a public tracker that rule stops being a convention and starts being an incident. |
Sketch directories are named for the issue number rather than a local `NN`, so `sketch-42` belongs to the question at `#42`. Same rules otherwise: throwaway, one command to run, never merged.
Research subagents work the same way with one substitution: the brief carries the **issue URL** instead of a file path, and the instruction is to post findings as a comment opening `## Evidence` via `gh issue comment` — and to decide nothing. `## Answer` and the close are still written by the session that fired it.
---
## The trail footer
Identical in shape; the inputs come from the query instead of the checklist.
- **Heading** — `🧭 <idea> · <Status> · <closed>/<total> cleared`, where `<total>` excludes anything labelled `pathfinder:out-of-scope`.
- **Glyph order** — sub-issue order, the same order Chart Step 6 created them in.
- **Glyphs** — `●` closed as completed · `◉` open and assigned to you · `○` open, unassigned, unblocked · `⊘` open with `blocked_by > 0` · `⊝` closed as not planned.
- **Named legend** — names, never `#numbers`. `(blocked:<name>)` names the blocker rather than numbering it, since there is no stable `NN` to point at.
- **Confidence** — the plain-English line, from the map body's `**Confidence:**`.
Form A's resume command changes, because there is no local path to resume from:
```
NEXT STEP · start a new conversation and run:
`/plan2code-0-pathfinder https://github.com/<owner>/<repo>/issues/<map-number>`
```
Form B is unchanged — a turn that asks the human something still says `WAITING ON YOU`, still names the outstanding probes, and still emits no resume command.
---
## Handoff (The Clearing Gate)
The gate's four dimensions, the 18/25 bar, the hard caps, and the honesty rules are unchanged. Only the preflight and the plumbing differ.
**Preflight, GitHub form:**
| # | Check |
|---|---|
| 1 | Reconcile pass run (above) |
| 2 | Zero open sub-issues — `gh api .../sub_issues --jq '[.[] \| select(.state=="open")] \| length'` returns `0` |
| 3 | `## Not yet specified` in the map body is empty |
| 4 | Every completed question has an `## Answer` comment carrying a `**Gist:**` |
| 5 | `ls specs/<idea>/` shows no existing `PLAN-DRAFT-*.md` |
| 6 | No `## Answer` defers a choice to "whoever implements this" |
**The draft is written to local disk**, at `specs/<idea>/PLAN-DRAFT-<YYYYMMDD>.md`, with the byte-exact status line `**Status:** Phase 3 Complete - Resume at Phase 4`. This is not a preference. `/plan2code-1-plan` discovers its input with `ls specs/`; it has no notion of an issue tracker, and a draft that exists only as an issue is a draft the rest of plan2code cannot see. Create `specs/<idea>/` if charting never needed it.
Three substitutions inside the template:
| Local | GitHub |
|---|---|
| `**Planning record:** specs/<idea>/pathfinder/map.md` | `**Planning record:** <map issue URL>` |
| `[Export format](./pathfinder/questions/01-export-format.md)` | `[Export format](https://github.com/o/r/issues/42)` |
| `## Out of scope` copied line for line, only the link prefix changing | Copied line for line, links already absolute — **nothing changes at all** |
Everything else — the mapping table, Section 5 left empty, the two load-bearing `**Next:**` bullets, no scrapable confidence numbers — is unchanged.
**Freeze the map:**
1. Set `**Status:** Cleared` in the map issue body.
2. Bump `**Updated:**`.
3. Add `**Plan:** specs/<idea>/PLAN-DRAFT-<YYYYMMDD>.md` under the status line.
4. `gh issue close <map-number> --reason completed`.
5. **Close nothing else, delete nothing, edit no answers.** Every question issue stays exactly as it is — it is the rationale record behind the plan.
Closing the map is the one addition over local mode, and it earns its place: `gh issue list --label pathfinder:map --state open` then reads as *the efforts still being charted*, which is the question a person scanning the tracker actually has.
A later session that opens a map issue reading `Status: Cleared` must not resume it. Point at the PLAN-DRAFT and `/plan2code-1-plan`, and stop. A redrawn destination is a fresh effort with a fresh map issue.
---
## Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| `Not Found` from a `sub_issues` or `dependencies` POST | An issue *number* was passed where a database *id* is required | `gh api repos/o/r/issues/<n> --jq .id`, retry with that |
| The wrong issue got attached | A number from another repo happened to be a valid id | Detach, re-fetch `.id` from the right repo, re-attach |
| Frontier is empty but open questions remain | Every one is blocked, or every one is assigned | Name the chain and stop — or, if the assignees are stale claims from your own dead sessions, reconcile first |
| A blocked question never unblocks | Its blocker was closed as `not planned` | Stranded. Re-frame to drop the dependency and cut the edge, or rule it out too |
| Two sessions resolved the same question | The claim was written after the work, not before | Claim is the *first* write. Merge the two answers into one comment, keep one close |
| A map edit lost someone's line | The body was edited from a stale copy read earlier in the session | Re-read the body immediately before every map write |
| `/plan2code-1-plan` finds nothing to resume | The PLAN-DRAFT was posted as an issue instead of written to `specs/<idea>/` | Write the file. The draft is always local |
| Secrets in the tracker | A `legwork` answer pasted a credential | Rotate the credential first, then delete the comment. Editing it is not enough — GitHub keeps the edit history |
@@ -0,0 +1,402 @@
# Grilling Playbook
> Part of plan2code-0-pathfinder — loaded once per session, used by Chart and Work alike.
Grilling is how a `grill · HITL` question resolves, how Chart Step 2 names the destination, and how Chart Step 4 maps the frontier. It is also the fallback for any question whose type gives you no better route. The output of a grill is a decision in the human's own words — never a decision you made on their behalf.
**Probe ≠ question file.** A *probe* is one turn of the interrogation; a *question file* is one `questions/NN-<slug>.md` on the map. Batching applies to probes only. **One question file per session still holds** — resolving three question files in one sitting is not what this is.
## The interview protocol
**Grilling is batched.** Put up to **three independent probes** to the human per turn, then wait. Never more than three, and never two probes where one's wording depends on the other's answer.
The old rule here was one probe per turn. It was safe and it was unusably slow: a charted map carries a dozen open questions, and a decision that costs one round trip per probe is a decision the human abandons half-finished. Batching is the default now; the discipline moved from *ask one* to *prove they are independent, then send three*.
| Rule | Why |
|---|---|
| Up to 3 probes per turn, never more | Past three the human skims, and a skimmed answer is worse than none. Three is a ceiling, not a quota — send two if only two are independent. |
| Only batch mutually independent probes | The independence test below. A probe whose wording or recommendation shifts based on another probe's answer waits for the next turn. |
| Reach for the structured question tool first | It is the intended channel, not the leftover bin. Shape the batch so it fits — three probes, plain headers, options a description can carry — and fall back to prose only when the detail test genuinely trips. |
| A probe that needs detail goes in prose, never in options | The detail test below. Batching buys round trips; it must never buy them by shrinking a decision to fit a picker. |
| Wait for the whole batch before sending the next | Their answers reshape what comes next. Pre-writing turn 2 wastes it. |
| Recommend an answer with every probe | A bare question makes the human do all the work. A recommendation gives them something to push against, which is faster and sharper. |
| Write it in plain English; keep the technical word only where that word IS the decision | A probe the human has to decode is a probe they answer approximately. See *Say it in plain English*. |
| Track the batch; re-ask what came back unanswered | This is what buys the batch. Dropped probes going unnoticed was the entire case for asking one at a time. |
| Walk one branch of the decision tree at a time | Batch across the branch's width, never down its depth. Do not ask about export scheduling before you know whether exports exist. |
| Do not act until they confirm shared understanding | Recap, get the confirmation, then write. No `## Answer` before the confirmation. |
| Never write implementation code during a grill | Grilling produces decisions. If you feel the pull to build, you have reached the edge of the map — say so and hand off. |
Shape of a single probe, batched or not:
```
Question — one decision, stated so it can be answered in a sentence.
Why it — one line: what it blocks, what breaks if it goes the other way.
matters
Recommend — your pick, with the reason. One line.
Options — the genuine alternatives, if there are more than two.
```
Worked probe, from a grill on `[Export format](./questions/04-export-format.md)`:
> **Q:** When a custodian exports a conversation that includes a 40 MB video attachment, does the export bundle the file or link to it?
> **Why it matters:** Bundling sets the size ceiling on an export job and decides whether exports can stream; linking makes the export useless once retention expires the blob.
> **Recommendation:** Bundle, with a per-job size cap of 2 GB and an automatic split into part files above that — reviewers open exports offline in tools that cannot follow links.
> **Alternatives:** Link-only (smaller, breaks offline); hybrid by MIME type (two code paths, two failure modes).
Batching does not shrink a probe. Three probes means three of these, each with its own why-it-matters and its own recommendation. Three bare questions in a list is not a batch, it is a form to fill in.
### Say it in plain English
Every probe gets read once, by a busy human, in a terminal. Write it the way you would say it out loud to a colleague who knows the product but has never read this skill. Plain is not vague — the two failures are opposite and both cost you the decision: woolly wording gets a woolly answer, dense wording gets a guessed one.
**Default to everyday words.** Short sentences, concrete nouns. "What should happen when the export is too big to email?" beats "what are the failure semantics of the artifact delivery path under a size-limit violation?" Same decision, one of them answerable on the first read.
**Keep the technical term where that term IS the decision.** A format name, a real file path, a column name, a version number, a limit, a `## Glossary` term the map already settled — those are load-bearing, and softening them makes the probe unanswerable. `.eml`-in-a-ZIP versus NDJSON *is* the choice; "normal email files versus one big machine-readable stream" is the gloss you put beside it, never the replacement for it.
> The test: would swapping the term for a plain phrase lose information the human needs in order to choose? Lose information — keep the term and gloss it. Lose nothing — cut it.
**Gloss an unavoidable term once, inline, then use it freely:** "…stored under Object Lock (S3's write-once mode — once it is set, even we cannot delete early)." Once per session, not once per probe. Re-explaining a term to the person who owns the system is its own insult.
**Never put Pathfinder's machinery in front of the human.** They are deciding something about their product; the vocabulary below is internal bookkeeping and buys them nothing:
| Do not say | Say |
|---|---|
| "this is a `grill · HITL`" | "this one is yours to call" |
| "the frontier holds two takeable questions" | "two things we can decide right now" |
| "graduating this out of the fog" | "this is sharp enough to write down as a real question now" |
| "the sacrificial boundary" | "name one thing people would assume is included that you are willing to cut" |
| "shall I set `Locked: yes`?" | "worth recording why we picked this, so nobody re-opens it in six months?" |
| "Q3 is blocked by 02" | "the export format question has to land before this one" |
| "this batch trips the detail test" | nothing — that call is yours, not theirs |
**Refer to questions by name, never by number** — "[Export format](./questions/04-export-format.md)", not "04". The number means something to the file system and to nobody else.
**No metaphor where a fact fits.** Maps, fog, and trails belong in the footer and the mascot. Inside a probe they cost a translation step: "three things here are still undecided" beats "the fog is thick in this quarter of the map."
The same discipline covers everything else the human reads — the recap turn in *Landing the grill*, the option labels and descriptions in the structured tool, the sketch probes, and the HITL checklists in the resolution playbook. Plain in the question, precise in the term that carries the decision.
### The independence test
A probe may join the current batch only if **all three** hold:
| Test | Fails when |
|---|---|
| Its wording would not change under any answer to another probe in the batch | "How do we name the part files?" reads differently if the format turns out to be a single stream |
| Its recommendation would not change either | You would recommend a 2 GB cap under ZIP parts and no cap under NDJSON |
| It does not presuppose another probe's answer | "How often do scheduled exports run?" assumes scheduled exports exist |
In doubt, hold it back. A held probe costs one extra round trip. A dependent probe sent early costs a wrong answer recorded as a decision, and you will not find out until the plan contradicts itself.
Independent probes are usually the ones that came from **different areas** — data, interface, security, operations, testing. Dependent probes are usually consecutive steps down one thread.
### Delivering a batch: choosing the channel
Two channels — the environment's structured question tool, or numbered Q blocks in prose. **Choose before you write a word, and choose per batch, not per probe.** One channel per turn: a batch split across a tool popup and a loose prose question loses the prose half every time, because the human answers in the tool and never scrolls back.
**The tool is where you start.** Assemble the batch for it — three probes, a plain two-or-three-word header each, alternatives a sentence or two of description can carry — and only then run the detail test to see whether anything forces you out. Prose is the exception you fall back to, not the safe default you retreat to. Two things make the tool worth the effort: a skipped probe comes back *visibly* skipped, and a picker is answerable in one pass by a human who has thirty seconds. Neither survives the move to prose.
The two failure directions are opposite and both real. Retreating to prose out of caution costs you the visible skip and the fast reply. Forcing a genuinely gnarly decision into a picker costs you the reasoning, which is worse. The detail test below is where that line sits — run it honestly in both directions.
**Whichever channel you pick, the turn closes with the waiting footer.** A turn that sends a batch is a **Form B turn** in `trail.md`: the Trail Footer under it names the outstanding probes after `WAITING ON YOU` and carries **no** resume command. Emitting "start a NEW conversation" above an unanswered batch tells the human to leave the session you are sitting in — they walk, and the batch you built to save round trips costs you the whole decision instead. Same for the recap turn below, which is also waiting on them.
#### The detail test — the only things that force you out of the tool
Numbered Q blocks are **required, not merely permitted**, if *any* probe in the batch trips *any* row below. One tripping probe downgrades the whole batch.
These four rows are the whole list. Nothing else forces prose — not a long question, not a hard decision, not a `Locked: yes`, not your discomfort with the widget.
| Trip | Looks like | Not this |
|---|---|---|
| The answer must be composed, not picked | "Name one thing a reasonable person would assume is in scope that you are willing to cut." There is no option set, because inventing one puts words in their mouth. | A decision with genuine named alternatives, however weighty. Write the options. |
| The probe needs an artifact inline to be answerable | A state table, a fake request/response pair, an ASCII UI, a worked example with real numbers — effectively every `sketch` probe | A probe that merely *mentions* a file path, a format, or a number. Those go in the question text. |
| An option cannot be conveyed even in its description | Each alternative needs a worked paragraph before it means anything — a migration path, a failure sequence, a schema | An alternative that needs one or two sentences of trade-off. That is what the description field is for. |
| The alternatives themselves are unknown to you | You cannot name the losing options at all, because the frame is theirs — a contract, an old incident, an org politics fact | You can name them but cannot say why each loses. Name them, recommend one, and let the recap turn supply the reasoning. |
Three things that look like trips and are not:
- **One label bundling several decisions** — "authentic counts, one-use per attempt, restored on death" is three answers wearing one coat. The fix is to **split it into separate probes**, not to write prose. Three separated probes is exactly one batch.
- **Two probes colliding on the tool's short header limit** (16 characters in Claude Code) — `Export scope` twice is unanswerable, but the fix is to rename them (`Date range`, `Who can run`) or to hold one for the next turn. Reword before you retreat.
- **A hybrid is possible** — the free-text escape hatch takes "the header from B with the list from C" fine. Trip only when you can already predict the answer *will* be a composition, which is the first row.
**`Locked: yes` on its own does not trip the test.** A lock's `## Answer` owes every alternative and the reason each lost — but if *you* can already name the alternatives, you have written the options, and the recap turn turns the pick into words the human said. A lock trips only on the fourth row, where you cannot name them at all. Treating every lock as an automatic downgrade sends almost every MODE B decision worth grilling to prose, which defeats the point — MODE B is exactly where a claimed question already has named alternatives and the tool earns its keep.
**Nothing tripped? Use the structured tool.** Not "may" — do. It is the intended channel, and it is where the visible-skip guarantee behind the partial-answer discipline below comes from.
**Never reshape a probe to fit the tool.** Reaching for the tool first is not licence to shrink a decision into it. The failure mode is not that the tool rejects a gnarly probe — it is that it *accepts* one. You compress a decision with real texture into three tidy options, the human clicks the least-bad one, and you have recorded a decision with no reasoning behind it. That answer cannot satisfy `## Answer`'s obligation to name what was rejected and why, and nobody finds out until handoff, when the PLAN-DRAFT's Architecture section turns out to have nothing to say. Splitting a bundled probe or renaming a colliding header is reshaping the *batch* and is always right. Cutting a real alternative, or thinning a description until the trade-off disappears, is reshaping the *decision* and is always wrong. When the honest choice is between paragraphs and dishonest options, write the paragraphs.
#### The structured tool
The default channel, and the one you build the batch for. One question object per probe, up to three in a single call:
- **Header** — the decision in two or three plain words (`Export format`, `Size cap`). Not a type, not a marker, not a number.
- **Question** — the probe, with its why-it-matters. This is prose and it is not rationed; the same sentences you would have written in a Q block go here.
- **Options** — the genuine alternatives, each described by its trade-off, with the recommended one named as such in its description. Two to four; the free-text escape hatch covers the rest. Label plainly, then let the description carry the precise term: `One file per message` labelling the `.eml`-in-a-ZIP option, with `.eml` named in the description.
A short *label* is not a short *decision*. The label is a handle — `Fixed tick count` — and the description carries the trade-off that makes it choosable. A probe only trips the third detail-test row when even that description cannot hold the option. A label bundling several independent answers is not that row — it is a probe that wants splitting.
**A click is a decision, not a sentence.** The HITL rule wants an `## Answer` traceable to something the human actually said, and a selected option label is thin evidence on its own. What makes tool-delivered answers legitimate is the recap turn in *Landing the grill* — you play the choices back in prose and they confirm or correct in their own words. Never skip the recap on the grounds that the tool already captured the answer; the tool captured the *pick*, and the recap captures the *agreement*.
**If a reply comes back thinner than the decision** — a bare click on something you now realise carries weight — do not paper over it. Fold the why into the recap turn as one more probe before writing the `## Answer`.
#### Numbered Q blocks
The mandatory channel for anything the detail test catches, and the fallback anywhere the structured tool does not exist. Give each probe the room the tool would have denied it.
**Copy this shape exactly.** The blank lines are load-bearing, not decoration:
````markdown
Three independent decisions are open. Answer in any order, skip any you want to punt — "Q2: b, Q3: the hybrid" is a perfectly good reply.
---
**Q1 — Export format**
When a custodian exports a year of a channel, what do they get back?
*Why it matters:* fixes the size ceiling, decides whether exports can stream, and determines whether the review vendors can ingest without a conversion step.
*Recommendation:* **(a)** — both vendors named in Codebase context read `.eml` natively.
- **a)** ZIP of one `.eml` per message, plus `manifest.csv`
- **b)** A single NDJSON stream — compact and streamable, but nobody downstream parses it
- **c)** PST — what Legal asked for by name, but single-writer with a ~50 GB ceiling
---
**Q2 — What's not included**
Name one thing a reasonable person would assume is part of this that you are willing to say is NOT part of it.
*Why it matters:* this becomes the first thing written down as out of scope, and every later "is that in or out?" call is measured against it. A destination nobody has excluded anything from has not been thought about.
*Recommendation:* none — this one is yours. If nothing comes to mind, I will offer two candidates and you reject one.
*(free text — no options on this one)*
````
Number them, keep the numbers stable across turns and sessions, and say out loud that partial answers are welcome — the invitation is what makes the skip visible instead of silent.
#### Formatting rules for a Q block
A batch is only worth sending if the human can read it. These are mechanical, and getting them wrong turns three careful probes into one unreadable paragraph:
| Rule | Why |
|---|---|
| **A blank line between every element** — the `**Qn — Name**` line, the question, *Why it matters*, *Recommendation*, and the option list | Markdown joins consecutive non-blank lines into a single paragraph. Without blank lines the entire batch renders as a wall of text and the human skims it, which is the failure the three-probe cap exists to prevent. |
| **Never hard-wrap a sentence across source lines** | The wrap is invisible to the renderer, so it buys nothing and costs you the paragraph break. Write each sentence as one logical line however long it is; the terminal wraps it. |
| **Options are a bullet list, one option per bullet** — `- **a)**` | Indented continuation lines are the specific thing that collapses: under four spaces the indent is stripped, at four or more it becomes a code block. A bullet list survives every renderer and keeps the options scannable. |
| **`---` between probes** | Three probes run together is one block of text. The rule gives the eye a stop and makes "answer Q2 and Q3" easy to aim at. |
| **The question itself gets its own line, not a run-on with the heading** | `**Q1 — Export format.** When a custodian…` buries the decision inside a paragraph. Name it, break, then ask it. |
| **Never use spaces to convey structure** | Whatever hierarchy you indent by hand disappears on render. Structure comes from blank lines, bullets, and bold — nothing else. |
The same applies to the recap turn in *Landing the grill*: it is prose the human has to check line by line, so give each recapped decision its own bullet.
### When answers come back partial
Assume they will. The human answers two and drops one, and the dropped one is often the hardest and most valuable.
1. **Diff what came back against what you sent.** Skipped, answered with "Other: skip", or silently omitted all count as unanswered.
2. **Lead the next turn with the unanswered probes**, at their original numbers, restated in full. Not "you missed Q3" — the whole probe again, with its recommendation, because they have lost the context by now. Unanswered probes come *before* any new probe, and they count against the cap of three.
3. **Skipped twice, stop pushing.** Record it under `## Evidence` as an open probe with your recommendation verbatim, then either narrow it into something answerable or spin it out — a fresh question file if you can phrase it sharply, a `## Not yet specified` line if you cannot.
4. **Never promote your own recommendation into `## Answer`.** A probe the human declined twice is unanswered, not decided. Writing it up as decided is self-answering a HITL question, which breaks the skill.
Never let a dropped probe fall off the end of the session unrecorded.
## Facts you look up, decisions you ask
This is the single rule that keeps a grill from feeling like an interrogation. If a **fact** can be found by exploring the environment — filesystem, codebase, tools, docs, config, git history, the map's own resolved questions — go find it. The **decisions** are the human's; put each one to them and wait.
| You look it up | You ask |
|---|---|
| Which Postgres version the app runs against | Whether the export index is allowed to add a new table |
| Whether `ExportJob` already has a `status` column | What states that column is allowed to hold |
| How the current retention sweep is scheduled | Whether exports must survive a retention sweep |
| Whether the repo uses Vitest or Jest | Whether these paths get unit tests, integration tests, or neither |
| What the S3 bucket lifecycle rule is today | Whether we are allowed to change it |
| What `AGENTS.md` says about naming conventions | Anything `AGENTS.md` does not already answer |
Look first, then ask. A question that opens with "I checked `src/export/job.ts` — it already has a `status` enum with `queued | running | failed`. Does a partial success need a fourth state?" is worth three of "how should export status work?"
**Never ask what `AGENTS.md`, the map's `## Ground rules`, or a resolved question already answers.** Re-asking a settled decision reopens it by accident and costs you the human's trust for the rest of the session.
**When the lookup is expensive**, that is not a grill — it is a `research · AFK` or `legwork` question. Say so, note it, and keep grilling the decisions you can still put to the human.
## The HITL rule, stated hard
An agent that answers its own grill has broken the skill.
A `grill · HITL` question resolves **only** through live exchange with the human. Not from the codebase, not from a plausible default, not from "the obvious industry standard," not from what you would have picked. The whole value of the question is that a human with context you do not have chose one branch over another.
Signs you are about to break it:
- You wrote a recommendation and then wrote the `## Answer` without a reply in between.
- You wrote "assuming the user would want X" anywhere.
- You resolved a question in a session where the human said nothing but "go".
- The `## Answer` contains no sentence traceable to something the human actually said.
**When the human goes unreachable mid-grill:**
1. The claim **stays**. `State: claimed` and `Claimed:` are left exactly as they are.
2. Append to `## Evidence` — never `## Answer` — the exchange so far: the probes already answered, **every probe of the last batch still outstanding**, and your recommendation for each, verbatim.
3. End the session. Report the question by name and say it is mid-grill and waiting on the human. This one IS a session end, so the Trail Footer takes **Form A** — the pathed resume command, not `WAITING ON YOU`.
4. Invent nothing. No provisional answer, no "pending confirmation" answer, no default recorded as a decision.
The next session picks the claim back up and re-sends the outstanding batch. If the session was truly abandoned rather than paused, Work Step 2 reconciliation resets it to `open` on its own — that is its job, not yours.
## The four disciplines
Run all four continuously during any grill. They are not stages; they fire whenever the trigger appears in what the human just said.
### Challenge against the glossary
When a term conflicts with the map's `## Glossary`, call it out immediately, mid-sentence if necessary.
> "The glossary defines **Export** as a completed archive file delivered to a custodian. You just used it for the background job that builds one. Which do you mean — or do we need a second term?"
### Sharpen fuzzy or overloaded language
When a term is vague or carries two meanings, propose a precise canonical term and get a ruling.
> "You keep saying 'account.' Sometimes you mean the organization paying us, sometimes the individual login. Those are a **Customer** and a **User** and they have different retention rules. Which one owns the export quota?"
### Stress-test relationships with concrete invented scenarios
Do not ask abstractly whether a relationship holds. Invent a specific scenario that probes the edge and force a precise boundary.
> "A custodian leaves the company on the 3rd. Their retention policy expires their messages on the 5th. Legal opens a hold on the 4th. On the 6th, does the export still contain those messages?"
Invent the numbers, names, and dates. Vague scenarios get vague answers.
### Cross-reference claims against the actual code
When the human states how something works, check whether the code agrees, and surface contradictions instead of quietly picking a side.
> "You said retention deletes rows. `RetentionSweep.run()` sets `deleted_at` and leaves the row in place — a soft delete. Which is the behavior we are designing against?"
A contradiction is a finding, not an embarrassment. Surface it in the same turn you found it.
## The glossary
Domain modeling would write a `CONTEXT.md`. Pathfinder does **not** — plan2code owns `AGENTS.md`, and a competing root glossary file would collide with it. The map's `## Glossary` section is the one place resolved terms live.
Entry format — one row in the map's `## Glossary` table: the term, a one-or-two-sentence definition, and the rejected synonyms in the Avoid column:
```
| Term | Meaning here | Avoid |
|---|---|---|
| Export | A completed, immutable archive file delivered to a custodian. Always the artifact, never the process that produces it. | download, extract, dump |
| Export Job | The background unit of work that produces an Export. Has states; an Export does not. | export run |
| Custodian | The person whose communications an Export contains. Not necessarily the person who requested it. | user, owner, subject |
```
Rules:
| Rule | Detail |
|---|---|
| Update inline, never defer | The moment a term is resolved, write it into `## Glossary` and save. A term you meant to add at the end of the session is a term you lost. |
| Be opinionated | When several words compete, pick one and put the rest in the Avoid column. A glossary that lists synonyms as equals has decided nothing. |
| Keep definitions tight | One or two sentences. Define what it **is**, not what it does. |
| Only project-specific terms | "Retention Policy" belongs. "Timeout", "retry", "DTO" do not, however heavily the project uses them. |
| It is a glossary and nothing else | No implementation details, no open questions, no scratch notes, no decisions. Decisions live in question files. |
| Group under sub-bullets only when clusters emerge | A flat list is fine for one cohesive area. |
Every term you resolve is a term no later session re-litigates. That is the whole return on the discipline.
## The Lock test
Domain modeling would offer an ADR here. Pathfinder sets `Locked: yes` on the question instead — the decision record and the question are the same file.
Offer `Locked: yes` **only** when all three hold:
| Test | Meaning | Fails when |
|---|---|---|
| Hard to reverse | Changing your mind later costs real time or migration | You could flip it in an afternoon |
| Surprising without context | A future reader will ask "why did they do it this way?" | It is the obvious choice anyone would make |
| A real trade-off | There were genuine alternatives and one was picked for specific reasons | There was only ever one option |
If any one is missing, skip it. An easy-to-reverse decision will just get reversed. An unsurprising one raises no questions to answer. One with no alternative records nothing beyond "we did the obvious thing."
**What earns a lock:** architectural shape ("the export index is a materialized view, not a table"); integration patterns between components ("retention and export communicate by event, never by direct call"); technology choices carrying lock-in (database, message bus, auth provider — not every library, just the ones that would take a quarter to swap); boundary and scope decisions, including the explicit no-s; deliberate deviations from the obvious path ("hand-written SQL here, not the ORM, because the ORM cannot express the retention join"); constraints invisible in the code ("no cross-region replication — the data residency contract forbids it"); and non-obvious rejections ("we considered and rejected GraphQL, for reasons someone will otherwise re-propose in six months").
**What `Locked: yes` obliges the `## Answer` to contain:**
1. The decision itself, in one or two sentences.
2. **Every alternative genuinely considered, and why each was rejected.** This is the part that makes a lock worth having. "We rejected X" with no reason is not a lock.
3. The consequences a later reader would not guess.
4. The one-line `**Gist:**` that Work Step 7 requires, same as any answer.
Offer it, do not impose it: *"This one looks hard to reverse and the reasoning will not be obvious in six months. Lock it?"* The human decides.
**Where locks go at handoff:** the handoff playbook lifts every `Locked: yes` question into the PLAN-DRAFT's **Architecture** section, and their rejected alternatives and standing constraints into **Assumptions**. Unlocked answers still inform the draft, but locks are the ones that survive verbatim into planning. Grill them harder for that reason.
## The one grill every map must resolve, and the lens applied to all of them
### 1. Testing posture — `grill · HITL`
Chart Step 6 requires this question. It exists because `/plan2code-1-plan` Phase 1 asks for exactly three things and stalls without them: testing types, whether tests run after each phase, and the coverage target. A map that clears without answering them hands the human a plan session that immediately re-asks.
The first three probes below pass the independence test against each other — none reads differently under another's answer — so **send all three as one batch**. This is the canonical worked example of a full batch, and it is the canonical case for the structured tool: every one of the three has named alternatives you can already write, the answers are fixed literals rather than prose, and nothing in the batch trips the detail test.
| Probe | Recommend by default |
|---|---|
| Which types are in scope — unit, integration, E2E, or none? | Unit plus integration; E2E only where a real browser or real broker is the only honest test |
| Does the suite run after each implementation phase, or once at the end? | `Run after each phase` — work that cannot be verified when it lands cannot be signed off |
| Coverage target: critical paths, moderate (~60-80%), or comprehensive (>80%)? | Critical paths, named explicitly, rather than a percentage nobody defends |
| What is deliberately not tested, and why? | **Fourth probe, and it does not ride in the batch.** It has no option set — the explicit no-s have to be composed — and it reads differently once the types are settled. Fold it into the recap turn, where you are already waiting on them. |
| What is already there — runner, fixtures, CI wiring? | **Not a probe.** Look it up before you send the batch, and cite it in the probes above |
Record the answer in the exact literals `/plan2code-2-document` string-matches: the types; `Run after each phase` or `Dedicated phase only`; and `Critical paths`, `Moderate (~60-80%)`, or `Comprehensive (>80%)`. Not prose — a paraphrase matches no branch downstream.
### 2. Test seams and verifiability — a lens, not a question file
For each major decision on the map, ask the same question: **how will anyone know it works?** A decision nobody can verify is a decision that silently rots.
This is not a question of its own and never gets a file or an `NN`. Run it inside whatever grill is claimed; anything it surfaces that needs deciding separately becomes a new question at Work Step 8.
These six all interrogate one decision from different sides, so they batch cleanly — but they ride along inside the claimed grill rather than owning a turn. Fold the two or three that bite into the batch you were already sending; never spend a whole turn on all six.
| Probe | What it flushes out |
|---|---|
| What observable behavior changes if this decision is implemented correctly? | Decisions with no observable effect — usually a sign the question was about implementation, not design |
| What is the cheapest thing that fails when it breaks? | The seam. If the answer is "a customer complains," there is no seam yet |
| Can this be tested without a live third-party account, a real S3 bucket, a wall-clock sleep? | Untestable-by-construction designs, while they are still cheap to change |
| Where does the boundary go so a test can stand at it — an interface, a queue, an HTTP edge, a pure function? | The seam the plan will need to name |
| What does the failure look like in production — log line, metric, alert, dead-letter queue? | Verifiability after ship, not just in CI |
| If we get this wrong, how long before we find out? | Decisions that need a canary or a feature flag rather than a test |
If a decision survives all six with no answer, it is not ready to leave the map. Either re-frame the question, or add the seam as a constraint in the `## Answer` so the plan inherits it.
## Anti-patterns
| Failure mode | What it looks like | Fix |
|---|---|---|
| Batching **dependent** probes | "What format, what do we name the part files, and how big is a part?" | Only the first is independent. Send it; hold the other two — they are unanswerable until format lands. |
| Drip-feeding one probe at a time | Twelve open questions on the map, one probe per response, the human gives up on session four | Batch up to three independent probes. On a charted map the human's round trips are the scarce resource, not your token budget. |
| Losing a probe the human skipped | Sent three, got two back, moved on and never mentioned the third | Diff the batch. Lead the next turn with what came back empty, restated in full. |
| A batch of naked questions | Three one-liners with no recommendations and no why-it-matters | Every probe in a batch carries its own recommendation and its own stake. Otherwise you have offloaded the thinking, not the round trips. |
| Sending a batch as a wall of text | Three probes hard-wrapped across source lines with their options indented, all collapsing into one paragraph on render | Blank line between every element, options as a bullet list, `---` between probes. The three-probe cap exists so the human reads all three; an unreadable batch gets skimmed, and a skimmed answer is worse than none. |
| Flattening a gnarly probe into a picker | An architectural decision reduced to three option labels because the batch was already going through the structured tool | Run the detail test. One tripping probe sends the whole batch to numbered Q blocks. A clicked option records no reasoning, and the `## Answer` needs reasoning. |
| Defaulting to prose when nothing tripped | A clean three-probe batch written as Q blocks "to be safe", or downgraded just because the question will be `Locked: yes` | The detail test is a test, not a preference, and it is four rows long. Nothing tripped means the tool: you gain visibly skipped probes, and the recap turn still captures the reasoning a lock needs. |
| Retreating to prose over a fixable batch | Two probes collided on the 16-character header, or one option label was bundling three answers, so the whole batch went to Q blocks | Neither is a detail-test trip. Rename the headers; split the bundled probe. Reshape the batch, never the decision. |
| Sending four probes because the tool accepts four | A fourth probe added to a clean batch of three because there was room in the call | The cap is three regardless of what the environment allows. The ceiling is the human's attention, not the tool's schema. |
| Grilling in Pathfinder's own vocabulary | "The frontier has one takeable `grill · HITL` — shall we graduate 02 out of the fog and lock it?" | Plain English. Markers, types, `NN` numbers, and fog are your bookkeeping; the human is deciding about their product. |
| Plain-washing the load-bearing term | "Do you want the friendly file or the compact one?" where the real choice is `.eml`-in-a-ZIP versus NDJSON | Plain wording, precise nouns. Name the formats and gloss them; a decision made on a euphemism cannot be written into `## Answer`. |
| Asking what the codebase already answers | "Do you use Postgres or MySQL?" | Look. Every avoidable question spends trust you need for the hard ones. |
| Accepting a vague answer and moving on | "Handle it sensibly" → recorded as the decision | Push once more, concretely: "Sensibly meaning we drop the attachment, or fail the whole job?" A vague answer is not an answer. |
| Leading the human to your preferred answer | "You'd want Postgres here, right?" | Recommend openly, then present the real alternatives with their real merits. A recommendation invites a fight; a leading question suppresses one. |
| Grilling past the decision into implementation | "Should the retry helper take a callback or return a promise?" | That is the plan's job, or the implementer's. Stop at the decision. The pull to keep going is the edge of the map. |
| Drifting off the claimed question | Claimed `[Export format](./questions/04-export-format.md)`, forty minutes later deep in auth | Name the drift out loud, capture the new thread as a fresh question or as a line in `## Not yet specified`, and return. One question _file_ per session. |
| Self-answering a HITL question | An `## Answer` with no words the human said | Delete it. Reopen the question. See the HITL rule. |
| Recording the decision but not the rejections | "We chose event-driven." | Rejections are half the record — and mandatory when `Locked: yes`. Ask what else was on the table before you close. |
| Grilling a fact | "How long does the retention sweep take?" | If it is measurable, measure it — or make it a `research · AFK` question. Do not make the human guess at their own system. |
| Letting the glossary go stale | Three terms resolved, none written down | Write each one the moment it lands. Deferring loses them. |
| Closing without confirmation | Answer written straight after the last reply | Recap the whole chain of decisions, get the explicit confirm, then write. |
## Landing the grill
When the branch is walked out. **Steps 1-3 are ONE turn, not three** — recap, confirmation request, and lock offer go out together, because a lock offer sent after a separate confirmation costs a round trip to ask a yes/no the human could have answered alongside the recap.
1. Recap the decisions in order, in the human's own terms, using glossary vocabulary. Include anything a structured-tool reply left implicit, so the confirmation covers the reasoning and not just the picks.
2. Ask for the confirmation. Do not skip this — the recap is where the human catches the one thing you misheard, and where a clicked option becomes words they said.
3. Apply the Lock test in the same message. Offer, do not impose.
4. Write `## Answer` per Work Step 7: the decision, what was rejected and why, consequences, and a one-line `**Gist:**`. Evidence, links, and transcript fragments go under `## Evidence`.
5. Anything the grill surfaced that belongs to a different question goes to the map — a fresh question if you can phrase it sharply, `## Not yet specified` if you cannot, `## Out of scope` if it sits past the destination.
@@ -0,0 +1,415 @@
# Handoff Playbook — Clearing the Map into a PLAN-DRAFT
Loaded at The Clearing Gate. Turns a cleared map into `specs/<idea>/PLAN-DRAFT-<YYYYMMDD>.md` that `/plan2code-1-plan` resumes from at Phase 4, then freezes `pathfinder/` as the rationale record.
**Backend note.** The scoring rubric, the hard caps, the honesty rules, the template, and the mapping table are the same either way — and the draft is written to local disk either way, because `/plan2code-1-plan` reads a file, not a tracker. On `**Backend:** github`, `github-issues.md` replaces only the preflight table, the question links (issue URLs, already absolute), and the freeze steps.
Nothing here is creative. The gate is scored, the mapping is fixed, the template is literal. Follow it exactly or the resuming plan session silently loses work.
## Preflight — before scoring anything
| # | Check | If it fails |
|---|---|---|
| 1 | Re-run the reconcile pass: read every file in `questions/`, rebuild every map marker from the files | Fix the map first. Markers are derived, never authored. |
| 2 | Zero `[ ]`, zero `[/]`, zero `[!]` rows in `## Question Checklist` | Not cleared. Return to the frontier. |
| 3 | `## Not yet specified` is empty | Not cleared. Graduate the fog into questions, or admit it is out of scope. |
| 4 | Every `[x]` row's file has a real `## Answer` with a `**Gist:**` | The file wins over the marker. Repair, then re-check. |
| 5 | `ls specs/<idea>/` shows no existing `PLAN-DRAFT-*.md` | One already exists — read it. Update it in place; never add a second dated draft. |
| 6 | The destination is reachable with nothing left to decide — no `## Answer` defers a choice to "whoever implements this" | Not cleared. Name the open decision and graduate it into a question. |
Only after all six pass do you score the four dimensions.
## The clearing-gate scoring rubric
Four dimensions, 0-25 each, scored against **written evidence in `questions/`** — never against your recollection of the conversation. The gate needs **every dimension at 18/25 or better**. There is no averaging: 25/25/25/14 fails.
### Band scale (applies to all four)
| Band | Meaning |
|---|---|
| 23-25 | Decided, written down, and consequences recorded. A developer could act on it without asking a question. |
| 18-22 | Decided and written down. Residual detail remains, but it is *specification* detail that Phase 5/6 settles — not a decision anyone still has to make. |
| 12-17 | A real decision is still open, or an answer exists with no evidence behind it. **Gate fails.** |
| 0-11 | The area was never charted. **Gate fails**, and the map was cleared prematurely. |
The 18-boundary is the honest line between *"needs designing"* and *"needs deciding"*. Pathfinder owns deciding. If someone still has to decide, you are not done.
### What each dimension scores
| Dimension | Scores | Raises it | Lowers it |
|---|---|---|---|
| **Requirements Clarity** | Are the requirements unambiguous? | Every resolved `grill` answer states the decision AND what was rejected; the testing-posture answer names types, cadence, and coverage | Answers phrased as preferences ("probably NDJSON") instead of decisions; a requirement that only exists in the map gist and not in a question file |
| **Technical Feasibility** | Do we know HOW to build each component? | Resolved `sketch` questions with a real artifact under `sketch-NN/`; `research` answers citing primary sources under `## Evidence`; `00-codebase-context.md` naming the actual files that change | A mechanism nobody has exercised in this codebase; a research answer whose `## Evidence` is empty or cites only a blog post |
| **Integration Points** | Are all external dependencies identified? | Every system named in `## Destination` has a resolved question touching it; auth, quota, and failure behavior named per integration | An integration mentioned in an answer but never questioned; a config store that was read but never written to during a sketch |
| **Risk Assessment** | Are blockers documented with mitigations? | Answers that record consequences; `Locked: yes` answers that say what breaks if reversed; ceilings with a stated behavior at the ceiling | A recorded limit with no decided behavior past it; a `Locked: yes` answer with no consequences section |
### Hard caps
A cap overrides your judgment. While a cap condition holds, the dimension **cannot** exceed 17, so the gate cannot pass.
| Cap | Condition |
|---|---|
| Requirements ≤ 17 | The testing-posture question is not `resolved`, or its answer omits any of types / cadence / coverage |
| Feasibility ≤ 17 | Any `research` question is `resolved` with an empty `## Evidence` |
| Integration ≤ 17 | A system named in `## Destination` has no resolved question touching it |
| Risk ≤ 17 | Any `Locked: yes` answer records no consequences |
### Honesty rules
- Score the **written record**, not the conversation. If the human agreed to something in a session and nobody wrote it into a `## Answer`, it does not exist and it does not earn points.
- A filled `## Answer` is not automatically 25. An answer that decides but records no consequences tops out around 20.
- Never round up to clear the gate. A 17 that "feels like an 18" is the exact case the gate exists to catch.
- Never move a decision to `## Out of scope` to raise a score. Out-of-scope is a scoping act with a reason; scope-cutting to pass a gate is score inflation with extra steps.
- If two dimensions are borderline, write the one-line justification for each score into the Session End report. Justifications that cannot be written are scores that cannot be defended.
### Worked example — idea `audit-log-s3-export`
Destination: *"A spec a developer can implement: nightly export of tenant audit logs to customer-owned S3 buckets, with a signed manifest per run."*
Resolved questions: [Codebase context](./questions/00-codebase-context.md), [Export format](./questions/01-export-format.md), [Scheduling model](./questions/02-scheduling-model.md), [Destination auth](./questions/03-destination-auth.md), [Retention and replay](./questions/04-retention-and-replay.md), [Testing posture](./questions/05-testing-posture.md), [Throughput ceiling](./questions/06-throughput-ceiling.md). Ruled out: [Failure notification](./questions/07-failure-notification.md).
**First scoring pass:**
| Dimension | Score | Justification against evidence |
|---|---|---|
| Requirements Clarity | 22/25 | Four `grill` answers state decisions and rejections (NDJSON chosen, CSV rejected for nested actor payloads). Testing posture settled: integration + unit, run after each phase, moderate coverage. Minus 3: the manifest's exact field list is unspecified — a Phase 5 spec detail, not an open decision. |
| Technical Feasibility | 21/25 | `sketch-02` uploaded a 400 MB multipart object to a real bucket end to end. `00-codebase-context.md` names `AuditExportJob` and the existing Hangfire registration as the extension points. Minus 4: no component in this codebase has ever assumed a cross-account IAM role; the pattern is documented in AWS docs cited under `## Evidence` but unexercised here. |
| Integration Points | 23/25 | Three integrations, each with a resolved question: S3 (Destination auth), Hangfire (Scheduling model), tenant config store (Codebase context). Auth and quota named per integration. Minus 2: the tenant config store's write path was read but never exercised by a sketch. |
| Risk Assessment | **16/25** | Consequences recorded on Export format and Destination auth. But Throughput ceiling establishes 200 MB per tenant per day at p95 and **records no decided behavior above it** — truncate, spill to the next run, or fail the run is still undecided. |
**Gate result: FAIL** on Risk Assessment (16 < 18), and the orchestrator third clearing condition — the destination reachable with nothing left to decide — fails with it: a decision is genuinely still open. Do not write a PLAN-DRAFT. Name the failure, graduate `questions/08-overflow-behavior.md` (`grill · HITL`, `Blocked by: none`) from the gap, and end the session on the frontier.
**Second scoring pass, one session later**, with [Overflow behavior](./questions/08-overflow-behavior.md) resolved (spill to the next run, alarm at three consecutive spills):
| Dimension | Score |
|---|---|
| Requirements Clarity | 22/25 |
| Technical Feasibility | 21/25 |
| Integration Points | 23/25 |
| Risk Assessment | 20/25 |
All four at 18 or better. Gate passes. Proceed to the mechanism.
## The mechanism, stated plainly
Pathfinder writes `specs/<idea>/PLAN-DRAFT-<YYYYMMDD>.md` with the header line:
`**Status:** Phase 3 Complete - Resume at Phase 4`
`/plan2code-1-plan` **already recognises that exact string.** Its "Check for Existing Progress" block, which runs before Phase 1, reads:
`- Status "Phase 3 Complete - Resume at Phase 4": Resume at Phase 4`
That is the whole handoff. The string is the contract, and it needs **zero changes to the planning skill** — pathfinder is impersonating the Large-project Context Checkpoint that 1-plan's own Phase 3 performs, which writes the same status for the same reason.
Consequences of that being a literal string match:
- Copy it byte for byte. Plain ASCII hyphen-minus surrounded by single spaces. An en dash, a colon, or "Phase 3 complete" in lower case breaks the match and 1-plan starts over at Phase 1 — throwing away every decision the map holds.
- It goes on its own `**Status:**` line in the header block, not buried in prose.
- The file must be named `PLAN-DRAFT-<YYYYMMDD>.md` and live directly in `specs/<idea>/`. `PLAN-DRAFT-*.md` is a reserved name inside `pathfinder/` — never write it there.
- Get the date from the shell (`date +%Y%m%d` in Bash, `Get-Date -Format yyyyMMdd` in PowerShell). Do not guess it.
**Discovery is shell-only.** `specs/` is gitignored, so Glob silently returns nothing and every downstream skill would report "no PLAN-DRAFT found". Use `ls specs/` and `ls specs/<idea>/` (Bash) or `Get-ChildItem specs/` (PowerShell) — the same rule 1-plan and `/plan2code-2-document` follow when they look for the file you are about to write.
**Do not append `## Planning Metrics` or any metrics comment.** Pathfinder is not a metered step. 1-plan's Phase 7 owns that block and will add it when it finishes the plan. Do not emit any of the loop's completion tokens listed in the skill's Rules anywhere under `specs/`.
## Map to PLAN-DRAFT mapping
Everything in the draft traces to something on the map. Nothing is invented at handoff time — if a section has no source, that is a gate failure you missed, not a paragraph to write from imagination.
| Source on the cleared map | Becomes |
|---|---|
| `## Destination` | Section 1 Executive Summary (2-3 sentences, present tense) and Section 7 Success Criteria (the destination restated as checkable outcomes) |
| Resolved `grill` answers describing behavior | Section 2.1 Functional Requirements, one `FR-N` per decided behavior |
| Resolved `grill` answers describing performance, security, scale, operability | Section 2.2 Non-Functional Requirements, one `NFR-N` each |
| `## Out of scope` | Section 2.3 Out of Scope — copied line for line, wording and order intact. The ONLY permitted change is the link prefix: `./questions/` becomes `./pathfinder/questions/`, because the draft sits one level above the map. Do not re-word, re-order, or summarise; a re-worded scope boundary is a re-litigated one. |
| The testing-posture question's answer | Section 2.4 Testing Strategy table (Types / Phase Testing / Coverage) |
| Resolved `research` answers and their `## Evidence` | Section 3 Tech Stack — the cited source becomes the Justification cell |
| `questions/00-codebase-context.md` | The `## System Context` section, and the components table in 4.3 |
| Answers with `Locked: yes` | Section 4 Architecture (4.1 Pattern rationale, 4.4 Data Model, 4.5 API Design) **and** Section 9 Assumptions — a locked decision is an assumption downstream work is allowed to rely on |
| Consequences recorded across all `## Answer` sections | Section 6 Risks and Mitigations — the consequence is the Risk, the decision that bounds it is the Mitigation |
| `## Ground rules` | The `AGENTS.md` line in `## System Context`; conventions the plan must not violate |
| The shape of the map (question count, integrations touched, components named) | The `## Scope Assessment` section |
| Resolved `sketch` questions and their artifacts | Section 3 Justification cells and Section 6 Mitigation cells ("proven by `pathfinder/sketch-02/`") |
| Every resolved question, by name | `## Pathfinder Provenance` |
| — | **Section 5 Implementation Phases stays empty.** 1-plan Phase 6 breaks the work into phases. Pathfinder decides; it does not slice. Leave the placeholder note in place and do not put implementation checkboxes there. |
Questions ruled `out-of-scope` never appear in Provenance and never become requirements. Their one line in `## Out of scope` is their only trace — that is the point of the marker.
## The PLAN-DRAFT template
Write this literally, substituting real content. Keep the section numbering exactly as shown — 1-plan and `/plan2code-2-document` both address sections by number.
````markdown
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
# Audit Log Export - Implementation Plan
**Created:** 2026-08-03
**Status:** Phase 3 Complete - Resume at Phase 4
**Charted by:** `/plan2code-0-pathfinder` over 9 sessions
**Planning record:** `specs/audit-log-s3-export/pathfinder/map.md` (no PLAN-CONVERSATION - this plan was charted, not conversed)
**Confidence (pathfinder):** Requirements-clarity 22/25 · Feasibility-technical 21/25 · Integration-points 23/25 · Risk-assessment 20/25
---
## 1. Executive Summary
Tenants can have their audit logs exported nightly to an S3 bucket they own, with a
signed manifest per run so they can prove completeness. Export runs on the existing
Hangfire schedule, writes NDJSON, and assumes a customer-provided cross-account IAM
role with an external ID. Runs that exceed the per-tenant daily ceiling spill into the
next run rather than truncating.
## 2. Requirements
### 2.1 Functional Requirements
- [ ] **FR-1:** Export each tenant's prior-day audit events as newline-delimited JSON, one object per event, UTF-8, no BOM
- [ ] **FR-2:** Write a per-run manifest listing object keys, event counts, byte counts, and a SHA-256 per object
- [ ] **FR-3:** Sign the manifest with the platform export key; publish the public key at a stable URL
- [ ] **FR-4:** Assume the tenant-configured IAM role with the tenant's external ID; never use platform-owned credentials against a customer bucket
- [ ] **FR-5:** Allow an operator to replay any run within a 7-day window without duplicating manifest sequence numbers
- [ ] **FR-6:** Spill events above the per-tenant daily ceiling into the next scheduled run, oldest first
- [ ] **FR-7:** Raise an alarm after three consecutive spilling runs for the same tenant
### 2.2 Non-Functional Requirements
- [ ] **NFR-1:** Sustain 200 MB per tenant per day at p95 without extending the nightly window past 04:00 UTC
- [ ] **NFR-2:** Never log tenant event bodies, bucket names, or assumed-role ARNs above debug level
- [ ] **NFR-3:** A failed run must leave no partial objects visible in the customer bucket
- [ ] **NFR-4:** Export must add no schema changes to the audit event write path
### 2.3 Out of Scope
<!-- copied line for line from pathfinder/map.md ## Out of scope; only the link prefix changes -->
- [Failure notification](./pathfinder/questions/07-failure-notification.md) — email/webhook delivery of run failures belongs to the platform alerting effort, not this export. The alarm in FR-7 is raised, not delivered.
- **On-demand export from the tenant UI** — the destination is the scheduled export. A user-triggered export is a separate effort with its own map.
- **Log formats other than NDJSON** — Parquet was raised and ruled past the destination.
### 2.4 Testing Strategy
| Aspect | Decision |
|---|---|
| Types | Unit + Integration |
| Phase Testing | Run after each phase |
| Coverage | Moderate (~60-80%) |
## System Context
**Project type:** Existing codebase — .NET 8 service, `src/Platform.Audit/`
| Aspect | Finding | Source |
|---|---|---|
| Entry points to change | `AuditExportJob`, registered in `HangfireStartup.ConfigureRecurringJobs()` | [Codebase context](./pathfinder/questions/00-codebase-context.md) |
| Existing patterns to follow | Jobs resolve tenant scope via `ITenantScopeFactory`; no job reads config directly | [Codebase context](./pathfinder/questions/00-codebase-context.md) |
| Integration surfaces | S3 (customer-owned), Hangfire scheduler, `TenantConfigStore` | [Destination auth](./pathfinder/questions/03-destination-auth.md), [Scheduling model](./pathfinder/questions/02-scheduling-model.md) |
| Technical debt in the path | `AuditQuery` materialises full result sets; streaming reader needed before FR-1 | [Codebase context](./pathfinder/questions/00-codebase-context.md) |
| System boundaries | Read-only against the audit store; writes only to customer buckets and the run-log table | [Retention and replay](./pathfinder/questions/04-retention-and-replay.md) |
| Conventions in force | `AGENTS.md` present and read; its logging and DI conventions govern | `pathfinder/map.md` ## Ground rules |
## Scope Assessment
**Assessment: Medium** — 11 requirements across 3 integrations, 5 components touched. No Large threshold is met.
| Indicator | Value |
|---|---|
| Requirements decided | 11 (7 FR + 4 NFR) |
| Components | 5 (`AuditExportJob`, `NdjsonWriter`, `ManifestSigner`, `S3RoleAssumer`, `ExportRunLog`) |
| Integrations | 3 (S3, Hangfire, `TenantConfigStore`) |
| Decisions charted | 8 resolved, 1 ruled out of scope |
## 3. Tech Stack
<!-- Phase 4 completes this table. Rows below are decided; do not re-open them. -->
| Category | Technology | Version | Justification |
|---|---|---|---|
| Serialization | `System.Text.Json` NDJSON writer | .NET 8 | [Export format](./pathfinder/questions/01-export-format.md) — no new dependency; CSV rejected for nested actor payloads |
| Object storage | `AWSSDK.S3` multipart upload | 3.7.x | [Throughput ceiling](./pathfinder/questions/06-throughput-ceiling.md) — proven in `pathfinder/sketch-02/` against a real bucket at 400 MB |
| Cross-account auth | STS `AssumeRole` + external ID | — | [Destination auth](./pathfinder/questions/03-destination-auth.md) — AWS confused-deputy guidance cited in that file's `## Evidence` |
| Scheduling | Existing Hangfire recurring job | in-repo | [Scheduling model](./pathfinder/questions/02-scheduling-model.md) — a new scheduler was rejected |
## 4. Architecture
### 4.1 Pattern
Pipeline inside the existing job host: query → stream → chunk → upload → manifest → sign.
Chosen because the audit store is the only source and the export is strictly one-way.
[Locked] A separate export microservice was rejected — see [Scheduling model](./pathfinder/questions/02-scheduling-model.md).
### 4.2 System Context Diagram
<!-- Phase 5 refines. Boundaries above are settled. -->
### 4.3 Components
| Component | Responsibility | Inputs | Outputs | Depends on |
|---|---|---|---|---|
| `AuditExportJob` | Orchestrates one tenant-run | Tenant id, run date | Run result | `TenantConfigStore` |
| `NdjsonWriter` | Streams events to chunked NDJSON | Event stream | Byte stream, counts | — |
| `S3RoleAssumer` | Assumes the tenant role, returns a scoped client | Role ARN, external ID | `IAmazonS3` | STS |
| `ManifestSigner` | Builds and signs the run manifest | Object metadata | Signed manifest | Platform export key |
| `ExportRunLog` | Records runs for replay and spill detection | Run result | Run rows | Platform DB |
### 4.4 Data Model
Manifest sequence numbers are per tenant, monotonic, and reused on replay.
[Locked] See [Retention and replay](./pathfinder/questions/04-retention-and-replay.md).
### 4.5 API Design
<!-- Phase 5 fills. No public API surface was decided during pathfinding. -->
## 5. Implementation Phases
<!-- Intentionally empty. Phase 6 of /plan2code-1-plan breaks the requirements
above into implementation phases. Pathfinder decides; it does not slice. -->
## 6. Risks and Mitigations
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Cross-account role assumption is unexercised in this codebase | Medium | High | Spike `S3RoleAssumer` against a second AWS account before any other component |
| A tenant exceeds the daily ceiling indefinitely | Medium | Medium | Spill oldest-first plus a three-run alarm — [Overflow behavior](./pathfinder/questions/08-overflow-behavior.md) |
| Partial objects visible after a failed run | Low | High | Upload to a run-scoped prefix, publish the manifest last — the manifest is the commit point |
| `AuditQuery` materialises full result sets | High | High | Streaming reader is a prerequisite, not an optimisation |
| Customer revokes the role mid-run | Low | Medium | Fail the run whole; replay window covers recovery |
## 7. Success Criteria
- [ ] A tenant with a configured role receives NDJSON and a signed manifest for the prior day, nightly
- [ ] The published public key verifies the manifest signature
- [ ] A 250 MB tenant-day completes without extending the window past 04:00 UTC
- [ ] A replay inside 7 days reproduces the run without a new sequence number
- [ ] A failed run leaves nothing visible in the customer bucket
## 8. Open Questions
<!-- 1-plan's template removes this section when empty; pathfinder keeps it with "None" so a
resuming session can see the map cleared clean, rather than that the section was forgotten. -->
None. The map cleared with zero open questions.
## 9. Assumptions
- Tenants can create an IAM role in their own account — [Destination auth](./pathfinder/questions/03-destination-auth.md) [Locked]
- The nightly Hangfire window remains available and is not contended by other jobs — [Scheduling model](./pathfinder/questions/02-scheduling-model.md) [Locked]
- Manifest sequence reuse on replay is acceptable to tenant compliance teams — [Retention and replay](./pathfinder/questions/04-retention-and-replay.md) [Locked]
- Audit events are immutable once written, so a replay reproduces byte-identical output
## Pathfinder Provenance
Charted over 9 sessions. Each requirement above traces to a decision below; open the
question for what was rejected, why, and what it costs.
| Question | Gist |
|---|---|
| [Codebase context](./pathfinder/questions/00-codebase-context.md) | `AuditExportJob` and `HangfireStartup` are the extension points; `AuditQuery` needs a streaming reader first |
| [Export format](./pathfinder/questions/01-export-format.md) | NDJSON with a per-run signed manifest; CSV rejected for nested actor payloads |
| [Scheduling model](./pathfinder/questions/02-scheduling-model.md) | Reuse the existing Hangfire recurring job; a dedicated export service was rejected |
| [Destination auth](./pathfinder/questions/03-destination-auth.md) | Customer-owned bucket via assumed role plus external ID; no platform-held customer credentials |
| [Retention and replay](./pathfinder/questions/04-retention-and-replay.md) | 7-day replay window, sequence numbers reused on replay |
| [Testing posture](./pathfinder/questions/05-testing-posture.md) | Unit + integration, run after each phase, moderate coverage |
| [Throughput ceiling](./pathfinder/questions/06-throughput-ceiling.md) | 200 MB per tenant per day at p95; multipart upload proven in `pathfinder/sketch-02/` |
| [Overflow behavior](./pathfinder/questions/08-overflow-behavior.md) | Spill oldest-first into the next run; alarm after three consecutive spills |
Ruled out of scope: [Failure notification](./pathfinder/questions/07-failure-notification.md) — recorded in 2.3.
---
**Next:** Resume with `/plan2code-1-plan` at Phase 4 (Tech Stack).
- **Phase 7 verification:** sections 1, 2, System Context and Scope Assessment are already settled — their source of truth is `specs/audit-log-s3-export/pathfinder/map.md`, not this conversation. Verify sections 3-7 only.
- **Phase 7:** replace THIS file in place. Do not create a second PLAN-DRAFT in this folder.
````
### Two things in that template that are not optional
**No scrapable confidence numbers anywhere in the file.** Write the confidence as `Requirements-clarity 22/25 · Feasibility-technical 21/25 · Integration-points 23/25 · Risk-assessment 20/25` — hyphenated dimension labels, sub-scores over 25, no total, no percent sign.
The reason is exact. When a PLAN-DRAFT carries no `METRICS_JSON` comment, the metrics collector falls back to scraping it by regex: an overall-confidence pattern that requires a literal `%`, and four breakdown patterns that match a bare `Requirements` / `Feasibility` / `Integration` / `Risk` followed directly by whitespace, a colon, or a pipe and then digits. **The breakdown patterns do not require a percent sign.** A pathfinder-written draft always lacks that comment until `/plan2code-1-plan` Phase 7 appends one, so both the percent sign *and* the bare dimension words have to be kept off the page — otherwise the pipeline records a planning-step confidence that no planning step ever produced. The hyphen in `Requirements-clarity` breaks the match; a table row reading `| Requirements | 11 |` does not, which is why the Scope Assessment row is labelled `Requirements decided`.
**The `**Next:**` footer must ship with both bullets.** Pathfinder cannot edit the planning skill, so those two instructions travel inside the artifact:
- *Without the verification bullet*, 1-plan's Phase 7 does exactly what it is told to do — "re-read conversation as source of truth" — finds a fresh conversation that starts at Phase 4 and contains no requirements discussion at all, concludes sections 1 and 2 are unsupported, and silently drops the requirements that N pathfinder sessions produced. The bullet redirects the source of truth for the settled sections to `map.md`.
- *Without the replace-in-place bullet*, 1-plan's Phase 7 creates `PLAN-DRAFT-<its own date>.md` alongside yours. `/plan2code-2-document` then finds two drafts in the folder, hits its "Multiple found: List all, ask which to document" branch, and asks the user to disambiguate between a pathfinder draft and a plan draft that partially supersedes it.
Never drop the footer to make the file tidier. It is load-bearing.
## System Context and Scope Assessment — why they buy you Phase 4
These two named, unnumbered sections are what make "Resume at Phase 4" legitimate rather than a shortcut. They stand in for the phases pathfinder already did the work of:
| Draft section | Satisfies | Because pathfinder already |
|---|---|---|
| Sections 1 and 2 (including 2.4 Testing Strategy) | 1-plan **Phase 1: Requirements Analysis** | Grilled every functional and non-functional decision, and always charted a testing-posture question — that question exists specifically so Phase 1's testing prompt is already answered |
| `## System Context` | 1-plan **Phase 2: System Context Examination** | Wrote `questions/00-codebase-context.md` at Chart Step 3: directory structure, key components verified against actual code, patterns and conventions, integration points, technical debt, boundaries — Phase 2's own checklist, item for item |
| `## Scope Assessment` | 1-plan **Phase 3: Scope Assessment** | Produced the counts Phase 3 measures — requirements, components, integrations — as a byproduct of charting. Map the totals onto Phase 3's Small / Medium / Large table and state the verdict |
Populate `## System Context` from `00-codebase-context.md` and nothing else. It is the one question guaranteed to exist on every map, it was resolved on the spot with the codebase open, and it is a `legwork · AFK` answer — factual, not preferential. Cite it in the Source column so a skeptical reader can check the finding against the file.
Populate `## Scope Assessment` from the shape of the map. Count resolved questions that produced requirements (not `00-codebase-context.md`, not out-of-scope ones), count distinct components named across the answers, count distinct external systems. Apply Phase 3's thresholds honestly: Large if **any** threshold is met. Score the counts, never the session count — a map can take nine sessions to clear and still be Medium, and the `Phase 3 Complete - Resume at Phase 4` status string works regardless of the verdict, so there is nothing to gain by inflating it. Pathfinder cannot count implementation phases (Section 5 is deliberately left empty), so assess on requirements, components, and integrations only.
If the charting session found `AGENTS.md` absent, say so in the `Conventions in force` row rather than leaving it blank. The plan session needs to know the conventions were never available, not guess that they were checked.
## Freeze the map
Once the PLAN-DRAFT is written and saved:
1. Set `**Status:** Cleared` in `map.md`.
2. Bump `**Updated:**` to today.
3. Add a plan pointer line under the status: `**Plan:** ../PLAN-DRAFT-20260803.md`.
4. Leave **everything** under `pathfinder/` exactly where it is — `map.md`, every file in `questions/`, every `sketch-NN/` directory.
**Never delete `pathfinder/`.** It is the rationale record behind the plan: what was decided, what was rejected, why, and what it costs. It sits in the same class as `PLAN-CONVERSATION-*.md` — the transcript a plan is defensible against — and `/plan2code-4-finalize` archives it alongside `PLAN-DRAFT.md` and `PLAN-CONVERSATION-*.md` into `specs--completed/<idea>/`. Deleting it turns every locked decision in the plan into an unexplained constraint six months from now.
Do not tidy it either. Do not collapse resolved questions into the map, do not prune `## Evidence`, do not remove sketch directories because the code is throwaway. The sketch is the proof behind a feasibility score.
**A later session that finds `Status: Cleared` must not resume work on it.** The map is finished; there is nothing left to decide inside it. Point at the PLAN-DRAFT and `/plan2code-1-plan`, and stop. If the destination has been redrawn — the scope grew, an out-of-scope item came back, the goal changed — that is a **fresh effort with a fresh map**, not a resumption: a new kebab-case idea name, a new `specs/<new-idea>/pathfinder/`, charting from Step 1. The frontier stops at the destination, so a new destination gets a new frontier. Reopening a cleared map silently invalidates the PLAN-DRAFT that was built from it, and nothing downstream would notice.
## What to tell the user
Session End for the cleared case reports six things, in this order:
| # | Report |
|---|---|
| 1 | The destination, restated — what the map was finding its way to, now reached |
| 2 | How many decisions were made, and the headline ones **by name** |
| 3 | What was ruled out of scope, and why (one line each) |
| 4 | The four confidence scores, with a one-line justification for any below 21 |
| 5 | The PLAN-DRAFT path |
| 6 | The next command — carried by the Trail Footer, which at `Cleared` routes to `/plan2code-1-plan` |
Example:
> **Destination reached:** a spec a developer can implement for nightly audit-log export to customer-owned S3 buckets with a signed manifest.
>
> **8 decisions made** across 9 sessions. The load-bearing ones: [Export format](./questions/01-export-format.md) settled on NDJSON with a signed manifest; [Destination auth](./questions/03-destination-auth.md) settled on assumed roles with an external ID so we never hold customer credentials; [Overflow behavior](./questions/08-overflow-behavior.md) settled on spilling rather than truncating.
>
> **Ruled out of scope:** failure-notification delivery (belongs to the platform alerting effort), on-demand export from the tenant UI (its own effort), Parquet output (past the destination).
>
> **Confidence:** Requirements-clarity 22/25 · Feasibility-technical 21/25 · Integration-points 23/25 · Risk-assessment 20/25. Feasibility is held at 21 because cross-account role assumption is unexercised in this codebase — it is a documented pattern, not a proven one here, and it is the first thing to spike.
>
> **Written:** `specs/audit-log-s3-export/PLAN-DRAFT-20260803.md`. It resumes planning at Phase 4 — sections 1, 2, System Context and Scope Assessment are already settled. The reasoning behind every one of them stays in `specs/audit-log-s3-export/pathfinder/`; do not delete it.
Then the closing block from the skill's Session End — the mascot with the message *The way is clear! Time to plan!* followed by the Trail Footer, whose trail shows every stop walked to the `` destination and whose command is `/plan2code-1-plan`.
**One branch.** If `## Ground rules` records `AGENTS.md` as **absent**, recommend `/plan2code-init` FIRST, and offer `questions/00-codebase-context.md` as its input:
> Before planning: this project has no `AGENTS.md`, and `/plan2code-1-plan` blocks on that. Run `/plan2code-init` first and attach `specs/audit-log-s3-export/pathfinder/questions/00-codebase-context.md` — the recon pass already established the structure, conventions, and integration points it asks for. Then `/plan2code-1-plan`.
Nothing to commit — `specs/` is gitignored. Say so once, then stop.
## Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| 1-plan starts at Phase 1 and re-asks for requirements | The status string does not match byte for byte | Compare against the quoted line above; watch for en dashes and casing |
| 1-plan cannot find the draft at all | Glob was used to discover `specs/` | Shell only: `ls specs/` |
| `/plan2code-2-document` asks which of two drafts to use | The replace-in-place bullet was dropped from the footer | Merge the two drafts into the pathfinder-dated one, delete the other, restore the footer |
| The finished plan is missing requirements the map decided | The verification bullet was dropped from the footer | Re-derive 2.1 and 2.2 from the resolved answers, restore the footer |
| Metrics report a planning confidence nobody scored | A percent sign, or a bare dimension word followed by a number, reached the file | Hyphenate the dimension labels and drop the percent sign |
| A locked decision in the plan has no visible reason | `pathfinder/` was deleted or pruned | Unrecoverable. This is why the freeze step exists |
| Gate passes but the first implementation session immediately hits an undecided question | A dimension was rounded up | The gate was the check. Score the written record, not the feeling |
@@ -0,0 +1,41 @@
# Questions & Map Format
> Part of plan2code-0-pathfinder — the on-disk format both modes share: directory layout, `NN` numbering, the question-file schema, `Type:` vocabulary, and the marker / blocking rules. The main file keeps only the marker legend and a layout gist; the authority is here.
>
> **This file describes the `local` backend only.** If `## Ground rules` says `**Backend:** github`, the equivalence table in `github-issues.md` replaces every rule below — there are no files, no `NN`, no schema lines, and no checklist.
## Layout
```
specs/<idea>/
├── pathfinder/
│ ├── map.md <- the index
│ ├── questions/NN-<slug>.md <- 00-codebase-context.md always exists
│ └── sketch-NN/ <- optional runnable sketch, throwaway
└── PLAN-DRAFT-<YYYYMMDD>.md <- written ONLY when the map clears
```
`questions/` is ground truth; `map.md` is a rebuildable index that gists and links. A filled `## Answer` beats any `State:` line. Detail lives in exactly one place — the question file.
## Numbering
`NN` is zero-padded from `00`, assigned in dependency order (blockers lower), **never reused or renumbered** — links and `Blocked by:` would rot silently. Next = max + 1. `00` is always `00-codebase-context.md`, never anything else. Gaps in the sequence are normal and harmless.
## The five schema lines
Each question file carries five contiguous `Key: value` lines after its H1 — NOT YAML frontmatter, no `---` delimiters:
| Line | Values |
|---|---|
| `Type:` | `grill · HITL` \| `research · AFK` \| `sketch · HITL` \| `legwork · HITL` \| `legwork · AFK` — one token, so type and mode cannot drift |
| `State:` | `open` \| `claimed` \| `resolved` \| `out-of-scope` |
| `Blocked by:` | `none` \| `02, 04` |
| `Claimed:` | `none` \| `<YYYY-MM-DD HH:mm>` |
| `Locked:` | `yes` only when hard to reverse AND surprising without context AND a real trade-off |
**Type meanings.** **grill** (default) — a decision only the human can make. **research** — a fact outside this directory gates it. **sketch** — the human needs something concrete to react to. **legwork** — manual work that must happen before a decision is possible.
## Markers and blocking
**Map markers**, rebuilt from the files every session: `[ ]` open — **these rows ARE the frontier** · `[/]` claimed · `[x]` resolved · `[!]` open but blocked · `[-]` out of scope.
**Unblocked** ⟺ every `NN` in `Blocked by:` is `resolved`. **Stranded:** a blocker gone `out-of-scope` never resolves — the question is not merely blocked. Re-frame its `## Question` to drop the dependency, or rule it out too. Never leave it sitting.
@@ -0,0 +1,522 @@
# Resolution Playbook
> Loaded at the top of MODE B. Work Step 6 routes here by `Type:`; Work Step 8 uses the fog procedure at the end.
>
> **Backend note.** Every resolution technique here is backend-independent — the type table, the primary-source rule, the sketch tiers, the checklist discipline, the `## Answer` anatomy. On `**Backend:** github`, `Type:` is a `pathfinder:<type>-<mode>` label, `## Answer` and `## Evidence` are comments rather than file sections, and a research subagent gets an issue URL instead of a path; see `github-issues.md`. Sketches stay on local disk regardless.
Every question resolves into the SAME shape — a filled `## Answer` plus whatever `## Evidence` backs it. The type only decides how you get there.
| `Type:` | Who drives | Parallel? | What "resolved" means |
|---|---|---|---|
| `research · AFK` | Subagent, alone | **Yes** — many at once | Facts found and cited; no decision made |
| `sketch · HITL` | Agent builds, human reacts | No | The human reacted and chose |
| `legwork · AFK` | Agent, alone | No | The work is done; resulting facts recorded |
| `legwork · HITL` | Human does, agent waits | No | The human confirmed it is done |
| `grill · HITL` | Human decides, agent interrogates | No — but probes batch, up to 3 per turn | The human said it in their own words |
**One question _file_ per session** holds for every row except `research`. HITL rows are never self-answered — an agent that writes its own `## Answer` on a `grill` has broken the skill.
Every file you create along the way — question files, sketch READMEs — opens with the banner:
`> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.`
---
## `research · AFK`
The only type an agent resolves alone, and the only type that may run several at once. Charting fires them in a batch at Chart Step 8; MODE B fires any that appear later the same way.
### Spin up a subagent
One subagent per research question. Do not read the docs yourself in the main session — the point is that the main session's context stays clean for the decision work.
The subagent's brief must carry, verbatim:
1. The absolute path of the question file it owns.
2. The `## Question` text.
3. The instruction to write into `## Evidence` of THAT file and nothing else — `## Answer` and `State:` are written by the session that fired it, at Work Step 7.
4. The primary-source rule below.
5. The no-deciding rule below.
### Primary sources only
Investigate against **primary sources** — official documentation, the library's own source code, the RFC or spec text, the first-party API reference, the vendor's own pricing page, the actual response from a live endpoint. Never a secondary write-up of them. A blog post, a Stack Overflow answer, or a model's recollection is a *lead*, not a source: follow every claim back to the source that owns it, and cite that.
| Claim about | Source that owns it |
|---|---|
| A library's behavior | That library's source or its own docs for the installed version |
| An HTTP API's shape | The vendor's API reference, or a real captured request/response |
| A file format | The format specification |
| A limit or quota | The vendor's own limits page, dated |
| This project's behavior | The code in this repo, by path |
If no primary source can be found, say so explicitly in `## Evidence` and mark the claim UNVERIFIED. An honest gap is worth more than a confident secondhand sentence — the gap becomes a `legwork` question (go run it and see) or a `grill` (the human decides under uncertainty).
**Version-pin everything.** "Node supports it" is not a finding. "Node 20.11 supports it; the project's `.nvmrc` pins 18.19, which does not" is.
### It does not decide
A research question's `## Answer` states the FACTS FOUND and what they imply. It never picks. The moment the findings enable a choice, that choice becomes a **separate `grill`** question, written at Work Step 8.
This split is the whole reason research is safe to run AFK. An agent that reads three storage options and then declares a winner has quietly made an architectural decision the human never saw.
### The dropped git branch
Upstream captures each prototype and research run on a throwaway git branch. **Pathfinder deliberately drops that.** `specs/` is already untracked scratch space, and plan2code never touches git during pathfinder. The question file IS the durable capture. Do not create branches, do not commit, do not mention branches in the `## Answer`.
### Worked example
`questions/04-object-lock-retention.md` after Work Step 7 recorded its subagent findings:
```markdown
> Pathfinder planning note - decisions, not implementation work. Archive with the spec; do not delete.
# Object Lock retention semantics for the archive bucket
Type: research · AFK
State: resolved
Blocked by: none
Claimed: 2026-08-03 14:02
Locked: no
## Question
If archived messages land in S3 under Object Lock, what retention modes exist,
which of them can be shortened or lifted after the fact, and what does each
require of the bucket at creation time? We need the facts before deciding
whether a 7-year hold can coexist with a customer's right-to-delete request.
## Answer
**Facts found**
1. Object Lock supports two retention modes: GOVERNANCE and COMPLIANCE.
2. Under GOVERNANCE, a principal holding `s3:BypassGovernanceRetention` can
shorten or remove a retention period on an object version.
3. Under COMPLIANCE, no principal — including the account root — can shorten
or remove it until the retain-until date passes.
4. Object Lock can only be enabled at bucket creation, and enabling it forces
versioning on permanently.
5. A legal hold is independent of retention mode, has no expiry, and is
toggled by any principal with `s3:PutObjectLegalHold`.
**What this implies**
- A right-to-delete request is satisfiable under GOVERNANCE and is NOT
satisfiable under COMPLIANCE before the retain-until date. These two
obligations cannot both be met on the same object version.
- Retrofitting Object Lock onto the existing `archive-prod` bucket is not
possible; it would require a new bucket and a full copy.
**Decision this enables (not made here)**
Which mode the archive uses, and how right-to-delete is honored under it,
is now a sharp question — see Work Step 8, spun off as a grill.
**Gist:** COMPLIANCE mode is irreversible even for root; GOVERNANCE is
bypassable — and Object Lock cannot be added to the existing bucket.
## Evidence
- Two modes, and the governance-bypass permission —
AWS S3 User Guide, "Object Lock overview", section "Retention modes".
https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lock-overview.html
(retrieved 2026-08-03)
- COMPLIANCE cannot be shortened by any user including the root user —
same page, "Compliance mode" paragraph, verbatim: "no user can overwrite or
delete the object version during the retention period."
- Enable-at-creation-only, and forced versioning —
AWS S3 User Guide, "Enabling Object Lock", first note block.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lock-configure.html
- Legal hold independence and no expiry —
same guide, "Legal holds" section.
- Existing bucket has Object Lock disabled — verified against this account:
`aws s3api get-object-lock-configuration --bucket archive-prod`
returns `ObjectLockConfigurationNotFoundError`. (run 2026-08-03)
- UNVERIFIED: whether our compliance counsel treats GOVERNANCE as sufficient
for the SEC 17a-4 attestation. No primary source exists for this — it is a
human judgment, not a fact. Belongs in a grill.
```
Note what the example does: every claim carries its own citation, a live command counts as a primary source, and the one thing that cannot be sourced is flagged rather than smoothed over.
---
## `sketch · HITL`
Raise the fidelity of the discussion by making something cheap and concrete for the human to react to. "How should this behave?" and "what should this look like?" produce vague answers in the abstract and sharp ones in front of an artifact.
Two tiers. **Start at Tier 1 every time.**
### Tier 1 — paper sketch (the default)
No executable code. You write a concrete thing into the question's `## Evidence`, show it to the human, and they react. Most sketch questions never need more than this.
Shapes that work:
| Shape | Use when the question is |
|---|---|
| Outline | "What are the steps, and in what order?" |
| State table | "What states exist and which transitions are legal?" |
| Worked example with real numbers | "Does this rule produce sane results?" |
| Fake request/response pair | "What should this API actually look like?" |
| ASCII UI | "What goes on this screen and what is primary?" |
| Decision table | "Under which conditions do we do which thing?" |
Rules for a paper sketch: use REAL-looking content, never placeholders. Real customer names, real message counts, real timestamps, real error strings. A table full of `foo` and `item 1` gets nodded at; a table with `retention_expired` in it gets argued with, and the argument is the point.
Present it, then ask the pointed questions the sketch actually opens — up to three, batched, in the same turn as the artifact. A sketch is the one place where several probes come free: the human has the whole picture in front of them, so a second and third question cost them almost nothing, and a state table with three questionable rows should not take three sessions.
Pointed is the discipline that survives batching. Never "thoughts?" — every probe names its row or element: "row 5 says a flagged message still archives when retention expires. Right, or does the hold pin it in place?" Sketches always trip the detail test's second row — the artifact has to sit inline, and a picker cannot carry it — so this is a numbered Q block below the sketch, never a structured tool call. This is the one place the tool-first rule is settled in advance; do not re-litigate it per sketch.
Follow the Q-block formatting rules in `grilling.md` for the probes below the artifact — blank line between every element, options as a bullet list, `---` between probes. A sketch batch is the easiest one to render as a wall of text, because the artifact above it already ate the human's attention.
Putting up a sketch is a **Form B turn** in `trail.md`: close it with `WAITING ON YOU` naming the probes, never a resume command. The human is meant to react to the artifact in this conversation, and a footer telling them to start a new one throws the sketch away.
#### Paper sketch example A — state table
For `questions/06-review-lifecycle.md`:
```markdown
| From | Event | To | Legal? |
|--------------|----------------------|--------------|-------------|
| ingested | policy match | flagged | yes |
| ingested | no match, 24h passes | archived | yes |
| flagged | reviewer clears | archived | yes |
| flagged | reviewer escalates | escalated | yes |
| flagged | retention expires | archived | ← QUESTION |
| escalated | case closed | archived | yes |
| escalated | retention expires | escalated | stays put |
| archived | legal hold applied | held | yes |
| held | hold released | archived | yes |
| held | retention expires | held | hold wins |
Open on this sketch: an item sitting in `flagged` when its retention window
expires. Row 5 currently drops it to `archived` unreviewed. The alternative is
that expiry cannot fire while a human review is outstanding — retention pauses.
```
#### Paper sketch example B — fake request/response pair
For `questions/09-export-job-api.md`:
```markdown
POST /v1/exports
{
"channel_ids": ["ms-teams-legal", "ms-teams-trading"],
"from": "2025-01-01T00:00:00Z",
"to": "2025-12-31T23:59:59Z",
"format": "eml",
"include_attachments": true
}
202 Accepted
{
"export_id": "exp_9fK2mQ",
"state": "queued",
"estimated_messages": 418377,
"estimated_bytes": 12884901888,
"poll_url": "/v1/exports/exp_9fK2mQ",
"expires_at": "2026-08-10T14:00:00Z"
}
GET /v1/exports/exp_9fK2mQ
200 OK
{
"export_id": "exp_9fK2mQ",
"state": "partial_failure",
"messages_written": 418202,
"messages_failed": 175,
"failure_manifest_url": "https://.../exp_9fK2mQ-failures.csv",
"download_urls": ["https://.../exp_9fK2mQ-part-001.zip", "..."]
}
Open on this sketch: `partial_failure` hands back a download plus a manifest of
what is missing. The alternative is all-or-nothing — 175 failures void the whole
12 GB export. Which does a compliance officer actually want at 4pm on a Friday?
```
#### Paper sketch example C — ASCII UI
For `questions/11-reviewer-queue-layout.md`:
```markdown
+----------------------------------------------------------------+
| Review Queue [ Mine 42 ] [ Team 318 ] [ Overdue 7 ] |
+---------------------------+------------------------------------+
| ! 2d K. Ondrusek | From: Kamil Ondrusek |
| "…move the block…" | To: trading-desk (14 members) |
| trading-desk | 2026-08-01 09:14 MS Teams |
|---------------------------| |
| 1d A. Whitfield | Policy hit: BLOCK-TRADE-LANGUAGE |
| "confirming size" | Confidence: 0.91 |
| trading-desk | |
|---------------------------| > can you move the block before |
| 4h R. Iyer | the close? size is 40k |
| "attached the deck" | |
| legal-general | [ Clear ] [ Escalate ] [ Hold ] |
+---------------------------+------------------------------------+
Open on this sketch: the policy hit and its confidence sit in the detail pane,
so the list gives no reason to pick one item over another beyond age. Should the
list rank by confidence instead of age, and show the rule name per row?
```
### Tier 2 — runnable sketch (available, and gated)
Every other plan2code step forbids writing code during planning. **Pathfinder is the one exception**, because some questions genuinely cannot be settled on paper: "does this state model actually hold once you push it through the ugly cases?", "what should this feel like?" A paper state table always looks fine. Driving it by hand for ninety seconds is where it falls over.
**All three gates must open before you write a line of code:**
1. The paper sketch was tried and did not settle it. Not skipped — tried. Say what the paper sketch failed to resolve.
2. There is an obvious way to run it in this project — an existing runtime and task runner. Do not add a package manager, language, or framework for a sketch.
3. The user says go. Ask explicitly: *"Paper didn't settle row 5. I can build a throwaway terminal app under `specs/audit-export/pathfinder/sketch-01/` that lets you drive the state machine by hand — about 60 lines, one command, deleted after. Go?"*
Any gate that stays shut: stay on paper, or convert the question to a `grill`.
#### Rules for a runnable sketch
1. **Throwaway from day one, and clearly marked.** Its README's first line is the banner, and its second says what question it exists to answer.
2. **It lives ONLY at `specs/<idea>/pathfinder/sketch-NN/`** — never in the project's own source tree, never beside the module it is sketching for. This is where pathfinder deliberately departs from upstream: upstream co-locates prototypes with the real code; pathfinder quarantines them, because `specs/` is gitignored scratch and the project tree is not.
3. **One command to run.** Print the exact command to the user. They must not have to remember a path or a flag. If the project has a task runner, use the runner's own idiom, but keep the entry point inside `sketch-NN/`.
4. **No persistence.** State lives in memory. Persistence is what the sketch is checking, not something it leans on. If the question is specifically about storage, use a local file named so its disposability is obvious.
5. **Skip all polish.** No tests, no error handling beyond what makes it run, no abstractions, no "we might want X later."
6. **Surface the full relevant state after every action** (logic) **or on every variant switch** (UI). The user must see the whole picture change, not a delta.
7. **Never merged.** The sketch is not lifted into the project. Only the validated decision it produced survives, in the `## Answer`. Reference the sketch directory from `## Evidence` so a later reader can re-run it, but the code is scaffolding, not output.
#### Pick the branch: logic or UI
| Question shape | Branch | Artifact |
|---|---|---|
| "Does this state model / data model / rule hold?" | **Logic** | One tiny interactive terminal app |
| "What should this look like?" | **UI** | Several radically different variations, switchable |
Getting this wrong wastes the entire sketch. If the question is genuinely ambiguous and the user is unreachable, default by what the question touches — a backend module or a rules engine points to logic, a page or component points to UI — and **state the assumption in the first lines of the sketch's README**, so the human can reject the framing before reading the code.
**Logic branch.** Build the smallest interactive terminal app that pushes the machine through the cases that are hard to reason about on paper. Keep the logic itself pure — a reducer, a state machine, or a small set of pure functions over a plain data type — with the terminal shell as a thin wrapper that imports it and never the reverse. Each frame: clear the screen, print the whole current state one field per line, then print the key legend, e.g. `[f] flag [c] clear [e] escalate [h] hold [t] advance clock 1d [q] quit`. Re-render the entire frame after every keystroke; never append to scrollback. The whole frame fits on one screen. The interesting moment is the user saying "wait, that shouldn't have been possible" — that is a bug in the *idea*, which is the entire point. Add actions on request; sketches evolve.
**UI branch.** Generate **several radically different variations side by side, switchable** — not one polished take. Default to 3, cap at 5. They must disagree about structure: different layout, different information hierarchy, different primary affordance. Three tweaked card grids is wallpaper, not a sketch. If two drafts come out similar, redo one with an explicit constraint against the shape they share. Switch by a URL search param plus a small floating bar (previous / current variant name / next), following whatever routing convention the project already uses — but with the files under `sketch-NN/`. Wire variants to stubbed data, never to real mutations; the question is what it should look like, not whether the backend works. The most valuable feedback is usually "I want the header from B with the list from C" — that hybrid IS the answer, and it goes in the `## Answer`.
#### After a runnable sketch
Record the verdict and the question it settled in `## Answer`. Record the sketch path, the run command, and what the user actually said while driving it in `## Evidence`. Leave the directory in place — it is gitignored scratch, it costs nothing, and the next session may want to re-run it. Never copy any of it into the project.
---
## `legwork · HITL` or `legwork · AFK`
The one type that DOES rather than decides. There is nothing here to research, sketch, or grill — a decision is simply blocked until some manual work happens.
Typical: provisioning access to a system, signing up for a service so its API can actually be judged, moving a data sample somewhere it can be looked at, requesting a sandbox tenant, reading the codebase (`questions/00-codebase-context.md` is always this type).
**It earns its place only by unblocking a decision, never by delivering the destination.**
### AFK or HITL
| Mode | When | How it resolves |
|---|---|---|
| `legwork · AFK` | The agent can do it with the tools it has — read the code, run a query, count rows, inspect a config | Do it, record the facts, resolve |
| `legwork · HITL` | It needs a human's hands, credentials, card, or signature | Hand over a precise numbered checklist and WAIT |
Drive it alone wherever you can. Do not hand a human a checklist for work you could have done yourself.
### The HITL checklist
Numbered, specific, and verifiable — every line names the exact place to click, the exact value to use, and what the human should see when it worked. Plain English throughout: exact names and values where they carry the work, no Pathfinder vocabulary anywhere (see *Say it in plain English* in the grilling playbook). No `- [ ]` checkboxes; question files never carry them.
```markdown
1. Go to https://console.vendor.example/settings/api and sign in with the
shared ops account (credentials in 1Password, item "Vendor Ops").
2. Create an API key named `pathfinder-eval-2026-08`. Scope it to
read-only — untick "Write" and "Admin".
3. Copy the key into 1Password as a NEW item named "Vendor Eval Key".
Do not paste it into this chat or into any file under specs/.
4. On the same page, note the "Rate limit" value shown for the key
and tell me the number.
5. Under Settings > Data, note whether "Historical backfill" is listed
as included or as a paid add-on, and tell me which.
Tell me when 1-5 are done, plus the two values from steps 4 and 5.
```
Then stop and wait. Do not guess the answers, do not proceed to the next question, do not mark it resolved on the assumption it went fine.
Handing over a checklist is a **Form B turn** in `trail.md` — close with `WAITING ON YOU` naming the checklist and the values you asked for, and no resume command. The human may be gone for hours, but the session is still theirs to come back to; only park it as a session end (Form A) once you are actually stopping.
### What its `## Answer` records
Two parts: **what was done**, and **the resulting facts later questions depend on**. Credentials locations (never the credentials), new URLs, row counts, version numbers, quota limits, table shapes, file paths.
```markdown
## Answer
**Done.** Read-only API key provisioned against the shared ops account and
stored in 1Password as "Vendor Eval Key". No key material is stored under specs/.
**Facts other questions depend on**
- Key location: 1Password item "Vendor Eval Key" (ops vault).
- Rate limit: 600 requests/minute per key, burst 1000.
- Historical backfill beyond 90 days is a paid add-on, not included in the
eval tier — so any evaluation against real 2024 traffic needs a purchase.
- Base URL for the eval tenant: https://eval-3f2.vendor.example/api/v2
(differs from the production host in their docs).
**Consequences** — the 90-day eval ceiling means the volume question cannot be
answered against real historical data on this tier; it has to be extrapolated
or the add-on has to be bought. That is a fresh decision, not one to make here.
**Gist:** Read-only eval key in 1Password; 600 rpm; history capped at 90 days
without a paid add-on.
```
### The guard
If the legwork turns out to BE the deliverable rather than an unblocker — you are migrating the data, not sampling it; you are building the integration, not evaluating it — **it is out of scope for pathfinder.** Stop. Say so plainly:
> "This has stopped being legwork that unblocks a decision and become the work itself. Pathfinder plans; it doesn't build. I'm ruling this out of scope and it belongs in planning."
Rule the question `out-of-scope`, mark the map row `[-]`, add one line to `## Out of scope` naming what it turned into, and hand off to `/plan2code-1-plan` for that piece. Never let pathfinder quietly become the implementation.
---
## `grill · HITL`
The default type: a decision only the human can make. Route to **the grilling playbook** — it owns the interrogation technique, the batching rules (up to three independent probes per turn), the detail test that decides whether the batch goes through the structured question tool (the default) or numbered Q blocks (the fallback), and the recommend-then-ask pattern.
Three things this playbook adds on top:
- **Zoom before you grill.** Work Step 5 already had you read the claimed question plus anything it references. Bring the resolved neighbors' gists into the first message so the human is not re-litigating settled ground.
- **Batch probes, not question files.** One question file per session is unchanged. A batch of three probes resolves ONE `questions/NN-*.md`; it is not licence to close three of them.
- **The human's own words.** A `grill` resolves only through live exchange. Never write the `## Answer` from what you inferred they would probably say — and a probe they skipped twice is unanswered, not decided.
A good grill `## Answer` contains four things:
1. **The decision** — stated flatly, in the human's terms, not hedged.
2. **What was rejected and why** — the alternatives that were live during the conversation, each with the reason it lost. This is the part that stops the decision from being reopened in three weeks.
3. **The consequences for other questions** — which open questions this constrains, which fog patches it just made sharp, which resolved answers it complicates. Name them; never number them in prose.
4. **A one-line bold Gist.**
---
## Writing the `## Answer`
Same anatomy for every type. Append it at Work Step 7; never edit `## Question` to match the answer.
| Part | Required | Content |
|---|---|---|
| The decision (or, for `research`, the facts found) | Always | What was settled, stated flatly |
| Rejected alternatives, with reasons | When alternatives existed | Each option that lost, and why |
| Consequences | Always | Effects on other questions, named not numbered |
| `**Gist:**` | Always | One line, last |
| `## Evidence` | When there is any | Sources, sketch paths, transcript quotes, commands run |
Worked example — `questions/03-export-format.md`:
```markdown
## Answer
**Decision.** Exports are written as one `.eml` file per message inside a ZIP,
with a top-level `manifest.csv` giving message id, channel, participants,
timestamp, SHA-256, and relative path. One ZIP per 2 GB, numbered `part-001`.
**Rejected**
- **Single NDJSON file.** Compact and trivially streamable, but the review
vendors named in Codebase context both ingest `.eml` natively and neither
parses NDJSON. Rejected because it moves the conversion cost onto the
customer's e-discovery team.
- **PST.** What the legal team asked for by name, but PST is a single-writer
format with a practical 50 GB ceiling and no first-party writer outside
Outlook. Rejected on the ceiling alone — the 2025 trading-desk export is
~12 GB and growing 40% year over year, so the ceiling is 3 years out.
- **One ZIP, no parts.** Rejected because S3 presigned downloads over 5 GB
fail on several corporate proxies the support team has already seen.
**Consequences**
- Makes [Export job API](./questions/09-export-job-api.md) sharper: the
response must return an ARRAY of download URLs, not one.
- Constrains [Integrity attestation](./questions/12-integrity-attestation.md) —
a per-message SHA-256 already exists in the manifest, so attestation can hang
off the manifest rather than needing a separate hash pass.
- Kills the "streaming export" fog bullet: parts and streaming are exclusive.
Removed from `## Not yet specified`.
**Gist:** One `.eml` per message in 2 GB ZIP parts, with a manifest.csv
carrying per-message SHA-256.
```
### The gist
The gist is what gets copied into the map row. It is not a summary of the answer — it is the one line that lets a future session decide, at a glance, whether to open the file.
- One line. Fits in a table row without wrapping twice.
- Says what was DECIDED, not what was discussed. "Chose ZIP parts" is weak; "One `.eml` per message in 2 GB ZIP parts" is judgeable.
- Carries the number or name that matters, if there is one.
- Never the full answer. If it needs a semicolon and a subordinate clause, cut it.
The map row it produces:
```markdown
- [x] [Export format](./questions/03-export-format.md) — one `.eml` per message in 2 GB ZIP parts, manifest.csv carries SHA-256
```
---
## Graduating the fog (Work Step 8)
`## Not yet specified` is the fog: in-scope questions you can *see* coming but could not phrase sharply when you wrote the map. Resolving a question clears the fog immediately ahead of it. Step 8 is where you collect what just became visible.
**The test is whether you can state the question precisely NOW — not whether you can answer it.** A question you cannot act on for weeks still gets a file, with `Blocked by:` filled in. A question you could answer this minute but cannot phrase without hand-waving stays fog.
### Procedure
1. **Re-read `## Not yet specified` in full.** Every bullet, every session, after every resolution. Not the ones you remember — you are assuming no memory of prior sessions, and the bullet that graduates is usually the one you forgot was there.
2. **Ask of each bullet: did the answer just make this sharp?** Can you now write a `## Question` paragraph a stranger could act on, without "we'll need to figure out" anywhere in it?
3. **If yes, write the question file NOW** — in this same session, before Work Step 9. Next `NN` = max existing + 1, never reused. Fill all five metadata lines, with `Blocked by:` wired in the same pass. Add its `[ ]` or `[!]` row to `## Question Checklist`.
4. **DELETE the bullet from `## Not yet specified`.** Immediately, in the same edit as writing the file.
5. **If no, leave the bullet alone** — untouched, not reworded into something that merely sounds sharper.
6. **Re-check the counts.** One fog patch may graduate into three questions, or into none. Both are normal. A patch that graduates into three was written at the right coarseness; a patch that graduates into exactly one every time was probably a question all along.
### The bullet you left behind
**A fog bullet still sitting in `## Not yet specified` after its question file exists is the single most common drift in this skill.** It is quiet — nothing errors, the map still renders — and it is corrosive:
- The Clearing Gate requires `## Not yet specified` to be EMPTY. A stale bullet blocks the gate forever, so the map never clears even when every question is resolved.
- A later session reads the bullet, does not recognize the question file as the same thing under a different phrasing, and writes a duplicate. Now two files hold half a decision each.
- It violates the rule that detail lives in exactly one place.
Detection is cheap: after writing any graduated question file, re-read `## Not yet specified` top to bottom and confirm the bullet is gone. If a bullet reads like a question you have already written a file for, delete the bullet — the file always wins.
### Ruling a question out of scope mid-work
The destination fixes the scope. When a resolution reveals that a question — the one you just claimed, or another on the map — sits past the destination:
1. Set that file's `State: out-of-scope`. Do NOT write a `## Answer`; there is no decision, only a scope boundary. Add one line under `## Question` saying why it is out.
2. Set its map row to `[-]`.
3. Add one line to `## Out of scope`: the name as a link, plus the reason.
4. Check for **stranded** questions — anything whose `Blocked by:` names it. A blocker that is `out-of-scope` will never be `resolved`, so the dependent is permanently blocked. Re-frame its `## Question` to drop the dependency, or rule it out too. Never leave it sitting.
Out-of-scope work never graduates back. The frontier stops at the destination. It returns only if the destination is redrawn, and then as a fresh effort with a fresh map.
```markdown
## Out of scope
- [Slack connector](./questions/08-slack-connector.md) — the destination names
MS Teams only; Slack is a separate effort with its own compliance posture.
```
### Re-framing or deleting a question the answer invalidated
An answer can also break questions that already exist. Three cases:
| What happened | Do this |
|---|---|
| The question still matters but is asked wrong | **Re-frame.** Rewrite `## Question` in place. Keep `NN`, keep the file, keep the links. Add one line noting which answer forced the re-frame. |
| The question no longer exists — the answer subsumed it | **Delete the file and its map row.** Add one line to the answering question's `## Answer` consequences saying what it absorbed. Do not renumber anything. |
| The question is now two questions | **Re-frame the original to the narrower half; write a new file at max+1 for the other.** Wire `Blocked by:` between them if one gates the other. |
Never leave a question standing that you know is wrong on the theory that a later session will notice. It will not — it assumes no memory, and a well-formed `## Question` reads as intentional.
Two hard constraints on all three cases: **`NN` is never reused and never renumbered** — links and `Blocked by:` lines would rot silently. And every re-frame or deletion is reflected in the map's `## Question Checklist` in the same edit, so the index never disagrees with the files.
@@ -0,0 +1,171 @@
# Trail Footer
> Loaded at the top of the skill. Read once; it applies to EVERY response in both modes. Defines the map visual and the pathed resume command that close each turn.
>
> **Backend note.** The glyphs, the two forms, and the discipline are identical either way. On `**Backend:** github` the inputs come from the sub-issue query rather than the checklist, and Form A's command is the map issue URL — see `github-issues.md`.
The trail is how a human with no memory of the last session sees, at a glance, how far the map has come and what is left — and copies the exact command to resume without hunting for a path that lives in gitignored `specs/`. It closes every response once a map exists.
It is **presentation only**. Rebuild it fresh from `map.md` each response; never write it to disk, and never let it emit a loop token (`TASK_COMPLETE`, `PHASE_COMPLETE`, and the rest). It is an index of an index — the question files remain ground truth.
---
## When it renders
| Situation | Footer? |
|---|---|
| `map.md` exists and the turn **ends the session** — Session End, a cleared map, a fully blocked frontier | **Yes**, both parts: the trail, then the resume command |
| `map.md` exists and the turn **asks the human something** — a probe batch, a sketch put up for reaction, a `legwork · HITL` checklist | **Yes**, but with the waiting form of Part 2. Never a resume command — see below |
| Intent Gate (Step 0), the no-fog off-ramp, or any route-and-stop before Step 5 | **No** — no map on disk yet, and no path to resume. There is nothing to draw. |
One idea per session, so there is only ever one trail. Draw the trail for the active idea and no other.
---
## The two parts
Always in this order, after everything else in the response:
1. **The trail** — the horizontal path plus its legend and confidence line. Identical on every turn.
2. **The next line** — either the pathed resume command or the waiting notice, chosen by turn type.
### Part 1 — the trail
```
🧭 audit-log-export · Working · 2/7 cleared
START ●━●━◉··○··○··⊘··⊝ ····⚑
● done · ◉ here · ○ open · ⊘ blocked · ⊝ out of scope · ⚑ destination · ~4 fog
1 Codebase context ✔ · 2 Export format ✔ · 3 Row-count ceiling ◀ here
4 Export authorization · 5 Testing posture · 6 Delivery channel (blocked:3) · 7 SIEM push (out of scope)
Confidence: solid, but Risk is borderline.
```
How each line is built, top to bottom:
- **Heading** — `🧭 <idea> · <Status> · <resolved>/<total> cleared`. `<Status>` is the map's `**Status:**` verbatim (`Charting` / `Working` / `Cleared`). `<resolved>` counts `[x]` rows; `<total>` counts every question row **except** out-of-scope `[-]` rows (a ruled-out question is off the route, not an unfinished stop).
- **The path** — one glyph per question row in `NN` order, left to right, from `START` to the `⚑` destination. Connectors carry meaning: solid `━` joins stops already walked (everything up to and including `◉ here`); dashed `··` joins stops still ahead. After the last question glyph, a fog stretch `····` then `⚑` — drop both if there is no fog and join straight to `⚑` with `━`.
- **The legend** — only the glyphs actually on this path, so a map with no blocked question does not advertise `⊘`. Append `~N fog` when `## Not yet specified` holds N bullets.
- **The named legend** — the same `NN` order, `<NN> <Name>` each, `·`-separated, wrapping across lines as needed. Tag each: `✔` resolved, `◀ here` the claimed one, `(blocked:NN)` with its blocker, `(out of scope)`. Names come straight from the checklist rows.
- **Confidence** — the map's four internal scores (Requirements / Feasibility / Integration / Risk, each out of 25) exist for the Clearing Gate, not the human. Never print the raw numbers or the `R·F·I·K` letters. Instead, reduce them to one plain-English line:
| Scores | Line |
|---|---|
| All four ≥ 20 | `Confidence: solid.` |
| All four ≥ 18, one or more sitting at 18-19 | `Confidence: solid, but <Dimension> is borderline.` (name every dimension in that range, comma-separated) |
| Any dimension < 18 | `Confidence: not yet — <Dimension(s)> still need work.` |
Omit the whole line if the map has no `**Confidence:**` line yet.
### Glyph reference
| Glyph | Map marker | Meaning |
|---|---|---|
| `●` | `[x]` | resolved — a stop already walked |
| `◉` | `[/]` | the claimed question — you are here |
| `○` | `[ ]` | open, on the frontier |
| `⊘` | `[!]` | open but blocked |
| `⊝` | `[-]` | ruled out of scope |
| `⚑` | — | the destination |
| `····` | — | the fog still between the last question and the destination |
### Part 2 — the next line
Part 2 answers exactly one question for the human: **is this turn over, or is it my move?** It has two forms, and the **turn type** picks between them — not the map's status.
#### Form A — the turn ends the session
```
NEXT STEP · start a new conversation and run:
`/plan2code-0-pathfinder specs/audit-log-export/pathfinder`
```
The path is always relative and always ends `/pathfinder` — that is exactly the argument this skill's Auto-Discovery resolves an idea from, so the human pastes it back with zero edits. Fill `<idea>` from the active map's directory; never leave the `<spec-folder>` placeholder in a rendered footer.
The command target follows the map's status, but the sentence is always the same shape — `start a new conversation and run:` followed by the command on its own line:
| Status | Command |
|---|---|
| `Charting` / `Working` | `/plan2code-0-pathfinder specs/<idea>/pathfinder` |
| `Cleared` | `/plan2code-1-plan` — point at the PLAN-DRAFT the Clearing Gate wrote |
| `Cleared`, but `## Ground rules` records `AGENTS.md` **absent** | `/plan2code-init` FIRST (offer `questions/00-codebase-context.md`), then `/plan2code-1-plan` |
#### Form B — the turn asks the human something
Any turn whose next move is theirs and happens **in this same conversation**: a batch of grill probes, a sketch put up for reaction, a `legwork · HITL` checklist, the Chart Step 2 destination grill, the Chart Step 4 frontier grill.
```
WAITING ON YOU · answer here, in this conversation:
Q1 Duration model · Q2 Contract while held · Q3 Placement & restore
```
Name every outstanding item at its stable number so a partial reply is cheap to give and a dropped probe is visible to both of you. `·`-separated on one line; one per line if the names run long. Never more items than the three-probe cap allows.
**Emit no resume command on a Form B turn.** There is nothing to resume — the session is alive and holding a claim. A resume command here reads as *we're done*, and the human either walks away mid-decision or burns the next turn asking what you meant. This is the single most common way the footer misfires, and it costs the batch the round trips batching was introduced to save.
One exception: `grilling.md`'s unreachable-human procedure. Parking a mid-grill question and stopping IS a session end — use Form A, and say in the body that the question is parked mid-grill with its batch outstanding.
---
## State-by-state examples
**Charting, mid-grill, no question claimed yet** (`◉` is omitted; the frontier head is the first `○`). The Chart Step 4 frontier grill is a Form B turn — the batch is above, so the footer says stay:
```
🧭 audit-log-export · Charting · 1/5 cleared
START ●··○··○··○··○ ····⚑
● done · ○ open · ⚑ destination · ~6 fog
1 Codebase context ✔ · 2 Export format · 3 Export authorization · 4 Testing posture · 5 Row-count ceiling
Confidence: not yet — Requirements, Feasibility, Integration, Risk still need work.
WAITING ON YOU · answer here, in this conversation:
Q1 Data · Q2 Surface · Q3 Permissions
```
**Working, a claimed question mid-grill** — probes are out, the claim is held, nothing is resolved yet. Same trail, Form B again. Note that Q2 is a re-ask carried over at its original number from a batch that came back partial:
```
🧭 audit-log-export · Working · 2/7 cleared
START ●━●━◉··○··○··⊘··⊝ ····⚑
● done · ◉ here · ○ open · ⊘ blocked · ⊝ out of scope · ⚑ destination · ~4 fog
1 Codebase context ✔ · 2 Export format ✔ · 3 Row-count ceiling ◀ here
4 Export authorization · 5 Testing posture · 6 Delivery channel (blocked:3) · 7 SIEM push (out of scope)
Confidence: not yet — Feasibility, Integration, Risk still need work.
WAITING ON YOU · answer here, in this conversation:
Q2 Ceiling behaviour past the cap (re-ask) · Q3 Who sees the truncation warning
```
**Frontier fully blocked** — report the chain in the body; the trail shows why nothing is takeable. Work Step 4 stops the session here, so Form A:
```
🧭 audit-log-export · Working · 5/7 cleared
START ●━●━●━●━●━⊘··⊘ ⚑
● done · ⊘ blocked · ⚑ destination
6 Delivery channel (blocked:2) · 7 Notification (blocked:6)
Confidence: not yet — Feasibility, Risk still need work.
NEXT STEP · start a new conversation and run:
`/plan2code-0-pathfinder specs/audit-log-export/pathfinder`
```
**Cleared** — every stop walked, fog empty, command hands off:
```
🧭 audit-log-export · Cleared · 8/8 cleared
START ●━●━●━●━●━●━●━●━⚑ arrived
Confidence: solid.
NEXT STEP · start a new conversation and run:
`/plan2code-1-plan`
```
At `Cleared` the named legend is optional — the destination is reached and the PLAN-DRAFT is the thing to point at. Keep the confidence line; the Clearing Gate leaned on it.
---
## Discipline
- **Never print a resume command on a turn that asks a question.** The footer must not tell the human to leave a conversation you are still waiting in. Before you write Part 2, ask whether the response above it ends with something for them to answer; if it does, Form B, no exceptions but the parked-grill one.
- **Alignment is not the point.** Glyphs sit in `NN` order and the legend names them in the same order; do not burn effort column-aligning numbers under waypoints across variable-width glyphs. Legibility over pixels.
- **Rebuild, never cache.** The markers come from the current `map.md`, which Work Step 2 has already reconciled against the question files this session. A footer that disagrees with the checklist above it means you drew from memory.
- **One trail.** Never render two ideas' trails, and never invent a stop the map does not list.
+151
View File
@@ -0,0 +1,151 @@
# 🧭 PATHFINDER MODE
Start all PATHFINDER MODE responses with '🧭 [PATHFINDER: Chart - Step X: Name]' or '🧭 [PATHFINDER: Work - Step X: Name]'.
## Role
Pathfinder, not architect. An idea has arrived too big or unclear to plan. Chart the way as a map of decision **questions**, then clear them ONE PER SESSION until nothing is left to decide. Hand off to `/plan2code-1-plan`.
Read references/grilling.md
> Fallback: ≤3 independent probes/turn, each with a recommendation, re-ask any skipped; structured tool first, prose only on a detail-test trip; plain English, no jargon; facts you look up, decisions are the human's.
## Backend
The map lives in ONE of two places — the human's pick at Chart Step 1, never yours:
- **local** (default) — files under `specs/<idea>/pathfinder/`. Private, gitignored, solo.
- **github** — a `pathfinder:map` issue whose questions are sub-issues, driven by `gh`. Shared, visible in the tracker UI, parallel.
Read references/github-issues.md — REQUIRED on `github`, skip it on `local`.
> Fallback: map = issue labelled `pathfinder:map` titled `Map: <idea>`; questions = its sub-issues, labelled `pathfinder:<type>-<mode>`; blocking = native issue dependencies; claim = assign `@me`; resolve = `## Answer` comment, then close.
Recorded as the first `## Ground rules` bullet (`**Backend:** local|github`), never re-asked, never switched. Either way the PLAN-DRAFT lands in local `specs/<idea>/` — downstream steps read files, not issues.
## Project Context
Load `./AGENTS.md` if it exists — its conventions govern; never re-ask what it answers. If missing, do NOT ask here; fold it into the Step 0 gate batch: *"No `AGENTS.md`. Pathfinder can chart without it. Continue, or run `plan2code-init` first?"* Record it in `## Ground rules` so no later session re-asks.
## Rules
- **Plan, don't do.** Every question resolves a DECISION. The pull to just build it is the edge of the map — hand off.
- **Confirm before creating anything.** No files, no issues, until the Intent Gate (Step 0) and backend pick (Step 1) return.
- **One question per session** (`research` excepted — parallel subagents).
- **Refer by name.** "[Export format](<link>)", never "02" or "#42" in prose. Bare ids belong on `Blocked by:` lines and in commands.
- **HITL questions are never self-answered.** Ask and wait. An agent that answers its own grill has broken the skill.
- **Questions are ground truth; the map is a rebuildable index.** A filled `## Answer` beats any state marker; detail lives in one place.
- **Never write implementation code** into the project. Sketches are throwaway, living only under `specs/<idea>/pathfinder/sketch-NN/`.
- **Reserved names — never create inside `pathfinder/`:** `overview.md`, `phase-<N>.md`, `PLAN-DRAFT-*.md`, `PLAN-CONVERSATION-*.md`.
- **Never emit the loop's completion tokens under `specs/`** — `TASK_COMPLETE`, `PHASE_COMPLETE`, `ALL_TASKS_COMPLETE`, `IMPLEMENTATION_COMPLETE`, `SPEC_COMPLETE`, `WORK_COMPLETE`. It scans for them.
- **No `- [ ]` checkboxes inside question files**, and no `METRICS_JSON` anywhere. Pathfinder is not a metered step.
## Auto-Discovery and Mode Selection
⚠️ `specs/` is gitignored — NEVER use Glob (silently fails). Shell only: `ls specs/` (Bash) or `Get-ChildItem specs/` (PS).
**Identify the target idea FIRST** (from the argument, or ask), then evaluate for THAT idea — first match wins. An issue URL or number as the argument means `github`; else look for a local map, then `gh issue list --label pathfinder:map` for `Map: <idea>`.
| Condition | Route |
|---|---|
| No map, but `specs/<idea>/overview.md` exists | Documented — offer `/plan2code-3-implement`. STOP |
| No map in either backend | MODE A, Step 0 (Intent Gate) |
| Map `**Status:** Charting` | MODE A, resume at Step 6 |
| Map `**Status:** Working` | MODE B |
| Map `**Status:** Cleared` | Point at the PLAN-DRAFT and `/plan2code-1-plan`. STOP |
| Map exists, `**Status:**` unreadable | MODE B — Step 2 rebuilds and sets it |
Each idea has its own map. Never chart two in one session.
## Questions
Read references/questions.md — the `local` format. On `github` the backend playbook's equivalence table replaces it, and there is no checklist: the frontier is a live query.
> Fallback (`local`): `map.md` indexes; `questions/NN-<slug>.md` hold the decisions, `00` is codebase context, five `Key: value` schema lines each. Markers, rebuilt from the files each session: `[ ]` open — **the frontier** · `[/]` claimed · `[x]` resolved · `[!]` blocked · `[-]` out of scope.
## MODE A: Chart
Read references/chart.md
> Fallback: confirm the outcome with the human FIRST; only then grill the destination, then breadth-first; write the map and one question per sharp decision.
0. `[Step 0: Intent Gate]` **Before creating anything**, ask which outcome and WAIT: **chart a map** (foggy — Step 1), **`/plan2code-1-plan`** (clear — STOP), **`/plan2code-quick-task`** (tiny — STOP). HITL, never self-select "chart".
1. `[Step 1: Name and backend]` Only after the gate returns "chart." Confirm the kebab-case idea name, then ask — HITL, never self-picked — **local files or GitHub Issues?** Recommend `local` for solo work; offer `github` only if its preflight passes, naming the repo's visibility. THEN the first write.
2. `[Step 2: Destination]` Grill until it is one or two lines. It fixes scope — settle it first.
3. `[Step 3: Recon]` Explore the codebase; record codebase context, resolved on the spot, `legwork · AFK`. On `github` hold it until Step 6 so a Step 4 off-ramp leaves no litter.
4. `[Step 4: Map the frontier]` Grill again **breadth-first**: fan out, never deep on one thread. Surface the open decisions and what is takeable now.
5. `[Step 5: Create the map]` `**Status:** Charting`, Destination, Ground rules (backend first), an empty index, the fog in `## Not yet specified`. Say once where it lives and who can see it.
6. `[Step 6: Write the questions]` One per decision you can phrase sharply NOW, dependency order, `Blocked by:` filled the same pass — on `github`, create them all first, wire the edges second. The rest stays fog. Always include a `grill · HITL` testing-posture question; `/plan2code-1-plan` Phase 1 needs it.
7. `[Step 7: Index]` Fill `## Question Checklist` from the files (`local` only). Set `**Status:** Working`.
8. `[Step 8: Fire research]` One subagent per `research` question, in parallel. Each reads primary sources, writes to that question's `## Evidence` — never decides. Then Session End.
**No fog at Step 4?** Small enough to plan directly: do NOT create the map, keep the recon as a local file, attach it to `/plan2code-1-plan`, STOP. Charting resolves nothing by hand — stop at Step 8.
## MODE B: Work
Read references/resolve.md
> Fallback: resolve by type — research reads sources, sketch makes something concrete, grill interviews the human, legwork does the manual work.
Assume NO memory of any prior session.
1. `[Step 1: Load]` Read the map whole. No question yet.
2. `[Step 2: Reconcile]` **Always.** Read every question. `## Answer` written but the state disagrees? The answer wins. Claimed with no `## Answer`? A crash: release it, say so. Rebuild every marker from the questions.
3. `[Step 3: Frontier]` Every question open, unclaimed, and unblocked. First in order.
4. `[Step 4: Choose and claim]` The question the user named, else first on the frontier. Mark it claimed on the question and the map, **saved before any work.** Frontier empty but questions remain? All blocked — report the chain, STOP. Stranded on an `out-of-scope` blocker? Re-frame or rule out, re-run Step 3. Nothing open? Go to The Clearing Gate.
5. `[Step 5: Zoom]` Read the claimed question in full, plus any closed question it references. Obey `## Ground rules`.
6. `[Step 6: Resolve]` Route by type per the resolve playbook. HITL needs the human's own words.
7. `[Step 7: Record]` Write `## Answer`: the decision, what was rejected and why, consequences, a one-line `**Gist:**`. Sources under `## Evidence`. Mark it resolved, index the gist on the map, bump `**Updated:**`.
8. `[Step 8: Graduate]` Fog now sharp? Write those questions, delete the graduated bullets. Past the destination? Rule it out of scope, one line in `## Out of scope`. Invalidated? Re-frame or rule out.
9. `[Step 9: Gate]` Run The Clearing Gate, then Session End.
## The Clearing Gate
Read references/handoff.md
> Fallback: write `specs/<idea>/PLAN-DRAFT-<YYYYMMDD>.md` from the map, Status `Phase 3 Complete - Resume at Phase 4`, then route to `/plan2code-1-plan`.
The map clears only when ALL hold:
1. Nothing open, claimed, or blocked
2. `## Not yet specified` is EMPTY
3. The destination is reachable with nothing left to decide
4. Every confidence dimension (Requirements, Feasibility, Integration, Risk) scores ≥ 18/25
Any failing: name it, keep working. All passing: follow the handoff playbook, set `**Status:** Cleared`, stop. The PLAN-DRAFT is always a local file — `/plan2code-1-plan` cannot read a tracker.
## Trail Footer
Read references/trail.md
> Fallback: once the map exists, close every response with a one-line path of markers (`●` done · `◉` here · `○` open · `⊘` blocked · `⊝` out of scope) from `START` to `⚑`, a numbered legend of question names, plus a plain-English confidence note.
Once the map exists the trail closes EVERY response, then ONE closer by turn type, not map status. Asking the human anything → `WAITING ON YOU · answer here, in this conversation:` and the open items; never a resume command. Ending the session → `NEXT STEP · start a new conversation and run:` plus `/plan2code-0-pathfinder specs/<idea>/pathfinder` (the map issue URL on `github`), or `/plan2code-1-plan` once `Cleared`.
## Session End
Report the question resolved (by name), its gist, what graduated from the fog, what's still open. Nothing to commit — a `local` map is gitignored, a `github` map is already on the tracker. Then the mascot, then the Trail Footer.
```
╭───╮
│ ★ │
│ ◡ │ One more decision down. The fog is thinner!
╰───╯
```
**When the map cleared**, the mascot says `The way is clear! Time to plan!` and the footer routes to `/plan2code-1-plan` — or `/plan2code-init` FIRST if `## Ground rules` records `AGENTS.md` absent.
## Abort / Recovery
| Issue | Action |
|---|---|
| Session stops mid-question, or the map drifted | Release the claim, note why. Work Step 2 repairs the map; the questions always win. |
| Frontier empty, fog remains | Not sharp yet. Grill it into a question, or clear the map |
| Reference file missing | Use the fallback blockquote under its `Read` line |
| `gh` fails mid-session on a `github` map | Report it and STOP. Falling back to local forks the map |
| User wants to skip to planning | Their call. Say what is undecided, route to `/plan2code-1-plan` |
## Learning Capture
If charting surfaced project-specific insights, suggest `/plan2code-init-update` to capture them in `AGENTS.md`.
+1 -1
View File
@@ -201,7 +201,7 @@ Report and resolve issues before continuing.
After revision is applied, optionally clean up with user confirmation:
1. Archive the previous PLAN-DRAFT: rename to `PLAN-DRAFT-<date>-prev.md` (preserves revision history)
2. Remove research or scratch files created during revision that are no longer needed
2. Remove scratch files created during revision that are no longer needed. NEVER remove `pathfinder/` — it is the decision record behind the plan
3. Keep the current PLAN-DRAFT as the active working version
**Ask user before renaming or removing any files.**
@@ -0,0 +1,72 @@
# Community Feedback Submission (STEP 6.5 detail)
> Loaded by `src/plan2code-4-finalize.md` STEP 6.5. This file has no character limit (see `AGENTS-architecture.md` Reference Files).
## 1. Generate a fresh `run_id`
Compute your own current timestamp and a freshly generated random 4-character hex string — do not reuse the example value shown below or anywhere in this spec. Format: `run-<YYYYMMDD>-<HHMMSS>-<4-hex-chars>` (e.g. `run-20260715-143000-a1b2`).
## 2. Assemble the payload
Read `PLAN-DRAFT-*.md`, `phase-*.md`, and `overview.md` in `specs--completed/<feature-name>/` (STEP 6 has already archived them there) and reason over their content to gather:
- `step1`: `final_confidence`, `confidence_breakdown` (`{requirements, feasibility, integration, risk}`), `clarification_rounds`, `tech_stack_revision_rounds`, `verification_gaps_found`, `functional_requirements_count`, `non_functional_requirements_count`, `risk_count`, `phase_count`
- `step2`: `total_tasks`, `phase_count`, `parallel_groups_identified`, `requirement_coverage_percent`, `verification_items_added`
- `step3`: `task_completion_rate`, `tasks_completed`, `tasks_total`, `blocker_count`
- `step4`: `completion_rate_at_audit`, `verification_failures_found`, `documentation_updates_needed`, `archival_succeeded` (safe to read now that Step 6 has run)
- `plan2code_version` — from `version.json`
- `prompt_versions_short` — first 12 characters of each of the 8 prompt file names' content (`plan`, `revise_plan`, `document`, `implement`, `finalize`, `init`, `init_update`, `quick_task`). You do not have `sha256File()` available — note these as best-effort/approximate if you cannot compute a real hash, or omit the field entirely if you cannot.
**Never include** `project.name` or any bulky arrays (e.g. `tasks_per_phase`).
Include the feedback collected at Step 5 as `user_feedback`: `overall_rating`, `rating_reason`, `what_went_well`, `what_went_poorly`.
Assemble the full nested JSON object matching this schema exactly:
```json
{
"schema_version": "1.0",
"run_id": "run-<YYYYMMDD>-<HHMMSS>-<4-hex>",
"plan2code_version": "<from version.json>",
"prompt_versions_short": { "plan": "...", "revise_plan": "...", "document": "...", "implement": "...", "finalize": "...", "init": "...", "init_update": "...", "quick_task": "..." },
"step1": { "final_confidence": 0, "confidence_breakdown": { "requirements": 0, "feasibility": 0, "integration": 0, "risk": 0 }, "clarification_rounds": 0, "tech_stack_revision_rounds": 0, "verification_gaps_found": 0, "functional_requirements_count": 0, "non_functional_requirements_count": 0, "risk_count": 0, "phase_count": 0 },
"step2": { "total_tasks": 0, "phase_count": 0, "parallel_groups_identified": 0, "requirement_coverage_percent": 0, "verification_items_added": 0 },
"step3": { "task_completion_rate": 0, "tasks_completed": 0, "tasks_total": 0, "blocker_count": 0 },
"step4": { "completion_rate_at_audit": 0, "verification_failures_found": 0, "documentation_updates_needed": 0, "archival_succeeded": true },
"user_feedback": { "overall_rating": 0, "rating_reason": "...", "what_went_well": "...", "what_went_poorly": "..." }
}
```
## 3. Render the GitHub Issue
- **Title:** `` `[Feedback] v<plan2code_version> — rating <N>/10` `` (e.g. `[Feedback] v1.15.3 — rating 8/10`)
- **Body:** a short human-readable markdown summary (version, rating, headline numbers such as completion rate and confidence), followed by the full payload as `<!-- METRICS_JSON {...} -->` (same HTML-comment convention used in Step 7's own summary block)
- **Label:** `community-feedback`
Estimate the combined URL-encoded size of `title` + `body` + `labels`. If it exceeds roughly 8KB, warn the user and offer to truncate the longest free-text `user_feedback` field(s) — starting with `rating_reason`, then `what_went_well`/`what_went_poorly` — before proceeding. This size limit only affects the browser/print fallback tiers (below), not the `gh` CLI tier.
## 4. Preview and approval gate
Display the exact rendered title, full body (including the `METRICS_JSON` block), and label(s) to the user, mirroring the Step 4 Documentation Review pattern:
```
╭───╮
│ ● │
│ ~ │ Ready to submit your feedback to the maintainer!
╰───╯
```
> Reply "approve" to proceed with submission, or "skip" to cancel.
Do NOT proceed to submission without an explicit "approve" reply. A "skip" or any non-approval reply cancels this sub-step entirely and proceeds to Step 7 with no submission.
## 5. Tiered submission
On approval, attempt each tier in order until one succeeds:
1. **Tier 1 (primary):** Attempt `gh issue create --repo jparkerweb/plan2code --title "<title>" --body "<body>" --label community-feedback` via your shell tool. If it succeeds, report the created issue URL to the user and stop.
2. **Tier 2 (secondary):** If `gh` is not installed or not authenticated (command fails), construct the URL `https://github.com/jparkerweb/plan2code/issues/new?title=<url-encoded title>&body=<url-encoded body>&labels=community-feedback` and attempt to open it in the user's default browser using the OS-appropriate command (`start "<url>"` on Windows, `open "<url>"` on macOS, `xdg-open "<url>"` on Linux). Tell the user they still need to click "Submit issue" themselves since they must be logged in.
3. **Tier 3 (tertiary):** If no browser can be opened (e.g. no shell tool access, headless/remote session), print the same URL from Tier 2 to the terminal/chat: "Please open this URL in your browser and click 'Submit issue' to share your feedback: `<url>`".
After any tier succeeds (or the user manually confirms Tier 3 submission), proceed to Step 7 as normal.
+17 -19
View File
@@ -203,9 +203,9 @@ Do NOT make documentation changes without user approval.
`🧹 [FINALIZATION STEP 5: User Feedback]`
**Objective:** Collect optional user feedback before archival.
**Objective:** Collect optional feedback before archival.
Ask the user: "Would you like to provide feedback on this workflow run? (optional)"
Ask the user: "Would you like to provide feedback on this run? (optional)" If no, skip to Step 6.
If yes, collect:
1. **Rating** (1-10): "How would you rate this workflow run overall?"
@@ -225,7 +225,7 @@ Append to `overview.md` (in the active spec directory):
| Went Poorly | [response] |
```
If the user declines, skip and proceed to Step 6 (Spec Cleanup).
Ask: "Submit this feedback + run metrics to the maintainer via GitHub? (optional)" Hold as "submission consent". If declined, skip to Step 6.
---
@@ -238,30 +238,27 @@ If the user declines, skip and proceed to Step 6 (Spec Cleanup).
**Confirm with user before moving files.**
1. Create: `specs--completed/<feature-name>/`
2. Move all files from `specs/<feature-name>/`:
2. Move all contents of `specs/<feature-name>/`:
- `overview.md` (with completion summary)
- All `phase-X.md` files
- `PLAN-DRAFT.md` (if present)
- `PLAN-CONVERSATION-*.md` (if present)
3. Remove any temporary research or scratch files not part of the final spec record
- `PLAN-DRAFT.md`, `PLAN-CONVERSATION-*.md`, `pathfinder/` (if present)
3. Remove temporary scratch files not part of the final spec record
4. Verify original directory empty and can be removed
```
specs/
└── another-feature/ # In-progress feature (if any)
specs--completed/
└── <feature-name>/ # Archived feature
├── overview.md # With completion summary
├── phase-1.md # All checkboxes [x]
├── phase-2.md
└── ...
```
**Keep folder name exactly as-is during archival.**
---
### STEP 6.5: Community Feedback Submission
`🧹 [FINALIZATION STEP 6.5: Community Feedback Submission]`
If Step 5 feedback/consent was declined, skip to Step 7. Otherwise assemble/preview/submit the payload:
Read references/community-feedback-submission.md
---
### STEP 7: Final Confirmation
`🧹 [FINALIZATION STEP 7: Final Confirmation]`
@@ -288,6 +285,7 @@ Replace METRICS_JSON values with actuals. `completion_rate_at_audit` = Y/Z as de
- [x] Step 4: Documentation Review
- [x] Step 5: User Feedback (Optional)
- [x] Step 6: Spec Cleanup
- [x] Step 6.5: Community Feedback Submission
- [x] Step 7: Final Confirmation
### Files Created/Modified During Finalization
+150
View File
@@ -0,0 +1,150 @@
# plan2code-handoff
Turn everything useful in the current conversation into a single, self-contained
handoff document that lets a *different* agent — a new session, a teammate's
session, or a subagent — resume the work without re-reading this transcript.
The reader of this document starts with **zero context**. They can see the repo
and can open files, but they cannot see this conversation. Write for them.
## The one hard rule: capture the next task, and confirm it with the user
Every handoff MUST end with a **Next task** that the incoming agent should start
on. This is the single most important part of the document — a handoff with a
vague or missing next step forces the reader to re-derive intent, which is
exactly what this skill exists to prevent.
Determine it like this:
1. **Try to infer it** from the conversation — the open TODO, the failing test,
the plan step you were mid-way through, the thing the user just asked for
next. Look at what's actually unfinished, not just the last message.
2. **Present it to the user for confirmation before writing the file.** If you
inferred a candidate, show it and ask them to confirm or correct it. If you
genuinely can't infer one, ask them to tell you what the next agent should do.
Use `AskUserQuestion` (offer your inferred task as the recommended option) or
a plain question — either is fine.
3. **Do not write the document until the user has confirmed or supplied the next
task.** This gate is mandatory even when your inference feels obviously
correct. The user's answer is the source of truth; your inference is only a
draft of it.
If the user passed a focus area as an argument, treat it as a strong signal for
the next task (and shape the whole document around it), but still confirm.
## Where to write it
Ask the user if they would like to save the file to the tempory directory of the user's OS (this should be the default) or to some other location like `./handoffs/` at the repo root. Filenames should have a timestamped filename so it's discoverable but doesn't collide with earlier handoffs:
```
<user-specified-path>/<YYYY-MM-DD-HHmm>-handoff.md
```
Get the timestamp from the shell rather than guessing — e.g. PowerShell
`Get-Date -Format 'yyyy-MM-dd-HHmm'`. Create the `<user-specified-path>/` directory if it
doesn't exist.
### Make sure you aren't leaking the file into version control
The handoff is working state for the next session, not a project artifact, so it
should stay out of commits and PRs. Don't assume it will — this skill may run in
any repo. Before (or right after) writing, check whether the path is ignored:
- Is this even a git repo? `git rev-parse --is-inside-work-tree` — if it errors,
there's nothing to ignore; skip this and just tell the user where the file is.
- Is the file ignored? `git check-ignore handoffs/` (exit 0 = ignored). This is
the reliable check — a repo may ignore `handoffs/` via a global or nested
`.gitignore`, so don't rely on grepping the root `.gitignore` alone.
If it is **not** ignored, do not silently modify the user's `.gitignore`. Tell
them the file would be tracked by git and offer to add a `handoffs/` line to
`.gitignore` — let them decide. Some users may want handoffs committed so
teammates get them; that's a legitimate choice, so present it, don't force it.
## If this touched a plan2code spec
`specs/` is gitignored — Glob/Grep and file search silently skip it; use a shell
listing instead: `ls specs/` (bash) or `Get-ChildItem specs/` (PowerShell). If the
conversation worked inside `specs/<feature>/`, confirm the exact state before
writing:
- Which `phase-X.md` is in progress, and whether its `- [ ]` tasks are still
unchecked (checkboxes are ground truth, not the overview's Phase Checklist).
- Cite that file and its checkbox state directly in **Current state** and
**Key files & pointers**, instead of relying on conversation memory alone.
- Let **Suggested skills** name the specific next pipeline command
(`/plan2code-3-implement` to keep implementing the phase,
`/plan2code-4-finalize` once all phases are checked) — but only as a
suggestion; the confirmed **Next task** above still governs what the reader
does first.
No `specs/` activity this session? Skip this section entirely.
## What to include
Keep it tight and high-signal. Prefer pointers over prose: this repo already
records a lot (plan specs, the loop's NDJSON logs, git history, diffs), so
**reference those by path or URL instead of copying them in**. The reader can
open a file; they can't open your memory.
Use this structure:
```markdown
# Handoff — <short title of the work>
<!-- written <timestamp> -->
## Next task
<the confirmed next task — concrete and actionable, e.g.
"Implement Step 3 of specs/<name>.md: wire the aggregator into cli.ts, then
run `npm run build` in plan2code-metrics/ and fix the two failing tests.">
## Goal / why
<13 sentences: what the user is ultimately trying to achieve, so the reader
can make good judgment calls the instructions don't cover.>
## Current state
<Where things stand right now. What's done, what's in progress, what's broken.
Name the branch. Point at the plan/spec file(s) by path rather than restating
them. Note anything half-applied or left uncommitted.>
## Key files & pointers
<Bulleted paths the reader will need, each with a one-line "why". Include plan
specs, the files you were editing, relevant logs (e.g. .plan2code-loop NDJSON),
and any PR/issue URLs.>
## Gotchas & decisions
<Non-obvious things learned this session: a constraint (e.g. the 11k-char limit
on src/plan2code-*.md), a decision made and why, a dead end already ruled out,
a command that must be run a specific way. Save the reader from re-discovering
these the hard way.>
## Suggested skills
<Which skills the next agent should use, and when — e.g. /plan2code-3-implement
to continue a phase, /plan2code-review before finishing, /plan2code-4-finalize
to wrap up. Skip if none apply.>
## Verification
<How the reader confirms their work: exact test/build commands, what "done"
looks like.>
```
Adapt the sections to the work — drop any that would be empty rather than
padding them. **Next task** is the only section that is never optional.
## Strip sensitive data
Before writing, remove credentials, API tokens, passwords, and personal
identifiers. If a secret is load-bearing for the next step, reference *where* it
lives (env var name, secret manager entry) rather than its value.
## After writing
Tell the user the path you wrote to and give a one-line summary of the confirmed
next task, so they know what the incoming agent will start on. Mention that a
fresh session can be pointed at the file to resume the work.
If your ignore check above found the file is **not** gitignored (or the repo has
no `.gitignore`, or it isn't a git repo at all), say so plainly here — e.g. "note:
`handoffs/` isn't gitignored in this repo, so this file will show up in `git
status` and could be committed" — and offer to add the ignore line. Never leave
the user unaware that the handoff might ride along into a commit.
@@ -0,0 +1,65 @@
# Step 7 — AI Agent File Sync
Check for other AI agent config files and offer to replace them with AGENTS.md
references, so AGENTS.md stays the single source of truth.
## Files to Detect
| File | Reference Path |
|------|----------------|
| `CLAUDE.md` (root) | `./AGENTS.md` |
| `GEMINI.md` (root) | `./AGENTS.md` |
| `.cursorrules` (root) | `./AGENTS.md` |
| `.github/copilot-instructions.md` | `../AGENTS.md` |
| `.cursor/rules/*.md` | `../../AGENTS.md` |
| `.windsurf/rules/*.md` | `../../AGENTS.md` |
**No files found:** Skip silently, end workflow.
## If Files Found
```
+---+
| o |
| ~ | Found other AI agent configs!
+---+
```
> Found AI config files that could reference AGENTS.md:
>
> | File | Size |
> |------|------|
> | `CLAUDE.md` | 45 lines |
>
> Replace with AGENTS.md references?
> - **Yes** - Update all
> - **Select** - Choose specific (numbered list)
> - **No** - Keep as-is
**Warning** for files >10 lines: "[file] has custom content that will be replaced."
## CLAUDE.md Template
CLAUDE.md gets a special template because Claude Code auto-loads it — the `CRITICAL — MANDATORY FIRST STEP` directive ensures AGENTS.md is always read:
```markdown
# CLAUDE.md
**CRITICAL — MANDATORY FIRST STEP: You MUST read [AGENTS.md](./AGENTS.md) before responding to ANY user message, including simple questions. Do NOT skip this step regardless of how trivial the request appears. No exceptions.**
See AGENTS.md for full project documentation: commands, architecture, environment, testing, deployment, and .agents-docs/ section details.
This file exists for Claude Code auto-loading. All AI coding agents should reference AGENTS.md.
```
## Reference Template (all other files)
Use title and path from the detection table:
```markdown
# [Title]
See [AGENTS.md]([Path]) for full project documentation: commands, architecture, environment, testing, deployment, and .agents-docs/ section details.
```
**For directory configs** (`.cursor/rules/`, `.windsurf/rules/`): Delete existing `.md` files, create single `reference.md`.
+3 -62
View File
@@ -185,70 +185,11 @@ If yes, return to Step 3. If done, proceed to Step 7.
## Step 7: AI Agent File Sync
Check for other AI agent config files and offer to replace with AGENTS.md references.
Check for other AI agent config files (`CLAUDE.md`, `GEMINI.md`, `.cursorrules`, `.github/copilot-instructions.md`, `.cursor/rules/*.md`, `.windsurf/rules/*.md`) and offer to replace them with AGENTS.md references. **No files found:** skip silently, end workflow.
### Files to Detect
Read references/ai-agent-file-sync.md
| File | Reference Path |
|------|----------------|
| `CLAUDE.md` (root) | `./AGENTS.md` |
| `GEMINI.md` (root) | `./AGENTS.md` |
| `.cursorrules` (root) | `./AGENTS.md` |
| `.github/copilot-instructions.md` | `../AGENTS.md` |
| `.cursor/rules/*.md` | `../../AGENTS.md` |
| `.windsurf/rules/*.md` | `../../AGENTS.md` |
**No files found:** Skip silently, end workflow.
### If Files Found
```
o o
\ /
+---+
| o |
| ~ | Found other AI agent configs!
+---+
```
> Found AI config files that could reference AGENTS.md:
>
> | File | Size |
> |------|------|
> | `CLAUDE.md` | 45 lines |
>
> Replace with AGENTS.md references?
> - **Yes** - Update all
> - **Select** - Choose specific (numbered list)
> - **No** - Keep as-is
**Warning** for files >10 lines: "[file] has custom content that will be replaced."
### CLAUDE.md Template
CLAUDE.md gets a special template because Claude Code auto-loads it — the `CRITICAL — MANDATORY FIRST STEP` directive ensures AGENTS.md is always read:
```markdown
# CLAUDE.md
**CRITICAL — MANDATORY FIRST STEP: You MUST read [AGENTS.md](./AGENTS.md) before responding to ANY user message, including simple questions. Do NOT skip this step regardless of how trivial the request appears. No exceptions.**
See AGENTS.md for full project documentation: commands, architecture, environment, testing, deployment, and .agents-docs/ section details.
This file exists for Claude Code auto-loading. All AI coding agents should reference AGENTS.md.
```
### Reference Template (all other files)
Use title and path from the detection table:
```markdown
# [Title]
See [AGENTS.md]([Path]) for full project documentation: commands, architecture, environment, testing, deployment, and .agents-docs/ section details.
```
**For directory configs** (`.cursor/rules/`, `.windsurf/rules/`): Delete existing `.md` files, create single `reference.md`.
> Fallback: for each detected file, offer Yes/Select/No to replace it with a pointer to AGENTS.md (warn when a file >10 lines has custom content that would be replaced). `CLAUDE.md` gets a special template opening with a `CRITICAL — MANDATORY FIRST STEP` directive to always read `AGENTS.md` (Claude Code auto-loads it); all other files get a short "See AGENTS.md for full project documentation" pointer using the correct relative path (`./`, `../`, or `../../` by location). For directory configs (`.cursor/rules/`, `.windsurf/rules/`), delete existing `.md` files and create a single `reference.md`.
---
+2 -2
View File
@@ -10,7 +10,7 @@ Start all CREATE AGENTS MODE responses with '💡'
╰───╯
```
Analyze this codebase and create `AGENTS.md` to guide future AI coding agents (Claude Code, Codex, Gemini CLI, Devin, Zed, etc.).
Analyze this codebase and create `AGENTS.md` to guide future AI coding agents (Claude Code, Codex, Devin, Zed, etc.).
## Content
@@ -31,7 +31,7 @@ Prefix the file with:
```
# AGENTS.md
This file provides guidance to AI coding agents like Claude Code (claude.ai/code), Cursor AI, Codex, Gemini CLI, GitHub Copilot, Devin, Zed, and other AI coding assistants when working with code in this repository.
This file provides guidance to AI coding agents like Claude Code (claude.ai/code), Cursor AI, Codex, GitHub Copilot, Devin, Zed, and other AI coding assistants when working with code in this repository.
```
---
@@ -0,0 +1,18 @@
# Review — Session End Next-Step Routing
Loaded at the end of a review session to suggest what genuinely helps next.
Principles to reason from, not a lookup table — adapt; when a case doesn't fit
cleanly, say what you verified and ask.
- **Plan2Code Workflow Pipeline:** `/plan2code-1-plan``PLAN-*` files · `/plan2code-2-document``overview.md` + `phase-*.md` (the "spec docs") in `specs/<feature>/` · `/plan2code-3-implement` → checks off phase tasks, one phase per run · `/plan2code-4-finalize` → archives to `specs--completed/`.
- **Find specs (any OS/shell):** `specs/` is gitignored, and search tools (Glob/Grep/project search) skip gitignored paths on many platforms — an empty search result is not evidence either way. Check with a terminal listing: `ls specs/<feature>/` (bash/zsh) · `Get-ChildItem specs/<feature>` (PowerShell) · `dir specs\<feature>` (cmd). Feature dir unknown? List `specs/` first. Command errors? Try another shell's form, then read the expected files directly — file reads see gitignored paths. Conclude "no specs" only after a terminal listing or a failed direct read.
- **Reconcile three signals:** session context (what this conversation was doing — a fresh session may have none), user intent (what they asked reviewed), the disk check above. Disk wins on state; context wins on intent and on disk silence; no context → intent + disk decide.
- **plan2code artifact reviewed (a plan, the spec docs, phases) — route on the reviewed feature's own `specs/<feature>/` state (another feature's specs prove nothing here), to the earliest unmet stage, suggesting only a step whose input exists:**
- `PLAN-*` without `overview.md``/plan2code-2-document`
- `overview.md` without `phase-*.md``/plan2code-2-document`
- Unchecked `- [ ]` tasks in `phase-*.md``/plan2code-3-implement` (checkboxes are ground truth; flag overview conflicts)
- All phase tasks checked → `/plan2code-4-finalize`
- Archived spec → pipeline complete; summary only
- **Anything else** (code/PRs, logs, docs, tickets, emails, a codebase): no pipeline step — close with the summary; add a next action only if the review makes one obvious and actionable.
- **Gates:** unresolved Criticals → fixing them (H/A/S) is the next step. Signals the rules above can't reconcile, or multiple candidate specs → ask one targeted question.
- **Output:** "Next (NEW conversation): `/plan2code-<step>` — [why + how you know]"; otherwise "Review complete -- [summary]."
+7 -7
View File
@@ -141,7 +141,7 @@ Run every finding through the false-positive detection shortcuts before presenti
> **Totals:** X Critical · Y Warning · Z Suggestion
**Fix options:** Reply `H` (high-priority), `A` (all), or `S 1,3,5` (specific findings).
**Fix options:** Reply `H` to fix Critical + Warning findings, `A` to fix all findings, or `S 1,3,5` to fix specific findings.
**Zero findings:** Skip table and fix options. Dimension Coverage justifies each clean dimension.
@@ -155,16 +155,16 @@ Run every finding through the false-positive detection shortcuts before presenti
## Session End
Render this section only when the session is over: fix options resolved (fixes applied and verified, or user declined) or the review had zero findings. Step 5 ends at fix options — stop there and wait.
Work summary — tell user: scope reviewed, findings count by severity (Critical/Warning/Suggestion), fixes applied, unresolved findings.
After approval, select template based on project state:
**Next step** — suggest what genuinely helps next. Principles to reason from, not a lookup table — adapt; when a case doesn't fit cleanly, say what you verified and ask.
**Pre-condition check:** Before suggesting implement or finalize, verify that `overview.md` AND at least one `phase-*.md` file exist in the specs directory. If they do NOT exist, the document step has not been run yet.
Read references/session-end.md
> Fallback: route plan2code artifacts (plan/spec docs/phases) on the reviewed feature's own `specs/<feature>/` state to the earliest unmet pipeline stage whose input exists — `PLAN-*` or `overview.md` without `phase-*.md` → `/plan2code-2-document`; unchecked `- [ ]` phase tasks → `/plan2code-3-implement`; all checked → `/plan2code-4-finalize`; archived → complete, summary only. Verify `specs/` on disk with a terminal `ls`/`Get-ChildItem` — it's gitignored, so search tools miss it and an empty result proves nothing. Non-pipeline artifacts (code/PRs/docs/logs) → summary only. Unresolved Criticals → fixing them (H/A/S) is the next step. Ambiguous or multiple candidate specs → ask one targeted question. Output: "Next (NEW conversation): `/plan2code-<step>` — [why + how you know]"; else "Review complete -- [summary]."
- **Specs dir exists but NO overview.md / phase-*.md files:** "Next: generate implementation docs. NEW conversation: `/plan2code-2-document`"
- **Specs + overview.md + phase files + more phases:** "Next: Phase X. NEW conversation: `/plan2code-3-implement`"
- **Specs + overview.md + phase files + all complete:** "All complete! NEW conversation: `/plan2code-4-finalize`"
- **Standalone:** "Review complete -- [summary]."
- **Commit** (code changes): `git add [files] && git commit -m "fix: [desc]" -m "<JIRA>" -m "AI Assisted"` -- derive JIRA from branch.
```
+37 -5
View File
@@ -17,12 +17,22 @@ A persistent three-line status bar for Claude Code that displays model info, pro
```
╭─╮ Sonnet 4.5 | Medium │ plan2code │ main │ +12 -3
│★│ ┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄
╰─╯ 2m ($0.18) │ ▰▰▱▱▱▱▱▱▱▱▱▱ 18% (36k) │ 88k in · 3k out
╰─╯ 2m ($0.18) │ ▰▰▱▱▱▱▱▱▱▱▱▱ 18% │ 88k in · 3k out
```
**Git worktree** — session running in a linked worktree at `C:\git\plan2code-user-auth`:
```
╭─╮ Opus 4.6 | High │ plan2code ⑂ │ feature/user-auth │ +12 -3
│★│ ┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄
╰─╯ 3h 5m ($4.62) │ ▰▰▰▰▰▱▱▱▱▱▱▱ 42% (84k) │ 5h: 28% · 7d: 61%
```
**Line 1:** Planny icon | Model name + reasoning effort | Project name | Git branch | Uncommitted changes
**Line 2:** Planny icon | Gray separator
**Line 3:** Planny icon | Session duration + cost | Context window bar | Usage (rate limits or token counts)
**Line 3:** Planny icon | Session duration + cost | Context window bar (+ absolute tokens used, Pro/Max/Teams only) | Usage (rate limits or token counts)
The context bar's absolute token count `(84k)` only appears on Pro/Max/Teams (rate-limit) accounts, where it's the only place total context tokens are shown. On Enterprise/Bedrock/Vertex/PAYG accounts, it's suppressed because the `in`/`out` token counts already cover it.
## Installation
@@ -65,6 +75,7 @@ Edit `~/.claude/statusline-config.json`:
"model": true,
"effort": true,
"project": true,
"worktree": true,
"branch": true,
"contextBar": true,
"planUsage": true,
@@ -83,9 +94,10 @@ Edit `~/.claude/statusline-config.json`:
| `items.model` | `true` | Show model name (Opus, Sonnet, Haiku) |
| `items.effort` | `true` | Append the current reasoning effort level (Low/Medium/High/XHigh/Max) to the model segment, e.g. `Sonnet 5 \| High`. Reads `effort.level` from stdin; hidden when the current model doesn't support an effort parameter (field absent from stdin). |
| `items.project` | `true` | Show project directory name |
| `items.worktree` | `true` | Detect [git worktrees](https://git-scm.com/docs/git-worktree) and render the project segment as `repo ⑂ worktree` instead of just the worktree's directory name. See [Worktree Display](#worktree-display). Requires `items.project`. |
| `items.branch` | `true` | Show current git branch (falls back to short SHA when detached HEAD) |
| `items.contextBar` | `true` | Show context window usage bar with percentage |
| `items.contextTokens` | `true` | Append raw tokens used in the context window next to the percentage, e.g. `42% (84k)`. Reads `context_window.total_input_tokens` (the same input-token count `used_percentage` is derived from — excludes output tokens). Requires `items.contextBar` to also be enabled. |
| `items.contextTokens` | `true` | Append raw tokens used in the context window next to the percentage, e.g. `42% (84k)`. Reads `context_window.total_input_tokens` (the same input-token count `used_percentage` is derived from — excludes output tokens). Requires `items.contextBar` to also be enabled. Suppressed automatically on Enterprise/Bedrock/Vertex/PAYG accounts, where the `in`/`out` token usage segment already shows the same number. |
| `items.planUsage` | `true` | Show usage info: rate limits (Pro/Max) or token counts (Bedrock/Vertex/PAYG) |
| `items.linesChanged` | `true` | Show uncommitted git diff stats (+added -removed) |
| `items.duration` | `true` | Show session duration |
@@ -98,8 +110,28 @@ Malformed config falls back to defaults silently. Type errors on individual keys
| Platform | Status | Notes |
|----------|--------|-------|
| Claude Code | Supported | Full feature set via `settings.json` registration |
| Copilot CLI | Planned | Deferred CLI lacks status line script support as of v0.0.421 |
| Others | Not supported | Codex CLI, Gemini CLI have built-in status, not scriptable |
| Others | Not yet supported | CLI lacks status line script support |
## Worktree Display
With `items.worktree` disabled, a session in a linked worktree shows only that worktree's directory name (`plan2code-user-auth`) — the underlying repository identity is lost, and the session looks like an unrelated project.
When enabled, the project segment becomes `repo ⑂ worktree`:
| Worktree directory | Branch | Rendered | Why |
|--------------------|--------|----------|-----|
| `plan2code` (primary checkout) | `main` | `plan2code` | Not a linked worktree — unchanged |
| `plan2code-spike` | `spike-thing` | `plan2code ⑂ spike` | Repo prefix stripped from the worktree name |
| `plan2code-user-auth` | `feature/user-auth` | `plan2code ⑂` | Name is redundant with the visible branch — collapses to the bare marker |
| `plan2code-user-auth` | `feature/user-auth`, `items.branch: false` | `plan2code ⑂ user-auth` | No branch shown, so the name is kept |
Details:
- **Detection** — a single `git rev-parse --git-dir --git-common-dir`. The two paths differ only inside a linked worktree. The true repo name comes from `--git-common-dir` (its parent directory, or the `<name>.git` basename for a bare main repo), so it is correct regardless of how the worktree directory was named.
- **Prefix stripping** — a leading repo name followed by `-`, `_`, or `.` is removed, so `plan2code-user-auth` reads as `user-auth`.
- **Redundancy collapse** — worktree directories usually mirror their branch. The name is dropped (leaving ``) when it matches the branch after normalizing case and `/ _ . -` separators, compared against both the full branch and its trailing segment so type prefixes like `feature/` don't defeat the match. This only applies when the branch is actually displayed.
- **Color** — the repo keeps its bold silver; the `` marker and worktree name render in amber so a non-primary checkout is obvious at a glance. Monochrome output is `repo ⑂ worktree`.
- **Cost** — one extra `git` invocation, subject to the same 1.5s timeout, and skipped entirely outside git repos.
## Usage Display
@@ -6,6 +6,7 @@
"model": true,
"effort": true,
"project": true,
"worktree": true,
"branch": true,
"contextBar": true,
"contextTokens": true,
+96 -9
View File
@@ -28,6 +28,7 @@ const DEFAULTS = {
model: true,
effort: true,
project: true,
worktree: true,
branch: true,
contextBar: true,
contextTokens: true,
@@ -90,6 +91,7 @@ const C = {
// Plan2Code brand
teal: rgb(190, 192, 200), // Silver — repo
sky: rgb(120, 195, 255), // Light sky blue — branch
amber: rgb(235, 170, 95), // Amber — linked worktree marker
muted: rgb(40, 120, 200), // Mid blue — model
label: rgb(130, 135, 150), // Mid gray — field labels (5h, 7d)
// Bar & quota thresholds
@@ -150,6 +152,54 @@ function getGitBranch(stdinData) {
return sha || '';
}
// Linked worktrees have their own .git dir under the main repo's
// .git/worktrees/<name>, while --git-common-dir still points at the main repo's
// .git. Equal paths mean we're in the primary checkout.
function getWorktreeInfo(cwd) {
const raw = gitExec(['rev-parse', '--git-dir', '--git-common-dir'], cwd);
if (!raw) return null;
const [gitDir, commonDir] = raw.split('\n').map((p) => path.resolve(cwd, p.trim()));
if (!gitDir || !commonDir) return null;
const samePath = process.platform === 'win32'
? gitDir.toLowerCase() === commonDir.toLowerCase()
: gitDir === commonDir;
if (samePath) return null;
// commonDir is normally <repo>/.git; bare repos use <name>.git directly
const commonBase = path.basename(commonDir);
const repoName = commonBase === '.git'
? path.basename(path.dirname(commonDir))
: commonBase.replace(/\.git$/i, '');
const toplevel = gitExec(['rev-parse', '--show-toplevel'], cwd);
const worktreeName = path.basename(toplevel || cwd);
if (!repoName || !worktreeName) return null;
return { repoName, worktreeName };
}
// "plan2code-user-auth" in repo "plan2code" reads as just "user-auth"
function stripRepoPrefix(worktreeName, repoName) {
const lowerWt = worktreeName.toLowerCase();
const lowerRepo = repoName.toLowerCase();
if (!lowerWt.startsWith(lowerRepo) || lowerWt === lowerRepo) return worktreeName;
const rest = worktreeName.slice(repoName.length);
return /^[-_.]/.test(rest) ? rest.slice(1) : worktreeName;
}
const normalizeName = (s) => s.toLowerCase().replace(/[/_.\s-]+/g, '-').replace(/^-|-$/g, '');
// Worktree dirs usually mirror their branch (plan2code-user-auth / feature/user-auth).
// Compare against both the full branch and its trailing segment so the common
// type prefixes (feature/, bugfix/, ...) don't defeat the match.
function isRedundantWithBranch(worktreeName, branch) {
if (!branch) return false;
const candidates = new Set([normalizeName(branch), normalizeName(branch.split('/').pop())]);
return candidates.has(normalizeName(worktreeName));
}
// ============================================================================
// Formatters
// ============================================================================
@@ -192,6 +242,34 @@ function getProjectName(stdinData) {
return path.basename(projectDir);
}
// Returns the project segment: "plan2code" normally, "plan2code ⑂ spike" in a
// linked worktree — the name collapsing to a bare ⑂ when the branch says it already.
function formatProject(stdinData, config, branch) {
const project = getProjectName(stdinData);
if (!project) return null;
const plain = (text) => (config.color ? `${C.bold}${C.teal}${text}${C.reset}` : text);
if (!config.items.worktree) return plain(project);
const cwd = stdinData?.workspace?.project_dir || process.cwd();
if (!isGitRepo(cwd)) return plain(project);
const info = getWorktreeInfo(cwd);
if (!info) return plain(project);
const shortName = stripRepoPrefix(info.worktreeName, info.repoName);
// Only dedupe against a branch the user can actually see — otherwise the
// worktree identity would vanish entirely.
const visibleBranch = config.items.branch ? branch : '';
const redundant = isRedundantWithBranch(shortName, visibleBranch)
|| isRedundantWithBranch(info.worktreeName, visibleBranch)
|| normalizeName(shortName) === normalizeName(info.repoName);
const marker = redundant ? '⑂' : `${shortName}`;
if (!config.color) return `${info.repoName} ${marker}`;
return `${C.bold}${C.teal}${info.repoName}${C.reset} ${C.amber}${marker}${C.reset}`;
}
function calculateContextPercent(stdinData, config) {
const cw = stdinData?.context_window;
if (!cw) return null;
@@ -204,7 +282,7 @@ function calculateContextPercent(stdinData, config) {
return Math.round(Math.min(100, (rawUsedTokens / usableTokens) * 100));
}
function formatContextBar(percent, tokens, config) {
function formatContextBar(percent, tokens, config, suppressTokens) {
if (percent == null) return null;
const clamped = Math.max(0, Math.min(100, percent));
@@ -212,7 +290,8 @@ function formatContextBar(percent, tokens, config) {
const filledStr = '▰'.repeat(filled);
const emptyStr = '▱'.repeat(CONTEXT_BAR_LEN - filled);
const tokenStr = config.items.contextTokens && tokens != null ? ` (${formatTokenCount(tokens)})` : '';
const showTokens = config.items.contextTokens && !suppressTokens && tokens != null;
const tokenStr = showTokens ? ` (${formatTokenCount(tokens)})` : '';
if (!config.color) return `${filledStr}${emptyStr} ${clamped}%${tokenStr}`;
@@ -339,14 +418,16 @@ function formatLine1(stdinData, config) {
}
}
// Branch is resolved first — the project segment dedupes the worktree name against it
const branch = config.items.branch || config.items.worktree ? getGitBranch(stdinData) : '';
if (config.items.project) {
const project = getProjectName(stdinData);
if (project) segments.push(config.color ? `${C.bold}${C.teal}${project}${C.reset}` : project);
const project = formatProject(stdinData, config, branch);
if (project) segments.push(project);
}
if (config.items.branch) {
const branch = getGitBranch(stdinData);
if (branch) segments.push(config.color ? `${C.sky}${branch}${C.reset}` : branch);
if (config.items.branch && branch) {
segments.push(config.color ? `${C.sky}${branch}${C.reset}` : branch);
}
if (config.items.linesChanged) {
@@ -365,15 +446,21 @@ function formatLine2(stdinData, config) {
if (dur) segments.push(dur);
}
// Token usage ("X in / Y out") already surfaces total_input_tokens — suppress the
// context bar's duplicate (XXk) in that case. Rate-limit accounts show no absolute
// tokens elsewhere, so the bar's (XXk) stays as the only source there.
const rateLimits = config.items.planUsage ? formatRateLimits(stdinData, config) : null;
const tokenUsage = config.items.planUsage && !rateLimits ? formatTokenUsage(stdinData, config) : null;
if (config.items.contextBar) {
const percent = calculateContextPercent(stdinData, config);
const tokens = stdinData?.context_window?.total_input_tokens;
const bar = formatContextBar(percent, tokens, config);
const bar = formatContextBar(percent, tokens, config, !!tokenUsage);
if (bar) segments.push(bar);
}
if (config.items.planUsage) {
const usage = formatRateLimits(stdinData, config) || formatTokenUsage(stdinData, config);
const usage = rateLimits || tokenUsage;
if (usage) segments.push(usage);
}
+3 -4
View File
@@ -1,6 +1,6 @@
{
"name": "Plan2Code",
"version": "1.15.3",
"version": "2.1.0",
"description": "A structured 4-step workflow methodology for AI-assisted software development",
"keywords": [
"ai",
@@ -16,8 +16,7 @@
"type": "git",
"url": "https://github.com/jparkerweb/plan2code"
},
"homepage": "https://plan2code.jparkerweb.com",
"releaseDate": "2026-06-28",
"homepage": "https://github.com/jparkerweb/plan2code",
"releaseDate": "2026-08-08",
"mode": "utility"
}