Bot now acts as an authentic QA agent instead of rubber-stamping every
question and step.
- intelligent-responder answers AskUserQuestion via LLM using current
observations (tools used, files created, errors) instead of keyword
matching; auto-responder removed.
- evaluator runs after each step with step-specific criteria
(prompts/evaluation-criteria.ts), produces 0-100 score plus
strengths/weaknesses/critical issues. Avg <60 blocks finalization.
- observation-collector captures tool_use, file writes, errors, and
question reasoning; session-runner falls back to scraping tool_use
blocks when canUseTool doesn't fire and dedupes both sources.
- Bot writes BOT-EVALUATION.md and BOT-NOTES.md for metrics analysis.
- Idea generator: 12 categories instead of CLI/web-app coin flip,
stronger seed adherence, less developer-tool bias.
- Init step now writes a minimal AGENTS.md stub instead of running
the full init skill, avoiding hallucinated architecture before plan.
- Implement step uses Read/Write/Edit/Glob/Grep directly instead of
a Skill sub-session that produced no visible tool observations.
- bin: add --help, accept --idea="value" form, strip surrounding
quotes; evaluator maxTurns 3 -> 30; warn when parser misses SCORE.
Add the autonomous workflow test runner (plan2code-bot) that uses the
Claude Agent SDK to run plan2code end-to-end. Includes --resume flag
to continue incomplete runs from saved state, and automatic state file
cleanup on successful completion.