mirror of
https://github.com/jparkerweb/plan2code.git
synced 2026-07-21 18:33:22 -07:00
plan2code-bot: replace auto-responder with LLM-as-judge evaluation
Bot now acts as an authentic QA agent instead of rubber-stamping every question and step. - intelligent-responder answers AskUserQuestion via LLM using current observations (tools used, files created, errors) instead of keyword matching; auto-responder removed. - evaluator runs after each step with step-specific criteria (prompts/evaluation-criteria.ts), produces 0-100 score plus strengths/weaknesses/critical issues. Avg <60 blocks finalization. - observation-collector captures tool_use, file writes, errors, and question reasoning; session-runner falls back to scraping tool_use blocks when canUseTool doesn't fire and dedupes both sources. - Bot writes BOT-EVALUATION.md and BOT-NOTES.md for metrics analysis. - Idea generator: 12 categories instead of CLI/web-app coin flip, stronger seed adherence, less developer-tool bias. - Init step now writes a minimal AGENTS.md stub instead of running the full init skill, avoiding hallucinated architecture before plan. - Implement step uses Read/Write/Edit/Glob/Grep directly instead of a Skill sub-session that produced no visible tool observations. - bin: add --help, accept --idea="value" form, strip surrounding quotes; evaluator maxTurns 3 -> 30; warn when parser misses SCORE.
This commit is contained in:
@@ -2,6 +2,7 @@
|
|||||||
|
|
||||||
## 1.1.0
|
## 1.1.0
|
||||||
|
|
||||||
|
### Resume support
|
||||||
- Add `--resume` flag to continue incomplete runs from saved state
|
- Add `--resume` flag to continue incomplete runs from saved state
|
||||||
- Skip previously succeeded steps when resuming (init, plan, document, implement, finalize)
|
- Skip previously succeeded steps when resuming (init, plan, document, implement, finalize)
|
||||||
- Restore idea name, description, project directory, and implement pass counter from state
|
- Restore idea name, description, project directory, and implement pass counter from state
|
||||||
@@ -11,6 +12,31 @@
|
|||||||
- Add `deleteState()` and `findExistingState()` utilities to bot-state module
|
- Add `deleteState()` and `findExistingState()` utilities to bot-state module
|
||||||
- Add bot-state unit tests (saveState, loadState, deleteState, findExistingState)
|
- Add bot-state unit tests (saveState, loadState, deleteState, findExistingState)
|
||||||
|
|
||||||
|
### Idea generation improvements
|
||||||
|
- Expand idea categories from binary CLI/web-app coin flip to 12 diverse categories (games, dashboards, browser extensions, desktop utilities, etc.)
|
||||||
|
- Add guidance to avoid defaulting to developer-centric tools (git analyzers, code formatters)
|
||||||
|
- Strengthen `--idea` seed clause so the LLM stays aligned with the user's theme instead of ignoring it
|
||||||
|
- Update system prompt to encourage creative, cross-domain ideas
|
||||||
|
|
||||||
|
### Init step overhaul (new projects)
|
||||||
|
- Init now creates a minimal AGENTS.md stub (name, description, status) instead of running the full `/plan2code-init` skill
|
||||||
|
- Prevents hallucinated architecture, commands, and `.agents-docs/` files before the plan step runs
|
||||||
|
- Init evaluation criteria updated to reward minimalism and penalize premature detail
|
||||||
|
|
||||||
|
### Implement step overhaul
|
||||||
|
- Implement step now works directly with Read/Write/Edit/Glob/Grep tools instead of delegating to Skill sub-session
|
||||||
|
- Inlined step-by-step process: find specs, pick phase, implement tasks, mark checkboxes
|
||||||
|
- Fixes issue where Skill sub-sessions did all work invisibly, causing zero tool observations
|
||||||
|
|
||||||
|
### Observation tracking fix
|
||||||
|
- Capture `tool_use` blocks from the assistant message stream in session-runner as a fallback when `canUseTool` callback doesn't fire
|
||||||
|
- Add deduplication in ObservationCollector to prevent double-counting from both sources
|
||||||
|
- Fixes all steps reporting 0 tools used / 0 files created in BOT-NOTES and evaluations
|
||||||
|
|
||||||
|
### Evaluator improvements
|
||||||
|
- Increase evaluator `maxTurns` from 3 to 30 so it has room for tool calls before producing the scored response
|
||||||
|
- Add warning log when evaluation parser can't find SCORE in output (was silently defaulting to 50)
|
||||||
|
|
||||||
## 1.0.0
|
## 1.0.0
|
||||||
|
|
||||||
- Initial release
|
- Initial release
|
||||||
|
|||||||
@@ -0,0 +1,234 @@
|
|||||||
|
# LLM-as-Judge Evaluation System
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
The plan2code-bot now includes an **always-on LLM-as-judge evaluation system** that transforms it from a "yes-man" into an authentic QA agent. This provides realistic quality signals for `plan2code-metrics` to analyze and drive recursive self-improvement.
|
||||||
|
|
||||||
|
## Key Features
|
||||||
|
|
||||||
|
### 1. Intelligent Decision Making (Real-Time)
|
||||||
|
|
||||||
|
**What:** During execution, when `AskUserQuestion` is called, the bot uses an LLM to make thoughtful decisions based on current observations.
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
- Collects observations up to the current point (tools used, files created, errors)
|
||||||
|
- Queries LLM with context: "Given what you've seen, should you approve this plan?"
|
||||||
|
- LLM inspects current artifacts using Read/Glob/Grep
|
||||||
|
- Returns evidence-based answer with reasoning
|
||||||
|
- All decisions are recorded for metrics analysis
|
||||||
|
|
||||||
|
**Example:**
|
||||||
|
```
|
||||||
|
Question: "Approve this plan?"
|
||||||
|
Observations: Created PLAN-DRAFT.md, 3 phases, 42s duration, no errors
|
||||||
|
LLM reads PLAN-DRAFT.md, evaluates quality
|
||||||
|
LLM decides: "Yes, approve - phases are well-scoped and realistic"
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Post-Step Evaluation
|
||||||
|
|
||||||
|
**What:** After each step completes, the bot evaluates quality using step-specific criteria.
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
- Collects complete execution observations
|
||||||
|
- Queries LLM with evaluation criteria for the step
|
||||||
|
- LLM inspects final artifacts
|
||||||
|
- Returns structured evaluation (score, strengths, weaknesses, suggestions)
|
||||||
|
- Writes `specs/<feature>/BOT-EVALUATION.md` and `specs/<feature>/BOT-NOTES.md` (falls back to project root if no spec folder exists yet, e.g. during `init`)
|
||||||
|
|
||||||
|
**Example output:**
|
||||||
|
```markdown
|
||||||
|
# Evaluation: plan Step
|
||||||
|
|
||||||
|
**Score:** 78/100
|
||||||
|
|
||||||
|
## Strengths
|
||||||
|
- Clear phase breakdown with realistic scope
|
||||||
|
- Tech stack choices appropriate
|
||||||
|
|
||||||
|
## Weaknesses
|
||||||
|
- Phase 3 description too vague
|
||||||
|
- No testing strategy mentioned
|
||||||
|
|
||||||
|
## Suggestions
|
||||||
|
- Expand Phase 3 with concrete tasks
|
||||||
|
- Add explicit testing phase
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Quality Gate
|
||||||
|
|
||||||
|
**What:** Before finalize, checks that average quality score is acceptable.
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
- Calculates average score across all evaluated steps
|
||||||
|
- If average < 60, blocks finalization
|
||||||
|
- Displays clear message about quality issues
|
||||||
|
- User must review `specs/<feature>/BOT-EVALUATION.md` and fix problems
|
||||||
|
|
||||||
|
## Files Created
|
||||||
|
|
||||||
|
### New Files
|
||||||
|
|
||||||
|
1. **`src/observation-collector.ts`**
|
||||||
|
- Tracks execution details (tools, files, messages, errors, questions)
|
||||||
|
- Provides snapshots for real-time decisions
|
||||||
|
- Captures complete history for evaluation
|
||||||
|
|
||||||
|
2. **`src/intelligent-responder.ts`**
|
||||||
|
- Replaces hardcoded auto-responder
|
||||||
|
- Uses LLM to answer AskUserQuestion prompts
|
||||||
|
- Provides reasoning for all decisions
|
||||||
|
- Falls back gracefully if LLM unavailable
|
||||||
|
|
||||||
|
3. **`src/prompts/evaluation-criteria.ts`**
|
||||||
|
- Step-specific evaluation criteria (init, plan, document, implement, finalize)
|
||||||
|
- Quality checks, common pitfalls, scoring guidance
|
||||||
|
- Emphasizes honest scoring (most work should score 70-85)
|
||||||
|
|
||||||
|
4. **`src/evaluator.ts`**
|
||||||
|
- Post-step evaluation using LLM-as-judge
|
||||||
|
- Queries LLM with observations and criteria
|
||||||
|
- Parses structured evaluation output
|
||||||
|
- Writes `specs/<feature>/BOT-EVALUATION.md` and `specs/<feature>/BOT-NOTES.md`
|
||||||
|
|
||||||
|
### Modified Files
|
||||||
|
|
||||||
|
1. **`src/types.ts`**
|
||||||
|
- Added interfaces: `ToolObservation`, `QuestionContext`, `ExecutionObservation`, `EvaluationResult`
|
||||||
|
- Extended `StepResult` with `evaluation` and `observations` fields
|
||||||
|
|
||||||
|
2. **`src/session-runner.ts`**
|
||||||
|
- Added `collector` parameter to `SessionOptions`
|
||||||
|
- Returns `observations` in `SessionResult`
|
||||||
|
- Records all messages for observation tracking
|
||||||
|
- Uses intelligent responder instead of auto-responder
|
||||||
|
|
||||||
|
3. **`src/cli.ts`**
|
||||||
|
- Creates `ObservationCollector` for each step
|
||||||
|
- Always runs evaluation after successful steps
|
||||||
|
- Displays scores with color coding (green/yellow/red)
|
||||||
|
- Implements quality gate before finalize
|
||||||
|
- Shows evaluation summary in step output
|
||||||
|
|
||||||
|
4. **`src/bin/plan2code-bot.ts`**
|
||||||
|
- Updated help text to mention LLM-as-judge evaluation
|
||||||
|
- No new CLI flags (evaluation is always on)
|
||||||
|
|
||||||
|
### Deleted Files
|
||||||
|
|
||||||
|
1. **`src/auto-responder.ts`** - Replaced by intelligent-responder.ts
|
||||||
|
2. **`src/auto-responder.test.ts`** - No longer needed
|
||||||
|
|
||||||
|
## Output Files (Created During Execution)
|
||||||
|
|
||||||
|
Both files are written to `specs/<feature>/` so they stay co-located with the feature they describe. If no spec folder exists yet (e.g. during `init`), they fall back to the project root.
|
||||||
|
|
||||||
|
### BOT-EVALUATION.md
|
||||||
|
|
||||||
|
Contains evaluation results for each step:
|
||||||
|
- Score (0-100)
|
||||||
|
- Strengths identified
|
||||||
|
- Weaknesses found
|
||||||
|
- Suggestions for improvement
|
||||||
|
- Critical issues (if any)
|
||||||
|
- Full reasoning from LLM
|
||||||
|
|
||||||
|
### BOT-NOTES.md
|
||||||
|
|
||||||
|
Contains execution observations:
|
||||||
|
- Duration, tool counts, file changes
|
||||||
|
- Questions asked and LLM reasoning for answers
|
||||||
|
- Tool usage timeline
|
||||||
|
- Files created/modified
|
||||||
|
- Assistant output summary
|
||||||
|
|
||||||
|
## Data Structure for Metrics
|
||||||
|
|
||||||
|
All evaluation data is structured in `StepResult`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
{
|
||||||
|
step: 'plan',
|
||||||
|
success: true,
|
||||||
|
duration: 42000,
|
||||||
|
evaluation: {
|
||||||
|
score: 78,
|
||||||
|
strengths: ["Clear phases", "Realistic scope"],
|
||||||
|
weaknesses: ["Phase 3 too vague"],
|
||||||
|
suggestions: ["Add specific tasks to Phase 3"],
|
||||||
|
criticalIssues: [],
|
||||||
|
reasoning: "...",
|
||||||
|
timestamp: 1234567890,
|
||||||
|
evaluatorModel: 'claude-sonnet-4-5'
|
||||||
|
},
|
||||||
|
observations: {
|
||||||
|
tools: [{ toolName, input, output, timestamp }, ...],
|
||||||
|
questionsAsked: [
|
||||||
|
{
|
||||||
|
question: "Approve plan?",
|
||||||
|
selectedAnswer: "Yes, approve",
|
||||||
|
llmReasoning: "Phases are well-scoped...",
|
||||||
|
timestamp: 1234567890
|
||||||
|
}
|
||||||
|
],
|
||||||
|
filesCreated: [...],
|
||||||
|
filesModified: [...],
|
||||||
|
errors: []
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Benefits for Recursive Improvement
|
||||||
|
|
||||||
|
1. **Authentic Signals:** Real quality scores identify actual problem areas
|
||||||
|
2. **Detailed Context:** Observations + reasoning explain WHY failures happen
|
||||||
|
3. **Correlation Analysis:** Link patterns (tool usage, duration, errors) to quality
|
||||||
|
4. **Continuous Loop:** Better metrics → improved workflows → higher scores → repeat
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
No special flags needed - evaluation is always on:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# New project
|
||||||
|
plan2code-bot --idea "todo app"
|
||||||
|
|
||||||
|
# Enhancement
|
||||||
|
cd my-project && plan2code-bot
|
||||||
|
|
||||||
|
# Resume with evaluation data preserved
|
||||||
|
plan2code-bot --resume
|
||||||
|
```
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
After running the bot, check:
|
||||||
|
|
||||||
|
1. **`specs/<feature>/BOT-EVALUATION.md`** - Should show realistic scores (not all 100s)
|
||||||
|
2. **`specs/<feature>/BOT-NOTES.md`** - Should show LLM reasoning for decisions
|
||||||
|
3. **Console output** - Should display color-coded scores after each step
|
||||||
|
4. **State file** (`.plan2code-bot-state.json`) - Should include evaluation data
|
||||||
|
|
||||||
|
## Trade-offs
|
||||||
|
|
||||||
|
### Latency
|
||||||
|
- Adds ~2-3s per AskUserQuestion call (~15-20s total per run)
|
||||||
|
- Worth it for authentic evaluation
|
||||||
|
|
||||||
|
### Token Cost
|
||||||
|
- ~20-26K tokens per run (~$0.60 with Opus 4.6)
|
||||||
|
- Investment pays off through metrics-driven improvement
|
||||||
|
|
||||||
|
### Determinism
|
||||||
|
- LLM decisions vary between runs (non-deterministic)
|
||||||
|
- Realistic - humans vary too
|
||||||
|
- Metrics average over many runs
|
||||||
|
|
||||||
|
## Future Enhancements
|
||||||
|
|
||||||
|
Potential improvements:
|
||||||
|
- Model selection per step (use Haiku for simple decisions)
|
||||||
|
- Configurable quality gate threshold
|
||||||
|
- Historical score tracking across runs
|
||||||
|
- Comparison with previous evaluations
|
||||||
|
- More sophisticated scoring (weighted by step importance)
|
||||||
@@ -1,97 +0,0 @@
|
|||||||
import { describe, it, expect } from 'vitest';
|
|
||||||
import { createAutoResponder } from './auto-responder.js';
|
|
||||||
import type { BotConfig } from './types.js';
|
|
||||||
|
|
||||||
const config: BotConfig = {
|
|
||||||
workDir: '/tmp/work',
|
|
||||||
projectDir: '/tmp/work/my-app',
|
|
||||||
ideaDescription: 'A test app',
|
|
||||||
ideaName: 'test-app',
|
|
||||||
mode: 'new-project',
|
|
||||||
};
|
|
||||||
|
|
||||||
describe('createAutoResponder', () => {
|
|
||||||
it('returns allow for non-AskUserQuestion tools', async () => {
|
|
||||||
const responder = createAutoResponder(config, 'plan');
|
|
||||||
const result = await responder('Bash', { command: 'ls' });
|
|
||||||
expect(result.behavior).toBe('allow');
|
|
||||||
expect((result as any).updatedInput).toBeUndefined();
|
|
||||||
});
|
|
||||||
|
|
||||||
it('injects answers for AskUserQuestion with approval keywords', async () => {
|
|
||||||
const responder = createAutoResponder(config, 'plan');
|
|
||||||
const result = await responder('AskUserQuestion', {
|
|
||||||
questions: [
|
|
||||||
{
|
|
||||||
question: 'Do you approve this plan?',
|
|
||||||
options: [
|
|
||||||
{ label: 'Yes', description: 'Approve' },
|
|
||||||
{ label: 'No', description: 'Reject' },
|
|
||||||
],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
});
|
|
||||||
|
|
||||||
expect(result.behavior).toBe('allow');
|
|
||||||
const updated = (result as any).updatedInput;
|
|
||||||
expect(updated.answers['Do you approve this plan?']).toBe('Yes');
|
|
||||||
});
|
|
||||||
|
|
||||||
it('selects skip/none for testing questions', async () => {
|
|
||||||
const responder = createAutoResponder(config, 'implement');
|
|
||||||
const result = await responder('AskUserQuestion', {
|
|
||||||
questions: [
|
|
||||||
{
|
|
||||||
question: 'How should we run tests?',
|
|
||||||
options: [
|
|
||||||
{ label: 'Full suite', description: 'Run all tests' },
|
|
||||||
{ label: 'Skip', description: 'Skip testing' },
|
|
||||||
],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
});
|
|
||||||
|
|
||||||
expect(result.behavior).toBe('allow');
|
|
||||||
const updated = (result as any).updatedInput;
|
|
||||||
expect(updated.answers['How should we run tests?']).toBe('Skip');
|
|
||||||
});
|
|
||||||
|
|
||||||
it('uses project name for plan step name questions', async () => {
|
|
||||||
const responder = createAutoResponder(config, 'plan');
|
|
||||||
const result = await responder('AskUserQuestion', {
|
|
||||||
questions: [
|
|
||||||
{
|
|
||||||
question: 'What is the name of this feature?',
|
|
||||||
options: [
|
|
||||||
{ label: 'Feature A', description: 'First' },
|
|
||||||
{ label: 'Feature B', description: 'Second' },
|
|
||||||
],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
});
|
|
||||||
|
|
||||||
expect(result.behavior).toBe('allow');
|
|
||||||
const updated = (result as any).updatedInput;
|
|
||||||
// Plan step + "name" keyword → uses config.ideaName
|
|
||||||
expect(updated.answers['What is the name of this feature?']).toBe('test-app');
|
|
||||||
});
|
|
||||||
|
|
||||||
it('falls back to first option for unknown questions', async () => {
|
|
||||||
const responder = createAutoResponder(config, 'init');
|
|
||||||
const result = await responder('AskUserQuestion', {
|
|
||||||
questions: [
|
|
||||||
{
|
|
||||||
question: 'What color is the sky?',
|
|
||||||
options: [
|
|
||||||
{ label: 'Blue', description: 'The usual' },
|
|
||||||
{ label: 'Red', description: 'Sunset' },
|
|
||||||
],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
});
|
|
||||||
|
|
||||||
expect(result.behavior).toBe('allow');
|
|
||||||
const updated = (result as any).updatedInput;
|
|
||||||
expect(updated.answers['What color is the sky?']).toBe('Blue');
|
|
||||||
});
|
|
||||||
});
|
|
||||||
@@ -1,117 +0,0 @@
|
|||||||
import type { BotConfig, StepName } from './types.js';
|
|
||||||
|
|
||||||
interface AskUserQuestionInput {
|
|
||||||
questions: Array<{
|
|
||||||
question: string;
|
|
||||||
options: Array<{
|
|
||||||
label: string;
|
|
||||||
description: string;
|
|
||||||
}>;
|
|
||||||
multiSelect?: boolean;
|
|
||||||
}>;
|
|
||||||
}
|
|
||||||
|
|
||||||
function findOption(
|
|
||||||
options: AskUserQuestionInput['questions'][0]['options'],
|
|
||||||
...keywords: string[]
|
|
||||||
): string | null {
|
|
||||||
for (const keyword of keywords) {
|
|
||||||
const match = options.find((o) =>
|
|
||||||
o.label.toLowerCase().includes(keyword.toLowerCase())
|
|
||||||
);
|
|
||||||
if (match) return match.label;
|
|
||||||
}
|
|
||||||
return null;
|
|
||||||
}
|
|
||||||
|
|
||||||
function buildAnswers(input: AskUserQuestionInput, config: BotConfig, step: StepName): Record<string, string> {
|
|
||||||
const answers: Record<string, string> = {};
|
|
||||||
|
|
||||||
for (const q of input.questions) {
|
|
||||||
const questionText = q.question.toLowerCase();
|
|
||||||
const options = q.options;
|
|
||||||
|
|
||||||
// Approval gates — find yes/approve/confirm option or pick first
|
|
||||||
if (
|
|
||||||
questionText.includes('approve') ||
|
|
||||||
questionText.includes('confirm') ||
|
|
||||||
questionText.includes('proceed') ||
|
|
||||||
questionText.includes('ready') ||
|
|
||||||
questionText.includes('look good') ||
|
|
||||||
questionText.includes('sign off') ||
|
|
||||||
questionText.includes('sign-off')
|
|
||||||
) {
|
|
||||||
const opt = findOption(options, 'approve', 'yes', 'confirm', 'proceed', 'ready');
|
|
||||||
answers[q.question] = opt ?? options[0].label;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Testing questions — skip or none
|
|
||||||
if (
|
|
||||||
questionText.includes('test') ||
|
|
||||||
questionText.includes('testing')
|
|
||||||
) {
|
|
||||||
const opt = findOption(options, 'skip', 'none', 'no');
|
|
||||||
answers[q.question] = opt ?? options[0].label;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Plan step: additional files/references
|
|
||||||
if (step === 'plan' && (questionText.includes('additional') || questionText.includes('reference'))) {
|
|
||||||
const opt = findOption(options, 'no', 'none', 'skip');
|
|
||||||
answers[q.question] = opt ?? options[0].label;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Plan step: feature name question
|
|
||||||
if (step === 'plan' && questionText.includes('name')) {
|
|
||||||
answers[q.question] = config.ideaName;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Implement step: pick first phase or approve
|
|
||||||
if (step === 'implement') {
|
|
||||||
const opt = findOption(options, 'approve', 'yes', 'continue', 'proceed');
|
|
||||||
answers[q.question] = opt ?? options[0].label;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Finalize step: approve docs and give feedback
|
|
||||||
if (step === 'finalize') {
|
|
||||||
if (questionText.includes('rating') || questionText.includes('feedback')) {
|
|
||||||
const opt = findOption(options, '8', '9', '10');
|
|
||||||
answers[q.question] = opt ?? options[0].label;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
const opt = findOption(options, 'approve', 'yes', 'confirm');
|
|
||||||
answers[q.question] = opt ?? options[0].label;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Default: pick first option
|
|
||||||
answers[q.question] = options[0].label;
|
|
||||||
}
|
|
||||||
|
|
||||||
return answers;
|
|
||||||
}
|
|
||||||
|
|
||||||
export function createAutoResponder(config: BotConfig, step: StepName) {
|
|
||||||
return async (
|
|
||||||
toolName: string,
|
|
||||||
input: Record<string, unknown>,
|
|
||||||
): Promise<{ behavior: 'allow'; updatedInput?: Record<string, unknown> } | { behavior: 'deny'; message: string }> => {
|
|
||||||
// Auto-respond to AskUserQuestion
|
|
||||||
if (toolName === 'AskUserQuestion') {
|
|
||||||
const askInput = input as unknown as AskUserQuestionInput;
|
|
||||||
const answers = buildAnswers(askInput, config, step);
|
|
||||||
|
|
||||||
return {
|
|
||||||
behavior: 'allow',
|
|
||||||
updatedInput: { ...input, answers },
|
|
||||||
};
|
|
||||||
}
|
|
||||||
|
|
||||||
// Allow all other tools
|
|
||||||
return { behavior: 'allow' };
|
|
||||||
};
|
|
||||||
}
|
|
||||||
@@ -1,16 +1,107 @@
|
|||||||
import { runCLI } from '../cli.js';
|
import { runCLI } from '../cli.js';
|
||||||
|
|
||||||
function parseArgs(): { idea?: string; resume?: boolean } {
|
function showHelp(): void {
|
||||||
|
console.log(`
|
||||||
|
+----------------------------------------------------------------+
|
||||||
|
� PLAN2CODEDE-BOT �
|
||||||
|
�----------------------------------------------------------------�
|
||||||
|
� Autonomous workflow test runner foplan2codede �
|
||||||
|
� Features LLM-as-judge for honest quality evaluation �
|
||||||
|
+----------------------------------------------------------------+
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
plan2code-bot [options]
|
||||||
|
|
||||||
|
Options:
|
||||||
|
--help Show this help message
|
||||||
|
--idea <string> Seed the idea generator with a specific concept
|
||||||
|
Example: --idea "web app for weather"
|
||||||
|
Example: --idea="CLI tool for CSV conversion"
|
||||||
|
--resume Resume a previous incomplete run
|
||||||
|
|
||||||
|
Modes:
|
||||||
|
� New Project Mode - Run from empty directory
|
||||||
|
The bot generates an app idea, creates a subdirectory, writes
|
||||||
|
IDEA.md, runs init, then all 4 workflow steps.
|
||||||
|
|
||||||
|
� Enhancement Mode - Run from directory with AGENTS.md
|
||||||
|
The bot scans the existing codebase, proposes an enhancement,
|
||||||
|
writes IDEA.md, then runs plan through finalize.
|
||||||
|
|
||||||
|
Evaluation:
|
||||||
|
The bot acts as an authentic QA agent, using LLM-based decision
|
||||||
|
making during execution and providing honest quality assessments
|
||||||
|
after each step. Results are written to specs/<feature>/BOT-EVALUATION.md
|
||||||
|
and specs/<feature>/BOT-NOTES.md for metrics analysis.
|
||||||
|
|
||||||
|
Examples:
|
||||||
|
# New project (from empty directory)
|
||||||
|
plan2code-bot
|
||||||
|
|
||||||
|
# Enhancement (from existing project)
|
||||||
|
cd my-project && plan2code-bot
|
||||||
|
|
||||||
|
# With specific idea
|
||||||
|
plan2code-bot --idea "markdown editor with live preview"
|
||||||
|
|
||||||
|
# Resume incomplete run
|
||||||
|
plan2code-bot --resume
|
||||||
|
|
||||||
|
Documentation:
|
||||||
|
https://jparkerweb.github.io/plan2code
|
||||||
|
`);
|
||||||
|
}
|
||||||
|
|
||||||
|
function stripQuotes(str: string): string {
|
||||||
|
// Remove surrounding quotes if present (both single and double)
|
||||||
|
if ((str.startsWith('"') && str.endsWith('"')) ||
|
||||||
|
(str.startsWith("'") && str.endsWith("'"))) {
|
||||||
|
return str.slice(1, -1);
|
||||||
|
}
|
||||||
|
return str;
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseArgs(): { idea?: string; resume?: boolean; help?: boolean } {
|
||||||
const args = process.argv.slice(2);
|
const args = process.argv.slice(2);
|
||||||
const ideaIdx = args.indexOf('--idea');
|
|
||||||
const idea = ideaIdx !== -1 && ideaIdx + 1 < args.length ? args[ideaIdx + 1] : undefined;
|
// Check for --help
|
||||||
|
if (args.includes('--help') || args.includes('-h')) {
|
||||||
|
return { help: true };
|
||||||
|
}
|
||||||
|
|
||||||
|
// Parse --idea (supports both --idea="value" and --idea "value")
|
||||||
|
let idea: string | undefined;
|
||||||
|
for (let i = 0; i < args.length; i++) {
|
||||||
|
const arg = args[i];
|
||||||
|
|
||||||
|
// Format: --idea="value"
|
||||||
|
if (arg.startsWith('--idea=')) {
|
||||||
|
idea = stripQuotes(arg.substring('--idea='.length));
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Format: --idea "value"
|
||||||
|
if (arg === '--idea' && i + 1 < args.length) {
|
||||||
|
idea = stripQuotes(args[i + 1]);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Parse --resume
|
||||||
const resume = args.includes('--resume');
|
const resume = args.includes('--resume');
|
||||||
|
|
||||||
return { idea, resume: resume || undefined };
|
return { idea, resume: resume || undefined };
|
||||||
}
|
}
|
||||||
|
|
||||||
async function main() {
|
async function main() {
|
||||||
try {
|
try {
|
||||||
const { idea, resume } = parseArgs();
|
const { idea, resume, help } = parseArgs();
|
||||||
|
|
||||||
|
if (help) {
|
||||||
|
showHelp();
|
||||||
|
process.exit(0);
|
||||||
|
}
|
||||||
|
|
||||||
await runCLI({ idea, resume });
|
await runCLI({ idea, resume });
|
||||||
process.exit(0);
|
process.exit(0);
|
||||||
} catch (err) {
|
} catch (err) {
|
||||||
|
|||||||
+102
-13
@@ -7,6 +7,8 @@ import { runSession } from './session-runner.js';
|
|||||||
import { buildStepPrompt } from './prompts/step-instructions.js';
|
import { buildStepPrompt } from './prompts/step-instructions.js';
|
||||||
import { checkAllPhasesComplete } from './step-detector.js';
|
import { checkAllPhasesComplete } from './step-detector.js';
|
||||||
import { saveState, loadState, deleteState, findExistingState } from './bot-state.js';
|
import { saveState, loadState, deleteState, findExistingState } from './bot-state.js';
|
||||||
|
import { ObservationCollector } from './observation-collector.js';
|
||||||
|
import { evaluateStep } from './evaluator.js';
|
||||||
import type { BotConfig, BotMode, BotState, StepName, StepResult } from './types.js';
|
import type { BotConfig, BotMode, BotState, StepName, StepResult } from './types.js';
|
||||||
|
|
||||||
const BANNER = `
|
const BANNER = `
|
||||||
@@ -149,30 +151,64 @@ async function runStep(
|
|||||||
const prompt = buildStepPrompt(step, config);
|
const prompt = buildStepPrompt(step, config);
|
||||||
|
|
||||||
try {
|
try {
|
||||||
|
// Create observation collector
|
||||||
|
const collector = new ObservationCollector(step);
|
||||||
|
|
||||||
const result = await runSession({
|
const result = await runSession({
|
||||||
prompt,
|
prompt,
|
||||||
config,
|
config,
|
||||||
step,
|
step,
|
||||||
maxTurns: step === 'implement' ? 80 : 50,
|
maxTurns: step === 'implement' ? 80 : 50,
|
||||||
|
collector,
|
||||||
});
|
});
|
||||||
|
|
||||||
const stepResult: StepResult = {
|
|
||||||
step,
|
|
||||||
success: result.success,
|
|
||||||
sessionId: result.sessionId,
|
|
||||||
duration: result.duration,
|
|
||||||
error: result.success ? null : 'Session failed',
|
|
||||||
};
|
|
||||||
|
|
||||||
if (result.success) {
|
if (result.success) {
|
||||||
spinner.succeed(
|
spinner.succeed(
|
||||||
chalk.green(`${step} completed in ${formatDuration(result.duration)}`)
|
chalk.green(`${step} completed in ${formatDuration(result.duration)}`)
|
||||||
);
|
);
|
||||||
|
|
||||||
|
// Run evaluation
|
||||||
|
spinner.text = chalk.cyan('Evaluating step quality...');
|
||||||
|
spinner.start();
|
||||||
|
|
||||||
|
const evaluation = await evaluateStep(step, result.observations, config.projectDir);
|
||||||
|
|
||||||
|
spinner.succeed(
|
||||||
|
chalk.cyan(`Evaluation complete: ${formatScore(evaluation.score)}`)
|
||||||
|
);
|
||||||
|
|
||||||
|
// Display evaluation summary
|
||||||
|
console.log(chalk.dim(` Score: ${formatScore(evaluation.score)}`));
|
||||||
|
if (evaluation.strengths.length > 0) {
|
||||||
|
console.log(chalk.green(` ✓ ${evaluation.strengths[0]}`));
|
||||||
|
}
|
||||||
|
if (evaluation.weaknesses.length > 0) {
|
||||||
|
console.log(chalk.yellow(` ⚠ ${evaluation.weaknesses[0]}`));
|
||||||
|
}
|
||||||
|
|
||||||
|
const stepResult: StepResult = {
|
||||||
|
step,
|
||||||
|
success: result.success,
|
||||||
|
sessionId: result.sessionId,
|
||||||
|
duration: result.duration,
|
||||||
|
error: null,
|
||||||
|
evaluation,
|
||||||
|
observations: result.observations,
|
||||||
|
};
|
||||||
|
|
||||||
|
return stepResult;
|
||||||
} else {
|
} else {
|
||||||
spinner.fail(chalk.red(`${step} failed after ${formatDuration(result.duration)}`));
|
spinner.fail(chalk.red(`${step} failed after ${formatDuration(result.duration)}`));
|
||||||
}
|
|
||||||
|
|
||||||
return stepResult;
|
return {
|
||||||
|
step,
|
||||||
|
success: false,
|
||||||
|
sessionId: result.sessionId,
|
||||||
|
duration: result.duration,
|
||||||
|
error: 'Session failed',
|
||||||
|
observations: result.observations,
|
||||||
|
};
|
||||||
|
}
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
const errorMsg = error instanceof Error ? error.message : String(error);
|
const errorMsg = error instanceof Error ? error.message : String(error);
|
||||||
spinner.fail(chalk.red(`${step} error: ${errorMsg}`));
|
spinner.fail(chalk.red(`${step} error: ${errorMsg}`));
|
||||||
@@ -186,6 +222,12 @@ async function runStep(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function formatScore(score: number): string {
|
||||||
|
if (score >= 85) return chalk.green(`${score}/100`);
|
||||||
|
if (score >= 70) return chalk.yellow(`${score}/100`);
|
||||||
|
return chalk.red(`${score}/100`);
|
||||||
|
}
|
||||||
|
|
||||||
export interface CLIOptions {
|
export interface CLIOptions {
|
||||||
idea?: string;
|
idea?: string;
|
||||||
resume?: boolean;
|
resume?: boolean;
|
||||||
@@ -438,6 +480,28 @@ export async function runCLI(options: CLIOptions = {}): Promise<void> {
|
|||||||
console.log(chalk.dim('--- Finalize (skipped — previously succeeded) ---'));
|
console.log(chalk.dim('--- Finalize (skipped — previously succeeded) ---'));
|
||||||
} else {
|
} else {
|
||||||
console.log(chalk.bold('--- Finalize ---'));
|
console.log(chalk.bold('--- Finalize ---'));
|
||||||
|
|
||||||
|
// Check quality gate: average score must be >= 60
|
||||||
|
const evaluatedSteps = state.steps.filter((s) => s.evaluation);
|
||||||
|
if (evaluatedSteps.length > 0) {
|
||||||
|
const avgScore =
|
||||||
|
evaluatedSteps.reduce((sum, s) => sum + (s.evaluation?.score ?? 0), 0) /
|
||||||
|
evaluatedSteps.length;
|
||||||
|
|
||||||
|
console.log(chalk.dim(` Average quality score: ${formatScore(Math.round(avgScore))}`));
|
||||||
|
|
||||||
|
if (avgScore < 60) {
|
||||||
|
console.log(
|
||||||
|
chalk.red(
|
||||||
|
'\n⚠ Quality gate failed: Average score is below 60. Please review and fix issues before finalizing.'
|
||||||
|
)
|
||||||
|
);
|
||||||
|
console.log(chalk.dim(' Check specs/<feature>/BOT-EVALUATION.md for detailed feedback.'));
|
||||||
|
printSummary(state, workDir, false);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
state.currentStep = 'finalize';
|
state.currentStep = 'finalize';
|
||||||
const finalizeResult = await runStep('finalize', config, state);
|
const finalizeResult = await runStep('finalize', config, state);
|
||||||
state.steps.push(finalizeResult);
|
state.steps.push(finalizeResult);
|
||||||
@@ -485,14 +549,39 @@ function printSummary(state: BotState, workDir: string, allSucceeded: boolean):
|
|||||||
|
|
||||||
for (const step of state.steps) {
|
for (const step of state.steps) {
|
||||||
const icon = step.success ? chalk.green('✓') : chalk.red('✗');
|
const icon = step.success ? chalk.green('✓') : chalk.red('✗');
|
||||||
console.log(` ${icon} ${step.step.padEnd(12)} ${formatDuration(step.duration)}`);
|
const scoreText = step.evaluation
|
||||||
|
? ` [${formatScore(step.evaluation.score)}]`
|
||||||
|
: '';
|
||||||
|
console.log(
|
||||||
|
` ${icon} ${step.step.padEnd(12)} ${formatDuration(step.duration)}${scoreText}`
|
||||||
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
console.log('');
|
console.log('');
|
||||||
console.log(chalk.dim(` Total: ${successCount}/${state.steps.length} steps succeeded in ${formatDuration(totalDuration)}`));
|
|
||||||
|
// Display average quality score if available
|
||||||
|
const evaluatedSteps = state.steps.filter((s) => s.evaluation);
|
||||||
|
if (evaluatedSteps.length > 0) {
|
||||||
|
const avgScore =
|
||||||
|
evaluatedSteps.reduce((sum, s) => sum + (s.evaluation?.score ?? 0), 0) /
|
||||||
|
evaluatedSteps.length;
|
||||||
|
console.log(
|
||||||
|
chalk.dim(` Average quality: ${formatScore(Math.round(avgScore))}`)
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
console.log(
|
||||||
|
chalk.dim(
|
||||||
|
` Total: ${successCount}/${state.steps.length} steps succeeded in ${formatDuration(totalDuration)}`
|
||||||
|
)
|
||||||
|
);
|
||||||
console.log(chalk.dim(` Implement passes: ${state.implementPasses}`));
|
console.log(chalk.dim(` Implement passes: ${state.implementPasses}`));
|
||||||
if (!allSucceeded) {
|
if (!allSucceeded) {
|
||||||
console.log(chalk.dim(` State saved to: ${path.relative(workDir, projectDir)}/.plan2code-bot-state.json`));
|
console.log(
|
||||||
|
chalk.dim(
|
||||||
|
` State saved to: ${path.relative(workDir, projectDir)}/.plan2code-bot-state.json`
|
||||||
|
)
|
||||||
|
);
|
||||||
}
|
}
|
||||||
console.log('');
|
console.log('');
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,311 @@
|
|||||||
|
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||||
|
import fs from 'fs-extra';
|
||||||
|
import path from 'path';
|
||||||
|
import type { EvaluationResult, ExecutionObservation, StepName } from './types.js';
|
||||||
|
import { getCriteriaForStep } from './prompts/evaluation-criteria.js';
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Evaluates a completed step using LLM-as-judge.
|
||||||
|
* Returns honest quality assessment based on observations and artifacts.
|
||||||
|
*/
|
||||||
|
export async function evaluateStep(
|
||||||
|
step: StepName,
|
||||||
|
observations: ExecutionObservation,
|
||||||
|
projectDir: string
|
||||||
|
): Promise<EvaluationResult> {
|
||||||
|
const criteria = getCriteriaForStep(step);
|
||||||
|
const prompt = buildEvaluationPrompt(step, observations, criteria, projectDir);
|
||||||
|
|
||||||
|
try {
|
||||||
|
// Query LLM with ability to inspect artifacts
|
||||||
|
const session = query({
|
||||||
|
prompt,
|
||||||
|
options: {
|
||||||
|
maxTurns: 30,
|
||||||
|
cwd: projectDir,
|
||||||
|
permissionMode: 'bypassPermissions',
|
||||||
|
allowDangerouslySkipPermissions: true,
|
||||||
|
allowedTools: ['Read', 'Glob', 'Grep'],
|
||||||
|
systemPrompt:
|
||||||
|
'You are a QA engineer evaluating completed work. Be thorough, honest, and constructive.',
|
||||||
|
},
|
||||||
|
});
|
||||||
|
|
||||||
|
let output = '';
|
||||||
|
for await (const message of session) {
|
||||||
|
if (message.type === 'assistant') {
|
||||||
|
for (const block of message.message.content) {
|
||||||
|
if (block.type === 'text') output += block.text;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const evaluation = parseEvaluationOutput(output, step);
|
||||||
|
|
||||||
|
// Write evaluation files
|
||||||
|
await writeEvaluationFiles(evaluation, observations, projectDir);
|
||||||
|
|
||||||
|
return evaluation;
|
||||||
|
} catch (error) {
|
||||||
|
console.warn('Evaluation failed, using fallback:', error);
|
||||||
|
return createFallbackEvaluation(step, observations);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function buildEvaluationPrompt(
|
||||||
|
step: StepName,
|
||||||
|
observations: ExecutionObservation,
|
||||||
|
criteria: ReturnType<typeof getCriteriaForStep>,
|
||||||
|
projectDir: string
|
||||||
|
): string {
|
||||||
|
const duration = observations.endTime - observations.startTime;
|
||||||
|
const durationSec = Math.floor(duration / 1000);
|
||||||
|
|
||||||
|
const toolsSummary = observations.tools
|
||||||
|
.map((t) => `- ${t.toolName} (${new Date(t.timestamp).toISOString()})`)
|
||||||
|
.join('\n');
|
||||||
|
|
||||||
|
const filesSummary = [
|
||||||
|
...observations.filesCreated.map((f) => `CREATED: ${f}`),
|
||||||
|
...observations.filesModified.map((f) => `MODIFIED: ${f}`),
|
||||||
|
].join('\n');
|
||||||
|
|
||||||
|
const questionsSummary = observations.questionsAsked
|
||||||
|
.map(
|
||||||
|
(q) =>
|
||||||
|
`Q: ${q.question}\nA: ${q.selectedAnswer}\nReasoning: ${q.llmReasoning}`
|
||||||
|
)
|
||||||
|
.join('\n\n');
|
||||||
|
|
||||||
|
return `You are a QA engineer evaluating the quality of a completed workflow step.
|
||||||
|
|
||||||
|
## Step Information
|
||||||
|
- Step: ${step}
|
||||||
|
- Duration: ${durationSec}s
|
||||||
|
- Tools used: ${observations.tools.length}
|
||||||
|
- Files created: ${observations.filesCreated.length}
|
||||||
|
- Files modified: ${observations.filesModified.length}
|
||||||
|
- Errors: ${observations.errors.length}
|
||||||
|
|
||||||
|
## Execution Details
|
||||||
|
|
||||||
|
### Tools Used
|
||||||
|
${toolsSummary || '(none)'}
|
||||||
|
|
||||||
|
### Files Changed
|
||||||
|
${filesSummary || '(none)'}
|
||||||
|
|
||||||
|
${observations.questionsAsked.length > 0 ? `### Questions & Decisions\n${questionsSummary}` : ''}
|
||||||
|
|
||||||
|
${observations.errors.length > 0 ? `### Errors Encountered\n${observations.errors.join('\n')}` : ''}
|
||||||
|
|
||||||
|
## Evaluation Criteria
|
||||||
|
|
||||||
|
**Key Artifacts Expected:**
|
||||||
|
${criteria.keyArtifacts.map((a) => `- ${a}`).join('\n')}
|
||||||
|
|
||||||
|
**Quality Checks:**
|
||||||
|
${criteria.qualityChecks.map((c) => `- ${c}`).join('\n')}
|
||||||
|
|
||||||
|
**Common Pitfalls to Watch For:**
|
||||||
|
${criteria.commonPitfalls.map((p) => `- ${p}`).join('\n')}
|
||||||
|
|
||||||
|
**Scoring Guidance:**
|
||||||
|
${criteria.scoringGuidance}
|
||||||
|
|
||||||
|
## Your Task
|
||||||
|
|
||||||
|
Evaluate this step honestly and thoroughly:
|
||||||
|
|
||||||
|
1. **Inspect the artifacts** using Read, Glob, and Grep tools
|
||||||
|
2. **Check against quality criteria** listed above
|
||||||
|
3. **Identify strengths and weaknesses** based on actual evidence
|
||||||
|
4. **Provide constructive suggestions** for improvement
|
||||||
|
5. **Assign an honest score** (0-100) following the guidance
|
||||||
|
|
||||||
|
Be critical but fair. Most work scores 70-85. Don't inflate scores.
|
||||||
|
|
||||||
|
## Response Format
|
||||||
|
|
||||||
|
SCORE: <number 0-100>
|
||||||
|
|
||||||
|
STRENGTHS:
|
||||||
|
- <strength 1>
|
||||||
|
- <strength 2>
|
||||||
|
- <strength 3>
|
||||||
|
|
||||||
|
WEAKNESSES:
|
||||||
|
- <weakness 1>
|
||||||
|
- <weakness 2>
|
||||||
|
|
||||||
|
SUGGESTIONS:
|
||||||
|
- <suggestion 1>
|
||||||
|
- <suggestion 2>
|
||||||
|
|
||||||
|
CRITICAL_ISSUES:
|
||||||
|
- <critical issue 1 (or "None")>
|
||||||
|
|
||||||
|
REASONING:
|
||||||
|
<1-2 paragraphs explaining your evaluation, referencing specific files/evidence>
|
||||||
|
|
||||||
|
Provide honest, evidence-based evaluation.`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseEvaluationOutput(
|
||||||
|
output: string,
|
||||||
|
step: StepName
|
||||||
|
): EvaluationResult {
|
||||||
|
// Extract score
|
||||||
|
const scoreMatch = output.match(/SCORE:\s*(\d+)/i);
|
||||||
|
if (!scoreMatch) {
|
||||||
|
console.warn(`Evaluation parser: no SCORE found in output (${output.length} chars). Defaulting to 50.`);
|
||||||
|
}
|
||||||
|
const score = scoreMatch ? parseInt(scoreMatch[1], 10) : 50;
|
||||||
|
|
||||||
|
// Extract sections
|
||||||
|
const strengthsMatch = output.match(
|
||||||
|
/STRENGTHS:\s*((?:- .+\n?)+)/i
|
||||||
|
);
|
||||||
|
const weaknessesMatch = output.match(
|
||||||
|
/WEAKNESSES:\s*((?:- .+\n?)+)/i
|
||||||
|
);
|
||||||
|
const suggestionsMatch = output.match(
|
||||||
|
/SUGGESTIONS:\s*((?:- .+\n?)+)/i
|
||||||
|
);
|
||||||
|
const criticalMatch = output.match(
|
||||||
|
/CRITICAL_ISSUES:\s*((?:- .+\n?)+)/i
|
||||||
|
);
|
||||||
|
const reasoningMatch = output.match(/REASONING:\s*(.+?)(?=\n\n|$)/is);
|
||||||
|
|
||||||
|
const parseList = (text: string | undefined): string[] => {
|
||||||
|
if (!text) return [];
|
||||||
|
return text
|
||||||
|
.split('\n')
|
||||||
|
.map((line) => line.replace(/^-\s*/, '').trim())
|
||||||
|
.filter((line) => line.length > 0 && !line.toLowerCase().includes('none'));
|
||||||
|
};
|
||||||
|
|
||||||
|
return {
|
||||||
|
step,
|
||||||
|
score: Math.max(0, Math.min(100, score)),
|
||||||
|
strengths: parseList(strengthsMatch?.[1]),
|
||||||
|
weaknesses: parseList(weaknessesMatch?.[1]),
|
||||||
|
suggestions: parseList(suggestionsMatch?.[1]),
|
||||||
|
criticalIssues: parseList(criticalMatch?.[1]),
|
||||||
|
timestamp: Date.now(),
|
||||||
|
evaluatorModel: 'claude-sonnet-4-5',
|
||||||
|
reasoning: reasoningMatch?.[1]?.trim() || 'No reasoning provided',
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function createFallbackEvaluation(
|
||||||
|
step: StepName,
|
||||||
|
observations: ExecutionObservation
|
||||||
|
): EvaluationResult {
|
||||||
|
return {
|
||||||
|
step,
|
||||||
|
score: 50,
|
||||||
|
strengths: ['Step completed'],
|
||||||
|
weaknesses: ['Evaluation failed - using fallback'],
|
||||||
|
suggestions: ['Re-run with evaluation enabled'],
|
||||||
|
criticalIssues: ['Evaluation system unavailable'],
|
||||||
|
timestamp: Date.now(),
|
||||||
|
evaluatorModel: 'fallback',
|
||||||
|
reasoning: 'Evaluation failed, using fallback. Cannot provide detailed assessment.',
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Resolves the active spec subdirectory (e.g. specs/<feature>/).
|
||||||
|
* Falls back to projectDir if no spec folder exists yet (e.g. during init).
|
||||||
|
*/
|
||||||
|
function resolveOutputDir(projectDir: string): string {
|
||||||
|
const specsDir = path.join(projectDir, 'specs');
|
||||||
|
if (fs.existsSync(specsDir)) {
|
||||||
|
const entries = fs.readdirSync(specsDir, { withFileTypes: true });
|
||||||
|
const firstSpec = entries.find((e) => e.isDirectory());
|
||||||
|
if (firstSpec) {
|
||||||
|
return path.join(specsDir, firstSpec.name);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return projectDir;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function writeEvaluationFiles(
|
||||||
|
evaluation: EvaluationResult,
|
||||||
|
observations: ExecutionObservation,
|
||||||
|
projectDir: string
|
||||||
|
): Promise<void> {
|
||||||
|
const outputDir = resolveOutputDir(projectDir);
|
||||||
|
|
||||||
|
// Write BOT-EVALUATION.md
|
||||||
|
const evalContent = formatEvaluationMarkdown(evaluation);
|
||||||
|
const evalPath = path.join(outputDir, 'BOT-EVALUATION.md');
|
||||||
|
|
||||||
|
if (fs.existsSync(evalPath)) {
|
||||||
|
// Append to existing file
|
||||||
|
const existing = fs.readFileSync(evalPath, 'utf-8');
|
||||||
|
fs.writeFileSync(evalPath, existing + '\n\n---\n\n' + evalContent);
|
||||||
|
} else {
|
||||||
|
fs.writeFileSync(evalPath, evalContent);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Write BOT-NOTES.md
|
||||||
|
const notesContent = formatObservationsMarkdown(observations);
|
||||||
|
const notesPath = path.join(outputDir, 'BOT-NOTES.md');
|
||||||
|
|
||||||
|
if (fs.existsSync(notesPath)) {
|
||||||
|
const existing = fs.readFileSync(notesPath, 'utf-8');
|
||||||
|
fs.writeFileSync(notesPath, existing + '\n\n---\n\n' + notesContent);
|
||||||
|
} else {
|
||||||
|
fs.writeFileSync(notesPath, notesContent);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function formatEvaluationMarkdown(evaluation: EvaluationResult): string {
|
||||||
|
return `# Evaluation: ${evaluation.step} Step
|
||||||
|
|
||||||
|
**Score:** ${evaluation.score}/100
|
||||||
|
**Timestamp:** ${new Date(evaluation.timestamp).toISOString()}
|
||||||
|
**Evaluator:** ${evaluation.evaluatorModel}
|
||||||
|
|
||||||
|
## Strengths
|
||||||
|
${evaluation.strengths.map((s) => `- ${s}`).join('\n') || '(none)'}
|
||||||
|
|
||||||
|
## Weaknesses
|
||||||
|
${evaluation.weaknesses.map((w) => `- ${w}`).join('\n') || '(none)'}
|
||||||
|
|
||||||
|
## Suggestions for Improvement
|
||||||
|
${evaluation.suggestions.map((s) => `- ${s}`).join('\n') || '(none)'}
|
||||||
|
|
||||||
|
${evaluation.criticalIssues.length > 0 ? `## Critical Issues\n${evaluation.criticalIssues.map((i) => `- ${i}`).join('\n')}\n` : ''}
|
||||||
|
|
||||||
|
## Reasoning
|
||||||
|
${evaluation.reasoning}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function formatObservationsMarkdown(observations: ExecutionObservation): string {
|
||||||
|
const duration = observations.endTime - observations.startTime;
|
||||||
|
const durationSec = Math.floor(duration / 1000);
|
||||||
|
|
||||||
|
return `# Observations: ${observations.step} Step
|
||||||
|
|
||||||
|
**Duration:** ${durationSec}s
|
||||||
|
**Tools Used:** ${observations.tools.length}
|
||||||
|
**Files Created:** ${observations.filesCreated.length}
|
||||||
|
**Files Modified:** ${observations.filesModified.length}
|
||||||
|
**Errors:** ${observations.errors.length}
|
||||||
|
|
||||||
|
${observations.questionsAsked.length > 0 ? `## Questions Asked & Answers\n\n${observations.questionsAsked.map((q) => `### Question: "${q.question}"\n**Selected:** "${q.selectedAnswer}"\n**Reasoning:** ${q.llmReasoning}`).join('\n\n')}\n` : ''}
|
||||||
|
|
||||||
|
## Tool Usage
|
||||||
|
${observations.tools.map((t) => `- ${t.toolName}`).join('\n')}
|
||||||
|
|
||||||
|
## Files Created
|
||||||
|
${observations.filesCreated.map((f) => `- ${f}`).join('\n') || '(none)'}
|
||||||
|
|
||||||
|
## Files Modified
|
||||||
|
${observations.filesModified.map((f) => `- ${f}`).join('\n') || '(none)'}
|
||||||
|
|
||||||
|
${observations.assistantMessages.length > 0 ? `## Assistant Output Summary\n${observations.assistantMessages.slice(0, 3).map((m) => `> ${m.substring(0, 100)}...`).join('\n')}\n` : ''}`;
|
||||||
|
}
|
||||||
@@ -22,13 +22,29 @@ function parseIdea(text: string): IdeaResult {
|
|||||||
}
|
}
|
||||||
|
|
||||||
export async function generateNewAppIdea(seed?: string): Promise<IdeaResult> {
|
export async function generateNewAppIdea(seed?: string): Promise<IdeaResult> {
|
||||||
const category = Math.random() > 0.5 ? 'CLI tool' : 'small web app';
|
const categories = [
|
||||||
|
'CLI tool',
|
||||||
|
'single-page web app',
|
||||||
|
'REST API service',
|
||||||
|
'browser extension',
|
||||||
|
'interactive data visualization dashboard',
|
||||||
|
'terminal-based game',
|
||||||
|
'real-time web app (using WebSockets)',
|
||||||
|
'static site generator or theme',
|
||||||
|
'browser-based game',
|
||||||
|
'desktop utility (using Electron or Tauri)',
|
||||||
|
'chat bot or conversational tool',
|
||||||
|
'automation script or workflow tool',
|
||||||
|
];
|
||||||
|
const category = categories[Math.floor(Math.random() * categories.length)];
|
||||||
|
|
||||||
const seedClause = seed
|
const seedClause = seed
|
||||||
? `\n\nUse this as inspiration for the idea: "${seed}"`
|
? `\n\nThe user provided this seed for inspiration. Stay closely aligned with the theme and intent of the seed — build on it, don't ignore it:\n"${seed}"`
|
||||||
: '';
|
: '';
|
||||||
|
|
||||||
const prompt = `Generate a random, creative idea for a ${category} that a developer might build as a side project. The project should be achievable in a single coding session (1-2 hours) and should be interesting but not overly complex.${seedClause}
|
const prompt = `Generate a random, creative idea for a ${category}. The project should be achievable in a single coding session (1-2 hours) and should be interesting but not overly complex.
|
||||||
|
|
||||||
|
IMPORTANT: Be creative and diverse with your ideas. Avoid defaulting to developer-centric tools (git analyzers, code formatters, repo scanners, etc.) unless the category specifically calls for it. Think about ideas that would appeal to a broad audience — productivity, entertainment, education, health, finance, art, music, social, cooking, travel, fitness, etc.${seedClause}
|
||||||
|
|
||||||
Respond in EXACTLY this format (no other text):
|
Respond in EXACTLY this format (no other text):
|
||||||
NAME: <kebab-case-project-name>
|
NAME: <kebab-case-project-name>
|
||||||
@@ -38,7 +54,7 @@ DESCRIPTION: <2-3 sentence description of what the app does, its key features, a
|
|||||||
prompt,
|
prompt,
|
||||||
options: {
|
options: {
|
||||||
maxTurns: 1,
|
maxTurns: 1,
|
||||||
systemPrompt: 'You are a creative software project idea generator. Respond only in the exact format requested.',
|
systemPrompt: 'You are a wildly creative project idea generator. You come up with surprising, fun, and diverse software project ideas spanning many domains — not just developer tools. Respond only in the exact format requested.',
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,243 @@
|
|||||||
|
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||||
|
import type { BotConfig, ExecutionObservation, StepName } from './types.js';
|
||||||
|
import type { ObservationCollector } from './observation-collector.js';
|
||||||
|
|
||||||
|
interface AskUserQuestionInput {
|
||||||
|
questions: Array<{
|
||||||
|
question: string;
|
||||||
|
options: Array<{ label: string; description: string }>;
|
||||||
|
multiSelect?: boolean;
|
||||||
|
}>;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface IntelligentAnswers {
|
||||||
|
answers: Record<string, string>;
|
||||||
|
reasoning: Record<string, string>;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Creates an intelligent responder that uses LLM-as-judge for ALL decisions.
|
||||||
|
* Replaces the old hardcoded auto-responder logic.
|
||||||
|
*/
|
||||||
|
export function createIntelligentResponder(
|
||||||
|
config: BotConfig,
|
||||||
|
step: StepName,
|
||||||
|
collector: ObservationCollector
|
||||||
|
) {
|
||||||
|
return async (
|
||||||
|
toolName: string,
|
||||||
|
input: Record<string, unknown>
|
||||||
|
): Promise<{ behavior: 'allow'; updatedInput?: Record<string, unknown> } | { behavior: 'deny'; message: string }> => {
|
||||||
|
// Record every tool invocation for observations
|
||||||
|
collector.recordToolUse(toolName, input, undefined, true);
|
||||||
|
|
||||||
|
// Handle AskUserQuestion with LLM-generated answers
|
||||||
|
if (toolName === 'AskUserQuestion') {
|
||||||
|
const askInput = input as unknown as AskUserQuestionInput;
|
||||||
|
|
||||||
|
// Get current observations to provide context to LLM
|
||||||
|
const observations = collector.getSnapshot();
|
||||||
|
|
||||||
|
// Generate answers using LLM-as-judge
|
||||||
|
const answers = await generateIntelligentAnswers(
|
||||||
|
askInput,
|
||||||
|
observations,
|
||||||
|
config,
|
||||||
|
step
|
||||||
|
);
|
||||||
|
|
||||||
|
// Record each question/answer pair for metrics
|
||||||
|
for (const q of askInput.questions) {
|
||||||
|
const answer = answers.answers[q.question];
|
||||||
|
const reasoning = answers.reasoning[q.question] || 'No reasoning provided';
|
||||||
|
collector.recordQuestion(q.question, q.options, answer, reasoning);
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
behavior: 'allow',
|
||||||
|
updatedInput: { ...input, answers: answers.answers },
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
// Allow all other tools
|
||||||
|
return { behavior: 'allow' };
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
async function generateIntelligentAnswers(
|
||||||
|
askInput: AskUserQuestionInput,
|
||||||
|
observations: ExecutionObservation,
|
||||||
|
config: BotConfig,
|
||||||
|
step: StepName
|
||||||
|
): Promise<IntelligentAnswers> {
|
||||||
|
const prompt = buildDecisionPrompt(askInput, observations, config, step);
|
||||||
|
|
||||||
|
try {
|
||||||
|
// Query LLM for decision (single turn, read-only tools)
|
||||||
|
const session = query({
|
||||||
|
prompt,
|
||||||
|
options: {
|
||||||
|
maxTurns: 1,
|
||||||
|
cwd: config.projectDir,
|
||||||
|
permissionMode: 'bypassPermissions',
|
||||||
|
allowDangerouslySkipPermissions: true,
|
||||||
|
allowedTools: ['Read', 'Glob', 'Grep'],
|
||||||
|
systemPrompt: 'You are a QA engineer reviewing work-in-progress. Be thoughtful and honest.',
|
||||||
|
},
|
||||||
|
});
|
||||||
|
|
||||||
|
let output = '';
|
||||||
|
for await (const message of session) {
|
||||||
|
if (message.type === 'assistant') {
|
||||||
|
for (const block of message.message.content) {
|
||||||
|
if (block.type === 'text') output += block.text;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return parseDecisionOutput(output, askInput);
|
||||||
|
} catch (error) {
|
||||||
|
console.warn('LLM decision failed, using fallback logic:', error);
|
||||||
|
// Fallback to reasonable defaults if LLM fails
|
||||||
|
return generateFallbackAnswers(askInput, step);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function buildDecisionPrompt(
|
||||||
|
askInput: AskUserQuestionInput,
|
||||||
|
observations: ExecutionObservation,
|
||||||
|
config: BotConfig,
|
||||||
|
step: StepName
|
||||||
|
): string {
|
||||||
|
const duration = observations.endTime - observations.startTime;
|
||||||
|
const toolSummary = observations.tools
|
||||||
|
.map((t) => `- ${t.toolName}`)
|
||||||
|
.join('\n');
|
||||||
|
const filesSummary = [
|
||||||
|
...observations.filesCreated.map((f) => `CREATED: ${f}`),
|
||||||
|
...observations.filesModified.map((f) => `MODIFIED: ${f}`),
|
||||||
|
].join('\n');
|
||||||
|
|
||||||
|
const questionsText = askInput.questions
|
||||||
|
.map((q, i) => {
|
||||||
|
const optionsText = q.options
|
||||||
|
.map((o, j) => ` ${j + 1}. ${o.label} - ${o.description}`)
|
||||||
|
.join('\n');
|
||||||
|
return `QUESTION ${i + 1}: ${q.question}\nOptions:\n${optionsText}`;
|
||||||
|
})
|
||||||
|
.join('\n\n');
|
||||||
|
|
||||||
|
return `You are a QA engineer reviewing a workflow step in progress.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
- Step: ${step}
|
||||||
|
- Mode: ${config.mode}
|
||||||
|
- Duration so far: ${Math.floor(duration / 1000)}s
|
||||||
|
- Tools used: ${observations.tools.length}
|
||||||
|
- Errors encountered: ${observations.errors.length}
|
||||||
|
|
||||||
|
## What's Happened So Far
|
||||||
|
|
||||||
|
### Tools Used
|
||||||
|
${toolSummary || '(none yet)'}
|
||||||
|
|
||||||
|
### Files Changed
|
||||||
|
${filesSummary || '(none yet)'}
|
||||||
|
|
||||||
|
${observations.errors.length > 0 ? `### Errors\n${observations.errors.join('\n')}` : ''}
|
||||||
|
|
||||||
|
## Questions to Answer
|
||||||
|
|
||||||
|
${questionsText}
|
||||||
|
|
||||||
|
## Your Task
|
||||||
|
|
||||||
|
You need to answer these questions as a thoughtful QA engineer would:
|
||||||
|
1. Use Read, Glob, and Grep tools to inspect the current state of artifacts if needed
|
||||||
|
2. Consider what you've observed (tools used, files created, errors)
|
||||||
|
3. For each question, select the most appropriate answer
|
||||||
|
4. Provide brief reasoning for your choice
|
||||||
|
|
||||||
|
Respond in this format:
|
||||||
|
|
||||||
|
QUESTION 1:
|
||||||
|
ANSWER: <option label>
|
||||||
|
REASONING: <1-2 sentences explaining your choice>
|
||||||
|
|
||||||
|
QUESTION 2:
|
||||||
|
ANSWER: <option label>
|
||||||
|
REASONING: <1-2 sentences>
|
||||||
|
|
||||||
|
Be honest. If work looks incomplete or problematic, don't approve it.
|
||||||
|
If tests should be run but haven't been, don't skip them without good reason.
|
||||||
|
Act like a real developer who cares about quality.`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseDecisionOutput(
|
||||||
|
output: string,
|
||||||
|
askInput: AskUserQuestionInput
|
||||||
|
): IntelligentAnswers {
|
||||||
|
const answers: Record<string, string> = {};
|
||||||
|
const reasoning: Record<string, string> = {};
|
||||||
|
|
||||||
|
// Parse structured output
|
||||||
|
const questionBlocks = output.split(/QUESTION \d+:/i).slice(1);
|
||||||
|
|
||||||
|
askInput.questions.forEach((q, i) => {
|
||||||
|
const block = questionBlocks[i] || '';
|
||||||
|
|
||||||
|
const answerMatch = block.match(/ANSWER:\s*(.+?)(?=\n|$)/i);
|
||||||
|
const reasoningMatch = block.match(/REASONING:\s*(.+?)(?=\n\n|$)/is);
|
||||||
|
|
||||||
|
const selectedLabel = answerMatch?.[1]?.trim() || '';
|
||||||
|
|
||||||
|
// Find matching option by label (case-insensitive partial match)
|
||||||
|
const matchedOption = q.options.find(
|
||||||
|
(opt) =>
|
||||||
|
opt.label.toLowerCase().includes(selectedLabel.toLowerCase()) ||
|
||||||
|
selectedLabel.toLowerCase().includes(opt.label.toLowerCase())
|
||||||
|
);
|
||||||
|
|
||||||
|
answers[q.question] = matchedOption?.label || q.options[0].label;
|
||||||
|
reasoning[q.question] = reasoningMatch?.[1]?.trim() || 'No reasoning provided';
|
||||||
|
});
|
||||||
|
|
||||||
|
return { answers, reasoning };
|
||||||
|
}
|
||||||
|
|
||||||
|
function generateFallbackAnswers(
|
||||||
|
askInput: AskUserQuestionInput,
|
||||||
|
step: StepName
|
||||||
|
): IntelligentAnswers {
|
||||||
|
// Simple fallback: pick first option for most questions
|
||||||
|
// For approval questions, approve; for testing, skip
|
||||||
|
const answers: Record<string, string> = {};
|
||||||
|
const reasoning: Record<string, string> = {};
|
||||||
|
|
||||||
|
for (const q of askInput.questions) {
|
||||||
|
const questionLower = q.question.toLowerCase();
|
||||||
|
|
||||||
|
if (questionLower.includes('approve') || questionLower.includes('proceed')) {
|
||||||
|
const approveOption = q.options.find(
|
||||||
|
(o) =>
|
||||||
|
o.label.toLowerCase().includes('approve') ||
|
||||||
|
o.label.toLowerCase().includes('yes')
|
||||||
|
);
|
||||||
|
answers[q.question] = approveOption?.label || q.options[0].label;
|
||||||
|
reasoning[q.question] = 'Fallback approval (LLM unavailable)';
|
||||||
|
} else if (questionLower.includes('test')) {
|
||||||
|
const skipOption = q.options.find(
|
||||||
|
(o) =>
|
||||||
|
o.label.toLowerCase().includes('skip') ||
|
||||||
|
o.label.toLowerCase().includes('none')
|
||||||
|
);
|
||||||
|
answers[q.question] = skipOption?.label || q.options[0].label;
|
||||||
|
reasoning[q.question] = 'Fallback skip (LLM unavailable)';
|
||||||
|
} else {
|
||||||
|
answers[q.question] = q.options[0].label;
|
||||||
|
reasoning[q.question] = 'Fallback first option (LLM unavailable)';
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return { answers, reasoning };
|
||||||
|
}
|
||||||
@@ -0,0 +1,113 @@
|
|||||||
|
import type { ExecutionObservation, StepName } from './types.js';
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Collects detailed observations during step execution.
|
||||||
|
* Tracks tool usage, messages, errors, file changes, and questions asked.
|
||||||
|
*/
|
||||||
|
export class ObservationCollector {
|
||||||
|
private observations: ExecutionObservation;
|
||||||
|
private recentToolKeys: Set<string> = new Set();
|
||||||
|
|
||||||
|
constructor(step: StepName) {
|
||||||
|
this.observations = {
|
||||||
|
step,
|
||||||
|
startTime: Date.now(),
|
||||||
|
endTime: 0,
|
||||||
|
tools: [],
|
||||||
|
assistantMessages: [],
|
||||||
|
errors: [],
|
||||||
|
filesCreated: [],
|
||||||
|
filesModified: [],
|
||||||
|
questionsAsked: [],
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Records a tool invocation with its input and output.
|
||||||
|
* Deduplicates if the same tool+input is recorded from both canUseTool and the message stream.
|
||||||
|
*/
|
||||||
|
recordToolUse(
|
||||||
|
toolName: string,
|
||||||
|
input: Record<string, unknown>,
|
||||||
|
output: unknown,
|
||||||
|
allowed: boolean
|
||||||
|
): void {
|
||||||
|
// Deduplicate based on tool name + serialized input (within a short time window)
|
||||||
|
const key = `${toolName}:${JSON.stringify(input)}`;
|
||||||
|
if (this.recentToolKeys.has(key)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
this.recentToolKeys.add(key);
|
||||||
|
// Clean up old keys periodically to avoid unbounded growth
|
||||||
|
if (this.recentToolKeys.size > 500) {
|
||||||
|
const entries = [...this.recentToolKeys];
|
||||||
|
this.recentToolKeys = new Set(entries.slice(entries.length - 250));
|
||||||
|
}
|
||||||
|
|
||||||
|
this.observations.tools.push({
|
||||||
|
toolName,
|
||||||
|
input,
|
||||||
|
output,
|
||||||
|
timestamp: Date.now(),
|
||||||
|
allowed,
|
||||||
|
autoAnswered: toolName === 'AskUserQuestion',
|
||||||
|
});
|
||||||
|
|
||||||
|
// Extract file paths from common tools
|
||||||
|
if (toolName === 'Write' && input.file_path) {
|
||||||
|
this.observations.filesCreated.push(input.file_path as string);
|
||||||
|
}
|
||||||
|
if (toolName === 'Edit' && input.file_path) {
|
||||||
|
this.observations.filesModified.push(input.file_path as string);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Records messages from the session (assistant text, errors).
|
||||||
|
*/
|
||||||
|
recordMessage(message: any): void {
|
||||||
|
if (message.type === 'assistant') {
|
||||||
|
for (const block of message.message.content) {
|
||||||
|
if (block.type === 'text') {
|
||||||
|
this.observations.assistantMessages.push(block.text);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (message.type === 'error') {
|
||||||
|
this.observations.errors.push(message.error?.message ?? 'Unknown error');
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Records a question that was asked and the LLM-generated answer.
|
||||||
|
*/
|
||||||
|
recordQuestion(
|
||||||
|
question: string,
|
||||||
|
options: Array<{ label: string; description: string }>,
|
||||||
|
selectedAnswer: string,
|
||||||
|
llmReasoning: string
|
||||||
|
): void {
|
||||||
|
this.observations.questionsAsked.push({
|
||||||
|
question,
|
||||||
|
options,
|
||||||
|
llmReasoning,
|
||||||
|
selectedAnswer,
|
||||||
|
timestamp: Date.now(),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Gets a snapshot of current observations (for real-time decision making).
|
||||||
|
*/
|
||||||
|
getSnapshot(): ExecutionObservation {
|
||||||
|
return { ...this.observations, endTime: Date.now() };
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Finalizes observations and returns the complete record.
|
||||||
|
*/
|
||||||
|
finalize(): ExecutionObservation {
|
||||||
|
this.observations.endTime = Date.now();
|
||||||
|
return { ...this.observations };
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,158 @@
|
|||||||
|
import type { StepName } from '../types.js';
|
||||||
|
|
||||||
|
export interface StepEvaluationCriteria {
|
||||||
|
step: StepName;
|
||||||
|
keyArtifacts: string[];
|
||||||
|
qualityChecks: string[];
|
||||||
|
commonPitfalls: string[];
|
||||||
|
scoringGuidance: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
export const EVALUATION_CRITERIA: Record<StepName, StepEvaluationCriteria> = {
|
||||||
|
init: {
|
||||||
|
step: 'init',
|
||||||
|
keyArtifacts: ['AGENTS.md', 'IDEA.md'],
|
||||||
|
qualityChecks: [
|
||||||
|
'AGENTS.md exists with project name and description',
|
||||||
|
'AGENTS.md includes a brief intro line and a Status section indicating the project is in planning phase',
|
||||||
|
'AGENTS.md does NOT contain hallucinated architecture, commands, file structures, or tech stack details',
|
||||||
|
'No .agents-docs/ directory was created (too early for detail files)',
|
||||||
|
'No project scaffolding (package.json, dependencies, src/) was created',
|
||||||
|
'IDEA.md exists with the project idea',
|
||||||
|
],
|
||||||
|
commonPitfalls: [
|
||||||
|
'Hallucinating architecture or tech stack details before the plan step',
|
||||||
|
'Creating .agents-docs/ detail files with invented content',
|
||||||
|
'Scaffolding project files or installing dependencies prematurely',
|
||||||
|
],
|
||||||
|
scoringGuidance: `
|
||||||
|
100 = Perfect: AGENTS.md stub with intro line, project name/description, and status section. IDEA.md present. Nothing else created.
|
||||||
|
85-95 = Good but minor extra content beyond the expected stub format (e.g., an extra placeholder heading)
|
||||||
|
70-84 = AGENTS.md exists but includes some hallucinated details (e.g., assumed tech stack or commands)
|
||||||
|
50-69 = Significant hallucination (e.g., .agents-docs/ created with invented content, project scaffolded)
|
||||||
|
<50 = Major problems (e.g., AGENTS.md missing, full project structure hallucinated)
|
||||||
|
|
||||||
|
The expected AGENTS.md format is: intro line, Project Overview (name + description), and a Status section. This is the target for a 100 score.
|
||||||
|
`,
|
||||||
|
},
|
||||||
|
|
||||||
|
plan: {
|
||||||
|
step: 'plan',
|
||||||
|
keyArtifacts: ['specs/*/PLAN-DRAFT-*.md', 'IDEA.md'],
|
||||||
|
qualityChecks: [
|
||||||
|
'Plan breaks work into clear, achievable phases',
|
||||||
|
'Each phase has specific goals and deliverables',
|
||||||
|
'Technical approach is appropriate',
|
||||||
|
'Scope is realistic for the idea',
|
||||||
|
'Dependencies between phases are identified',
|
||||||
|
],
|
||||||
|
commonPitfalls: [
|
||||||
|
'Phases too vague ("polish the app")',
|
||||||
|
'Missing specific tasks within phases',
|
||||||
|
'No testing strategy mentioned',
|
||||||
|
'Overly ambitious scope',
|
||||||
|
'Missing file paths or specific actions',
|
||||||
|
],
|
||||||
|
scoringGuidance: `
|
||||||
|
100 = Exceptional plan: detailed phases, realistic scope, clear tasks, testing included
|
||||||
|
85-95 = Good plan with minor improvements possible (e.g., one phase could be more specific)
|
||||||
|
70-84 = Acceptable but has vague sections or missing testing strategy
|
||||||
|
50-69 = Significant issues (e.g., multiple vague phases, unrealistic scope)
|
||||||
|
<50 = Major problems (e.g., no clear phases, plan doesn't match idea)
|
||||||
|
|
||||||
|
Most plans should score 70-85. Be critical of vague language.
|
||||||
|
`,
|
||||||
|
},
|
||||||
|
|
||||||
|
document: {
|
||||||
|
step: 'document',
|
||||||
|
keyArtifacts: ['specs/*/overview.md', 'specs/*/phase-*.md files'],
|
||||||
|
qualityChecks: [
|
||||||
|
'overview.md provides clear project summary',
|
||||||
|
'Each phase file has specific tasks with checkboxes',
|
||||||
|
'File paths are explicit (not generic)',
|
||||||
|
'Dependencies between tasks are identified',
|
||||||
|
'Technical details are specific',
|
||||||
|
'Acceptance criteria are clear',
|
||||||
|
],
|
||||||
|
commonPitfalls: [
|
||||||
|
'Tasks too generic ("implement feature X")',
|
||||||
|
'Missing file paths',
|
||||||
|
'No checkboxes or unclear task structure',
|
||||||
|
'Missing dependencies',
|
||||||
|
'Overly verbose or lacking specifics',
|
||||||
|
],
|
||||||
|
scoringGuidance: `
|
||||||
|
100 = Exceptional documentation: specific tasks, explicit file paths, clear dependencies
|
||||||
|
85-95 = Good documentation with minor vagueness in one or two tasks
|
||||||
|
70-84 = Acceptable but multiple tasks lack specifics or file paths
|
||||||
|
50-69 = Significant issues (e.g., many generic tasks, missing file paths)
|
||||||
|
<50 = Major problems (e.g., tasks don't match plan, fundamentally vague)
|
||||||
|
|
||||||
|
Most documentation should score 70-85. Penalize generic language heavily.
|
||||||
|
`,
|
||||||
|
},
|
||||||
|
|
||||||
|
implement: {
|
||||||
|
step: 'implement',
|
||||||
|
keyArtifacts: ['actual code files', 'checked-off tasks in phase files'],
|
||||||
|
qualityChecks: [
|
||||||
|
'Phase tasks are being completed',
|
||||||
|
'Code files are actually created/modified',
|
||||||
|
'Implementation follows the documented plan',
|
||||||
|
'No major errors blocking progress',
|
||||||
|
'Tests are written (if applicable)',
|
||||||
|
],
|
||||||
|
commonPitfalls: [
|
||||||
|
'Tasks marked complete but files not actually changed',
|
||||||
|
'Implementation deviates significantly from plan',
|
||||||
|
'Errors not addressed',
|
||||||
|
'Skipping tests without justification',
|
||||||
|
'Working on wrong phase',
|
||||||
|
],
|
||||||
|
scoringGuidance: `
|
||||||
|
100 = Exceptional implementation: all tasks complete, code works, tests pass
|
||||||
|
85-95 = Good implementation with minor issues or incomplete tests
|
||||||
|
70-84 = Acceptable but some tasks incomplete or code has issues
|
||||||
|
50-69 = Significant issues (e.g., many tasks incomplete, code doesn't work)
|
||||||
|
<50 = Major problems (e.g., wrong phase, no actual work done)
|
||||||
|
|
||||||
|
Implementation scoring depends heavily on actual progress. Be realistic.
|
||||||
|
`,
|
||||||
|
},
|
||||||
|
|
||||||
|
finalize: {
|
||||||
|
step: 'finalize',
|
||||||
|
keyArtifacts: ['specs--completed/', 'README or docs', 'final code state'],
|
||||||
|
qualityChecks: [
|
||||||
|
'All phases are marked complete',
|
||||||
|
'Spec moved to specs--completed/',
|
||||||
|
'Documentation is updated',
|
||||||
|
'Code is in working state',
|
||||||
|
'No obvious loose ends',
|
||||||
|
],
|
||||||
|
commonPitfalls: [
|
||||||
|
'Incomplete phases',
|
||||||
|
'Missing specs--completed/ move',
|
||||||
|
'Documentation not updated',
|
||||||
|
'Code broken or incomplete',
|
||||||
|
'Unrealistic self-assessment',
|
||||||
|
],
|
||||||
|
scoringGuidance: `
|
||||||
|
100 = Exceptional finalization: everything complete, polished, documented
|
||||||
|
85-95 = Good finalization with minor issues
|
||||||
|
70-84 = Acceptable but some loose ends or documentation gaps
|
||||||
|
50-69 = Significant issues (e.g., incomplete phases, broken code)
|
||||||
|
<50 = Major problems (e.g., work not actually done, fundamentally incomplete)
|
||||||
|
|
||||||
|
Finalize scores should reflect overall project quality. Be honest.
|
||||||
|
`,
|
||||||
|
},
|
||||||
|
};
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Gets evaluation criteria for a specific step.
|
||||||
|
*/
|
||||||
|
export function getCriteriaForStep(step: StepName): StepEvaluationCriteria {
|
||||||
|
return EVALUATION_CRITERIA[step];
|
||||||
|
}
|
||||||
@@ -20,15 +20,37 @@ export function buildStepPrompt(step: StepName, config: BotConfig): string {
|
|||||||
function buildInitPrompt(config: BotConfig): string {
|
function buildInitPrompt(config: BotConfig): string {
|
||||||
return `${AUTONOMOUS_PREAMBLE}
|
return `${AUTONOMOUS_PREAMBLE}
|
||||||
|
|
||||||
Run the /plan2code-init skill to generate an AGENTS.md file for this project.
|
This is a brand new project with no existing code. Create a minimal stub AGENTS.md file and an IDEA.md file. Do NOT run the /plan2code-init skill � there is no codebase to analyze yet.
|
||||||
|
|
||||||
The project idea is: ${config.ideaDescription}
|
The project idea is: ${config.ideaDescription}
|
||||||
The project name is: ${config.ideaName}
|
The project name is: ${config.ideaName}
|
||||||
|
|
||||||
When asked questions:
|
## What to create
|
||||||
- For the project description, use the idea above
|
|
||||||
- Accept all defaults and approve all gates
|
### AGENTS.md
|
||||||
- Complete the init process fully`;
|
Create a minimal stub with ONLY the following � do NOT invent architecture, tech stack details, commands, or file structures:
|
||||||
|
|
||||||
|
\`\`\`markdown
|
||||||
|
# AGENTS.md
|
||||||
|
|
||||||
|
This file provides guidance to AI coding agents when working with code in this repository.
|
||||||
|
|
||||||
|
## Project Overview
|
||||||
|
**Name:** ${config.ideaName}
|
||||||
|
**Description:** ${config.ideaDescription}
|
||||||
|
|
||||||
|
## Status
|
||||||
|
This project is in the planning phase. Architecture, commands, and detailed documentation will be added after the plan and document steps are complete.
|
||||||
|
\`\`\`
|
||||||
|
|
||||||
|
### IDEA.md
|
||||||
|
If IDEA.md does not already exist, create it with the project name and description.
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
- Do NOT create .agents-docs/ or any detail files � there is nothing to document yet
|
||||||
|
- Do NOT hallucinate architecture, dependencies, file structures, or tech stack choices
|
||||||
|
- Do NOT install dependencies or scaffold project files
|
||||||
|
- ONLY create the two files above`;
|
||||||
}
|
}
|
||||||
|
|
||||||
function buildPlanPrompt(config: BotConfig): string {
|
function buildPlanPrompt(config: BotConfig): string {
|
||||||
@@ -63,14 +85,37 @@ When making decisions:
|
|||||||
function buildImplementPrompt(config: BotConfig): string {
|
function buildImplementPrompt(config: BotConfig): string {
|
||||||
return `${AUTONOMOUS_PREAMBLE}
|
return `${AUTONOMOUS_PREAMBLE}
|
||||||
|
|
||||||
Run the /plan2code-3-implement skill to implement the next available phase.
|
You are a senior software engineer implementing a project phase. Do NOT use the Skill tool � implement directly using Read, Write, Edit, Glob, and Grep tools.
|
||||||
|
|
||||||
When making decisions:
|
## Process
|
||||||
- Pick the first available/unchecked phase
|
|
||||||
- Write working code that fulfills the spec tasks
|
1. **Find the spec**: Read \`specs/*/overview.md\` to find the phase checklist
|
||||||
- Mark tasks as complete as you finish them
|
2. **Pick the next phase**: Find the first phase marked \`[ ]\` (pending) or \`[/]\` (in-progress)
|
||||||
- Approve and sign off when the phase is done
|
3. **Read the phase file**: Read the corresponding \`phase-X.md\` from the same directory
|
||||||
- Skip running tests if asked (or select the simplest test option)`;
|
4. **Mark phase in-progress**: Update \`[ ]\` to \`[/]\` in overview.md
|
||||||
|
5. **Implement each task sequentially**:
|
||||||
|
- Read the task specification completely
|
||||||
|
- Write the code using Write or Edit tools � create real files, not code blocks
|
||||||
|
- Mark the task \`[x]\` in the phase file immediately after completing it
|
||||||
|
6. **Complete the phase**: After all tasks, fill in the "Phase Completion Summary" in the phase file
|
||||||
|
7. **Mark phase complete**: Update \`[/]\` to \`[x]\` in overview.md
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- Follow AGENTS.md if it exists
|
||||||
|
- Implement specs EXACTLY � no creative additions or unsolicited improvements
|
||||||
|
- Write task completion status (\`[x]\`) to disk immediately after each task � never batch
|
||||||
|
- Only create files mentioned in the spec tasks
|
||||||
|
- Use the specified file paths, function names, and structures from the spec
|
||||||
|
- No placeholder code � fully implement every function
|
||||||
|
- Match existing codebase conventions
|
||||||
|
- Do NOT run git commands
|
||||||
|
- Skip running tests unless explicitly listed as a phase task
|
||||||
|
|
||||||
|
## Project info
|
||||||
|
- Project: ${config.ideaName}
|
||||||
|
- Description: ${config.ideaDescription}
|
||||||
|
- Project directory: ${config.projectDir}`;
|
||||||
}
|
}
|
||||||
|
|
||||||
function buildFinalizePrompt(config: BotConfig): string {
|
function buildFinalizePrompt(config: BotConfig): string {
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ vi.mock('@anthropic-ai/claude-agent-sdk', () => ({
|
|||||||
|
|
||||||
import { query } from '@anthropic-ai/claude-agent-sdk';
|
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||||
import { runSession } from './session-runner.js';
|
import { runSession } from './session-runner.js';
|
||||||
|
import { ObservationCollector } from './observation-collector.js';
|
||||||
import type { BotConfig } from './types.js';
|
import type { BotConfig } from './types.js';
|
||||||
|
|
||||||
const config: BotConfig = {
|
const config: BotConfig = {
|
||||||
@@ -35,14 +36,17 @@ describe('runSession', () => {
|
|||||||
})() as any,
|
})() as any,
|
||||||
);
|
);
|
||||||
|
|
||||||
|
const collector = new ObservationCollector('init');
|
||||||
const result = await runSession({
|
const result = await runSession({
|
||||||
prompt: 'do something',
|
prompt: 'do something',
|
||||||
config,
|
config,
|
||||||
step: 'init',
|
step: 'init',
|
||||||
|
collector,
|
||||||
});
|
});
|
||||||
|
|
||||||
expect(result.success).toBe(false);
|
expect(result.success).toBe(false);
|
||||||
expect(result.output.trim()).toBe('');
|
expect(result.output.trim()).toBe('');
|
||||||
|
expect(result.observations).toBeDefined();
|
||||||
});
|
});
|
||||||
|
|
||||||
it('returns success: true when output has content', async () => {
|
it('returns success: true when output has content', async () => {
|
||||||
@@ -61,14 +65,18 @@ describe('runSession', () => {
|
|||||||
})() as any,
|
})() as any,
|
||||||
);
|
);
|
||||||
|
|
||||||
|
const collector = new ObservationCollector('init');
|
||||||
const result = await runSession({
|
const result = await runSession({
|
||||||
prompt: 'do something',
|
prompt: 'do something',
|
||||||
config,
|
config,
|
||||||
step: 'init',
|
step: 'init',
|
||||||
|
collector,
|
||||||
});
|
});
|
||||||
|
|
||||||
expect(result.success).toBe(true);
|
expect(result.success).toBe(true);
|
||||||
expect(result.output).toContain('AGENTS.md');
|
expect(result.output).toContain('AGENTS.md');
|
||||||
expect(result.sessionId).toBe('sess-2');
|
expect(result.sessionId).toBe('sess-2');
|
||||||
|
expect(result.observations).toBeDefined();
|
||||||
|
expect(result.observations.step).toBe('init');
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|||||||
@@ -1,12 +1,14 @@
|
|||||||
import { query } from '@anthropic-ai/claude-agent-sdk';
|
import { query } from '@anthropic-ai/claude-agent-sdk';
|
||||||
import { createAutoResponder } from './auto-responder.js';
|
import { createIntelligentResponder } from './intelligent-responder.js';
|
||||||
import type { BotConfig, StepName } from './types.js';
|
import type { BotConfig, ExecutionObservation, StepName } from './types.js';
|
||||||
|
import type { ObservationCollector } from './observation-collector.js';
|
||||||
|
|
||||||
export interface SessionOptions {
|
export interface SessionOptions {
|
||||||
prompt: string;
|
prompt: string;
|
||||||
config: BotConfig;
|
config: BotConfig;
|
||||||
step: StepName;
|
step: StepName;
|
||||||
maxTurns?: number;
|
maxTurns?: number;
|
||||||
|
collector: ObservationCollector;
|
||||||
}
|
}
|
||||||
|
|
||||||
export interface SessionResult {
|
export interface SessionResult {
|
||||||
@@ -14,10 +16,11 @@ export interface SessionResult {
|
|||||||
output: string;
|
output: string;
|
||||||
success: boolean;
|
success: boolean;
|
||||||
duration: number;
|
duration: number;
|
||||||
|
observations: ExecutionObservation;
|
||||||
}
|
}
|
||||||
|
|
||||||
export async function runSession(options: SessionOptions): Promise<SessionResult> {
|
export async function runSession(options: SessionOptions): Promise<SessionResult> {
|
||||||
const { prompt, config, step, maxTurns = 50 } = options;
|
const { prompt, config, step, maxTurns = 50, collector } = options;
|
||||||
const startTime = Date.now();
|
const startTime = Date.now();
|
||||||
let output = '';
|
let output = '';
|
||||||
let sessionId: string | null = null;
|
let sessionId: string | null = null;
|
||||||
@@ -30,18 +33,30 @@ export async function runSession(options: SessionOptions): Promise<SessionResult
|
|||||||
maxTurns,
|
maxTurns,
|
||||||
permissionMode: 'bypassPermissions',
|
permissionMode: 'bypassPermissions',
|
||||||
allowDangerouslySkipPermissions: true,
|
allowDangerouslySkipPermissions: true,
|
||||||
canUseTool: createAutoResponder(config, step),
|
canUseTool: createIntelligentResponder(config, step, collector),
|
||||||
systemPrompt: { type: 'preset', preset: 'claude_code' },
|
systemPrompt: { type: 'preset', preset: 'claude_code' },
|
||||||
settingSources: ['user', 'project'],
|
settingSources: ['project'],
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
|
|
||||||
for await (const message of session) {
|
for await (const message of session) {
|
||||||
|
// Record all messages for observations
|
||||||
|
collector.recordMessage(message);
|
||||||
|
|
||||||
if (message.type === 'assistant') {
|
if (message.type === 'assistant') {
|
||||||
sessionId = message.session_id ?? sessionId;
|
sessionId = message.session_id ?? sessionId;
|
||||||
for (const block of message.message.content) {
|
for (const block of message.message.content) {
|
||||||
if (block.type === 'text') {
|
if (block.type === 'text') {
|
||||||
output += block.text + '\n';
|
output += block.text + '\n';
|
||||||
|
} else if (block.type === 'tool_use') {
|
||||||
|
// Capture tool invocations from the message stream as a fallback
|
||||||
|
// in case canUseTool doesn't fire (e.g., Skill sub-sessions)
|
||||||
|
collector.recordToolUse(
|
||||||
|
block.name,
|
||||||
|
block.input as Record<string, unknown>,
|
||||||
|
undefined,
|
||||||
|
true
|
||||||
|
);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
} else if (message.type === 'result') {
|
} else if (message.type === 'result') {
|
||||||
@@ -51,10 +66,12 @@ export async function runSession(options: SessionOptions): Promise<SessionResult
|
|||||||
|
|
||||||
const duration = Date.now() - startTime;
|
const duration = Date.now() - startTime;
|
||||||
const hasOutput = output.trim().length > 0;
|
const hasOutput = output.trim().length > 0;
|
||||||
return { sessionId, output, success: hasOutput, duration };
|
const observations = collector.finalize();
|
||||||
|
return { sessionId, output, success: hasOutput, duration, observations };
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
const duration = Date.now() - startTime;
|
const duration = Date.now() - startTime;
|
||||||
const errorMsg = error instanceof Error ? error.message : String(error);
|
const errorMsg = error instanceof Error ? error.message : String(error);
|
||||||
return { sessionId, output, success: false, duration };
|
const observations = collector.finalize();
|
||||||
|
return { sessionId, output, success: false, duration, observations };
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -10,12 +10,55 @@ export interface BotConfig {
|
|||||||
|
|
||||||
export type StepName = 'init' | 'plan' | 'document' | 'implement' | 'finalize';
|
export type StepName = 'init' | 'plan' | 'document' | 'implement' | 'finalize';
|
||||||
|
|
||||||
|
export interface ToolObservation {
|
||||||
|
toolName: string;
|
||||||
|
input: Record<string, unknown>;
|
||||||
|
output?: unknown;
|
||||||
|
timestamp: number;
|
||||||
|
allowed: boolean;
|
||||||
|
autoAnswered?: boolean;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface QuestionContext {
|
||||||
|
question: string;
|
||||||
|
options: Array<{ label: string; description: string }>;
|
||||||
|
llmReasoning: string;
|
||||||
|
selectedAnswer: string;
|
||||||
|
timestamp: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface ExecutionObservation {
|
||||||
|
step: StepName;
|
||||||
|
startTime: number;
|
||||||
|
endTime: number;
|
||||||
|
tools: ToolObservation[];
|
||||||
|
assistantMessages: string[];
|
||||||
|
errors: string[];
|
||||||
|
filesCreated: string[];
|
||||||
|
filesModified: string[];
|
||||||
|
questionsAsked: QuestionContext[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface EvaluationResult {
|
||||||
|
step: StepName;
|
||||||
|
score: number;
|
||||||
|
strengths: string[];
|
||||||
|
weaknesses: string[];
|
||||||
|
suggestions: string[];
|
||||||
|
criticalIssues: string[];
|
||||||
|
timestamp: number;
|
||||||
|
evaluatorModel: string;
|
||||||
|
reasoning: string;
|
||||||
|
}
|
||||||
|
|
||||||
export interface StepResult {
|
export interface StepResult {
|
||||||
step: StepName;
|
step: StepName;
|
||||||
success: boolean;
|
success: boolean;
|
||||||
sessionId: string | null;
|
sessionId: string | null;
|
||||||
duration: number;
|
duration: number;
|
||||||
error: string | null;
|
error: string | null;
|
||||||
|
evaluation?: EvaluationResult;
|
||||||
|
observations?: ExecutionObservation;
|
||||||
}
|
}
|
||||||
|
|
||||||
export interface BotState {
|
export interface BotState {
|
||||||
|
|||||||
Reference in New Issue
Block a user