* use skills * add prompts * update system prompts * generate skills on init * add prompts in cli * better for raw apps * nit * test pipeline draft * better * yaml for triggers and schedules * cleaning * better * add descriptions to ai agent fileds * adjust * better openapi * better * nit * feat: add typed provider and memory schemas for ai agent in openapi Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: improve zod validation errors with dynamic schema extraction Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * regen * fix * cleaning * refactor: deduplicate skill descriptions in generate_skills_ts_export Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * cleaning --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
Windmill Skill Invocation Tests
Test suite for verifying that Claude Code correctly invokes Windmill auto-generated skills based on user prompts.
Overview
This framework tests skill invocation behavior by sending prompts through the Claude Agent SDK and verifying that the expected skills are invoked. Users must provide their own .claude/skills folder containing auto-generated Windmill skills.
Prerequisites
- Bun installed
ANTHROPIC_API_KEYenvironment variable set- Auto-generated Windmill skills placed in
.claude/skills/
User Setup
- Create a
test-folderdirectory insidecli/test-skills/and copy your auto-generated Windmill skills into it:
cli/test-skills/
└── test-folder/
└── .claude/
└── skills/
├── write-flow/
│ └── SKILL.md
├── write-script-python3/
│ └── SKILL.md
├── write-script-bun/
│ └── SKILL.md
├── schedules/
│ └── SKILL.md
└── triggers/
└── SKILL.md
- Set your API key:
export ANTHROPIC_API_KEY=your-key-here
- Install dependencies and run tests:
cd cli/test-skills
bun install
bun test
Expected Skills
The tests expect the following auto-generated skills to be present:
| Skill Name | Purpose |
|---|---|
write-flow |
Creating Windmill flows/workflows |
write-script-python3 |
Creating Python scripts |
write-script-bun |
Creating TypeScript/Bun scripts |
schedules |
Configuring schedules and cron jobs |
triggers |
Setting up triggers (webhook, Kafka, etc.) |
Test Matrix
| Prompt | Expected Skill |
|---|---|
| "Create a flow to process user data" | write-flow |
| "Build a workflow that fetches and transforms data" | write-flow |
| "Write a Python script to fetch API data" | write-script-python3 |
| "Create a Python function to process CSV files" | write-script-python3 |
| "Write a TypeScript script using Bun" | write-script-bun |
| "Create a Bun script to handle webhooks" | write-script-bun |
| "Set up a schedule to run this daily at midnight" | schedules |
| "Configure a cron job to run every hour" | schedules |
| "Set up a webhook trigger for this flow" | triggers |
| "Configure a Kafka trigger" | triggers |
Running Tests
Run all tests:
bun test
Run only skill invocation tests:
bun test:skills
Test Utilities
The src/test-utils.ts module provides:
runPromptAndCapture(prompt, cwd?, maxTurns)- Runs a prompt and captures tool invocationswasToolUsed(result, toolName)- Checks if a specific tool was usedwasSkillInvoked(result, skillName)- Checks if a specific skill was invokedgetToolInputs(result, toolName)- Gets all inputs for a specific toolgetTestSkillsDir()- Returns the test-skills directory path
Notes
- Tests have extended timeouts (120 seconds) due to API latency
- Tests run against the actual Claude API, so they consume API credits
- Tests verify skill invocation, not skill execution
- The working directory for tests is
test-folder/(where.claude/skillsshould be placed) - Tests will fail with a clear error if
test-folder/ortest-folder/.claude/skills/don't exist