1.4 KiB
Failing Tests
This file tracks benchmark cases that still fail or need follow-up validation.
Flow
-
flow-test6-ai-agent-toolsLatest failing run:ai_evals/results/2026-04-09T11-25-24.107Z__flowIssues: final output does not include the actions or tool-result details the prompt asks foropen_support_ticketcontains a syntax bug -
flow-test7-simple-modificationLatest failing run:ai_evals/results/2026-04-09T11-25-24.107Z__flowIssues:validate_datawas added, but the failure behavior still does not match the requested contractsave_resultsthrows instead of returning a graceful structured result -
flow-test11-preprocessor-and-failure-handlerLatest failing run:ai_evals/results/2026-04-09T11-25-24.107Z__flowIssues: the model creates regularpreprocessorandfailuremodules it does not use Windmill's special top-levelpreprocessor_moduleandfailure_module
Needs Reconfirmation
flow-test4-order-processing-loopFull-suite failing run:ai_evals/results/2026-04-09T11-25-24.107Z__flowFollow-up passing run after prompt improvement:ai_evals/results/2026-04-09T13-29-15.877Z__flowNote: this case failed on invalidbranchonedownstream result access it passed after adding explicit branch-output guidance to the flow prompt rerun the full flow suite to confirm the fix holds in the broader benchmark