fix(agent-pod): force judge to actually call tools instead of narrating
Observed live: phase-judge (on homelab-reasoning) wrote a full page of 'I should check X, then Y' reasoning, declared VERDICT: PASS, and showed the touch command as a fenced code block in its own text -- never ran git diff, never wrote the result file, never touched the sentinel. Coordinator timed out waiting on a file that was never going to appear.
This commit is contained in:
@@ -107,6 +107,22 @@ data:
|
|||||||
"Read and follow ~/.pi/agent/skills/karpathy-guidelines/SKILL.md and " +
|
"Read and follow ~/.pi/agent/skills/karpathy-guidelines/SKILL.md and " +
|
||||||
"~/.pi/agent/skills/caveman/SKILL.md as hard rules for this entire task, before anything else. ";
|
"~/.pi/agent/skills/caveman/SKILL.md as hard rules for this entire task, before anything else. ";
|
||||||
|
|
||||||
|
// judge (routed to homelab-reasoning) has been observed narrating an
|
||||||
|
// entire review in prose -- "I should run git diff, then check X..." --
|
||||||
|
// and then writing a verdict based on that narration without ever calling
|
||||||
|
// a real tool. Live example: a phase-judge call produced a page of
|
||||||
|
// "I would check..." reasoning, declared VERDICT: PASS, and showed the
|
||||||
|
// touch command as a fenced code block IN ITS OWN TEXT rather than
|
||||||
|
// executing it. Coordinator just timed out waiting on a sentinel that was
|
||||||
|
// never going to appear, since nothing was ever actually run. Spelled out
|
||||||
|
// explicitly since "use the judge skill" alone apparently isn't enough to
|
||||||
|
// rule this out.
|
||||||
|
const REQUIRE_REAL_TOOL_CALLS =
|
||||||
|
"Do not narrate what you would check -- actually run the commands via a real tool call and read their real " +
|
||||||
|
"output before writing anything. A verdict based on describing checks instead of executing them is invalid. " +
|
||||||
|
"Writing the verdict file and touching the sentinel are themselves tool calls you must execute, not text to " +
|
||||||
|
"display in your response. ";
|
||||||
|
|
||||||
function parseVerdictLine(text, label) {
|
function parseVerdictLine(text, label) {
|
||||||
if (!text) return null;
|
if (!text) return null;
|
||||||
const re = new RegExp(`${label}:\\s*(\\w+)`, "i");
|
const re = new RegExp(`${label}:\\s*(\\w+)`, "i");
|
||||||
@@ -331,9 +347,9 @@ data:
|
|||||||
|
|
||||||
function judgePrompt(task, baseBranch, attempt, cwd) {
|
function judgePrompt(task, baseBranch, attempt, cwd) {
|
||||||
return (
|
return (
|
||||||
`Use the judge skill to review the diff against ${baseBranch}...HEAD for task ${task}. Write your verdict to ` +
|
`${REQUIRE_REAL_TOOL_CALLS}Use the judge skill to review the diff against ${baseBranch}...HEAD for task ${task}. ` +
|
||||||
`${path.join(cwd, `.task-result-${task}`)} as a single "VERDICT: PASS" or "VERDICT: FAIL" line plus one line of rationale, then ` +
|
`Write your verdict to ${path.join(cwd, `.task-result-${task}`)} as a single "VERDICT: PASS" or "VERDICT: FAIL" ` +
|
||||||
`run: touch ${path.join(cwd, `.stage-done-${task}-judge-${attempt}`)}`
|
`line plus one line of rationale, then run: touch ${path.join(cwd, `.stage-done-${task}-judge-${attempt}`)}`
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -694,7 +710,7 @@ data:
|
|||||||
cwd,
|
cwd,
|
||||||
repoId,
|
repoId,
|
||||||
"judge",
|
"judge",
|
||||||
`Use the judge skill to review the full phase diff for phase ${phaseLabel} against ` +
|
`${REQUIRE_REAL_TOOL_CALLS}Use the judge skill to review the full phase diff for phase ${phaseLabel} against ` +
|
||||||
`${baseBranch}...HEAD (covers every task in this phase: ${[...phaseTaskIds].join(", ")}). Every ` +
|
`${baseBranch}...HEAD (covers every task in this phase: ${[...phaseTaskIds].join(", ")}). Every ` +
|
||||||
`individual task already passed its own judge review -- your job here is different: confirm the ` +
|
`individual task already passed its own judge review -- your job here is different: confirm the ` +
|
||||||
`tasks integrate correctly as one coherent narrative, and that real integration tests (not just ` +
|
`tasks integrate correctly as one coherent narrative, and that real integration tests (not just ` +
|
||||||
|
|||||||
Reference in New Issue
Block a user