Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
e43b82f
[Improve] Check finished coding turns against the request in about a …
mrubens Sep 21, 2026
430e0e3
Run the completion check only with a hosted judgment model and addres…
mrubens Sep 21, 2026
cbad542
Keep late follow-ups, mid-check requests, and untracked noise files f…
mrubens Sep 21, 2026
6eb4f0b
Hold a follow-up that lands mid-check until the closing turn completes
mrubens Sep 21, 2026
c225920
Hold follow-ups for the whole closing path, not just the check itself
mrubens Sep 21, 2026
c03eee3
Cover what the judge pass weighed: checklist, validation evidence, pr…
mrubens Sep 21, 2026
9995d78
Count earlier validation runs only while the code they covered is unc…
mrubens Sep 21, 2026
32e842b
Treat a validation run as stale once source is edited after it
mrubens Sep 21, 2026
4b28144
Count shell rewrites of source as edits that make an earlier validati…
mrubens Sep 21, 2026
e4607c0
Only file-rewriting forms of git checkout, restore, and reset make a …
mrubens Sep 21, 2026
2d1b787
Moving to an existing branch makes a validation run stale; creating a…
mrubens Sep 21, 2026
918fb46
Decide from checkout and switch arguments whether the working tree wa…
mrubens Sep 21, 2026
971b6bf
See through git global options when deciding whether a command replac…
mrubens Sep 21, 2026
cf725fb
Drop validation runs that predate the last source edit instead of ask…
mrubens Sep 21, 2026
6f15a40
Catch an unsupported validation claim reliably with its own measured …
mrubens Sep 21, 2026
5b21bae
Decide whether a validation run is stale from the code itself, not fr…
mrubens Sep 21, 2026
b13e475
Fingerprint every changed path, falling back to size and mtime past t…
mrubens Sep 21, 2026
91249ec
Keep parentheses in the code fingerprint
mrubens Sep 21, 2026
c0f73ae
Fingerprint files past the caps by their bytes rather than size and m…
mrubens Sep 21, 2026
8559bef
Stream large files into the fingerprint instead of buffering them
mrubens Sep 21, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

77 changes: 77 additions & 0 deletions apps/api/src/handlers/tasks/checkTaskCompletion.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
import type { Context } from 'hono';

import { evaluateTaskCompletionGate } from '@roomote/cloud-agents/server';
import { db, eq, taskRuns } from '@roomote/db/server';
import {
taskCompletionCheckRequestSchema,
type TaskCompletionCheckResponse,
} from '@roomote/types';

import type { Variables } from '../../types';
import type { McpAuth } from '../mcp/middleware';
import { isRunTokenContext } from '../mcp/proxy-utils';
import { logHandlerError } from '../utils';

/**
* The sandbox harness calls this when a coding turn ends with a changed diff.
* The sandbox supplies only what lives there (the diff and the agent's
* report); what was asked is read from the transcript server-side, and the
* judgment model key never leaves the API. Every non-verdict outcome is
* `skipped` so the harness completes the turn normally.
*/
export async function checkTaskCompletion(
c: Context<{ Variables: Variables & { mcpAuth: McpAuth } }>,
): Promise<Response> {
const auth = c.get('mcpAuth').authContext;

if (!isRunTokenContext(auth)) {
return c.json({ error: 'Completion checks require a task run token' }, 403);
}

const runId = Number(c.req.param('runId'));

if (!Number.isInteger(runId) || runId <= 0) {
return c.json({ error: 'Invalid task run id' }, 400);
}

if (auth.runId !== runId) {
return c.json(
{ error: 'Task run token does not match requested task run' },
403,
);
}

const parsed = taskCompletionCheckRequestSchema.safeParse(
await c.req.json().catch(() => null),
);

if (!parsed.success) {
return c.json(
{ error: 'Invalid completion check', issues: parsed.error.issues },
400,
);
}

try {
const taskRun = await db.query.taskRuns.findFirst({
columns: { taskId: true, actingUserId: true },
where: eq(taskRuns.id, runId),
});

if (!taskRun) {
return c.json({ error: 'Task run not found' }, 404);
}

const result: TaskCompletionCheckResponse =
await evaluateTaskCompletionGate({
taskId: taskRun.taskId,
userId: taskRun.actingUserId,
check: parsed.data,
});

return c.json(result);
} catch (error) {
logHandlerError('checkTaskCompletion', error);
return c.json({ error: 'Failed to check task completion' }, 500);
}
}
2 changes: 2 additions & 0 deletions apps/api/src/handlers/tasks/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ import { manageSourceControl } from './manageSourceControl';
import { updateTaskModelSelection } from './updateModelSelection';
import { listTaskModels } from './listModels';
import { saveTaskMemory } from './saveTaskMemory';
import { checkTaskCompletion } from './checkTaskCompletion';
import { recordAutomationResult } from './recordAutomationResult';
import { updatePersonalization } from './updatePersonalization';

Expand All @@ -43,4 +44,5 @@ tasksRouter.post('/:taskId/task_suggestions', submitTaskSuggestions);
tasksRouter.post('/:taskId/automation_result', recordAutomationResult);
tasksRouter.post('/:taskId/mcp_recommendations', submitMcpRecommendations);
tasksRouter.post('/runs/:runId/memory', saveTaskMemory);
tasksRouter.post('/runs/:runId/completion_check', checkTaskCompletion);
tasksRouter.post('/runs/:runId/personalization', updatePersonalization);
4 changes: 4 additions & 0 deletions apps/docs/models.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -344,6 +344,10 @@ When a judgment model is on, Roomote asks it first for:
- choosing the forum tag when Roomote opens a thread in a Discord forum channel
- whether a settled session or task turn holds something worth saving to
[Memory](/memory) that the agent did not record itself
- checking a finished coding turn against the evidence: the request, the
agent's checklist, the diff, and the commands it ran. This check has no
helper-model fallback; without a judgment model the agent runs its own slower
review pass instead. See [Tasks](/tasks)

Roomote acts on the judgment model only when it is confident. When it is
unsure, unavailable, or returns an error, Roomote keeps the behavior it has
Expand Down
13 changes: 13 additions & 0 deletions apps/docs/tasks.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -171,6 +171,19 @@ video. Roomote skips visual proof when it would not add useful evidence. Judge
the evidence against the stated result rather than requiring a recording for
every change.

When a [judgment model](/models#judgment-model) is configured and a task turn
ends with code changes, Roomote automatically holds the agent's closing report
against the evidence: what you asked for, the agent's own checklist, the diff,
and the shell commands it actually ran. It looks for a requested item or
checklist item left undone without explanation, a claimed change the diff does
not contain, a claimed test or build result the recorded commands do not
support, code shipped with no validation and no reason, an interface change
with visual proof waved off, a plain defect in the changed lines, and leftovers
such as debug logging or a disabled test. If anything is found, the agent gets
one chance to fix it or explain before the task reports back. The check takes
about a second and does not replace pull request review. Without a judgment
model, the agent runs a slower review pass of its own before delivery instead.

Visual proof should use genuine application, authentication, database, and
backend state when practical. When an artifact instead uses disclosed simulated
state to make a UI reachable, treat it as evidence only for the rendered
Expand Down
32 changes: 31 additions & 1 deletion apps/worker/src/run-task/agent-home.ts
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,7 @@ import {
resolveOpenRouterVariantModelAlias,
toBedrockMantleRuntimeModelId,
OPENCODE_ARCHITECT_AGENT,
TASK_COMPLETION_GATE_ENV_VAR,
TASK_MODEL_CONTEXT_WINDOWS_ENV_VAR_NAME,
TASK_MODEL_COSTS_ENV_VAR_NAME,
CREDENTIAL_EGRESS_METHODS,
Expand Down Expand Up @@ -1403,6 +1404,33 @@ function createJudgeModelInstructions(): string {
].join('\n');
}

/**
* With the turn-end completion check active, the platform already compares
* the request, the report, and the diff, so the judge is left with the one
* thing that check cannot do: open proof images.
*/
function createProofOnlyJudgeModelInstructions(): string {
return [
`A hidden OpenCode \`${ROOMOTE_OPENCODE_JUDGE_AGENT_NAME}\` subagent is configured for visual-proof checks only.`,
'',
'When `R_VISION_MODEL` is configured, the judge runs on that vision model so it can open proof screenshots directly. Otherwise it falls back to the active coding model for the task.',
'',
`Delegate one focused pass to the \`${ROOMOTE_OPENCODE_JUDGE_AGENT_NAME}\` subagent with the Task tool only when a pre-delivery \`capture-visual-proof\` step for this shipped change kept screenshots or keyframes. When that step kept no images (a no-op, not-applicable, unnecessary, or blocked result), or the workflow required no proof step, do not spawn the judge. Whether the work matches the request is checked by the platform automatically when your turn ends: it compares the request, your closing report, and the diff, and sends you a follow-up only if something needs another look.`,
'',
'Include in the judge brief: the plan or requested outcome, the validation results, the proof report verbatim, the path `/tmp/capture-visual-proof/diff-at-start.patch`, and the local paths of every kept screenshot and keyframe so the judge can open them.',
'',
'Treat the judge as a narrow proof check. Ask it to open the kept screenshot and keyframe images and verify them against the plan and shipped change, to treat weak, mismatched, or falsely claimed proof as a gap, and to report any undisclosed source drift between the proof snapshot and the shipped diff. Do not ask for an open-ended repo review.',
'',
'Do not spawn the judge subagent when the current task is itself a pull-request or workspace code review (`review-code`, PR review, or PR re-review). Those workflows are already the review pass and must produce findings directly.',
'',
'Keep judge tool use minimal and targeted. Prefer the supplied diff and proof evidence, and only read extra files to resolve a specific ambiguity or verify an obvious risk.',
'',
'Treat the judge response as review input for the parent workflow. If judge-driven fixes change repository files, re-run the `capture-visual-proof` step once for the updated shipped change, replace prior proof evidence with that latest result, then run one more focused judge pass against the refreshed diff and refreshed proof result before delivery. Keep orchestration, code changes, and final user-facing decisions in the parent agent.',
'',
"Do not paste the judge's full output into chat or any user-facing reply. The judge verdict is internal review material; surface at most a brief, parent-authored summary of the actionable outcome (what was fixed or what still needs attention), never the raw review dump.",
].join('\n');
}

/**
* PR review/re-review tasks already run as the code-reviewer on the review
* model. Configuring and instructing a nested judge pass would double the
Expand Down Expand Up @@ -2026,7 +2054,9 @@ export function generateOpenCodeConfig({
);
fs.writeFileSync(
judgeModelInstructionsPath,
createJudgeModelInstructions(),
runtimeEnv[TASK_COMPLETION_GATE_ENV_VAR] === 'true'
? createProofOnlyJudgeModelInstructions()
: createJudgeModelInstructions(),
'utf8',
);
instructions.push(judgeModelInstructionsPath);
Expand Down
Loading
Loading