Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 11 additions & 7 deletions apps/docs/tasks.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -171,17 +171,21 @@ video. Roomote skips visual proof when it would not add useful evidence. Judge
the evidence against the stated result rather than requiring a recording for
every change.

When a [judgment model](/models#judgment-model) is configured and a task turn
ends with code changes, Roomote automatically holds the agent's closing report
against the evidence: what you asked for, the agent's own checklist, the diff,
and the shell commands it actually ran. It looks for a requested item or
When a [judgment model](/models#judgment-model) is configured, Roomote
automatically holds the agent's work against the evidence at the moments that
matter: before the agent reports to you, before it pushes or opens a pull
request, and when its turn ends with code changes. The evidence is what you
asked for, the agent's own checklist, the diff, and the shell commands it
actually ran. It looks for a requested item or
checklist item left undone without explanation, a claimed change the diff does
not contain, a claimed test or build result the recorded commands do not
support, code shipped with no validation and no reason, an interface change
with visual proof waved off, a plain defect in the changed lines, and leftovers
such as debug logging or a disabled test. If anything is found, the agent gets
one chance to fix it or explain before the task reports back. The check takes
about a second and does not replace pull request review. Without a judgment
such as debug logging or a disabled test. If anything is found, the report or
push is held once with the reasons, so the agent fixes or explains before a
person sees the claim or the code ships; a second attempt for the same work
always goes through. The check takes about a second and does not replace pull
request review. Without a judgment
model, the agent runs a slower review pass of its own before delivery instead.

Visual proof should use genuine application, authentication, database, and
Expand Down

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

14 changes: 13 additions & 1 deletion apps/worker/src/run-task/agent-home.ts
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,7 @@ import { SLACK_STOP_HOOK_SCRIPT } from './slack-stop-hook-script';
import { OPENCODE_SLACK_HOOKS_PLUGIN_SCRIPT } from './opencode-slack-hooks-plugin-script';
import { OPENCODE_CHATGPT_GATEWAY_PLUGIN_SCRIPT } from './opencode-chatgpt-gateway-plugin-script';
import { OPENCODE_TOOL_SAFETY_PLUGIN_SCRIPT } from './opencode-tool-safety-plugin-script';
import { OPENCODE_COMPLETION_GATE_PLUGIN_SCRIPT } from './opencode-completion-gate-plugin-script';
import { resolveOpenCodeModelSelection } from './opencode-model';
import {
getRepoLocalSkillInvocations,
Expand Down Expand Up @@ -187,6 +188,8 @@ const ROOMOTE_OPENCODE_CHATGPT_GATEWAY_PLUGIN_FILE_NAME =
'roomote-chatgpt-gateway.js';

const ROOMOTE_OPENCODE_TOOL_SAFETY_PLUGIN_FILE_NAME = 'roomote-tool-safety.js';
const ROOMOTE_OPENCODE_COMPLETION_GATE_PLUGIN_FILE_NAME =
'roomote-completion-gate.js';

const ROOMOTE_OPENCODE_IDENTITY_PLUGIN_FILE_NAME = 'roomote-identity.js';

Expand Down Expand Up @@ -967,6 +970,10 @@ function writeOpenCodeManagedFiles(openCodeConfigDir: string): void {
pluginsDir,
ROOMOTE_OPENCODE_IDENTITY_PLUGIN_FILE_NAME,
);
const completionGatePluginPath = path.join(
pluginsDir,
ROOMOTE_OPENCODE_COMPLETION_GATE_PLUGIN_FILE_NAME,
);
const silenceHookPath = path.join(
openCodeConfigDir,
ROOMOTE_OPENCODE_SLACK_SILENCE_HOOK_FILE_NAME,
Expand All @@ -989,6 +996,11 @@ function writeOpenCodeManagedFiles(openCodeConfigDir: string): void {
'utf8',
);
fs.writeFileSync(identityPluginPath, OPENCODE_IDENTITY_PLUGIN_SCRIPT, 'utf8');
fs.writeFileSync(
completionGatePluginPath,
OPENCODE_COMPLETION_GATE_PLUGIN_SCRIPT,
'utf8',
);
fs.writeFileSync(silenceHookPath, SLACK_SILENCE_HOOK_SCRIPT, 'utf8');
fs.writeFileSync(stopHookPath, SLACK_STOP_HOOK_SCRIPT, 'utf8');
fs.chmodSync(silenceHookPath, 0o755);
Expand Down Expand Up @@ -1416,7 +1428,7 @@ function createProofOnlyJudgeModelInstructions(): string {
'',
'When `R_VISION_MODEL` is configured, the judge runs on that vision model so it can open proof screenshots directly. Otherwise it falls back to the active coding model for the task.',
'',
`Delegate one focused pass to the \`${ROOMOTE_OPENCODE_JUDGE_AGENT_NAME}\` subagent with the Task tool only when a pre-delivery \`capture-visual-proof\` step for this shipped change kept screenshots or keyframes. When that step kept no images (a no-op, not-applicable, unnecessary, or blocked result), or the workflow required no proof step, do not spawn the judge. Whether the work matches the request is checked by the platform automatically when your turn ends: it compares the request, your closing report, and the diff, and sends you a follow-up only if something needs another look.`,
`Delegate one focused pass to the \`${ROOMOTE_OPENCODE_JUDGE_AGENT_NAME}\` subagent with the Task tool only when a pre-delivery \`capture-visual-proof\` step for this shipped change kept screenshots or keyframes. When that step kept no images (a no-op, not-applicable, unnecessary, or blocked result), or the workflow required no proof step, do not spawn the judge. Whether the work matches the request is checked by the platform automatically: before you report to a person, before you push or open a pull request, and when your turn ends. It compares the request, your report, the commands you ran, and the diff. If something needs another look, the tool call fails once with the reasons (fix or explain, then call it again) or you get a follow-up after the turn.`,
'',
'Include in the judge brief: the plan or requested outcome, the validation results, the proof report verbatim, the path `/tmp/capture-visual-proof/diff-at-start.patch`, and the local paths of every kept screenshot and keyframe so the judge can open them.',
'',
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,90 @@
/**
* OpenCode plugin: before a tool call that reports to a person or ships the
* work, ask the sandbox server to run the completion check. A denial fails
* the tool call with the reasons, which the agent reads as the tool result.
* Everything else, including any failure to reach the server, lets the call
* through. Written into the OpenCode plugins directory by agent-home.
*/
export const OPENCODE_COMPLETION_GATE_PLUGIN_SCRIPT = `const SANDBOX_SERVER_URL =
process.env.ROOMOTE_SANDBOX_SERVER_URL || 'http://127.0.0.1:4200';
const CHECK_TIMEOUT_MS = 20_000;
// A cheap pre-filter; the sandbox server makes the real classification.
const CANDIDATE_TOOL = /report_to_parent_session|send_chat_reply|send_chat_message|manage_source_control|^(bash|shell)$/;
const CANDIDATE_COMMAND = /\\bgit\\b[^|;&\\n]*\\bpush\\b|\\bgh\\s+pr\\b|\\bglab\\s+mr\\b/;

function isCandidate(tool, args) {
if (!CANDIDATE_TOOL.test(tool)) {
return false;
}

if (tool === 'bash' || tool === 'shell') {
return CANDIDATE_COMMAND.test(String(args?.command ?? ''));
}

return true;
}

async function checkCompletionBeforeTool(tool, args) {
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), CHECK_TIMEOUT_MS);

try {
const response = await fetch(
SANDBOX_SERVER_URL + '/trpc/commands.checkCompletionBeforeTool',
{
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: 'Bearer ' + process.env.ROOMOTE_CLOUD_TOKEN,
},
// superjson envelope, as the sandbox server's tRPC expects.
body: JSON.stringify({ json: { tool, args } }),
signal: controller.signal,
},
);

if (!response.ok) {
return { allowed: true };
}

const body = await response.json();
return body?.result?.data?.json ?? { allowed: true };
} finally {
clearTimeout(timer);
}
}

export const RoomoteOpenCodeCompletionGate = async () => ({
'tool.execute.before': async (input, context) => {
if (process.env.ROOMOTE_COMPLETION_GATE !== 'true') {
return;
}

const tool = typeof input?.tool === 'string' ? input.tool : '';
const args = context?.args ?? input?.args;

if (!isCandidate(tool, args)) {
return;
}

let decision;

try {
decision = await checkCompletionBeforeTool(tool, args);
} catch (error) {
process.stderr.write(
'WARN [CompletionGate] check before ' +
tool +
' failed; allowing the call: ' +
(error instanceof Error ? error.message : String(error)) +
'\\n',
);
return;
}

if (decision && decision.allowed === false && decision.reason) {
throw new Error(decision.reason);
}
},
});
`;
8 changes: 8 additions & 0 deletions apps/worker/src/sandbox-server/lib/harness.ts
Original file line number Diff line number Diff line change
Expand Up @@ -314,6 +314,14 @@ export interface Harness extends EventEmitter<HarnessEvents> {
* Returns an unsubscribe function.
*/
subscribeRuntimeOutput(listener: (event: AcpMessage) => void): () => void;
/**
* Run the completion check before a tool call that reports to a person or
* ships the work. Harnesses without the check allow every call.
*/
checkCompletionBeforeTool?(input: {
tool: string;
args?: unknown;
}): Promise<{ allowed: boolean; reason?: string }>;

/**
* Subscribe to persisted Roomote runtime envelope events.
Expand Down
Loading
Loading