Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ core ← db ← cli
- **@agentops/core** — Domain types, policy engine, scoring algorithm, and builder functions for all entities. No external dependencies. All types use `readonly` and branded ID types (RunId, JobId, SessionId, etc.) for type safety.
- **@agentops/db** — SQLite persistence via Drizzle ORM + better-sqlite3. Sixteen tables: `runs`, `policies`, `policy_results`, `run_metrics`, `jobs`, `sessions`, `events`, `locks`, `users`, `api_tokens`, `auth_sessions`, `device_codes`, `webhooks`, `webhook_deliveries`, `audit_log`, `user_budgets`. Complex fields stored as JSON columns. DB defaults to `~/.agentops/agentops.db` (override with `AGENTOPS_DB_PATH`). WAL mode with `foreign_keys = ON` — deletes of parent rows must cascade children first (see `deletePolicy`, `deleteOldRuns`).
- **@agentops/cli** — CLI entry point (`agentops`). Commands: `init`, `serve`, `setup`, `hook`, `login`, `doctor`, `user`, `admin`, `cleanup`, `run`, `policy`, `report`, `wrap`, `watch`, `link`, `pr`, `job`, `session`, `events`, `lock`, `dispatch`. Supports `--json` output and `--db-path` override. `init` bootstraps the DB (`--seed` for sample data, `--seed-policies` for the starter policy set, `--clean` to reset). `serve` starts the dashboard server (`--port` to override 3000). `setup` configures Claude Code hooks (`--global`, `--uninstall`, `--dry-run`). `hook` handles Claude Code hook events (session-start, pre-tool-use, post-tool-use, user-prompt-submit, stop, subagent-stop, session-end) — reads JSON from stdin, manages state under `~/.agentops/state/`, evaluates policies in real-time, can block risky tool calls (exit code 2), and reads cost/token usage from the Claude Code transcript (deduped by `message.id`; unknown models warn on stderr instead of pricing at $0). `login` runs the device-flow auth against a dashboard; `doctor` diagnoses a local install. Helper modules: `format.ts` (output formatting), `git.ts` (git integration), `github.ts` (GitHub API), `transcript.ts` (usage/cost from Claude Code transcripts), `pricing` lives in core. Note: `job`, `lock`, `dispatch`, `wrap`, and `watch` operate on orchestration machinery that Claude Code hooks never populate — treat them as experimental/vestigial.
- **@agentops/sdk** — Lightweight HTTP client for agent runtimes to talk to the AgentOps server. Depends only on `@agentops/core` for types. Uses native `fetch`. Provides `AgentOpsClient` class (via `createClient()` factory) with methods: `createSession`, `startRun`, `reportAction`, `reportArtifact`, `reportMetrics`, `checkPolicy`, `heartbeat`, `completeRun`, `failRun`. Also exports `PolicyMiddleware` for pre-flight policy checks before actions. Throws typed `AgentOpsError` with status codes.
- **@agentops/sdk** — Lightweight HTTP client for agent runtimes to talk to the AgentOps server. Depends only on `@agentops/core` for types. Uses native `fetch`. Provides `AgentOpsClient` class (via `createClient()` factory) with methods: `createSession`, `startRun`, `reportAction`, `reportArtifact`, `reportMetrics`, `checkPolicy`, `heartbeat`, `completeRun`, `failRun`, `terminateSession`. Also exports `PolicyMiddleware` for pre-flight policy checks before actions. Throws typed `AgentOpsError` with status codes.
- **@agentops/web** — Next.js 16 App Router dashboard with React 19, Tailwind CSS 4. API routes under `src/app/api/` organized by resource (runs, sessions, policies, events, analytics, admin, stats, sdk, auth, budgets, webhooks). Jobs, Locks, and Coordination pages were removed from the dashboard because hooks do not populate this data. Sidebar nav: Runs | Sessions | Events | Analytics | Usage | Policies | Settings. **Every API route is authenticated** (`src/lib/auth.ts`): bearer tokens or session cookies, `admin`/`member` roles, members are scoped to their own runs/sessions (`resolveViewScope`), non-owners get 404 (not 403) to avoid ID enumeration, mutations additionally pass `checkSameOrigin` CSRF checks — new routes must follow this pattern (regression tests live in `api/__tests__/auth-gaps.test.ts`). Inbound SDK routes under `/api/sdk/` use bearer auth + per-token rate limits. The Usage page (`/usage`) always shows the local hook-captured rollup (incl. Bedrock-vs-direct backend split and per-user attribution); when `ANTHROPIC_ADMIN_API_KEY` is set it additionally shows org-wide cost/tokens via `/api/admin/{status,cost,analytics}`, which normalize the Anthropic Admin API's reports (RFC 3339 `starting_at`, paginated, amounts are decimal-string cents) server-side. Server components access SQLite directly via a singleton lazy-loaded DB instance (`src/lib/db.ts`). Uses `serverExternalPackages: ["better-sqlite3"]` in next.config.ts. Path alias: `@/*` → `./src/*`. Dark theme by default.

### Core Domain Concepts
Expand Down
30 changes: 27 additions & 3 deletions packages/sdk/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,9 +20,33 @@ npm install @agentops/sdk

```ts
import { AgentOpsClient } from "@agentops/sdk";

const client = new AgentOpsClient({ baseUrl: "http://localhost:3000" });
await client.startSession({ agentId: "my-agent" });
import { createAgentId, AgentRole } from "@agentops/core";

const client = new AgentOpsClient({ baseUrl: "http://localhost:3000", apiKey: "ao_..." });

const { sessionId } = await client.createSession({ agentId: "my-agent" });
const { runId } = await client.startRun({
goal: { humanReadable: "Fix the bug", structured: { type: "bugfix", description: "", parameters: {} } },
agents: [{ id: createAgentId("my-agent"), model: "claude-opus-4-6", role: AgentRole.Implementer }],
environment: { repo: "acme/api", branch: "main", permissions: [], sandbox: { enabled: false, isolationLevel: "none" } },
sessionId,
});

// Metrics reports are partial updates — omitted fields keep their stored values.
await client.reportMetrics(runId, {
tokenUsage: { input: 100, output: 50, total: 150 },
costUsd: 0.05,
backend: "bedrock", // optional Bedrock/Anthropic spend attribution
});

// Test results reported at completion drive the correctness score and
// merge recommendation.
const { recommendation } = await client.completeRun(runId, {
testResults: [{ name: "unit", passed: true, duration: 12, message: "" }],
confidenceScore: 0.9,
});

await client.terminateSession(sessionId);
```

## License
Expand Down
68 changes: 67 additions & 1 deletion packages/sdk/src/__tests__/client.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ import {
createRunId,
createSessionId,
createAgentId,
createArtifactId,
AgentRole,
} from "@agentops/core";
import { AgentOpsClient, AgentOpsError, createClient } from "../client.js";
Expand Down Expand Up @@ -123,6 +124,13 @@ describe("AgentOpsClient", () => {

expect(result.runId).toBe(testRunId);
expect(result.status).toBe("running");

// The declared agents must actually go over the wire — the server
// persists them onto the run (request-side contract).
const sentBody = JSON.parse(mockFetch.mock.calls[0]![1].body as string);
expect(sentBody.agents).toEqual([
{ id: "agent-1", model: "claude-opus-4-6", role: "implementer" },
]);
});
});

Expand Down Expand Up @@ -175,6 +183,20 @@ describe("AgentOpsClient", () => {

expect(result.runId).toBe(testRunId);
});

it("sends backend and byModel for Bedrock spend attribution", async () => {
mockFetch.mockResolvedValueOnce(jsonResponse<ReportMetricsResponse>({ runId: testRunId }));

await client.reportMetrics(testRunId, {
costUsd: 2.5,
backend: "bedrock",
byModel: { "us.anthropic.claude-opus-4-7-v1:0": 2.5 },
});

const sentBody = JSON.parse(mockFetch.mock.calls[0]![1].body as string);
expect(sentBody.backend).toBe("bedrock");
expect(sentBody.byModel).toEqual({ "us.anthropic.claude-opus-4-7-v1:0": 2.5 });
});
});

describe("checkPolicy", () => {
Expand Down Expand Up @@ -248,7 +270,7 @@ describe("AgentOpsClient", () => {
expect(result.runId).toBe(testRunId);
});

it("sends result when provided", async () => {
it("sends result when provided (deprecated, still transmitted)", async () => {
mockFetch.mockResolvedValueOnce(
jsonResponse({ runId: testRunId } as unknown as CompleteRunResponse),
);
Expand All @@ -258,6 +280,50 @@ describe("AgentOpsClient", () => {
const sentBody = JSON.parse(mockFetch.mock.calls[0]![1].body as string);
expect(sentBody.result).toBe("Bug fixed successfully");
});

it("sends testResults / policyChecks / confidenceScore / artifacts", async () => {
mockFetch.mockResolvedValueOnce(
jsonResponse({ runId: testRunId } as unknown as CompleteRunResponse),
);

await client.completeRun(testRunId, {
testResults: [{ name: "unit", passed: false, duration: 5, message: "boom" }],
policyChecks: [],
confidenceScore: 0.8,
artifacts: [
{
id: createArtifactId("artifact_1"),
diffs: [],
logs: [],
testOutputs: ["1 failed"],
reports: [],
},
],
});

const sentBody = JSON.parse(mockFetch.mock.calls[0]![1].body as string);
expect(sentBody.testResults).toEqual([
{ name: "unit", passed: false, duration: 5, message: "boom" },
]);
expect(sentBody.policyChecks).toEqual([]);
expect(sentBody.confidenceScore).toBe(0.8);
expect(sentBody.artifacts).toHaveLength(1);
expect(sentBody.artifacts[0].id).toBe("artifact_1");
});
});

describe("terminateSession", () => {
it("posts to /api/sdk/sessions/:id/terminate and returns { status }", async () => {
mockFetch.mockResolvedValueOnce(jsonResponse({ status: "terminated" }));

const result = await client.terminateSession(testSessionId);

expect(result.status).toBe("terminated");
expect(mockFetch).toHaveBeenCalledWith(
`${BASE_URL}/api/sdk/sessions/${testSessionId}/terminate`,
expect.objectContaining({ method: "POST" }),
);
});
});

describe("failRun", () => {
Expand Down
15 changes: 15 additions & 0 deletions packages/sdk/src/client.ts
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@ import type {
CheckPolicyRequest,
CheckPolicyResponse,
HeartbeatResponse,
TerminateSessionResponse,
CompleteRunRequest,
CompleteRunResponse,
FailRunRequest,
Expand Down Expand Up @@ -158,6 +159,20 @@ export class AgentOpsClient {
);
}

/**
* Gracefully end a session: archives its current run and marks it
* terminated. Without this, SDK-created sessions only end when the
* server's staleness reaper gives up on their heartbeats.
*/
async terminateSession(
sessionId: SessionId,
): Promise<TerminateSessionResponse> {
return this.request<TerminateSessionResponse>(
`/api/sdk/sessions/${sessionId}/terminate`,
{},
);
}

async completeRun(
runId: RunId,
result?: CompleteRunRequest,
Expand Down
1 change: 1 addition & 0 deletions packages/sdk/src/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ export type {
CheckPolicyResponse,
PolicyViolation,
HeartbeatResponse,
TerminateSessionResponse,
CompleteRunRequest,
CompleteRunResponse,
FailRunRequest,
Expand Down
29 changes: 29 additions & 0 deletions packages/sdk/src/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ import type {
Metrics,
Agent,
Environment,
Evaluation,
Goal,
ScoreCard,
MergeRecommendation,
Expand Down Expand Up @@ -47,11 +48,18 @@ export interface ReportArtifactRequest {
readonly reports?: Artifact["reports"];
}

// Partial-update semantics: every field is optional, and the server MERGES
// each report into the run's stored metrics — an omitted field preserves the
// previously reported value (it is never reset to zero).
export interface ReportMetricsRequest {
readonly tokenUsage?: Metrics["tokenUsage"];
readonly wallTimeMs?: number;
readonly costUsd?: number;
readonly flakeRate?: number;
/** Which API backend served the traffic ("anthropic" | "bedrock"). */
readonly backend?: Metrics["backend"];
/** Per-model cost in USD, keyed by raw model id (Bedrock ids stay namespaced). */
readonly byModel?: Metrics["byModel"];
}

export interface CheckPolicyRequest {
Expand All @@ -62,8 +70,25 @@ export interface CheckPolicyRequest {
readonly editedFiles?: ReadonlyArray<string>;
}

// Mirrors what /api/sdk/runs/[id]/complete actually consumes: the evaluation
// (test results, policy checks, confidence) plus any final artifacts to
// append. Without testResults the scorer has nothing to grade — correctness
// reads "not scored" and mutating runs can reach Merge with no tests run.
export interface CompleteRunRequest {
/**
* @deprecated The server has never read this field; it is silently
* dropped. Report outcome data via testResults / policyChecks /
* confidenceScore instead. Kept only so existing callers keep compiling.
*/
readonly result?: string;
/** Test outcomes for the run — these drive the correctness score. */
readonly testResults?: Evaluation["testResults"];
/** Client-side policy check outcomes to record on the evaluation. */
readonly policyChecks?: Evaluation["policyChecks"];
/** Agent self-reported confidence, 0-1. Defaults to 0 server-side. */
readonly confidenceScore?: number;
/** Final artifacts to append to the run before scoring. */
readonly artifacts?: ReadonlyArray<Artifact>;
}

export interface FailRunRequest {
Expand Down Expand Up @@ -118,6 +143,10 @@ export interface HeartbeatResponse {
readonly commands: ReadonlyArray<unknown>;
}

export interface TerminateSessionResponse {
readonly status: string;
}

export interface CompleteRunResponse {
readonly runId: RunId;
readonly score: ScoreCard;
Expand Down
114 changes: 114 additions & 0 deletions packages/web/src/app/api/sdk/__tests__/sdk-contract.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,7 @@ import { POST as reportMetricsRoute } from "@/app/api/sdk/runs/[id]/metrics/rout
import { POST as completeRunRoute } from "@/app/api/sdk/runs/[id]/complete/route";
import { POST as failRunRoute } from "@/app/api/sdk/runs/[id]/fail/route";
import { POST as heartbeatRoute } from "@/app/api/sdk/sessions/[id]/heartbeat/route";
import { POST as terminateSessionRoute } from "@/app/api/sdk/sessions/[id]/terminate/route";
import { POST as policyCheckRoute } from "@/app/api/sdk/policy/check/route";

let db: AgentOpsDb;
Expand Down Expand Up @@ -177,6 +178,21 @@ describe("SDK route ⇄ type contract", () => {
expect(Array.isArray(body.commands)).toBe(true);
});

it("terminateSession → { status } (TerminateSessionResponse)", async () => {
const { insertSession } = await import("@agentops/db");
const { createSession, activateSession } = await import("@agentops/core");
const session = { ...activateSession(createSession("agent", {})), userId: alice.user.id };
insertSession(db, session);
const res = await terminateSessionRoute(
reqFor(`/api/sdk/sessions/${session.id}/terminate`, {}),
withParams({ id: session.id as string }),
);
expect(res.status).toBe(200);
const body = await bodyOf(res);
expect(Object.keys(body)).toEqual(["status"]);
expect(body.status).toBe("terminated");
});

it("policy check (allow) → { decision, violations, warnings } (CheckPolicyResponse)", async () => {
const runId = seedRun();
const res = await policyCheckRoute(reqFor("/api/sdk/policy/check", { runId, toolName: "Read", toolInput: {} }));
Expand Down Expand Up @@ -212,3 +228,101 @@ describe("SDK route ⇄ type contract", () => {
expect(Array.isArray(body.warnings)).toBe(true);
});
});

// Request-side contract: bodies shaped exactly like the SDK's REQUEST types
// must be consumed by the routes, not silently dropped. This is the drift the
// response-only tests above couldn't see (StartRunRequest.agents was ignored,
// CompleteRunRequest couldn't carry testResults, ReportMetricsRequest
// replaced instead of merged).

describe("SDK request type ⇄ route contract", () => {
it("StartRunRequest.agents is persisted onto the run", async () => {
// Body shaped like StartRunRequest, agents included.
const res = await createRunRoute(
reqFor("/api/sdk/runs", {
goal: { humanReadable: "t", structured: { type: "t", description: "t", parameters: {} } },
agents: [
{ id: "agent-1", model: "claude-opus-4-6", role: "implementer" },
{ id: "agent-2", model: "claude-haiku-4-5", role: "reviewer" },
],
environment: { repo: "acme/test", branch: "main", permissions: [], sandbox: { enabled: false, isolationLevel: "none" } },
}),
);
expect(res.status).toBe(201);
const { runId } = (await bodyOf(res)) as { runId: string };
const { getRun } = await import("@agentops/db");
const { createRunId } = await import("@agentops/core");
const saved = getRun(db, createRunId(runId))!;
expect(saved.agents).toEqual([
{ id: "agent-1", model: "claude-opus-4-6", role: "implementer" },
{ id: "agent-2", model: "claude-haiku-4-5", role: "reviewer" },
]);
});

it("CompleteRunRequest.testResults reach the stored run AND drive scoring", async () => {
const runId = seedRun();
// Body shaped like CompleteRunRequest (post-fix): failing test included.
const res = await completeRunRoute(
reqFor(`/api/sdk/runs/${runId}/complete`, {
testResults: [
{ name: "passes", passed: true, duration: 10, message: "" },
{ name: "fails", passed: false, duration: 12, message: "assertion failed" },
],
policyChecks: [],
confidenceScore: 0.7,
artifacts: [
{ id: "artifact_final", diffs: ["+x"], logs: [], testOutputs: ["1 failed"], reports: [] },
],
}),
withParams({ id: runId }),
);
expect(res.status).toBe(200);
const body = await bodyOf(res);

// Scoring saw the tests: correctness is 1/2, not the "not scored" 1.0
// that let mutating runs reach Merge with zero tests run.
const score = body.score as { correctness: { score: number; rationale: string } };
expect(score.correctness.score).toBe(0.5);
expect(score.correctness.rationale).toContain("1/2 tests passing");

// Round-trip: the evaluation + artifacts landed on the stored run.
const { getRun } = await import("@agentops/db");
const { createRunId } = await import("@agentops/core");
const saved = getRun(db, createRunId(runId))!;
expect(saved.evaluations).toHaveLength(1);
expect(saved.evaluations[0]!.testResults).toHaveLength(2);
expect(saved.evaluations[0]!.confidenceScore).toBe(0.7);
expect(saved.artifacts.map((a) => a.id as string)).toContain("artifact_final");
});

it("ReportMetricsRequest is a partial update: omitted fields keep stored values", async () => {
const runId = seedRun();
// First report: tokens only.
await reportMetricsRoute(
reqFor(`/api/sdk/runs/${runId}/metrics`, {
tokenUsage: { input: 100, output: 50, total: 150 },
}),
withParams({ id: runId }),
);
// Second report: cost + backend attribution only (ReportMetricsRequest
// now declares backend/byModel so SDK clients can set Bedrock spend).
const res = await reportMetricsRoute(
reqFor(`/api/sdk/runs/${runId}/metrics`, {
costUsd: 2.5,
backend: "bedrock",
byModel: { "us.anthropic.claude-opus-4-7-v1:0": 2.5 },
}),
withParams({ id: runId }),
);
expect(res.status).toBe(200);

const { getRun } = await import("@agentops/db");
const { createRunId } = await import("@agentops/core");
const saved = getRun(db, createRunId(runId))!;
// Tokens from report #1 survived report #2 (previously zero-filled).
expect(saved.metrics.tokenUsage).toEqual({ input: 100, output: 50, total: 150 });
expect(saved.metrics.costUsd).toBe(2.5);
expect(saved.metrics.backend).toBe("bedrock");
expect(saved.metrics.byModel).toEqual({ "us.anthropic.claude-opus-4-7-v1:0": 2.5 });
});
});
Loading
Loading