Skip to content

feat: add schema-constrained structured outputs and recovery loop to Agent - #12990

Open
rautaditya2606 wants to merge 2 commits into
deepset-ai:mainfrom
rautaditya2606:feature/agent-response-schema
Open

rautaditya2606 wants to merge 2 commits into
deepset-ai:mainfrom
rautaditya2606:feature/agent-response-schema

Conversation

@rautaditya2606

Copy link
Copy Markdown
Contributor

Related Issues

Proposed Changes

Adds response_schema and max_schema_retries parameters to Agent, enabling schema-constrained output validation and automatic self-correction for terminal agent replies:

  • Provider-agnostic post-hoc validation: Validates final assistant text completions against a Pydantic BaseModel subclass or JSON Schema dictionary without coupling to provider-specific client parameters.
  • Automatic recovery loop: When the LLM generates invalid JSON or fails schema validation, the agent appends a targeted correction message (ChatMessage.from_user) containing the schema and error description, re-looping up to max_schema_retries times within the agent's overall max_agent_steps budget.
  • Fail-fast initialization: Validates JSON schema dictionaries at construction time via Draft202012Validator.check_schema, raising ValueError on malformed schemas before running.
  • Markdown fence handling: Pre-extracts JSON from markdown code blocks (json ... ) before validation to prevent unnecessary retries on formatted responses.
  • Terminal gating: Gated strictly on model_exit_reason == _EXIT_REASON_TEXT, ensuring schema validation never fires on unexecuted tool-call messages when no tools are present.
  • State & Pipeline typing: Registers "structured_output" in resolved_state_schema and output sockets (type: response_schema | None) while ensuring it is excluded from input sockets.
  • Serialization & clone: Added serialization/deserialization support for Pydantic type paths and JSON schema dictionaries in to_dict, from_dict, and clone.

How did you test it?

Added TestAgentStructuredOutput in test/components/agents/test_agent.py covering:

  • Init-time validation (type check, negative retries, reserved state key collision, and malformed schema check).
  • Clean Pydantic model and JSON Schema dictionary extraction.
  • Markdown-fenced JSON responses.
  • Recovery loop on invalid JSON string followed by valid response.
  • Recovery loop on schema mismatch followed by valid response.
  • Exceeded retry budget returning structured_output=None without unhandled exceptions.
  • max_agent_steps cutoff stopping recovery.
  • Gating check: tool call with empty tools does not trigger spurious schema validation.
  • run_async parity with sync run.
  • Backward compatibility when response_schema=None.
  • to_dict, from_dict, and clone serde roundtrips.

Quality checks run:

  • hatch run fmt (clean)
  • hatch run test:types haystack/components/agents/agent.py test/components/agents/test_agent.py (clean)
  • hatch -e test run pytest test/components/agents/test_agent.py (138 passed)

Notes for the reviewer

The validation and retry logic is completely isolated to haystack/components/agents/agent.py. When response_schema is omitted, the execution path and output dictionary remain byte-identical to existing behavior.

Note: when response_schema is a Pydantic model, structured_output holds a live BaseModel instance in agent state. This hasn't been tested against Agent's breakpoint/snapshot resume path — if that machinery serializes arbitrary state values mid-run, a Pydantic instance may not round-trip the same way plain dicts do. Dict-schema mode is unaffected. Flagging as untested rather than claiming full compatibility.

Checklist

@rautaditya2606
rautaditya2606 requested a review from a team as a code owner September 27, 2026 20:01
@rautaditya2606
rautaditya2606 requested review from anakin87 and a lite review from Copilot and removed request for a team September 27, 2026 20:01
@vercel

vercel Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

@rautaditya2606 is attempting to deploy a commit to the deepset Team on Vercel.

A member of the Team first needs to authorize it.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions github-actions Bot added topic:tests type:documentation Improvements on the docs labels Sep 27, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Coverage report

Click to see where and how coverage changed

FileStatementsMissingCoverageCoverage
(new stmts)
Lines missing
  haystack/components/agents
  agent.py 219, 227-231, 240-244, 1288
Project Total  

This report was generated by python-coverage-comment-action

@rautaditya2606
rautaditya2606 requested a lite review from Copilot September 27, 2026 20:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@anakin87 anakin87 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

See #12989 (comment)

No changes are needed from your side.

We'll follow up in the issue/here if we decide to implement this feature.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

topic:tests type:documentation Improvements on the docs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support schema-constrained structured outputs with automatic recovery in Agent

3 participants