Skip to content

Support MCP structuredContent and Artifacts in Tool Results #134

Description

@Asaf-prog

Support MCP structuredContent and Artifacts in Tool Results

Context

Follow-up to #133.

#133 fixes the model-facing text path for MCP tool results by correctly extracting text from MCP content blocks instead of stringifying the raw result.

That solves cases like:

{
  "content": [
    {
      "type": "text",
      "text": "Balance: $1,250"
    }
  ]
}

so the model receives:

Balance: $1,250

instead of a Python representation of the content block.

However, MCP tools may also return structured output in addition to text.

For example:

{
  "content": [
    {
      "type": "text",
      "text": "Found 2 invoices"
    }
  ],
  "structuredContent": {
    "invoices": [
      {
        "id": "INV-123",
        "amount": 500,
        "currency": "USD"
      },
      {
        "id": "INV-456",
        "amount": 800,
        "currency": "USD"
      }
    ]
  }
}

Today Extra normalizes the text content for the model, but the structured result is not preserved through the execution pipeline.


Problem

The current tool execution contract effectively reduces every successful tool result to:

str

Conceptually:

MCP result
   ↓
ToolMessage
   ↓
extract text
   ↓
str
   ↓
ToolExecutionManager
   ↓
model

This is sufficient for plain textual tools, but lossy for MCP tools that expose structured output.

For example:

{
  "content": [
    {
      "type": "text",
      "text": "Found 2 invoices"
    }
  ],
  "structuredContent": {
    "invoices": [...]
  }
}

becomes only:

Found 2 invoices

The actual invoice objects are discarded from Extra's runtime representation.

This limits several capabilities:

  • downstream agents cannot consume structured tool output directly;
  • orchestration code cannot reason over exact fields;
  • tool outputs cannot be reliably reused without re-parsing text;
  • evaluation and replay lose the original structured result;
  • UI components cannot render structured tool results;
  • future agent-to-agent composition has to fall back to text;
  • typed MCP tools lose much of the value of their output schema.

Goal

Preserve MCP tool results as structured runtime data while keeping the existing model-facing text behavior intact.

Extra should distinguish between:

model-facing text

and:

machine-readable structured result

instead of collapsing both into a single string.

Conceptually:

MCP Tool Result
      │
      ├── content
      │      ↓
      │   model-facing text
      │
      └── structuredContent / artifact
             ↓
        structured runtime result

Suggested Runtime Model

Instead of treating a successful tool result as only:

str

introduce a generic normalized result object.

Conceptually:

NormalizedToolResult(
    text="Found 2 invoices",
    structured={
        "invoices": [
            {"id": "INV-123", "amount": 500},
            {"id": "INV-456", "amount": 800},
        ]
    },
    artifact=None,
)

The exact naming/API should fit the existing runtime architecture.

The important point is that tool execution should preserve both:

text
structured data

independently.


MCP Example

Given an MCP tool result:

{
  "content": [
    {
      "type": "text",
      "text": "Found 2 invoices"
    }
  ],
  "structuredContent": {
    "invoices": [
      {
        "id": "INV-123",
        "amount": 500
      },
      {
        "id": "INV-456",
        "amount": 800
      }
    ]
  }
}

Extra should preserve something equivalent to:

NormalizedToolResult(
    text="Found 2 invoices",
    structured={
        "invoices": [
            {"id": "INV-123", "amount": 500},
            {"id": "INV-456", "amount": 800},
        ]
    },
)

The model may still initially receive:

Found 2 invoices

if that remains the current model contract.

But the structured value should remain available to the runtime.


Local Tool Compatibility

This should not become MCP-specific throughout the engine.

Local tools may also eventually return structured data.

For example:

def get_orders(...) -> dict:
    return {
        "orders": [...]
    }

The normalization boundary should therefore ideally be generic:

Provider result
      ↓
Tool result normalization
      ↓
NormalizedToolResult

rather than:

if MCP:
    special-case structuredContent

MCP-specific parsing belongs at the provider/adapter boundary.

The rest of Extra should consume a provider-agnostic result representation.


Relationship to ToolMessage

LangChain may expose MCP structured output through the tool result / ToolMessage artifact path.

The implementation should inspect the actual adapter contract rather than assuming all result data lives in:

ToolMessage.content

Conceptually:

LangChain ToolMessage
├── content
└── artifact

Extra should extract the MCP-specific structured result from the appropriate field and normalize it into Extra's own runtime representation.

The Extra domain/runtime layer should not depend on LangChain-specific result objects after normalization.


Persistence

The current ToolExecutionManager persists the final result as text for idempotent replay.

If structured results are introduced, replay must preserve the same logical result.

For example:

first execution
→ text + structuredContent

graph resumes
→ idempotency cache hit

expected:
→ same text + same structuredContent

It would be incorrect if:

first execution
→ structured output available

replayed execution
→ only text available

because execution behavior would depend on whether the call was executed live or restored from the ledger.

The persisted execution result should therefore evolve accordingly.


Hooks

The current transform_tool_result hook receives text.

We should explicitly decide how structured data interacts with hooks.

Possible direction:

ToolResultContext(
    result_text=...,
    structured_result=...,
)

or a normalized result object:

ToolResultContext(
    result=NormalizedToolResult(...)
)

The important invariant is that text transformations such as truncation must not accidentally destroy structured output unless that behavior is intentional.

For example:

truncate text to 5,000 chars

should not automatically imply:

discard structuredContent

Model Consumption

This issue should preserve structured content first.

It does not necessarily need to define the final long-term strategy for how every model consumes structured tool results.

Possible future approaches include:

Text serialization

Provide a deterministic JSON representation to the model:

{
  "invoices": [...]
}

Native structured tool messages

Use provider-supported structured tool result capabilities where available.

Selective exposure

Keep structured data in runtime state while sending only selected fields/text to the model.

The initial implementation should avoid locking Extra into one model-provider-specific strategy.


Artifacts

MCP / LangChain tool results may also carry artifacts in addition to text.

For example:

content
+
artifact

Artifacts may represent:

structuredContent
files
binary references
provider-specific metadata

Extra should preserve artifact metadata where it is meaningful and safe.

However, large binary payloads should not automatically be copied into:

  • model context;
  • logs;
  • traces;
  • execution persistence.

Prefer preserving references/metadata rather than blindly serializing arbitrary binary data.


Security and Data Handling

Structured tool output may contain sensitive business data.

Preserving it must not mean automatically logging it.

The existing runtime already avoids logging raw tool arguments/results in several paths.

The same principle should apply here.

Structured results must not be automatically emitted into:

logs
trace metadata
error messages

unless explicitly configured.

The normalized result should remain application/runtime data.


Failure Semantics

Malformed structured output should not corrupt an otherwise valid textual tool response.

For example:

valid text
+
invalid structuredContent

should have an explicitly defined policy.

Possible behavior:

preserve text
mark structured result unavailable
record controlled warning

rather than converting the entire successful tool execution into a failure.

If the MCP tool declares an output schema and the structured output violates it, validation behavior should be defined separately and explicitly.


Suggested Flow

MCP Server
    ↓
MCP result
├── content
└── structuredContent
    ↓
LangChain MCP adapter
    ↓
ToolMessage / artifact
    ↓
Extra provider normalization
    ↓
NormalizedToolResult
├── text
├── structured
└── artifact metadata
    ↓
Tool hooks
    ↓
Execution persistence
    ↓
Model / UI / orchestration

Scope for v1

  • Introduce a provider-agnostic structured tool-result representation.
  • Preserve model-facing text.
  • Preserve MCP structuredContent.
  • Preserve relevant artifact metadata.
  • Ensure the idempotency/execution ledger restores structured results.
  • Keep local string-returning tools fully backward compatible.
  • Expose structured output to runtime hooks.
  • Do not automatically log structured result values.
  • Add tests covering live execution and replay.

Tests

Text + structured content

Input:

{
  "content": [
    {
      "type": "text",
      "text": "Found 2 invoices"
    }
  ],
  "structuredContent": {
    "count": 2
  }
}

Expected:

text == "Found 2 invoices"
structured == {"count": 2}

Multiple text blocks + structured content

All text blocks remain normalized correctly while structured output is preserved unchanged.

Structured-only result

Define and test behavior when structured output exists but no useful text content exists.

For example:

{
  "structuredContent": {
    "balance": 1250
  }
}

Extra should not silently discard the result.

Local string tool

return "done"

continues to behave exactly as today.

Local structured tool

If supported by the normalization abstraction, verify structured Python output can also be preserved without MCP-specific code leaking into the runtime.

Idempotent replay

tool executed
→ structured result persisted

same tool call replayed
→ no provider call
→ identical normalized result restored

Hook transformation

A text transformation does not accidentally discard structured output.

Unsupported/non-text MCP block

Non-text content is handled deterministically and does not destroy accompanying structured data.


Acceptance Criteria

  • MCP structuredContent is no longer silently discarded.
  • Model-facing text behavior introduced in fix(engine): extract clean text from MCP tool results instead of stri… #133 remains unchanged.
  • The runtime no longer assumes every successful tool result is only a string.
  • Structured results are represented in a provider-agnostic Extra abstraction.
  • LangChain-specific result objects do not leak beyond the normalization boundary.
  • Structured results survive idempotent replay.
  • Hooks can access structured output without requiring text re-parsing.
  • Local string tools remain backward compatible.
  • Structured result values are not automatically logged.
  • Text and structured outputs can coexist independently.
  • Tests cover text-only, structured-only, mixed, replay, hooks, and local-tool compatibility.

Non-Goals

  • Full multimodal MCP support.
  • Rendering arbitrary artifacts in the widget.
  • Persisting large binary payloads.
  • Automatically injecting all structured output into every model context.
  • Defining a universal serialization strategy for every LLM provider.
  • Output-schema validation enforcement for all tools.

Those can build on top of the normalized result abstraction later.

Why This Matters

MCP tools are not only text generators.

A tool may return precise machine-readable data such as:

{
  "invoice_id": "INV-123",
  "amount": 500,
  "currency": "USD"
}

Reducing that result to:

"Found invoice INV-123"

throws away information the runtime already received.

Extra should preserve the original structured result so future layers can decide how to use it:

model
UI
agent orchestration
evaluation
replay
workflow logic

The core invariant should be:

A tool execution boundary may normalize provider-specific result formats, but it should not silently destroy structured information returned by the tool.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions