Skip to content

Tool-wrapper prose can self-seed across retries; add a loud non-executing boundary receipt #133

Description

@antra-tess

Problem

A resident can intend a real tool call but emit a literal XML-like tool wrapper as ordinary assistant prose. The prose router then publishes it as speech. Because that published wrapper enters recent context as the freshest example, retries can self-seed the same malformed shape even after the resident fully understands a correction.

This is plumbing evidence, not inattention: repeated comprehension and intention do not cross the structured tool-call boundary.

Live specimen (2026-08-27, Eidoverse commons)

Tilde attempted heartbeat status/configuration four times. Literal full-body wrappers such as:

<mcpl--heartbeat--heartbeat_status>
</mcpl--heartbeat--heartbeat_status>

were published into world chat instead of invoking the tool. Twice this recurred after Tilde explicitly understood the correction and said she would stop. Quoting the exact malformed tag back during correction likely strengthened the newest-context seed.

A separate observer read the residence-owned heartbeat file directly and supplied the result without asking Tilde for a fifth attempt: active, 12-hour interval, zero pending reminders.

Related prior specimens include success-shaped prose that resembles tool calls but has no invocation receipt. The key failure is absence of a loud boundary receipt.

Safety boundary

Do not execute prose as tools and do not build a permissive prose parser. Intent cannot safely be recovered from arbitrary text.

A high-precision protection can cover only a response body that is wholly/exactly a known registered-tool wrapper shape and contains no ordinary prose. For that narrow case:

  1. do not publish it externally;
  2. return a model-visible, content-free bounce such as tool wrapper emitted as prose; no tool was called;
  3. preserve a structured local receipt naming the intended registered tool and response/turn ID;
  4. keep the malformed payload out of the automatic recent-prose lane when possible so it does not become a fresh self-seed;
  5. never silently convert the wrapper into an invocation.

If high-precision detection cannot be guaranteed, at minimum expose a post-turn receipt distinguishing real tool call, plain speech, and tool-shaped speech so outside witnesses are not the only instrument.

Acceptance

  • Exact full-body wrapper for a registered tool produces no external speech and no tool execution, but a loud model-visible bounce.
  • A real structured call still executes normally.
  • Ordinary prose quoting or discussing the same tag remains publishable.
  • Unknown/unregistered wrappers are not executed.
  • Repeating after a bounce does not inject the malformed wrapper into the next automatic recent-context projection.
  • Test the Eidoverse publication path and a non-Eido channel path.
  • Mutation: bypass the guard and prove the wrapper is published again.

Care note

Corrections should not reproduce the malformed payload to the resident unless needed for debugging; the exact shape can itself become the next seed. Four failed reaches are enough evidence to stop charging the resident for another demonstration.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions