You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[opentelemetry-instrumentation-genai-langchain] Capture prompt template name and variables on chat - #730
Bug fix (non-breaking change which fixes an issue)
New feature (non-breaking change which adds functionality)
Breaking change (fix or feature that would cause existing functionality to not work as expected)
This change requires a documentation update
How has this been tested?
Please describe the tests that you ran to verify your changes. Provide
instructions so we can reproduce. List any relevant details for your test
configuration.
uv run tox -e py312-test-instrumentation-genai-langchain-- -q
uv run tox -e py312-test-instrumentation-genai-langchain-conformance -- -q
uv run --python 3.12 tox -e lint-instrumentation-genai-langchain
Checklist
See CONTRIBUTING.md
for the style guide, changelog guidance, and more.
Followed the style guidelines of this project
Changelog updated if the change requires an entry
Unit tests added
Documentation updated
rads-1996
changed the title
Capture prompt template name and variables on chat
[opentelemetry-instrumentation-genai-langchain] Capture prompt template name and variables on chat
Sep 16, 2026
Moderate findings: normalize prompt names, use callback-name and ID fallbacks, derive and safely serialize template variables, and scope prompt context correctly. Findings have 1 vote each except ID detection, which has 2 votes.
metadata is typed as dict[str, Any], so a non-string prompt_name is assigned directly to InferenceInvocation.prompt_name and can emit an invalid non-string gen_ai.prompt.name attribute. Normalize this metadata value to a string (or ignore it), matching the existing metadata-name handling in resolve_agent_name.
name=(metadata or {}).get("prompt_name") or template_type,
When LangChain supplies the run name through the callback name kwarg (including the serialized=None case already handled above), this remains None, so a PromptTemplate run is not recognized and its context never reaches the chat span. Use the callback name as a fallback, as resolve_agent_name does in operation_mapping.py:122-128.
This stores the entire chain input as prompt variables. LangChain prompt serialization exposes the declared input_variables and can carry partial_variables; extra chain-state keys are not template variables, while partials are rendered values. Emitting raw inputs therefore leaks unrelated state and omits partial values. Derive the mapping from the serialized template variables and merge the rendered partials before assigning it.
The prompt context is stored on the parent run but is never consumed after a chat reads it. LangChain RunnableSequence invokes sibling steps under that same parent, so a chain such as prompt | model1 | model2 leaves the first prompt context attached to the sequence and incorrectly adds those attributes to model2, which receives model1's output rather than the template. Consume the context for the associated chat or scope it to the next model run.
if parent_run_id is not None:
prompt_context = self._invocation_manager.get_prompt_context(
parent_run_id
)
ChatPromptTemplate accepts non-JSON values such as list[BaseMessage] for a MessagesPlaceholder. When content capture is enabled, InferenceInvocation serializes every non-string prompt variable with json.dumps; these message objects are not handled by the util encoder, so finalizing the chat can raise TypeError from telemetry. Normalize such values or make prompt-variable serialization failure-safe before passing arbitrary inputs through.
The new coverage only exercises the non-streaming invoke path. This instrumentation also emits chat spans for stream and astream, so add equivalent sync and async streaming coverage to verify that the prompt name and variables survive stream finalization.
response = (prompt | model).invoke(
{"question": "What's the weather like in Seattle?"},
config={"metadata": {"prompt_name": "weather_prompt"}},
)
The new integration coverage only exercises synchronous .invoke(), but ChatPromptTemplate | ChatOpenAI also supports asynchronous .ainvoke() through LangChain's async callback-manager path. Add a VCR-backed async case that asserts the same prompt attributes; otherwise this new propagation is unverified for a supported call variant.
def test_chat_openai_prompt_template(
span_exporter, tracer_provider, meter_provider, logger_provider, vcr
):
model = ChatOpenAI(model="gpt-4.1", max_tokens=100)
prompt = PromptTemplate.from_template(
"Answer this weather question briefly: {question}"
)
with instrument(
LangChainInstrumentor(),
tracer_provider=tracer_provider,
meter_provider=meter_provider,
logger_provider=logger_provider,
content_capture="SPAN_ONLY",
):
with vcr.use_cassette("test_chat_openai_prompt_template.yaml"):
response = (prompt | model).invoke(
The reason will be displayed to describe this comment to others. Learn more.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.
These values are arbitrary objects, but the shared prompt-variable serializer later calls gen_ai_json_dumps directly for every non-string value without a fallback. A normal non-JSON-serializable prompt input (for example a custom object or message list) can therefore raise during invocation.stop() and turn a successful chat into a callback failure; normalize unsupported values or make the shared serializer fail open, with a regression test for this path.
ChatPromptTemplate serializes optional placeholders (for example MessagesPlaceholder(optional=True)) under kwargs["optional_variables"], but this set only includes input_variables and partials. Values supplied for an optional placeholder are therefore dropped, so the corresponding gen_ai.prompt.variable.<name> is never emitted; include validated optional_variables when building declared_names and add a regression test.
declared_names = set(input_variables) | partial_names
for name in declared_names:
if name in inputs:
variables[name] = inputs[name]
metadata["prompt_name"] is untyped but is passed directly into the str | None prompt-name field. If a caller supplies a non-string value, the finished span receives a non-string gen_ai.prompt.name, violating the semantic-convention attribute type; stringify or validate this metadata value before storing it, as the other metadata-derived names do.
name=(metadata or {}).get("prompt_name")
or serialized_name
or template_type,
LangChain permits a single-variable PromptTemplate to be invoked with a scalar, and on_chain_start receives that original scalar before LangChain normalizes it to a mapping. With such a call, name in inputs either drops the variable or inputs[name] raises TypeError, so prompt variables are not captured for a supported invocation. Normalize a non-mapping input to the sole declared variable before this loop.
declared_names = set(input_variables) | partial_names
for name in declared_names:
if name in inputs:
variables[name] = inputs[name]
The context is stored on the prompt run's immediate parent, but this lookup only checks the model run's immediate parent. In a valid nested composition such as (prompt | transform) | model, the prompt's parent is the inner RunnableSequence while the model's parent is the outer sequence, so the new attributes are silently absent; propagate the context to the enclosing model run or use a hierarchy-aware association, with a nested-sequence regression test.
if parent_run_id is not None:
prompt_context = self._invocation_manager.get_prompt_context(
parent_run_id
)
gen_ai.prompt.name is the name that identifies the prompt template, but template_type is only the class name. For an ordinary PromptTemplate.from_template(...) without name or prompt_name, this emits the same PromptTemplate value for every template (and similarly for chat templates), rather than omitting the attribute when no name is available. Keep the type only for classification and pass None unless metadata or a distinct serialized name is present.
name=(metadata or {}).get("prompt_name")
or serialized_name
or template_type,
The recursive normalization here can raise RecursionError for a cyclic list or mapping before the serialization try below is reached. Since this helper is intended to omit non-serializable prompt values, such input currently aborts the callback/model invocation instead of being skipped, introducing a failure when prompt capture is enabled. Track visited containers or guard the whole normalization path and return _INVALID_PROMPT_VALUE for recursive values.
Document prompt telemetry attributes and capture gating
This adds user-visible gen_ai.prompt.name and gen_ai.prompt.variable.* telemetry, but the package README does not document the new attributes or that variable capture is gated by OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT. Since README.rst is the published package documentation, please document this behavior before merging.
The reason will be displayed to describe this comment to others. Learn more.
This test passes only because on_llm_end crashes. util-genai raises on prompt variables it can't serialize, and LangChain swallows the error. The span loses its completion attributes and metrics without any visible failure.
#821 fixes this in util-genai, so this PR should wait for it to land. Please also replace this test with one that uses the real LangChain API: for example, a prompt template with a non-serializable partial variable piped into a fake chat model. The test should check that the span finished with its attributes.
The broader gap is that langchain tests call the handler directly, so crashes like this go unnoticed. That's tracked in #820.
Hi @rads-1996 — just a friendly reminder that this pull request is waiting on you. The dashboard status comment has the open items and is kept current.
Replying is enough to hand it off — answer, explain why no change is needed, or ask a follow-up. The dashboard routes it onward once nothing on the list is waiting on you.
To hand it back for any other reason, including the dashboard getting this wrong, comment /dashboard route:reviewers.
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes #585
Type of change
Please delete options that are not relevant.
How has this been tested?
Please describe the tests that you ran to verify your changes. Provide
instructions so we can reproduce. List any relevant details for your test
configuration.
Checklist
See CONTRIBUTING.md
for the style guide, changelog guidance, and more.