Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
80 changes: 80 additions & 0 deletions docs/roadmap/06-think-passthrough.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# Think Block Passthrough

## Branch: `feature/think-passthrough`

## Problem

The OpenAI generator (openaiContentGenerator.ts:537-593) has a `filterThinkTags()`
method that **strips** `<think>...</think>` blocks from streaming output. For
non-streaming, it does `content.replace(/<think>[\s\S]*?<\/think>/g, '')` at
line 1355.

This means reasoning traces from models like DeepSeek-R1, QwQ, Qwen3-coder,
and others are silently discarded. Users running these models via LM Studio
can't see the chain-of-thought that makes them useful.

The Anthropic generator doesn't handle think blocks at all.

## What OpenClaude does

Wraps `thinking` blocks in `<thinking>` XML tags and passes them as text content.
Simple, but it at least surfaces the reasoning.

## Better approach for Delta

Rather than stripping or flattening, treat think blocks as a first-class part
type. This pairs naturally with the `DeltaPart` types from branch 5, but can
be implemented independently with the current type system.

## Plan

### Phase 1: Surface think blocks in streaming (OpenAI generator)

- Replace `filterThinkTags()` stripping logic with a **think block extractor**
- When a `<think>` tag opens: buffer content into a separate think accumulator
- When `</think>` closes: yield the buffered content as a think part
- Keep the stateful cross-chunk handling (it's well-implemented, just pointed
at the wrong outcome)

For the current `@google/genai` type system, represent think parts as:
```typescript
{ text: '', thought: true } // extend Part with metadata
```
Or use the `customMetadata` escape hatch on Part if available.

### Phase 2: Surface think blocks in non-streaming (OpenAI generator)

- Replace the regex strip at line 1355 with extraction
- Parse out all `<think>...</think>` blocks, emit as separate parts

### Phase 3: Handle Anthropic extended thinking

When/if we adopt the Anthropic SDK (branch 1), Claude's `thinking` content
blocks are returned natively — no XML parsing needed. Map them through.

### Phase 4: UI rendering

- `packages/cli/src/ui/` — render think blocks in a collapsible/dimmed style
- Show them by default (reasoning is the point), but respect a config flag
`showThinking: true|false` for users who want clean output
- Consider a `/thinking` toggle command

### Phase 5: Anthropic generator

- Handle `thinking` type content blocks in the streaming SSE parser
- Map to the same think part representation

## Files to modify
- `packages/core/src/core/openaiContentGenerator.ts` — replace filterThinkTags,
modify convertStreamChunkToGenAIFormat and convertToGenAIFormat
- `packages/core/src/core/anthropicContentGenerator.ts` — handle thinking blocks
- `packages/cli/src/ui/` — render think parts (identify exact component)
- `packages/cli/src/config/settingsSchema.ts` — add `showThinking` setting

## Files to create
- `packages/core/src/utils/thinkBlockParser.ts` — shared parser for XML think tags
- `packages/core/src/utils/thinkBlockParser.test.ts`

## Dependencies
- Independent of other branches, but aligns naturally with branch 5 (delta types
would include `{ type: 'thinking'; text: string }` as a first-class DeltaPart)
20 changes: 18 additions & 2 deletions packages/core/src/core/anthropicContentGenerator.ts
Original file line number Diff line number Diff line change
Expand Up @@ -213,6 +213,16 @@ export class AnthropicContentGenerator implements ContentGenerator {
FinishReason.FINISH_REASON_UNSPECIFIED,
responseId,
);
} else if (delta?.type === 'thinking_delta' && 'thinking' in delta) {
// Extended thinking chunk — surface as thought part
const thinkText = (delta as { thinking?: string }).thinking;
if (thinkText) {
yield this.makeStreamResponse(
[{ thought: true, text: thinkText }],
FinishReason.FINISH_REASON_UNSPECIFIED,
responseId,
);
}
} else if (delta?.type === 'input_json_delta' && 'partial_json' in delta) {
// Tool call argument fragment — append to the accumulator.
// We don't yield yet because the JSON is incomplete.
Expand Down Expand Up @@ -539,8 +549,14 @@ export class AnthropicContentGenerator implements ContentGenerator {
},
});
}
// thinking blocks are intentionally not surfaced yet — branch 6
// (feature/think-passthrough) will handle that
else if (block.type === 'thinking' && 'thinking' in block) {
// Anthropic's extended thinking — surface as a thought part so the
// Turn system renders it as DeltaEventType.Thought
const thinkText = (block as { thinking?: string }).thinking;
if (thinkText) {
parts.push({ thought: true, text: thinkText });
}
}
}

response.responseId = message.id;
Expand Down
36 changes: 23 additions & 13 deletions packages/core/src/core/openaiContentGenerator.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3504,8 +3504,8 @@ describe('OpenAIContentGenerator', () => {
});
});

describe('think tag filtering', () => {
it('should filter complete <think> blocks from non-streaming responses', async () => {
describe('think tag handling', () => {
it('should extract <think> blocks as thought parts from non-streaming responses', async () => {
const mockResponse = {
id: 'test-id',
object: 'chat.completion',
Expand Down Expand Up @@ -3533,11 +3533,13 @@ describe('OpenAIContentGenerator', () => {
};

const response = await generator.generateContent(request, 'test-prompt');
const text = response.candidates?.[0]?.content?.parts?.[0];
expect(text).toEqual({ text: 'Here is my answer.' });
const parts = response.candidates?.[0]?.content?.parts;
// Think content becomes a thought part, visible text is separate
expect(parts?.[0]).toEqual({ thought: true, text: 'I need to think about this...' });
expect(parts?.[1]).toEqual({ text: 'Here is my answer.' });
});

it('should filter <think> blocks from streaming responses', async () => {
it('should extract <think> blocks as thought parts from streaming responses', async () => {
const chunks = [
{
id: 'test-id',
Expand Down Expand Up @@ -3597,14 +3599,19 @@ describe('OpenAIContentGenerator', () => {
responses.push(response);
}

// Collect all text parts
const allText = responses
.flatMap((r) => r.candidates?.[0]?.content?.parts || [])
.filter((p) => 'text' in p && p.text)
// Collect all parts
const allParts = responses.flatMap((r) => r.candidates?.[0]?.content?.parts || []);

// Visible text should not include think content
const visibleText = allParts
.filter((p) => 'text' in p && p.text && !('thought' in p && p.thought))
.map((p) => ('text' in p ? p.text : ''))
.join('');
expect(visibleText).toBe('Hello world');

expect(allText).toBe('Hello world');
// Think content should be surfaced as thought parts
const thoughtParts = allParts.filter((p) => 'thought' in p && p.thought);
expect(thoughtParts.length).toBeGreaterThan(0);
});

it('should pass through text with no think tags unchanged', async () => {
Expand Down Expand Up @@ -3638,7 +3645,7 @@ describe('OpenAIContentGenerator', () => {
expect(text).toEqual({ text: 'Just a normal response with no tags.' });
});

it('should filter multiple <think> blocks from non-streaming responses', async () => {
it('should extract multiple <think> blocks as a combined thought part', async () => {
const mockResponse = {
id: 'test-id',
object: 'chat.completion',
Expand Down Expand Up @@ -3666,8 +3673,11 @@ describe('OpenAIContentGenerator', () => {
};

const response = await generator.generateContent(request, 'test-prompt');
const text = response.candidates?.[0]?.content?.parts?.[0];
expect(text).toEqual({ text: 'First part. Second part.' });
const parts = response.candidates?.[0]?.content?.parts;
// Thought part comes first with combined think content
expect(parts?.[0]).toEqual({ thought: true, text: 'thought 1\n\nthought 2' });
// Visible text follows
expect(parts?.[1]).toEqual({ text: 'First part. Second part.' });
});
});
});
98 changes: 22 additions & 76 deletions packages/core/src/core/openaiContentGenerator.ts
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ import { Config } from '../config/config.js';
import { openaiLogger } from '../utils/openaiLogger.js';
import { safeJsonParse } from '../utils/safeJsonParse.js';
import { normalizeSchemaForProvider, type ProviderType } from '../tools/schemaNormalizer.js';
import { extractThinkBlocks, StreamingThinkExtractor } from '../utils/thinkBlockParser.js';

// Extended types to support cache_control
interface ChatCompletionContentPartTextWithCache
Expand Down Expand Up @@ -104,8 +105,7 @@ export class OpenAIContentGenerator implements ContentGenerator {
arguments: string;
}
> = new Map();
private insideThinkBlock: boolean = false;
private thinkTagBuffer: string = '';
private thinkExtractor: StreamingThinkExtractor = new StreamingThinkExtractor();

constructor(
contentGeneratorConfig: ContentGeneratorConfig,
Expand Down Expand Up @@ -551,72 +551,15 @@ export class OpenAIContentGenerator implements ContentGenerator {
}
}

/**
* Filter <think>...</think> tags from streaming text chunks.
* Handles tags that span multiple chunks using stateful buffering.
*/
private filterThinkTags(text: string): string {
let result = '';
let i = 0;

while (i < text.length) {
if (this.insideThinkBlock) {
// Look for closing </think> tag
const closeIdx = text.indexOf('</think>', i);
if (closeIdx !== -1) {
// Found closing tag, skip everything up to and including it
this.insideThinkBlock = false;
i = closeIdx + '</think>'.length;
} else {
// No closing tag in this chunk, check if a partial </think> is at the end
// Buffer up to 8 chars (length of "</think>") to handle split tags
const tailLen = Math.min(text.length - i, 8);
const tail = text.substring(text.length - tailLen);
if ('</think>'.startsWith(tail) && tail.length < '</think>'.length) {
this.thinkTagBuffer = tail;
}
// Skip all remaining text (we're inside a think block)
break;
}
} else {
// Look for opening <think> tag
const openIdx = text.indexOf('<think>', i);
if (openIdx !== -1) {
// Emit text before the tag
result += text.substring(i, openIdx);
this.insideThinkBlock = true;
i = openIdx + '<think>'.length;
} else {
// Check if a partial <think> tag might be at the end of the chunk
let partialMatch = false;
for (let len = Math.min(7, text.length - i); len >= 1; len--) {
const candidate = text.substring(text.length - len);
if ('<think>'.startsWith(candidate)) {
// Potential partial opening tag at end of chunk
result += text.substring(i, text.length - len);
this.thinkTagBuffer = candidate;
partialMatch = true;
break;
}
}
if (!partialMatch) {
result += text.substring(i);
}
break;
}
}
}

return result;
}
// filterThinkTags removed — replaced by StreamingThinkExtractor in thinkBlockParser.ts
// which extracts think content as thought parts instead of discarding it.

private async *streamGenerator(
stream: AsyncIterable<OpenAI.Chat.ChatCompletionChunk>,
): AsyncGenerator<GenerateContentResponse> {
// Reset the accumulator for each new stream
// Reset accumulators for each new stream
this.streamingToolCalls.clear();
this.insideThinkBlock = false;
this.thinkTagBuffer = '';
this.thinkExtractor.reset();

for await (const chunk of stream) {
yield this.convertStreamChunkToGenAIFormat(chunk);
Expand Down Expand Up @@ -1374,11 +1317,15 @@ export class OpenAIContentGenerator implements ContentGenerator {

const parts: Part[] = [];

// Handle text content (filter out <think>...</think> tags from local models)
// Extract <think> blocks from local models as thought parts instead of stripping them.
// The Turn system surfaces these as DeltaEventType.Thought events in the UI.
if (choice.message.content) {
const filtered = choice.message.content.replace(/<think>[\s\S]*?<\/think>/g, '');
if (filtered) {
parts.push({ text: filtered });
const { text: visibleText, thinkContent } = extractThinkBlocks(choice.message.content);
if (thinkContent) {
parts.push({ thought: true, text: thinkContent });
}
if (visibleText) {
parts.push({ text: visibleText });
}
}

Expand Down Expand Up @@ -1462,18 +1409,17 @@ export class OpenAIContentGenerator implements ContentGenerator {
if (choice) {
const parts: Part[] = [];

// Handle text content (filter out <think>...</think> tags from local models)
// Extract <think> blocks from streaming content instead of stripping them.
// Completed think blocks become { thought: true } parts that the Turn
// system surfaces as DeltaEventType.Thought events.
if (choice.delta?.content) {
if (typeof choice.delta.content === 'string') {
// Prepend any buffered partial tag from previous chunk
let textToFilter = choice.delta.content;
if (this.thinkTagBuffer) {
textToFilter = this.thinkTagBuffer + textToFilter;
this.thinkTagBuffer = '';
const { visibleText, completedThink } = this.thinkExtractor.process(choice.delta.content);
if (completedThink) {
parts.push({ thought: true, text: completedThink });
}
const filtered = this.filterThinkTags(textToFilter);
if (filtered) {
parts.push({ text: filtered });
if (visibleText) {
parts.push({ text: visibleText });
}
}
}
Expand Down
Loading
Loading