Skip to content

fix: replay OpenAI Responses reasoning from encrypted content - #63

Merged
ibetitsmike merged 9 commits into
coder_2_33from
mike/openai-encrypted-reasoning
Oct 11, 2026
Merged

ibetitsmike merged 9 commits into
coder_2_33from
mike/openai-encrypted-reasoning

Conversation

@ibetitsmike

@ibetitsmike ibetitsmike commented Oct 2, 2026 •

Copy link
Copy Markdown

OpenAI Responses reasoning was only replayed as an item_reference, which needs store=true. With the default store=false it was dropped from later requests, so reasoning models lost their chain of thought between tool steps and turns.

This captures the completed reasoning item (ID, encrypted_content, summary) from response.output_item.done and from the Generate output, marks that metadata Finalized, and replays it as a full reasoning input item regardless of store, before its function, computer, and web search items. Completed summaries are kept verbatim, so an empty summary stays []. Unfinalized metadata (stream placeholders and rows persisted before this field) keeps the old behavior: an item reference with store=true, skipped otherwise.

Finalized metadata also records whether the source response was stored: the requested store, unless the response echoes store: false (as it does for zero data retention organizations). A following web_search_call item reference is replayed only when both the source and the destination request store; otherwise it is omitted, because a reference to an unstored item returns 404, while the reasoning is still replayed inline.

The follow-up request bodies of the six summary-thinking cassettes were updated offline to include the replayed reasoning items; they were not re-recorded live.

Targets coder_2_33. The branch merges in coder_2_33: #64 and #65 at 88cffb5, then #66 at d93660c. The only conflict was the assistant text replay, which keeps #66's phase and this branch's web search reference flag. Squash-merged as 6ba8681, which coder/coder#30298 pins.

Testing

  • go test ./..., go build ./..., go vet ./..., golangci-lint, and targeted race tests for providers/openai pass.
  • responses_reasoning_replay_test.go covers Generate and Stream capture, replay under store true and false, unfinalized fallbacks, and web search reference eligibility from stored and unstored sources. Red-green verified by disabling each gate.
  • Live behavior was verified through remote UAT of fix: replay stateless OpenAI reasoning across chat turns coder#30298 with real gpt-5-mini and gpt-5.4 requests: round 3 PASS at f8a25c7.

Review record

Xum, an AI coding agent, implemented, tested, and opened this PR on behalf of @ibetitsmike.

Defer synthetic user media until all contiguous tool-result messages have
been emitted. OpenAI rejects media inserted between replies to parallel
tool calls. Preserve accompanying text and the order of media payloads.

Cover both Chat Completions serializers with separate and grouped tool
results, two media results, and following conversation messages.

> Xum acted on Mike's behalf.
Capture encrypted_content, item ID and summary from the completed
reasoning output item (stream output_item.done and Generate output) and
mark that metadata Finalized. Replay finalized metadata with a blob as a
full reasoning input item regardless of store, keeping its position
before function, computer and web_search items. Unfinalized metadata
(stream placeholders and rows persisted before this field) keeps the
previous behavior: item_reference with store=true, skipped otherwise.

Update the recorded follow-up request bodies of the summary-thinking
cassettes to include the replayed reasoning items.
Empty summaries remain empty arrays rather than invented summary text entries.
…nses

Record whether the source response was stored on finalized reasoning
metadata and only replay a following web_search_call item reference
when it was. Reasoning from an unstored response is still replayed
inline from its encrypted content.
@ibetitsmike

Copy link
Copy Markdown
Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-11T10:20:21.781841Z d93660c Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. What shall we delve into next?

Reviewed commit: f8a25c79b3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@mafredri mafredri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Capturing the reasoning item from output_item.done and replaying it inline regardless of store is the right shape, and gating inline replay on Finalized keeps the output_item.added placeholders already stored in Coder chats away from the API. One concern, one check, two suggestions.

Suggestions:

  • The six summary-thinking cassettes (Azure gpt-5-mini, OpenAI gpt-5, OpenAI o4-mini) were edited offline, and the UAT through Coder ran gpt-5-mini and gpt-5.4. Could we re-record them, so each edited follow-up request has been accepted by the real API?
  • charmbracelet#407 replays the same item fields inline, but only for store=false, and has no Finalized gate, so it would also replay the .added placeholder blobs in rows persisted before the fix. Could we propose Finalized there, so the fork does not have to carry it?

Didn't review the tests in depth.

🤖 This review was automatically generated with Coder Agents.

Comment thread providers/openai/responses_language_model.go
Comment thread providers/openai/responses_language_model.go Outdated
OpenAI echoes store=false for zero data retention organizations even when
the request asked to store, so a later web_search_call item reference to
that response would 404. Record SourceStoreEnabled from the requested
store unless the response (Generate) or response.created event (Stream)
echoes store=false. An absent echo keeps the requested value, and an echo
never upgrades a requested false.
@ibetitsmike

Copy link
Copy Markdown
Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Keep it up!

Reviewed commit: d55a86791a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@kevinle128

Copy link
Copy Markdown

Thanks for the review and the pointer, @mafredri. I agree the Finalized gate is the right shape. Without it, metadata persisted from output_item.added before the fix would be replayed with partial encrypted_content.

I added it to charmbracelet#407 with the same field name and JSON tag (finalized,omitempty) as this PR, so metadata written by either side stays compatible. Generate and response.output_item.done set it, and stateless replay now skips unfinalized metadata.

charmbracelet#407 still replays inline only when store=false. Changing behavior for store=true callers seems like a call for the upstream maintainers, so I'll raise it there and reference this PR. SourceStoreEnabled and the web search reference rules stay out of charmbracelet#407 for now.

I re-recorded the four OpenAI summary-thinking cassettes (gpt-5, o4-mini) against the live API: the follow-up requests with inline encrypted reasoning return 200. The two Azure cassettes are still edited offline, because I don't have Azure access.

@linear-code

linear-code Bot commented Oct 6, 2026

Copy link
Copy Markdown

CODAGT-1350

@ibetitsmike
ibetitsmike changed the base branch from mike/mcp-tool-media-batch-order to coder_2_33 October 9, 2026 15:38
ibetitsmike added a commit to coder/coder that referenced this pull request Oct 9, 2026
coder/fantasy#63 now targets coder_2_33 and merges in #64 and #65, so the
pin carries the merged base plus the reasoning replay changes.
Picks up #66 (Responses message phase). The only conflict was the
assistant text replay in toResponsesPromptWithValidation: keep #66's
phase replay and this branch's canReferenceWebSearch flag.
@ibetitsmike

Copy link
Copy Markdown
Author

@codex review

Xum requested this review on behalf of @ibetitsmike.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Bravo.

Reviewed commit: d93660ce7f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@ibetitsmike
ibetitsmike merged commit 6ba8681 into coder_2_33 Oct 11, 2026
5 of 6 checks passed
ibetitsmike added a commit to coder/coder that referenced this pull request Oct 11, 2026
coder/fantasy#63 is merged into coder_2_33, which now also carries #66,
so the pin moves from the PR branch to the coder_2_33 head.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants