Skip to content

fix(workflow): exclude thoughts from typed node inputs - #7456

Open
BlueRaddish wants to merge 1 commit into
google:mainfrom
BlueRaddish:fix/oct8-workflow-answer-text
Open

BlueRaddish wants to merge 1 commit into
google:mainfrom
BlueRaddish:fix/oct8-workflow-answer-text

Conversation

@BlueRaddish

Copy link
Copy Markdown
Contributor

Link to Issue or Description of Change

Problem: A workflow node output can contain model thought parts followed by the final text or JSON answer. Typed Content conversion currently concatenates both. A downstream str function receives the thought text, while a BaseModel function fails validation even when the answer alone is valid JSON.

Reproduction: On main 7d8c9502f924ed5e7a1ab00dc1f5a603e858abc3, run a real Workflow through Runner with a producer returning Event(output=Content(parts=[Part(text="Let me think first.", thought=True), Part(text=answer)])), followed by a typed function. The string case receives Let me think first.hello; the structured case reports model_type validation failure before the function runs. No LLM is needed to reproduce the conversion defect.

Expected behavior: Typed values contain the answer text. Callers explicitly requesting types.Content still receive the full Content.

Solution: Reuse the existing thought-filtering extract_text_from_content helper in shared schema validation and FunctionNode argument conversion. Preserve existing warnings for discarded media parts and the handling of ordinary multipart text.

Testing Plan

Unit Tests:

  • Added regression coverage in the existing schema and FunctionNode test files.
  • Affected and adjacent test files pass locally: 176 passed on Python 3.13.14.
pytest tests/unittests/utils/test_schema_utils.py tests/unittests/utils/test_content_utils.py tests/unittests/workflow/test_function_node.py -q

The six regression cases fail before the fix and pass after it: string/structured schema conversion with both preserve_content settings, plus real workflow string and BaseModel inputs. Existing ordinary multipart-text, media warning, and schema conversion cases remain green. All native pre-commit hooks pass.

Manual End-to-End (E2E) Tests: The two new workflow cases execute real Workflow, FunctionNode, Runner, InMemorySessionService, and invariant checking through run_workflow. Before: thought text leaks into the string input and valid JSON input fails. After: the functions receive hello and _OutputModel(name='test', value=42), respectively. The setup is deterministic and performs no model or network calls.

Checklist

  • Read the contribution guide and performed a self-review.
  • Added tests that demonstrate the bug and exercise both affected conversion paths.
  • New and affected existing tests pass locally.
  • Exercised the change through the real workflow execution path.

Additional context

No public signatures, dependencies, or prompt behavior change. Full multi-version tox and the entire repository test suite were not run locally; the focused suite and native hooks above were run. Separate pending PRs #7430 (null preservation) and #6750 (JSON-mode serialization) touch schema utilities but do not address thought extraction; their actual diffs were checked. The earlier #7452 addresses debug printing and is independent of this workflow conversion fix.

Prepared with assistance from OpenAI Codex.

Use the existing answer-text helper for schema validation and function
argument coercion so thoughts do not invalidate JSON or enter string input.
Cover both routes through real workflow execution and focused schema tests.

Assisted-by: OpenAI Codex
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants