Repository navigation
Python: [Bug]: Distinct non-English memory topics silently share a file #9205
Description
Activity
- addedpythonUsage: [Issues, PRs], Target: PythonUsage: [Issues, PRs], Target: PythontriageUsage: [Issues], Target: All issues that still need to be triagedUsage: [Issues], Target: All issues that still need to be triaged
on Oct 8, 2026 - addedreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflowUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
on Oct 8, 2026 🤖 Automated triage reproduction notes (agent-authored — trust but verify)
Agent analysis
Repro:
python/packages/core/agent_framework/_harness/_memory.py::_slugify_topicat lines 141-143 andMemoryFileStore._topic_pathat lines 776-777 map distinct topics with identical ASCII projections to one file. Twowrite_memorycalls for旅行计划and饮食偏好merge both facts, anddelete_memory_topicfor the latter removes the former. Minimal repro: runtest_memory_context_provider_keps_distinct_non_ascii_topics_separateusing a temporaryMemoryFileStore.- Failing test:
python/packages/core/tests/core/test_harness_memory.py::test_memory_context_provider_keps_distinct_non_ascii_topics_separate - Files examined: python/packages/core/agent_framework/_harness/_memory.py, python/packages/core/tests/core/test_harness_memory.py, python/packages/core/pyproject.toml
- Tests run: test_memory_context_provider_keps_distinct_non_ascii_topics_separate (failed), test_memory_context_provider_tools_and_automation (passed)
- Reported version:
1.20.0 - Current version:
1.21.0
- Failing test:
- addedharness[Issues, PRs], Target: harness-level items[Issues, PRs], Target: harness-level itemsand removedtriageUsage: [Issues], Target: All issues that still need to be triagedUsage: [Issues], Target: All issues that still need to be triaged
on Oct 8, 2026 Reproduced on current main (declarative/core at c24ef0a). The collision is exactly where the triage notes point:
_slugify_topicstrips every non-[a-z0-9]character, so旅行计划and饮食偏好both fall through to thememory-topicfallback andcafé/cafèboth becomecaf. Store-level check with one owner/session:_slugify_topic("旅行计划") == _slugify_topic("饮食偏好") == "memory-topic" _slugify_topic("café") == _slugify_topic("cafè") == "caf" # two write_topic calls -> a single topics/memory-topic.md # get_topic("旅行计划") returns the second topic's memoriesSince the second write clobbers identity and
delete_memory_topicthen removes the shared file, this is silent data loss for any non-ASCII deployment.I'd like to take this. Plan: keep
_slugify_topic's readable output byte-identical for topics that are already safe (slug == normalized.lower(), so existing ASCII stems are untouched), and append a short sha256 digest of the normalized topic whenever the mapping discarded information, e.g.café->caf-4f8b…,旅行计划->memory-topic-9c21…. Same collision-resistance standard as the_storage_key_segmentdigest fallback already used for owner/source segments in this file. On top of that,get_topic/delete_topicfall back to the legacy lossy path when the new file is absent, andwrite_topicabsorbs a legacy-named file into the new name, so stores written by older versions stay readable and migrate on first rewrite instead of stranding. Regression tests: distinct CJK/Cyrillic/accented topics stay separate, deleting one leaves the other readable, legacy-name read-back, and slug idempotency throughMemoryTopicRecord/MemoryIndexEntryre-derivation.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNo status
Observed Behavior
Writing two distinct non-English topic names through
MemoryContextProvider'swrite_memorytool silently merges them into one file. For example,旅行计划(travel plans) and饮食偏好(food preferences) both resolve totopics/memory-topic.md. The second write returns the first topic's name with both facts in its memories. Deleting the second topic throughdelete_memory_topicthen removes the shared record, so the first topic can no longer be read.This occurs with one ordinary owner and session, using the public tools and real local Markdown files. It also reproduces with two Japanese topics, two Cyrillic topics, and
café/cafè. Ordinary distinct ASCII topics remain separate. A single non-English topic writes and reloads successfully, and writing the exact same topic twice correctly updates one record.Expected Behavior
Distinct supported human-readable topic names should not silently address the same record. Writing or deleting one of these topics should leave the other topic's facts intact. The public topic parameter is documented as a human-readable name and currently accepts these values without rejecting them; non-English memory selection is also explicitly supported by #7130.
Steps to Reproduce
MemoryContextProviderwith aMemoryFileStorefor one owner.write_memoryfor the two different topic names below.Minimal Reproduction
Error Messages and Stack Traces
No error is raised by either write. The ASCII pair produces two records followed by
first topic still exists. The Chinese pair produces one record named旅行计划, with both facts and slugmemory-topic, followed byfirst topic is missing.The experimental HARNESS warning is present. No model or external service is required.
Package Versions
Executed the complete tracked core package source from
ca936db37708539bd502240a48543f6094259ce5with cached dependencies. Currentmain91ab44faa4824a30247d00c341f3498ba46b1ca3has identical executable core package source; within this package,pyproject.tomlchanges the version from 1.20.0 to 1.21.0 and raises optionalallextra package requirements. Installed core distribution metadata is 1.20.0. A fresh installation of the 1.21.0 dependency set or optionalallextra has not been tested.Python Version
Python 3.14.3
Operating System
Windows
Regression
Unknown
Additional Context
_slugify_topic()keeps only ASCII letters and digits, usingmemory-topicif none remain._merge_memory()looks up the existing record by this slug and appends the new fact without checking whether it is a different topic name. Topic lookup and deletion use the same mapping.The original memory harness was added in #5613. #7130 fixes Unicode keyword extraction for selecting memories; it does not change topic filenames. Open #9153 preserves explicitly chosen custom slugs during reload while leaving default slug derivation unchanged. This report concerns collisions between automatically derived topic identities, with no custom slug supplied. #9197 concerns contributor session ID serialization, and #9079 concerns summary metadata parsing.
No product patch has been implemented. Please confirm the intended topic identity and compatibility approach before implementation. Already conflated records cannot be reliably separated from their stored contents alone; changing the filename mapping also needs an explicit decision for existing files and lookup by slug.
AI Assistance
AI-assisted issue analysis and reproduction tests.
Acknowledgements