Repository navigation
feat(core): report work units on indexing and embedding results - #1672
Conversation
sync_entity_vectors now returns the VectorSyncBatchResult the shared sync already produced instead of dropping it. EmbeddingIndexResult and EmbeddingIndexBatchResult gain chunks_embedded, taken from embedding_jobs_total: only chunks actually sent to the embedder count, so an unchanged re-sync reports zero. Cloud bills work units from this field. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF Signed-off-by: phernandez <paul@basicmachines.co>
FileIndexResult, IndexedEntity, and SyncedMarkdownFile gain indexed_bytes: the byte length of the stored content the indexer read and parsed (the loaded input bytes, before any frontmatter rewrite; 0 for a regular file indexed from metadata alone, and 0 when unchanged Markdown is not re-indexed). IndexFileJobResult carries it for processed files only; current, missing, failed, and superseded outcomes stay 0. IndexFileBatchJobResult and ProjectIndexCoordinatorResult expose the sum. Cloud bills work units from these counts; indexing behavior is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF Signed-off-by: phernandez <paul@basicmachines.co>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: eb9de24323
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| worker=lambda path: self._refresh_search_index( | ||
| prepared_entities[path], | ||
| entities_by_id[prepared_entities[path].entity_id], | ||
| indexed_bytes=_input_content_bytes(files[path]), |
There was a problem hiding this comment.
Zero bytes for stale resource no-ops
When _upsert_regular_file loses a note/resource classification race, its guarded exits return _PreparedEntity(refresh_search=False) without updating metadata or search state, but this unconditional argument still attaches the full payload length. build_index_file_batch_job_result then emits a processed result and aggregates those bytes, permanently over-reporting billable work for a stale writer that intentionally did nothing. Pass zero for these no-op prepared entities, or carry an explicit indexing outcome alongside refresh_search.
AGENTS.md reference: AGENTS.md:L162-L167
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in c73fc89: _refresh_search_index, which every path goes through, reports bytes only when the prepared entity refreshes, so the guarded no-op exits report 0. test_stale_resource_pass_preserves_newer_markdown_state now asserts indexed_bytes == 0 (it was 20 before the fix).
A regular-file pass that loses a note/resource race returns without updating metadata or search, but its result still carried the full payload length, so a stale writer that did nothing would be billed for it. The search refresh now reports bytes only when the prepared entity actually refreshes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF Signed-off-by: phernandez <paul@basicmachines.co>
Summary
Basic Memory Cloud is moving to per-tenant unit economics and will bill indexing and embeddings by the work a tenant caused (basic-memory-cloud#2373). Core now reports that work on its job results; it does not change how indexing or embedding runs.
Chunks embedded
sync_entity_vectorsreturnsVectorSyncBatchResultinstead ofNone(protocols,SearchRepositoryBase,SearchService). Clearing an entity's vectors returns an all-zero result.EmbeddingIndexResult.chunks_embeddedandEmbeddingIndexBatchResult.chunks_embeddedcome fromembedding_jobs_total: chunks actually sent to the embedder. Unchanged chunks that sync skips count 0.Bytes indexed
indexed_bytesonIndexedEntity,SyncedMarkdownFileandFileIndexResult: the stored bytes the indexer read and parsed (Markdown before any frontmatter rewrite; regular files the object read). An unchanged file returned without re-indexing reports 0.IndexFileJobResult.indexed_bytes(default 0): only a processed, non-superseded file reports bytes; current, missing, failed and superseded outcomes report 0.IndexFileBatchJobResult.indexed_bytesandProjectIndexCoordinatorResult.indexed_bytessum their file results.Tests
test_embedding_index_reports_only_chunks_actually_embedded(real SQLite, deterministic stub embedder): first sync > 0, unchanged re-sync 0.test-int/test_indexed_bytes_work_units.py(real DB): a processed file reportslen(content.encode())including Unicode; an unchanged file iscurrentand reports 0; the single-file path reports bytes.test_index_file_job_result_bills_bytes_only_for_current_content: superseded results report 0.Downstream: Cloud fills the new required fields at its next pin bump.
🤖 Generated with Claude Code
https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF