Repository navigation
fix(core): create vector storage at initialization, never at runtime - #1656
Conversation
Every embedding job builds a fresh search repository, so its per-instance "tables ready" flags start false and setup re-ran CREATE TABLE/INDEX IF NOT EXISTS on every job. CREATE INDEX IF NOT EXISTS takes a SHARE lock before it sees the index exists, so under concurrent writers it queued every chunk and embedding write behind it until queries hit the 30s timeout (one cloud tenant: ~28k embedding job timeouts in two days). Setup now reads the catalog first (to_regclass/pg_attribute, no table locks) and only runs DDL when the chunk table, the embeddings table, or their indexes are missing or out of date. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF Signed-off-by: phernandez <paul@basicmachines.co>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7d27cbafa9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| "SELECT to_regclass('search_vector_chunks') IS NOT NULL " | ||
| "AND to_regclass('idx_search_vector_chunks_project_entity') " | ||
| "IS NOT NULL" |
There was a problem hiding this comment.
Keep readiness probes in the current schema
When a deployment uses a layered search path such as tenant,public and the tenant schema is being bootstrapped while public already contains these objects, to_regclass searches the entire path, so this probe evaluates true for the public table and index. Unlike the previous unconditional CREATE TABLE/CREATE INDEX, which targets the first writable schema, setup now skips tenant-local creation and later unqualified vector writes fall through to the public tables, mixing search state across schema boundaries. Verify that both OIDs belong to current_schema() before taking the fast path.
AGENTS.md reference: AGENTS.md:L392-L396
Useful? React with 👍 / 👎.
to_regclass searches the whole search_path, so with a layered path like `tenant, public` the vector tables in public satisfied the probe and setup skipped creating the tenant schema's own. The probes now read pg_class for current_schema() only, matching where unqualified DDL and writes land. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF Signed-off-by: phernandez <paul@basicmachines.co>
Runtime code no longer issues DDL. Search and vector-sync paths only bind the vector adapter; PgVectorIndex upsert/search/delete no longer call initialize(). Storage is created by init_search_index() -> create_storage(), which local startup already calls and cloud will call once per tenant per process alongside migrations. A database missing vector storage now fails on its first query instead of creating tables mid-job. create_storage() still reads the catalog first, so initializing an already initialized database (every new worker) takes no locks on the tables. Tests that relied on runtime creation now initialize the database first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF Signed-off-by: phernandez <paul@basicmachines.co>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2b89294ab5
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| existing_dimensions = await self._existing_dimensions(session) | ||
| storage_missing = existing_dimensions is None | ||
| source_hash_missing = ( | ||
| existing_dimensions is not None and not await self._has_source_hash_column(session) |
There was a problem hiding this comment.
Keep mismatch checks in the target schema
With a layered search path such as tenant,public, when the tenant embedding table is absent but public.search_vector_embeddings has different dimensions or lacks source_hash, the new local _read_storage probe correctly takes the creation path, but _existing_dimensions() and _has_source_hash_column() still resolve the public table. The resulting mismatch reaches the unqualified DROP TABLE IF EXISTS search_vector_embeddings, which resolves to and deletes the public table rather than creating isolated tenant storage. Fresh evidence beyond the earlier comment is that the fast probe is now schema-scoped while these follow-up probes remain search-path-scoped; scope them to current_schema() as well.
Useful? React with 👍 / 👎.
create_storage() already reads pg_class/pg_attribute for current_schema(); the dimension and source_hash checks used separate lookups that searched the whole search_path, so a mismatched same-named table in another schema could be treated as ours and dropped. Reuse the schema-local read and delete the two path-wide helpers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF Signed-off-by: phernandez <paul@basicmachines.co>
Problem
Every embedding job builds a fresh
PostgresSearchRepository, and its first vector call ranCREATE EXTENSION/CREATE TABLE IF NOT EXISTS/CREATE INDEX IF NOT EXISTSfor the chunk and embedding tables.CREATE INDEX IF NOT EXISTStakes a SHARE lock on the table before it sees the index exists, so under concurrent writers every chunk/embedding write queued behind it. In production one tenant had ~28kindex_embeddings*jobs time out (asyncpg 30s) in two days;pg_stat_activityshowed 9 of theseCREATE INDEXstatements and ~50 writes waiting onrelationlocks.Fix: no DDL at runtime
_ensure_vector_tables, used by search and vector sync) only binds the vector adapter.PgVectorIndex.upsert/search/delete*no longer callinitialize();PgVectorIndex.initialize()is now a runtime no-op. A database without vector storage fails on its first query.init_search_index()→_create_vector_storage()→PgVectorIndex.create_storage()) creates the chunk and embedding tables. Local startup already calls this fromdb.run_migrations; cloud will call it from its once-per-tenant-per-process schema hook (companion cloud PR).create_storage()reads the catalog first (pg_class/pg_attribute/pg_extension, scoped tocurrent_schema()), so initializing an already-initialized database — every new worker — takes no locks on our tables. Only a fresh or outdated database runs DDL.Tests
ROW EXCLUSIVEon both vector tables, runtime binding and re-initialization both finish within 5s. Onmainthis times out, the production failure.search_path(tenant, public) with tables only in public still creates the tenant schema's own tables (Codex P1).init_search_index().just lint,just typecheckclean.🤖 Generated with Claude Code
https://claude.ai/code/session_01T8wjd6HrtSA2LN9ssC4NzF