Repository navigation
v0.9.17: tables ledger, oracle integrations, mship improvements - #8853
Merged
Merged
Conversation
* feat(integrations): add saved Ramp and Vanta credentials * fix(integrations): harden credential renewal and connection setup * fix(integrations): make credential renewal waits abortable * docs(integrations): document token credential availability registration * improvement(integrations): refresh Ramp branding and simplify credential forms
Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Bill Leoutsakos <bill@sim.ai>
…rface lint on deploy (#8803) * fix(workflows): lint unquoted string references in JSON fields and surface lint on deploy * fix(workflows): lint only the JSON fields the operation sends, and lint deploys as the acting user * fix(workflows): check JSON editors sent under a canonical param
* docs(library): update best-ai-agents-support-ticket-triage * Pi Babysit: address PR #8805 feedback --------- Co-authored-by: Sim Pi Agent <pi@sim.ai>
… 2026 (#8806) Co-authored-by: Sim Pi Agent <pi@sim.ai>
…ags (#8808) * chore(config): retire unused env vars and fully-rolled-out feature flags * chore(helm): bump chart version for removed TABLE_ROW_TTL value
* feat(oracledb): add Oracle Database integration * test(oracledb): remove retired block secret fields from masking inventory --------- Co-authored-by: Waleed Latif <walif6@gmail.com>
* fix(files): support arbitrary uploads and in-document links * fix(files): normalize previews and avoid oversized retries * fix(files): bound decoded text for encoded responses
…es a missing organization (#8809) * fix(billing): acknowledge Stripe webhooks whose subscription references a missing organization instead of retrying for days * test(billing): prove the not-yet-matching issuance branch by redelivery; drop the logging-only update catch * test(testing): a swapped price in the Stripe fake no longer reports the previous amount
* fix(db): retire unused keyword search projections * fix(db): retire keyword writers before schema push * fix(db): require approval before retiring push writers
* fix(google-ads): align reporting with current API access * fix(google-ads): preserve nullable inputs and redact retired credentials
* feat(dashboards): add empty-state graphic for the dashboard page Claude-Session: https://claude.ai/code/session_01J7A6CWpREr1jiPAdWyTQbA * improvement(dashboards): mark empty-state chart points as const Claude-Session: https://claude.ai/code/session_01J7A6CWpREr1jiPAdWyTQbA
…table definition row (#8765) * fix(tables): log row writes to the change log instead of locking the table definition row * fix(tables): lock the change log before the dev count reconcile, and tighten the held-row test * test(tables): name the renumbered change-log migration
* fix(tables): drop the per-table row-order lock from inserts Appends mint keys in a random slot after the last key, so concurrent appends never share a key and a batch stays contiguous. Positions are left best-effort and may repeat; the run dispatcher finishes a tied position before advancing its cursor. Positional inserts skip tied keys. Concurrent replaces stay serialized by the table's unique lock, which every replace already takes. * test(tables): order rows bytewise in the lockless-insert tests and stub the job queue locally
…ition source (#8815) * feat(analytics): attribute sign-ups and demo requests to their acquisition source - Record first and last marketing touch (campaign params, referring domain, landing path) in consent-gated first-party cookies on marketing and sign-in pages; attach them to user_created and the PostHog person - Capture $pageview on marketing routes, landing_demo_request_submitted, landing_demo_booked, external_sign_in_started, and email_type on identify - Add useCaptureWhenReady so view events captured on mount are no longer dropped before PostHog initializes - Include attribution in the demo-request sales notification - List the attribution cookies in the cookie policy * fix(analytics): cap attribution cookies at the consent grant, enforce stored field shapes, and lock sign-in buttons while pending
* feat(changelog): publish curated product updates * fix(changelog): harden feed rendering and media validation * improvement(changelog): present updates in a crawlable list * improvement(changelog): define the Sim editorial workflow Document scheduled discovery, durable candidate state, draft PR preparation, real-media review, and operational checks. Reuse the existing cached checkout for trusted Helm PR checks while preserving full history. * fix(changelog): preserve editorial decisions in the workflow plan * docs(changelog): refine product stories and define editorial quality * refactor(changelog): simplify presentation and validate model links
… leave-site prompt is answered (#8432) * fix(desktop-browser): ask the user before a page's alert, confirm, or leave-site prompt is answered * fix(desktop-browser): ask only about the user's own on-screen page, never the agent's * fix(desktop-browser): label frame dialogs with their own origin and cover reload leave prompts * fix(desktop-browser): preserve native dialog ownership and keyboard focus
…IPAA, GDPR, Data Residency, On-Prem, and VPC (#8822) Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
* fix(mcp): add directory verification compatibility * test(mcp): verify ownership challenge over HTTP
* fix(chat): send entitlements and agent mode from the public chat API `/api/v2/chat` (what `sim chat` uses) stopped sending `entitlements` and `mode` in #8208, so CLI and API chats got no entitlement-gated tools and every CLI service refused them with "CLI services require agent mode". The route now computes entitlements per turn like the workspace chat and sends `mode: 'agent'`. Its test still mocked the old entitlements function and listed `mode` as a forbidden legacy field; both are updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * feat(tests): workflow tests as a workspace resource Workflow tests are a workspace resource whose source is a plain vitest file, `tests/<name>.test.js`, owned by the test (workspace_files.context = 'test'). Sim creates a test's metadata with the tests tool and writes its cases with the file tools; every write is collected in the sandbox and refused if the file does not load. - Runner: test files run in the isolated-vm sandbox against draft or deployed workflows. `runWorkflow` executes real runs; `mockBlock`, `mockTool` (Agent tool calls) and `spyOnBlock` reach blocks in the tested workflow and in every child workflow it runs, matched by name as each workflow starts. `.mockSampleOutput()` builds outputs shaped like the real block or tool. `toMatchRubric` asks a model judge for pass or fail. - Runs record live per-case progress, the source hash, and the deployment of every workflow they ran, so results show as out of date once the test or a workflow changes. - UI: Tests page and test page (Edit / Split / Preview over the file, the preview a dashboard of the selected run), a test resource type in chat, and a Tests sidebar entry behind the `workflow-tests` flag. - Owned files never open as file tabs in chat: only workspace files and chat uploads do. - Migration 0400 adds workflow_test and workflow_test_run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * feat(tests): pick a run from a dropdown and open what each run ran against The test page shows one run at a time, chosen from a run picker with status dots and Draft / Out of date chips. Case statuses use the Badge status chip. Each ran-against entry records one execution, so a draft row opens the workflow snapshot from that run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): type errors and tests broken by workflow tests - Narrow the test principal to the kinds workflow_tests.run admits before handing it to executeWorkflow. - Select progress with the latest-run rows, guard file upsert ids in the tab filter, and set the sandbox Event polyfills through Reflect. - Cover the tests tool in the management tool contract, expect content writes to reach test files, and stub test availability in the payload test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): address review on redaction, staleness, and the harness - Redact each run's resolved secrets from what returns to the sandbox (output, errors, mocked tool inputs); mocked tools get only declared params. - Custom blocks no longer receive the consumer's test hooks. - Draft runs go stale when the draft changes; children a run calls are recorded in ran-against. - Test cases commit in the same transaction as the source file write. - Harness: runWorkflow is rejected in suite hooks, a timed-out case stops the file, and expect.assertions/hasAssertions are supported. - Insert run rows in one statement and start each run's clock with its file; check bans before each workflow run; restrict owned-file access to Copilot delegation; validate names in the tool contract. - Delete soft-deletes the test file and removes the chat tab; a finished run shows its own cases; polling at 3s on a separate read bucket; list error state; store reset; tests stay in the org Add Resource picker. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): reserve run slots, await pending assertions, keep dynamic tool args - Each test workflow run reserves and releases an execution slot. - A case waits for assertions it did not await and fails if one fails. - Mocked MCP and custom tools keep the arguments their schema declares. - A closed session refuses starts still awaiting their lookups. - Stable refresh callback; scroll fade on the results pane. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): keep test sources out of file tabs, refresh after runs, reopen tests - File-edit tool results mark a non-tab file `fileTab: false`, and the browser skips promoting it. - Idle test pages poll every 15s so runs started elsewhere appear; a Mothership run returns its tests as resource changes. - open_resource accepts test resources through an authorized read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): refresh test tabs after a run, keep saved edits successful - A finished Mothership run refreshes its tests instead of upserting tabs, so a test deleted mid-run does not come back. - A failed file-tab lookup after a saved edit opens no tab instead of reporting the edit as failed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): starter source imports every test helper Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * build(tests): add @vitest/expect and @vitest/spy for the sandbox bundle The vitest-expect sandbox bundle builds from these packages; rebuilt with the Reflect-based event polyfills. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * build(tests): tell knip the sandbox bundle uses @vitest/expect and @vitest/spy Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * feat(tests): name MCP and custom tool mocks by server and title MCP tool ids embed the server's database id, which changes when a server is re-added or a workspace is forked, so a stale mock silently stopped matching and the real server was called. Tests now name workspace tools the way the workspace does: mockTool({ mcp: 'Server', tool: 'name' }) resolved per run (failing on an unknown or ambiguous server), and mockTool({ customTool: 'Title' }) matched case- and space-insensitively. Raw mcp- and custom_ ids are rejected; built-in catalog ids are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): reserve the run name, hold Run for unsaved edits, cancel judges on close - A test named "run" collided with the static run endpoint, so its detail page got a 405; the name is now reserved. - Run is disabled while the open editor holds edits the server has not saved (including a refused save), so a run never uses the previous source. - toMatchRubric model calls are aborted when the sandbox run ends, so a stopped test no longer keeps calling or billing the judge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R * fix(tests): fail a run whose selected case names no test in the file A renamed or misspelled `only` path skipped every case, and the run was then saved as passing. The harness now rejects unknown names, so the run is recorded as an error with the names it could not find. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHGBgHrtePpi7KMrEfy41R --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(providers): remove default agent tool-call iteration cap * docs(providers): note no execution timeout when billing is disabled
* fix(desktop): ask before accessing local files * fix(desktop): remember folder permissions across chats * fix(desktop): cancel pending file consent and recheck access * fix(desktop): enumerate approved directory descriptors * fix(desktop): reject replaced directory listings * fix(desktop): share pending folder consent decisions * fix(desktop): complete file consent and add full file access * fix(desktop): revalidate folder consent and preserve exact identities
…rites (#8870) * fix(execution): tag deterministic admission rejections and throttle blocked-run logs Usage-limit, suspended-account, and missing-billing-account refusals now carry stable codes, so surfaces can tell a refusal that holds until a person acts from a transient one. Unattended surfaces can opt into throttleErrorLogs: each refusal of a workflow by the same gate records at most one execution error row per 15-minute window (Redis SET NX, in-process LRU without Redis, fail-open on Redis errors). A sender resending a refused delivery no longer writes an execution row, trace archive, and workspace_files row per attempt. A caller-supplied logging session is always completed. * fix(webhooks): acknowledge deterministic admission rejections for Telegram and Slack Telegram resends a non-2xx update until it succeeds or 24 hours pass, and Slack disables an app's event subscription once most deliveries fail, so answering a usage-limit refusal with 402 only loops. Providers opt in with acknowledgeAdmissionRejections and get an empty 200 with an ignored outcome; transient refusals (rate limit, concurrency, reservation outage) still fail so the sender retries. Generic webhooks keep the 402. Polling receives the raw refusal with its code. Webhook preprocessing throttles its error rows. * fix(webhooks): skip polls for over-limit payers and back off failing sources The poll orchestrator checks each workspace's payer once per tick and skips its webhooks while the payer is over its usage limit: nothing is fetched, marked seen, or counted as a failure, so items deliver once the payer is back under the limit and triggers are no longer auto-disabled over billing state. A webhook with consecutive poll failures waits 2^(n-1) minutes (capped at an hour) before the next fetch, and a source's Retry-After or FLOOD_WAIT_<n> is persisted and honored. RSS logs a source's 4xx once at warn, and an admission refusal mid-poll leaves the remaining items unseen. * fix(webhooks): stop Slack redelivering to deleted trigger paths and log them once A POST to a path with no webhook keeps its 404 but now carries x-slack-no-retry, sent unconditionally so it reveals nothing the 404 does not. The processor and route lines for an unknown path drop to debug, leaving the route handler's single client-error line. * fix(telegram): verify webhook deliveries with a per-webhook secret token Register a generated secret_token with setWebhook, store it as providerConfig.secretToken, and reject deliveries whose X-Telegram-Bot-Api-Secret-Token header does not match (401). Webhooks registered before this change have no stored secret and stay accepted until their next deploy registers one. A new registration reuses the active deployment's secret for the same bot so deliveries keep verifying during cutover. Drops the stale empty User-Agent warning: the proxy exempts webhook trigger routes from the empty-UA block. * refactor(webhooks): declare the Telegram admission opt-in on its handler * fix(execution): only treat billing and account refusals as deterministic An unreadable usage ledger fails closed as exceeded; it said nothing about the payer, yet it was tagged USAGE_LIMIT_EXCEEDED and acknowledged-and-dropped for Telegram and Slack. It now stays untagged and retryable. Reservation headroom denials clear as in-flight runs settle, so they leave the deterministic set too. The blocked-run log claim drops its in-process fallback: usage refusals only happen on hosted billing deployments, which run Redis, and without Redis every refusal records its row. A Redis failure logs at debug. * fix(webhooks): keep a fan-out target's retryable failure visible past a dropped refusal An acknowledged admission refusal counted as an acknowledgment, so a Slack or path fan-out answered 200 even when another target failed and needed the sender to retry. Dropped targets (block missing, acknowledged refusal) now answer 200 only when no other target failed. * fix(telegram): match the active bot through env-var token references Active rows store the bot token as authored, often a {{VAR}} reference, while subscription calls receive it resolved, so the comparison never matched: every deploy minted a fresh secret (a candidate that never activated left the bot rejected by the active row), and retiring an old version could delete the webhook the active version still used. Stored tokens are now resolved against the background webhook env before comparing. * fix(webhooks): skip polls only after a recorded refusal and back off source failures The per-tick payer pre-check read billing attribution and the usage ledger for every polled workspace, including healthy idle ones. A deterministic admission refusal of a polled event now records the workspace in Redis for five minutes, and the orchestrator skips only those workspaces in one MGET; healthy payers cost no billing reads, and billing-disabled deployments skip the mechanism. Backoff now follows source fetch failures only, tracked in providerConfig and stamped from the failed poll's start, so transient item-processing refusals never back a webhook off. Poll state keys are system-managed so deploy change detection ignores them. * refactor(webhooks): route every poller's source failures through one backoff Every poller's outer catch now records a source failure, so a failing Gmail, Outlook, IMAP, Drive, Sheets, Calendar, or HubSpot source backs off like RSS instead of only RSS. markWebhookSuccess clears the backoff in its existing reset write, the window uses the shared jittered backoff, and the orchestrator asks a boolean isPollBackedOff. Smaller cleanups: one isDroppedDispatch predicate for the Slack and path fan-outs, explicit precedence for a polled refusal's code, a typed RSS refusal error instead of a flag, Telegram resolves only the stored bot token, and PollOutcome lives with the polling types. * chore(webhooks): tighten poll comments and backoff tests Poll outcome docs describe what a skipped poll actually does, source failures keep logging the full error object as the pollers did before, the backoff table pins the clock past the poll start so it proves the window is anchored there, and the RSS rate-limit test drives a Retry-After header. * fix(webhooks): stop every poller's batch on a deterministic admission refusal Only RSS stopped at a refused item; Gmail, Outlook, and IMAP advanced their cursors past refused emails, and every poller counted the refusals toward auto-disable. A shared PollAdmissionRefusedError now leaves the idempotency callback, stops the batch, and returns skipped before any cursor update or failure count; items that already ran replay as idempotent no-ops. Source backoff goes back to RSS only, where the rate-limited feed was: the other pollers' fetch helpers do not carry status or Retry-After, so routing their failures through it would back off on a guess. The block-missing 404 also tells Slack not to redeliver. * fix(webhooks): never replay completed poll work or mask a retryable fan-out failure A poller stops its batch on a deterministic refusal only while nothing in the batch has completed; once an item has run, the refusal is an ordinary item failure, so the poller saves its completed work exactly as before and no completed event can replay after the idempotency window. In a multi-target delivery a missing block's no-retry 404 no longer stands in for a target that failed and needs the sender to retry. A source's Retry-After is counted from its answer rather than the poll's start. The RSS backoff keys are cleared by RSS's own state write, so other pollers' success path is unchanged, and two fields nothing reads are dropped. * fix(execution): drop the unreachable billing-account admission code A workspace without a billing account fails inside payer resolution and takes the retryable attribution-error path; the branch that tagged BILLING_ACCOUNT_REQUIRED only ran for an attribution with no actor, which system attribution never produces. The branch goes back to its staging form and the deterministic set keeps the usage limit and suspended accounts. * chore(webhooks): key blocked-run claims by gate and centralize the polling utils mock Two gates that fail without a code (a ban lookup error and a usage lookup error) no longer share one throttle claim, so neither hides the other's row. The polling utils module gets one central mock in @sim/testing, replacing the partial importOriginal mocks and the hand-rolled factory in the table trigger test. * fix(telegram): match the active bot with the env the caller resolved its token with Subscription creation resolves the incoming bot token with the deployer's env, while cleanup resolves with the background env; the active-row matcher now uses the same env as its caller, so a bot referenced through a personal variable still reuses the active secret and is not deleted from under the active deployment. Tests: the idempotency service gets one central mock in @sim/testing, used by every test that mocked it locally, and the new tests import single factory and mock files instead of the @sim/testing barrel. * chore(testing): stub IdempotencyService.createWebhookIdempotencyKey in the central mock
…ons they declare (#8877) * improvement(sim-cli): commands reach the API only through the operations they declare Hand-written commands called SimClient by path and declared nothing, so the command inventory could not mark one Mothership-unavailable when it calls a refused route. Every command now gets its client from callsOperations: calls name an operation, route and method come from the operation table, and the client is typed to the declared operations. * test(sim-cli): assert declared-operation routing at the transport * test(sim-cli): assert directory requests at the transport
…dary (#8872) * improvement(logs): log each execution failure once at its owning boundary A single external tool failure was logged at ERROR up to eight times as it propagated from the tool layer through the block executor, the engine, execution-core, and the trigger surfaces. Each boundary now calls logFailureOnce, which skips a failure an inner boundary already logged and sets severity by attribution: the author's input, configuration, or code at info, a third-party 4xx or 5xx at warn, and Sim's own faults (database, retryable setup, hosted-key rejection, anything unattributed) at error. The logged mark and attribution live in WeakMap/WeakSet side tables keyed by the thrown value, walked through the cause chain, and carried on the flattened tool failure's output so they survive a handler rebuilding a failed tool result as a new error. Run logs and trace spans are unchanged. * test(logs): pin the logged mark and attribution at the block executor Asserts the block executor's thrown error is already logged and attributed (database failure internal, user Stop user), switches the tool-boundary test to the central jsonResponse helper, and updates exact-argument log assertions for the added failureKind and executionId fields. * fix(logs): scope the logged mark to one propagation and attribute in-process failures Never mark a raw thrown value as logged: a persistent fault that rethrows one object (a rejected dynamic import, a memoized rejected promise) was logged once per process and then silenced everywhere. Only carriers created during the propagation (the flattened tool failure output, the block error, a handler's rebuilt error) carry the mark, and execution-level boundaries scope a raw value's mark to their execution id. Handlers that rebuild a failed tool result now call adoptToolFailure instead of relying on the result output being copied by reference. The child workflow and agent handlers no longer log failures the block executor logs with the block's and run's identity. In-process operations are attributed to Sim (4xx user, 5xx internal) rather than to a third party, only external HTTP failures log their response body, a missing required field is the only serializer refusal at info, and the size-limit and JSON-parse paths no longer log twice. * refactor(logs): build failure log metadata lazily and attribute author failures by type logFailureOnce takes its metadata as a thunk so a skipped boundary does no secret projection, adds the execution id itself, and returns the attribution it logged with. UserFailure attributes author-caused failures by type (missing required fields, boundary-safe custom block refusals, no start block) wherever they surface, replacing a try/catch at one serializer caller. One inheritFailureMarks carries both marks across a boundary that drops cause, the hosted-key check reuses classifyHostedKeyFailure, the Pi backends share one toolResultError, and the sandbox and Function route no longer log author code failures the tool boundary already logs. * test(logs): attribute the serializer refusal at its owner and pin block log redaction Moves the missing-required-fields attribution check to the serializer that throws it, drops an execution-core case whose internal row passed by default, and covers the block failure line's secret projection now that it, not the Agent handler, logs provider errors. Tightens comments the diff added. * improvement(logs): drop casts and test-only exports from the failure log Reads status and cause without casts, keeps the thrown missing-required-fields error's name and stack unchanged, ignores an empty execution id when scoping a logged mark, unexports wasFailureLogged and FailureKind (tests observe the outer boundary through logFailureOnce instead), and shrinks the explicit-any baseline for the Agent handler. * fix(logs): close the remaining double logs found in review Carries failure marks through the scrubbed Pi error and normalizes a non-Error node failure before logging so the engine sees its mark; drops the condition handler's and response-size handler's own error lines in favor of the owning boundary; logs MCP failures once and marks their output; passes the block id to the API and Function tool calls so the tool line carries it; attributes custom tool parameter validation and condition expression errors to the author; checks the whole cause chain for a retryable setup failure before honoring a mark; and restores the sandbox failure logs, whose adapters cannot yet tell a provider exception from a non-zero exit. * fix(logs): keep run identity on every remaining failure line Fixes the type error on the external failure body, attributes a provider error payload on a 2xx to the provider, and adds workflow, execution, and block ids to the MCP failure lines that replace the block executor's. The condition batch carries its failed result's marks and passes its block id. Pi GitHub calls carry no run identity, so their errors now carry only the tool layer's attribution and the block executor still logs them with ids. Restores the HITL notification warning, which covers soft failures the tool layer never logs, and attributes refused tool and proxy URLs to the author. * fix(logs): attribute a revoked OAuth credential to its owner A CredentialRevokedError anywhere in the cause chain classifies as a user failure, so the boundaries that log it after token resolution's WARN write INFO rather than ERROR. * fix(logs): keep one logged mark per execution and attribute unparseable responses A raw value logged at an execution boundary now remembers every execution that logged it (bounded), so overlapping runs sharing one persistent rejection each log it once. A 2xx body that is not JSON is the endpoint's failure, not Sim's.
* improvement(app): show branded wordmark while loading * fix(app): preload workspace data alongside branding
…data-typed next/script (#8875) * fix(docs): render the FAQ JSON-LD as a native script tag, and forbid data-typed next/script The FAQ component rendered its FAQPage JSON-LD through next/script with the default afterInteractive strategy, which returns null on the server and injects the tag from a client effect, so the served HTML never carried it. Render a native <script> the way #6763 fixed the other docs JSON-LD blocks; serializeJsonLd already escapes '<'. check:source-text now also fails on a next/script element in apps/** whose type is not JavaScript, naming the native-script fix. * fix(audits): match only the real type prop, fail closed on expression types, and scan MDX pages * fix(audits): recognize default-as next/script imports and skip MDX code examples * fix(audits): keep JSX template literals intact when skipping MDX inline code
…ode; fix the PPTX blend-group call they hid (#8876) * chore(audits): drop dead ESLint directives and the redundant check:dead-code script - Remove all 89 `eslint-disable` comments. Nothing runs ESLint (no eslint dependency or config in any workspace; Next 16 `next build` has no lint step), so they suppressed nothing. Reasons they carried are kept as plain `//` whys. - `check:comment-hygiene` now rejects `eslint-disable`/`eslint-enable` comments so they cannot return. - Remove `check:dead-code`: `check:unused-exports` already runs knip with the same config and a superset of its issue types, and `check:audits` skipped the alias. - Ignore unused exports in the generated `apps/docs/components/icons.tsx` (a verbatim copy of the sim icon set, whose export surface is ratcheted at the source) and shrink the unused-exports baseline by its 66 entries. * fix(pptx): stop calling the appended rect blend group as a function * docs(skills): name both ESLint directives the comment audit rejects * docs: name both ESLint directives in the comment rule * test(pptx): cover the rect gradient blend-group path
…adline (#8878) * fix(tools): stop provider timeout params from becoming the request deadline The request transport and the internal-operation path read params.timeout as a millisecond deadline for every tool. Twilio make_call, New Relic NRQL, Apify (3 tools), Daytona (2 tools), and Trigger.dev waitpoint tokens declare their own timeout param in seconds or as a duration, so a 60-second setting aborted the call after 60 ms. A declared timeout param is now the deadline only when the tool sets timeoutParamIsDeadline (http_request, firecrawl_map, firecrawl_parse); callers can still bound tools that declare none. No param ids change. Redis and Upstash coerced params with Number() inside tools.config.tool, which runs at serialization on the serialized params object, turning <Block.output> references into NaN. The coercions now run in tools.config.params. Guardrails: check-block-registry rejects coerced assignments to params inside an inline tools.config.tool; check-tool-param-reachability rejects a method param on a fixed-verb external tool, which the transport would send as the HTTP verb. * fix(audits): catch coercions under fallbacks in tools.config.tool, and clarify declared timeout params * fix(audits): treat every value-deriving expression as a selector coercion * fix(audits): reject compound assignments and increments on params in selectors
…through wrappers (#8885) - v1 admin: unauthorized/forbidden/notFound/badRequest/internalErrorResponse -> admin*Response; errorResponse module-local; conflict/notConfiguredResponse also admin-prefixed - files createErrorResponse -> createFileErrorResponse; workflows createErrorResponse -> createCodedErrorResponse (+ central mock) - import VFS segment codecs from @/lib/vfs/path; delete the mothership pass-throughs and the dead canonicalizeVfsPath - search-replace resource group key uses sortObjectKeysDeep instead of a localeCompare serializer
… discovers them (#8882) * chore(audits): name generated-artifact checks check:<x> so run-audits discovers them Rename tool-metadata, deployment-config, integration-catalog, docs, agent-stream-docs and docs-manifest checks to check:<x> (generators to generate:<x>, matching the existing generate:openapi/check:openapi pairs). run-audits now derives every audit from the check:* prefix, so the hand-kept EXTRA_AUDITS list is gone, and check:docs-manifest (local fs only, ~1s) joins check:audits instead of its own CI step and CLAUDE.md gate line. The <x>:check suffix stays only for checks this job cannot run (sibling copilot repo, Helm, per-machine covers). Add module headers to check-monorepo-boundaries, check-realtime-prune-graph and check-openapi saying what each protects and how to fix a finding. * docs(ship): regenerate tool metadata in Phase A; drop the actionlint comment with no command
…8884) * chore(naming): rename 41 baselined files to the file-name convention Kebab-case the guardrails validators and azure-blob destination, drop folder stutter in lib/logs, hooks/queries/oauth, emcn charts/chip/popover/tooltip/ tab-strip/calendar and workflow-renderer note/subflow/workflow-block, and rename _polyfills/_nodemailer. Updates every importer, vi.mock string, package exports/imports/sideEffects target and doc reference; fixes the anys and redundant non-null assertions in renamed files. check:file-names baseline 171 -> 130. * docs(workflow-renderer): point the handle-position comment at the renamed views
* docs(agents): align agent guidance with code and enforced checks - contracts: export the contract; export a schema or type only when imported - testing: mock-assertion ban matches test-audit; .dom.test convention; runner commands - imports: biome organizeImports owns order - stores: devtools for new stores; reset() + registerUserDataReset - add apps/sim/CLAUDE.md and (landing)/AGENTS.md symlinks; name db-migrate skill - skills: safeUrlPathSegment for path segments, hmacSha256Hex and bounded reads in trigger templates, hosted-key per_request/enabledWhen, stale identifiers fixed, pressure markers removed - CONTRIBUTING: commit types and format match history * docs(add-model): keep the check:/generate: script names from staging * docs(emcn-design-review): list only the Chip variants its props accept
) * chore(tests): tighten check:test-patterns and shrink its baseline - Exempt *.live.test.ts from global-remock and local-factory: live suites bind real boundaries (real Postgres via a re-mocked @sim/db). - Fail when a central mock vi.mocks an @/ module id that no longer resolves; repoint copilot-http to @/lib/mothership/request/http (renamed in #8208) and drop executor.mock's mocks of deleted modules. The check also scans vitest.setup.ts, whose mock and alias of the deleted @/stores/console/store are removed. - Gate local-helper on the workspace depending on @sim/testing. - redundant-hook inspects every afterEach statement and the leading run of a beforeEach (a later reset can deliberately discard setup's own calls). - New module-scope-stub rule: module-scope vi.stubGlobal/stubEnv/spyOn is undone before the first test; delete the six dead fetch stubs. - Rewrite @sim/testing @example barrel imports to direct paths. - Replace real sleeps with fake timers in file-doc-store and update-cost. - Convert table-constants and sim-search-connectors local factories to the central mocks; fold v1-route's subscription/rate-limiter duplicates into billingSubscriptionMock/rateLimiterMock. Baseline: 140 -> 92 entries. * fix(tests): catch exported and type-asserted module-scope stubs, and cover the stub and hook rules with fixtures
…face-neutral application imports (#8883) * chore(audits): check raw-route parseRequest contracts and enforce surface-neutral application imports check:route-verbs now resolves the contract each raw withRouteHandler route passes to parseRequest and checks the exported verb and path against it (322 sites across 258 files, previously unchecked). check:boundaries now bans next/server and app/api imports, and runtime route-contract, presenter and Copilot-handler imports, from application code. listSearchSources takes its cursor route from the adapter instead of importing the route contract. * fix(tests): pass cursorRoute in the org source-summary integration case * fix(audits): resolve aliased contract imports and export-list verbs in check:route-verbs Read a contract by the name its module exports, not the local alias, and read handler verbs from export lists (export { GET }, export { h as GET }) as well as inline exports. Both builder and raw parseRequest sites share the fix. * fix(audits): follow aliased parseRequest imports and relative application imports check:route-verbs matched raw parseRequest calls only by the literal name, so `import { parseRequest as parse }` left the contract unchecked; it now matches every local name bound to parseRequest from @/lib/api/server(/validation). check:boundaries applied the application rules only to @/ specifiers, so a relative import of a contract object, presenter, Copilot handler or app/api module passed; relative specifiers are now normalized to their @/ form first. * fix(audits): accept HEAD on a GET contract and treat empty import/export clauses as runtime edges * fix(audits): keep the HEAD-on-GET allowance off defineScimRoute and catch extension-suffixed presenter imports * fix(audits): follow exported non-verb helpers when resolving raw-route contracts
…blish, attest via actions/attest (#8873) * improvement(ci): halve the mobile job, narrow the desktop gate, skip proven checks on main, fix the top flake - e2e (mobile): Chromium and WebKit run as two processes against one app instead of back to back (each seeds its own fixtures and writes its own report), and every e2e group restores its own Turbopack dev cache. The job was the PR critical path at ~19.5 min: Chromium's pass (718 s median) carried the cold compile and WebKit waited it out (310 s). `suite setup` failures are now logged instead of only reaching the uploaded report. - dev-cache action: one mount shared by desktop-live and the e2e groups, keyed by app, event, fork, Next version and the hash of the file that sets the app's environment, so a cache is never read under NEXT_PUBLIC_* values it was not compiled with. http-e2e.sh stops apps with SIGINT so the cache write completes, drops a cache Turbopack reports as corrupt, and retries a startup it aborted once from an empty cache. - desktop-live-changes.sh: run the live desktop suite when a pull request touches a path it exercises instead of skipping only docs-like paths. Every genuine desktop-live failure since the layout change was on a PR the new rule still runs; pushes always run it. - ci.yml: a `proof` job lets a main push skip the checks when it is a merge whose tree is identical to its staging parent and a staging push run of that parent passed `checks / ci`; migrate gates on that proof instead. Any other shape or any error runs the full checks. - codeql, helm, desktop-e2e skip the staging -> main release PR (main's push and schedule still run them); helm never cancels a push mid-publish; trigger-promote on 2 vCPU; attestations via actions/attest with artifact-metadata: write. - stripe-sync-convergence: start each contender only after the parked transaction holds its locks, wait on a deadline rather than 200 polls, and release the transaction before failing, so a lost race no longer hangs the suite's teardown. It was the most frequent flake (9 red runs). - Integration shard weights refreshed from a recent run; /ship runs the affected workspaces' suites instead of the full suite locally. * ci: include the chat page in the desktop gate, scope release-PR skips to this repository, give a cache-retry boot the full deadline * ci: keep the release PR's CodeQL scan, require the main base for release-PR skips, gate desktop-live on workspace permissions * ci: keep the desktop live gate's skip-only-unrelated rule * ci: run the mobile browsers one after the other again * ci: drop the e2e Turbopack dev cache and keep /ship's full test gate * ci: drop the main proof job and the release-PR skips
* fix(chat): keep composer drafts across hydration - Adopt text typed into the server-rendered composer before hydration; the dropped input event left state empty and the next render wiped the text - Restore a saved draft after hydration instead of in the initial render, which mismatched the server HTML and regenerated the composer on every reload with a draft - Share one useHydrated hook across the composer, organization home, and theme toggle - Mobile E2E: wait for the draft to overflow its editor instead of measuring once, and check a saved draft returns after a reload without a hydration error * fix(chat): re-chipify a restored draft once skills load * fix(chat): keep the caret where the user typed before hydration
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Companion: simstudioai/mothership#656 (merge this first; then re-run sim-sync on #656)