-
-
Notifications
You must be signed in to change notification settings - Fork 1.4k
Comparing changes
Open a pull request
base repository: triggerdotdev/trigger.dev
base: v4.5.8
head repository: triggerdotdev/trigger.dev
compare: v4.5.9
- 20 commits
- 245 files changed
- 11 contributors
Commits on Jul 27, 2026
-
perf(run-ops-database): index BatchTaskRun for the batches list on th…
…e dedicated schema (#4396) ## Summary The batches list page orders by `(createdAt DESC, id DESC)`, which is why [#4361](#4361) added a matching index on `BatchTaskRun`. That index only landed in `@trigger.dev/database`. The dedicated run-ops database has its own migration history, so it never received the index. `BatchListPresenter` reads both databases and merges, so for environments whose batches live in the dedicated database the page kept falling back to a scan and in-memory sort, which is the exact behaviour #4361 set out to fix. ## Fix Adds the index to the run-ops schema with its own migration. `CREATE INDEX CONCURRENTLY IF NOT EXISTS`, so it is a no-op where the index already exists and still records its ledger row. The second half is the interesting part. Because the two packages own separate migration histories, a run-graph schema change has to be authored twice, and nothing made the miss visible: the run-ops status check truthfully reports "up to date" against its own history, so the apply step just skips. `schemaParity.test.ts` compares the physical shape of every model the run-ops schema declares against its counterpart in `@trigger.dev/database`: scalar fields with their attributes, plus `@@index`, `@@unique`, `@@id` and `@@map`. Relation navigation fields are excluded, since the run-ops schema deliberately drops relations that would cross a database boundary while keeping the scalar FK column. A field counts as a relation when its type resolves to a model name, which keeps enum-typed columns in scope. Two models are listed as run-ops-only: `CompletedWaitpoint` and `WaitpointRunConnection`, both explicit FK-free replacements for a control-plane implicit many-to-many, since an implicit m2m carries a foreign key that cannot resolve across databases. The test also asserts that exception list is exhaustive, so a new unpaired model fails rather than being silently skipped. Confirmed the guard actually fails: reverting the index turns `BatchTaskRun` red with the missing `@@index` named in the diff.
Configuration menu - View commit details
-
Copy full SHA for fc57643 - Browse repository at this point
Copy the full SHA fc57643View commit details -
ci: let the claude bot trigger the PR audit workflows (#4392)
## Summary PRs opened by the claude GitHub app fail both the agent instructions audit and the REVIEW.md drift audit before Claude gets a chance to run. `claude-code-action` refuses any actor whose account type is not `User` unless the actor is listed in `allowed_bots`: ``` Workflow initiated by non-human actor: claude (type: Bot). Add bot to allowed_bots list or use '*' to allow all bots. ``` So those PRs land with two permanently red checks and no audit coverage at all. Both workflows already allowlist Devin; this adds the claude app alongside it. ## Why this does not open the workflows up to outside contributors `allowed_bots` is only consulted for non-`User` actors. Humans, contributor or maintainer, take the separate write-permission path and are unaffected by what is in the list. Beyond that, both jobs are guarded by `github.event.pull_request.head.repo.full_name == github.repository`, so a fork PR skips the job entirely, and they trigger on `pull_request` rather than `pull_request_target`, so a fork-triggered run would get no API key and a read-only token anyway. The bot is named explicitly instead of using `"*"`, which would let every bot trigger these audits, dependabot's PR stream included.
Configuration menu - View commit details
-
Copy full SHA for 3ed4851 - Browse repository at this point
Copy the full SHA 3ed4851View commit details
Commits on Jul 28, 2026
-
fix(webapp): remove unused Electric sync trace routes (#4400)
<!-- ccr-slack-attribution --> _Requested by **Eric Allam** · [Slack thread](https://triggerdotdev.slack.com/archives/C0AU83M3136/p1785222101937829?thread_ts=1785207509.304669&cid=C0AU83M3136)_ Removes two dead Remix routes and the helpers only they used. `app/routes/sync.traces.runs.$traceId.ts` (`/sync/traces/runs/:traceId`) and `app/routes/sync.traces.$traceId.ts` (`/sync/traces/:traceId`) were added with the original ElectricSQL run page and lost their only consumers when the dashboard hooks that called them were deleted. Nothing in the repo references either route today. Also removed, because the deleted routes were their only callers: - `OtelTraceIdSchema`, `RESERVED_ELECTRIC_SHAPE_PARAMS`, `TraceScope`, `buildElectricTraceWhereClause` from `app/v3/electricShape.server.ts` (the file stays — `UNSAFE_REALTIME_TAG_CHARS` / `sanitizeRealtimeTagForSql` / `sanitizeRealtimeTagsForSql` are still used by `realtime.v1.runs.ts` and `realtimeClient.server.ts`) - the loader-specific cases in `apps/webapp/test/spanTraceRoutes.replicaLag.test.ts` and `internal-packages/run-store/src/runOpsStore.routesSpanTraceReadView.replicaLag.test.ts` `app/utils/longPollingFetch.ts` is untouched — `realtimeClient.server.ts` still uses it. `runOpsStore.ts` / `PostgresRunStore.ts` are untouched too; the unrouted-lookup mechanism there is generic and stays. As a plain code fact: the run lookup these loaders performed keyed on `TaskRun.traceId` alone, which is not an index-backed query shape. That is noted only as context for why the code is not worth keeping around unused. ### Judgement call worth a maintainer's opinion The request was specifically about `/sync/traces/runs/:traceId`, the route that looks up a run by `traceId`. This PR **also** deletes its sibling `/sync/traces/:traceId`. The reasoning: - both routes came in with the same ElectricSQL run-page work - both lost their only consumers in the same later commit - neither has any caller anywhere in the repo - they share the same helper module, so keeping one means keeping the helpers half-used If you would rather keep the sibling, reverting just that one file deletion is easy and does not affect the rest of this PR — say the word and I will restore it along with the helpers it needs. ## ✅ Checklist - [x] I have followed every step in the [contributing guide](https://github.com/triggerdotdev/trigger.dev/blob/main/CONTRIBUTING.md) - [x] The PR title follows the convention. - [x] I ran and tested the code works --- ## Testing Verification run locally from the repo root: | Command | Result | | --- | --- | | `pnpm run format` | clean, no changes produced | | `pnpm run lint:fix` | clean | | `pnpm run lint` | pass (exit 0, no findings) | | `pnpm run typecheck --filter webapp` | pass | | `pnpm run typecheck --filter @internal/run-store` | pass | A ripgrep sweep for `sync.traces`, `sync/traces`, `syncTraceRunsLoader`, `buildElectricTraceWhereClause`, `OtelTraceIdSchema` and `RESERVED_ELECTRIC_SHAPE_PARAMS` (excluding `node_modules`) returns zero hits. **Not fully verified:** both edited test files are testcontainers suites and need a Docker runtime, which was not available in my environment. I confirmed each file *collects* correctly with exactly the three intended remaining tests and no import errors — notably, dropping the `session.server` / `controlPlaneResolver.server` / `longPollingFetch` / `env.server` mocks does not break module loading for the surviving loaders. The assertions themselves then failed only on `Could not find a working container runtime strategy`. CI should be the real signal here. Per `apps/webapp/CLAUDE.md`, `pnpm run build --filter webapp` was deliberately not run. --- ## Changelog Removed two unused sync routes left over from the original ElectricSQL run page, along with the helpers and tests that existed only to serve them. No behaviour change — neither route had any caller. --- ## Screenshots _n/a — no user-visible surface changes._ 💯 --------- Co-authored-by: Claude <noreply@anthropic.com>
Configuration menu - View commit details
-
Copy full SHA for ec562c0 - Browse repository at this point
Copy the full SHA ec562c0View commit details -
feat(webapp): org-gated internal API origin in run env vars (#4366)
Adds an opt-in way for operators to route deployed runs' API traffic through a different origin than the public one, per organization. Set `INTERNAL_API_ORIGIN` on the webapp and enable the `internalApiOriginEnabled` feature flag (globally or per org, with the org override winning in both directions): deployed runs for enabled orgs then get `TRIGGER_API_URL` set to the internal origin instead of `API_ORIGIN`. Useful for gradually moving run traffic onto a private network path. ## Design The origin is resolved when an attempt starts, so flag changes take effect on the next attempt and roll back the same way, with no task redeploys. The org override is read fresh per attempt; the global default comes from the cached flags registry (a cold read fails safe to the public origin). When `INTERNAL_API_ORIGIN` is unset the flag is a no-op and no extra queries run, so existing deployments are unaffected. Dev runs always use the public origin, and `TRIGGER_STREAM_URL` remains unchanged.
Configuration menu - View commit details
-
Copy full SHA for 44eca4d - Browse repository at this point
Copy the full SHA 44eca4dView commit details -
Configuration menu - View commit details
-
Copy full SHA for 38bf82a - Browse repository at this point
Copy the full SHA 38bf82aView commit details -
docs(wait): separate compute billing from concurrency release (#4405)
The wait docs describe the 5 second compute-billing threshold as if it were also the suspension threshold. It isn't, and the gap is confusing when you're sizing a poll interval: - **Compute** stops being charged for any wait longer than 5 seconds. - **Concurrency** is only released once the machine has been snapshotted and shut down. For `wait.for` and `wait.until` that happens 60 seconds into the wait — a shorter wait stays `EXECUTING` and holds its concurrency slot for the whole wait, even though the compute is free. So `await wait.for({ seconds: 30 })` in a polling loop never releases its slot, which looks like a bug if the docs told you waits over 5 seconds checkpoint. ## Changes **`docs/snippets/paused-execution-free.mdx`** — rendered on `/wait`, `/wait-for` and `/wait-until`. Drops "we checkpoint and" from the billing sentence so it's purely about compute, then adds one paragraph for the concurrency half. **`docs/queue-concurrency.mdx`** — the "Waits and concurrency" section states flatly that waiting runs don't consume slots. Adds a short subsection for the time-based exception. **`docs/how-to-reduce-your-spend.mdx`** — "Waits longer than 5 seconds automatically checkpoint your task, meaning you don't pay for compute" → the compute claim only. Code comments follow, plus a pointer that waiting doesn't always free concurrency. **`docs/how-it-works.mdx`** — the Checkpoint-Resume walkthrough used `wait.for({ seconds: 30 })` as *the* example of a wait that suspends. Bumped to 5 minutes and noted the sub-60s exception. No behaviour change — docs only.Configuration menu - View commit details
-
Copy full SHA for 205bdc3 - Browse repository at this point
Copy the full SHA 205bdc3View commit details -
fix(webapp): restyle the leave and remove team member dialogs (#4411)
## Summary The confirmation dialog for leaving a team or removing a teammate was still built on the old `Alert` primitive: the entire question sat in the title, there was no header divider or `Esc` affordance, and the footer used small buttons pinned to the right. It now uses the standard `Dialog` layout the rest of the dashboard uses. The title is static ("Remove team member" / "Leave team"), the question moves into the body with the person's name and the organization highlighted, and the footer is a bordered row with medium Cancel and confirm buttons. A member who has not set a name is now identified by their email instead of "them". Verified against a local dashboard on both dialogs. Confirming a removal posts the member id, deletes the membership and shows the success toast. Cancel, `Esc`, and Enter while Cancel is focused all close the dialog without issuing a request, leaving the member in place. No release note needed: this is a visual restyle of an existing dialog with no behaviour change.Configuration menu - View commit details
-
Copy full SHA for 1e14e29 - Browse repository at this point
Copy the full SHA 1e14e29View commit details
Commits on Jul 29, 2026
-
fix(webapp): fade overflowing side menu selector labels (#4412)
Long organization, project, and environment names in the side menu were cut off mid-character. They now fade out at the right edge like the rest of the side menu items already did. ### Example of faded long names: <img width="246" height="200" alt="CleanShot 2026-07-28 at 22 59 24" src="https://github.com/user-attachments/assets/efa60b87-286f-4ab0-9d4e-490ef2de53e5" />
Configuration menu - View commit details
-
Copy full SHA for a11e5ff - Browse repository at this point
Copy the full SHA a11e5ffView commit details -
Configuration menu - View commit details
-
Copy full SHA for 15e160d - Browse repository at this point
Copy the full SHA 15e160dView commit details -
docs: give docs pages unique title tags and redirect stale pages (#4416)
## Summary Several docs pages rendered identical `<title>` tags, which weakens search indexing and makes results ambiguous. Each affected page now has a unique, descriptive title while keeping its existing sidebar label unchanged. Alongside the retitles: - Removed two stale build-system upgrade pages that were no longer in the navigation, with redirects to the current package upgrade guide. - Redirected the build-extensions group index to its overview page so the two URLs stop sharing a title. - Dropped a leftover orphaned API reference page (its old URL already redirects to the management overview). No links break: nothing in the docs points at the removed pages, and every redirect target exists.
Configuration menu - View commit details
-
Copy full SHA for 8d321f8 - Browse repository at this point
Copy the full SHA 8d321f8View commit details -
fix(webapp): don't apply an invite's role to an existing org member (#…
…4409) <!-- ccr-slack-attribution --> _Requested via [Slack thread](https://triggerdotdev.slack.com/archives/C097ZHVKZFA/p1785249693523749)_ ## Summary Accepting an old invitation could change the role of someone who was already in the organization. A long-pending invite can carry a lower role than the member has since been promoted to, so accepting it was a silent demotion. When the accepting user was the organization's only Owner, the role layer refused that demotion, and the refusal (an expected, protective outcome) was logged as an error. An invitation now only sets a role on a membership the accept actually created, and people who are already in an organization are skipped when invitations are sent. ## How `acceptInvite` already skipped the `OrgMember` create when it found an existing membership, but the `rbac.setUserRole` call below it was gated only on `invite.rbacRoleId`. It now also tracks whether this accept created the membership. A create that loses the unique-constraint race counts as pre-existing, since whichever flow won it owns that membership's role. Skipping existing members outright would regress one case: a member with no RBAC role at all would never receive the invitation's role. `ensureOrgMember` handles that with `healMissingRoleAssignment`, which fills in a null role but never overwrites a real one, so `assignInviteRbacRole` takes the same gate. An established role is never touched; an absent one is filled in. `assignInviteRbacRole` branches on the result's machine-readable `code` instead of logging every refusal at `error`. `last_owner` goes to `logger.info`, matching the two directory-sync role paths; everything else, including a refusal that carries no code, goes to `logger.warn`. The helper is best-effort and never throws, so no outcome it produces warrants `error`. No string matching on the error text is involved. `inviteMembers` resolves the organization's members by email and skips those addresses before creating invites. The invite table's `@@unique([organizationId, email])` only dedupes *pending invites*, so it could never catch this. ## Invite surfaces Skipping addresses means a batch can now come back empty, and neither caller handled that: - The dashboard action built its redirect from `invites[0].organization`, so a batch where every address was skipped threw a `TypeError` that reached the admin as a raw error string. It also reported the submitted count rather than the created one. It now names what it skipped ("No invitations sent: 1 already a member of this organization") and counts what it actually created. - The invites API derived `alreadyInvited` as "everything not created", so an existing member was reported as though they had already been invited. `inviteMembers` now returns the two groups separately and the endpoint reports `alreadyMembers` alongside `alreadyInvited`. ## Testing `apps/webapp/test/member.server.test.ts` passes 16/16 locally, up from 12. Getting there needed a harness fix. The `~/db.server` mock did not export `Prisma`, so any code reaching `PrismaNamespace.PrismaClientKnownRequestError` threw before it could branch, leaving every duplicate-key path in `member.server.ts` unreachable from tests. The mock now re-exports the real `Prisma`, and there is a case covering the pending-invite skip. New cases: the invite role is applied when the accept creates the membership; it is not applied when the member already has a role; it is applied when an existing member has no role assigned; the organization is still joined when the assignment is refused with `last_owner`; and `inviteMembers` reports members separately from pending invites. Forcing the gate off fails exactly the "already has a role" case, so the coverage is load-bearing. `pnpm run typecheck --filter webapp` and `oxfmt --check` both pass. ## Changelog Accepting an old invitation could change the role of someone who was already in the organization. An invitation now leaves an existing member's role untouched, people who are already in an organization are no longer sent invitations to it, and the invite form says which addresses it skipped instead of failing with an unhelpful error. --- ## ✅ Checklist - [x] I have followed every step in the [contributing guide](https://github.com/triggerdotdev/trigger.dev/blob/main/CONTRIBUTING.md) - [x] The PR title follows the convention. - [x] I ran and tested the code works. ## Screenshots No visual changes. The invite form's toast copy changes, as described above. --------- Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Matt Aitken <matt@mattaitken.com>
Configuration menu - View commit details
-
Copy full SHA for 639eaf6 - Browse repository at this point
Copy the full SHA 639eaf6View commit details -
feat(webapp,run-engine): queue metrics and health dashboard (#4131)
## Summary Three related changes, each independently gated: **Queue metrics and health.** Per-queue depth, throughput (enqueued, started, completed), concurrency, whether a queue is throttled, and scheduling delay (how long a run waits between becoming eligible and actually starting), plus a per concurrency-key breakdown for keyed queues. Collected from inside the run queue itself, stored in ClickHouse, and surfaced on the Queues list, a new per-queue detail page, the task pages, and the run inspector. The question it answers is "does this queue have enough concurrency to keep up, and if not, which key or which limit is the constraint". **Percent-based queue concurrency limits.** A queue's concurrency override can now be expressed as a percentage of the environment limit, stored as the source of truth and re-materialized whenever the environment limit changes. Absolute overrides above the environment limit are now **rejected with a 400** instead of being silently capped, which is a behavior change on `POST /api/v1/queues/:queue/concurrency/override`. **The `health` report.** A server-computed verdict on whether work is flowing, whether the runs that do start are healthy, and whether telemetry is fresh, rendered as text with sparklines. Available as `GET /api/v1/reports/:key`, `trigger report`, and the `get_report` MCP tool (plus a `report` MCP prompt, which shows up as a slash command in hosts that support prompts). With the flags off, the Queues page renders the pre-metrics component verbatim, nothing is emitted, and nothing is written to ClickHouse. ## Configuration Two independent gates, on purpose. Emission is global so data accrues for everyone before anyone can look at it; the view is per organization so it can be turned on for one org at a time without a deploy. **Runtime flags (no restart)** | Flag | Store | Gates | | --- | --- | --- | | `queue_metrics:enabled` | run-queue Redis key (`"1"`/`"0"`, off by default) | All emission, gauges and counters. Cached in-process for 10s with stale-while-revalidate, warmed eagerly at boot so the first op after a deploy is not dropped. | | `queue_metrics:gauge_sample_rate` | run-queue Redis key, `0..1` | Fraction of queue ops that emit a gauge. Counters are never sampled, so throughput stays exact at any rate. | | `queueMetricsUiEnabled` | feature-flag catalog: global `FeatureFlag` row, per-org `Organization.featureFlags` override wins | Whether an org sees the metrics view at all: the Queues list variant, the queue detail route, the built-in Queues dashboard, the concurrency-keys endpoint, and the metrics blocks on task pages and the run inspector. Off by default; a gated org gets a 404 on the detail route rather than an empty page. | Both Redis keys are readable and writable from `/admin/queue-metrics` (super-admin UI, with a live per-shard stream-health table) and `GET`/`POST /admin/api/v1/queue-metrics` (admin PAT). The admin surface uses its own Redis client, so it works on any instance regardless of whether that instance runs the emitter or the consumer. **Environment variables (boot time)** | Variable | Default | Notes | | --- | --- | --- | | `QUEUE_METRICS_EMIT_ENABLED` | `0` | Constructs the emitter and injects it into the run engine. Without it the run queue has no emitter at all. | | `QUEUE_METRICS_CONSUMER_ENABLED` | `0` | Boots the stream consumer on this instance. Independent of emission, so consumers can be sized separately from the API. | | `QUEUE_METRICS_STREAM_SHARD_COUNT` | `4` | Stream shards, hashed per queue. | | `QUEUE_METRICS_CONSUMER_BATCH_SIZE` | `1000` | Poll batch equals insert batch, so an ack can never outrun a write. | | `QUEUE_METRICS_REDIS_{HOST,PORT,USERNAME,PASSWORD,TLS_DISABLED}` | falls back to the run-queue Redis | Set `HOST` to move the metrics stream onto a dedicated instance so a metrics backlog cannot compete with the run queue for memory. Self-hosters can leave it unset and get a single-Redis deployment. | | `QUEUE_METRICS_COUNTER_STREAM_MAXLEN` | `2000000` shared, `8000000` dedicated | Bound on how much a stalled consumer can hold. The default is deliberately lower when the stream shares the queue-critical Redis. | | `QUEUE_METRICS_COUNTER_ODOMETER_TTL_SECONDS` | `604800` | TTL on the per-queue cumulative counter key, refreshed on every write, so only queues idle for the whole window are purged. | | `QUEUE_METRICS_MAX_QUEUE_NAMES_PER_ENV` | `1000` | Distinct queue names tracked per environment; overflow collapses into `__overflow__`. | | `QUEUE_METRICS_MAX_CONCURRENCY_KEYS_PER_QUEUE` | `10000` | Same idea one level down, per queue. | | `QUEUE_METRICS_GAUGE_SAMPLE_RATE` | `1` | Default for the live sample-rate key above. | | `QUEUE_METRICS_QUERY_TABLES_VISIBLE` | `0` | Lists the queue-metrics tables in the Query page, its schema docs, the schema API and the AI query context. Off keeps them unlisted while the feature is dark; a query naming them still runs either way. | | `QUEUE_METRICS_CLICKHOUSE_URL` | falls back to the shared wiring | Runs queue metrics on their own ClickHouse service: the consumer's inserts and every queue-metrics read go through it, so a metrics-heavy chart refresh never competes with runs-list or trace reads. Unset reproduces the previous split exactly (inserts on `CLICKHOUSE_URL`, reads on the query pool). | | `QUEUE_METRICS_CLICKHOUSE_READER_URL` | the write URL | Reader split, so the consumer's inserts can never land on a read endpoint. | | `QUEUE_METRICS_CLICKHOUSE_{KEEP_ALIVE_ENABLED,KEEP_ALIVE_IDLE_SOCKET_TTL_MS,MAX_OPEN_CONNECTIONS,LOG_LEVEL,COMPRESSION_REQUEST}` | `1`, unset, `10`, `info`, `1` | Pool tuning, matching the other per-workload ClickHouse clients. | Migrations to apply: ClickHouse `036_create_queue_metrics_v1.sql`, and a Postgres migration adding the nullable `TaskQueue.concurrencyLimitOverridePercent`. Both are additive. ## How collection works Queue operations produce two kinds of signal, and they have opposite failure modes, so they are handled differently. **Gauges** (queued, running, queue limit, env queued, env running, env limit, throttled, plus keys-with-backlog and worst-key wait on keyed queues) are read *inside* the same Redis script that performs the enqueue or dequeue, so the reading is atomic with the operation it describes rather than a racy follow-up read. The script returns them on its reply and the app forwards them to the stream. Gauges are sampled and drop-tolerant: they are aggregated with `max`, so a lost reading costs resolution, never correctness. **Counters** (enqueued, started, completed, plus nack and dead-lettered) are cumulative odometers. Each event increments a per-queue key on the metrics Redis and emits the absolute total, and ClickHouse takes the difference across buckets at read time. This is the important property of the design: a summed-delta counter undercounts permanently on any lost event, while a cumulative one self-heals, because the next surviving reading restates the whole total. Only bucket granularity can be lost, never the total. A queue returning after its odometer TTL expired restarts at 1 and reset detection handles it, which is safe precisely because expiry only spans a window with no activity. Both land on one sharded Redis stream. A consumer reads it with a consumer group, reclaims stale pending entries on a 15s interval rather than on every poll, maps one entry to one or two ClickHouse rows (whole-queue and, for keyed queues, per-key), and acks only after the insert lands. Each batch carries a dedup token derived from its stream-entry ids, and the target tables set `non_replicated_deduplication_window`, so a retried batch cannot double-count either the raw rows or the aggregates that hang off them. Consumer and emitter both emit OTel metrics (`queue_metrics.emitter.emitted`, `queue_metrics.consumer.{entries,rows_inserted,insert_errors,insert_duration,stream_depth,group_lag,pending,lag_unknown}`); stream depth and group lag are the two worth alerting on, and `lag_unknown` exists because Redis can report a null lag after a trim, which must not be read as zero. ## Storage and read path `queue_metrics_raw_v1` is a short landing table with a 6 hour TTL. Four aggregate tiers are materialized straight from raw, never cascaded off each other, each with a 30 day TTL: - `queue_metrics_v1`, 10 second buckets per queue, the default read path - `queue_metrics_5m_v1`, 5 minute buckets per queue, for wide ranges and cross-queue ranking - `env_metrics_v1`, 10 second buckets per environment, queue-independent so it stays cheap at any range - `queue_metrics_ck_v1`, 10 second buckets per concurrency key Every tier is an MV from raw because the counter states do not survive a cascade: their merge is order sensitive, so a `-MergeState` chain off the 10s table inflates the result, and the same property means an aggregate state may only be merged inside one queue. That constraint is now enforced by the query engine rather than by reviewer discipline: a column can declare a `mergeGroupKey`, and any query that references it without grouping by, or pinning to a single value of, every named key fails to compile with an actionable message. On the read side, TRQL gains three tables (`queue_metrics`, `env_metrics`, and a `queue_metrics_by_key` that is hidden from the editor, schema docs and schema API but still queryable, so per-key rows can never silently merge into a plain per-queue query), plus `deltaSumTimestampMerge` and `quantilesTDigestMerge`. Two schema-level optimizations ride along: a table can declare coarser rollups, so a query whose bucket interval is 5 minutes or wider is routed to the 5m table with no change to the query itself, and it can opt into the ClickHouse query cache with time bounds floored to a fixed grid, so the auto-refreshing dashboards actually share cache entries instead of missing on every tick. Both are caller-side substitutions, so the printer stays unaware of physical layout. All of this can also live on its own ClickHouse service. A table declares the pool its reads run on, the three queue-metrics tables name the dedicated one, and the ingestion consumer writes through the same client, so both directions move together with one env var and nothing else routes differently. The other engine change is opt-in gap filling: charts can request rows for empty buckets, where counters zero-fill and gauges carry forward. Grouped gauge series are densified per group and carried inside a partition, so a quiet queue's line holds its last value without bleeding another queue's value into it. ## Queue concurrency limits `concurrencyLimitOverridePercent` on `TaskQueue` is the source of truth when an override is set as a percentage; the absolute `concurrencyLimit` is materialized from it (floored, clamped to at least 1 so a percentage can never act as a pause, and never above the environment limit). Every path that changes an environment limit now recalculates the environment's percent-based overrides afterwards, outside the transaction, and pushes changed limits to the engine. The push is attempted even when the stored value did not change, so a previously failed sync self-heals rather than leaving the database and the engine diverged; paused queues are skipped so a recalculation cannot effectively unpause one. The API accepts exactly one of `concurrencyLimit` or `percent`, and the reject-instead-of-clamp change above means a request asking for more than the environment allows now fails loudly. The percent bound (greater than 0, at most 100) is defined once and shared by the zod schema, the dashboard mutation handler and the service, so the three cannot drift. The concurrency-keys table on a queue is now paginated against the ClickHouse per-key tier, ranked by peak backlog with the total on every row from a single scan, and only the keys on the current page are enriched with live counts from Redis. That replaces a hard top-50 cap with something whose cost is a function of page size rather than key cardinality. ## The health report `GET /api/v1/reports/:key?period=&format=markdown|ansi|json`. The verdict is computed on the server and is deterministic, not model-generated. Three independent analyzers run over one input snapshot: flow (is work moving, and if not, is the cause a limit, throttling, one bad queue, or dead-lettering), execution (are the runs that start succeeding, and at what latency), and liveness (how fresh is the telemetry). When telemetry is genuinely stale, the first two are forced to unknown and every actionable field is stripped, so no surface ever advises action off stale data. Authorization is per query table rather than a blanket query grant: a JWT must be scoped to every table the report reads (`runs`, `env_metrics`, `queue_metrics`), so a narrowly scoped token cannot pull a report that reads more than it was granted. `period` is validated as a shorthand with a 90 day ceiling at the edge. The report catalog is a registry of `{ load, interpret }` entries, so the next report is a new entry and no change to the route, the view model, the renderers, the CLI or the MCP tool. `trigger mcp` no longer launches the install wizard when stdout is a TTY, which fixed a real failure: hosts spawn the server over a PTY, so the wizard would open and the client would time out waiting for a server that never started. The wizard now needs `trigger mcp --install`. ## The part that is live regardless of every flag The enqueue and dequeue scripts now return a 2-tuple so a gauge reading can ride back on the reply. Every return site in the eight affected scripts is wrapped, and a `nil` original is converted to `false` on the way out, because a raw `nil` in the first slot would make Lua truncate the multi-bulk reply and silently drop the gauge on the throttled and empty-queue paths. The reply shape and the destructuring on the app side are exercised on every queue operation whether or not metrics are enabled, so that is the part of `run-engine` worth the closest review. One behavior fix in the same area: the scheduling-delay anchor is set only on a run's first entry into the queue. Anchoring it to trigger time on re-enqueues made waitpoint and checkpoint resumes report the entire wait as scheduling delay. Queue ordering is untouched, so a re-enqueued run keeps its position, and nacks deliberately keep the original anchor because a rolled-back dequeue is the same continuous wait. A pending-version promotion still anchors to trigger time, on purpose: that promotion is the run's first real entry into the queue, since the trigger deliberately held it back waiting for a worker version, and the TTL is armed at the same point for the same reason. The consequence is worth naming, because it is a judgement call: a run that waits on a deployment reports that wait as scheduling delay on its queue, which is time unrelated to queue capacity. ## Verification Unit and integration suites across the new package, the run queue, the mapping layer, the query engine and ClickHouse (including a test that applies migration 036 through the same splitter CI uses, and a regression test that inserts the same batch three times to prove the aggregates do not inflate). Beyond that, the whole path was driven end to end against a live stack with real runs: emitter to Redis stream to consumer to ClickHouse to the dashboards, for both the local dev path and the deployed path where a supervisor drives the dequeue, with assertions on exact counter reconstruction per queue and per concurrency key, throttling, environment saturation, scheduling delay, and a deliberate mid-stream reading drop to confirm the cumulative counters still reconstruct the correct total. The gated-off state was checked on every touched surface. The dedicated ClickHouse service was verified against a second, separately-schema'd instance: with it configured, the driven counters reconstruct exactly on the dedicated instance, the shared instance gains no rows for that window, a read through the query API returns the value that exists only on the dedicated instance, and a `runs` query still succeeds (it would fail outright if it were mis-routed to a service without that table). With the variable unset, the full suite passes unchanged. --------- Co-authored-by: Katia Bulatova <katia@trigger.dev> Co-authored-by: Katia Bulatova <katherine.bulatova@gmail.com> Co-authored-by: James Ritchie <james@trigger.dev>Configuration menu - View commit details
-
Copy full SHA for 4eb9292 - Browse repository at this point
Copy the full SHA 4eb9292View commit details -
feat(database,rbac): add multiple environment API key foundations (#4388
) Adds the storage model and authorization contracts needed for multiple environment API keys. Credentials are represented by hashed values, revocation and expiration state, and persisted effective scopes. The built-in authorization fallback exposes full-access policy preparation, while optional authorization extensions can supply additional presets and task-aware scope generation. This change does not create, display, or authenticate additional keys.
Configuration menu - View commit details
-
Copy full SHA for a81ad49 - Browse repository at this point
Copy the full SHA a81ad49View commit details -
Configuration menu - View commit details
-
Copy full SHA for 878c158 - Browse repository at this point
Copy the full SHA 878c158View commit details -
Configuration menu - View commit details
-
Copy full SHA for ed8f5e1 - Browse repository at this point
Copy the full SHA ed8f5e1View commit details -
Configuration menu - View commit details
-
Copy full SHA for a098171 - Browse repository at this point
Copy the full SHA a098171View commit details -
Configuration menu - View commit details
-
Copy full SHA for 8ebc8a4 - Browse repository at this point
Copy the full SHA 8ebc8a4View commit details -
Configuration menu - View commit details
-
Copy full SHA for 2f1734c - Browse repository at this point
Copy the full SHA 2f1734cView commit details
Commits on Jul 30, 2026
-
fix(webapp,clickhouse): stop invalid customer queries alerting, and i…
…solate Sentry scope per request (#4372) ## Summary A query sent to the query API with a typo in it, like a column name that does not exist, was being reported as a server error. That put customer SQL mistakes into our error alerting, where they made up almost all of the volume on one of our noisiest alerts, and it drowned out the failures that are actually ours to fix. This makes the level match who is at fault, and fixes two related problems found alongside it. ## Invalid queries are the caller's, not ours The query API route already got this right. It checks for `QueryError`, logs at warn, and returns a 400, with a comment saying the system handles it gracefully and no alert is needed. The layer underneath ignored that. `executeTSQL` logged every exception out of its catch block at error, including the compile failures the route was about to turn into a 400, and error-level logs are forwarded to error reporting. The TSQL package already draws the line we need: ```ts export class ExposedTSQLError extends BaseTSQLError { /** An exception that can be exposed to the user. */ } export class InternalTSQLError extends BaseTSQLError { /** An internal exception in the TSQL engine. */ } ``` `SyntaxError` and `QueryError` extend the first. So the catch block now branches on `ExposedTSQLError` and logs those at warn, keeping error for `InternalTSQLError` and anything unanticipated, which is a genuine compiler bug. ## SQL the caller wrote is their mistake, not ours The same asymmetry showed up one level down. A query that compiles fine can still be rejected by ClickHouse at execution, and most of those rejections mean the caller's SQL is wrong rather than that we generated something bad. This is where the volume actually is. Checking production, one error group alone, a missing `GROUP BY` on the public query API (`NOT_AN_AGGREGATE`), accounts for over a million events across hundreds of users. It is by far the largest error group in the project, and classifying only by resource limit would have left every one of those at error level. So rejections are split three ways in `ClickhouseClient`, which is the only place holding the parsed `ClickHouseError` and its symbolic type. By the time the error reaches `executeTSQL` it has been wrapped and the type is gone, and the type never appears in the message text, so it cannot be recovered by string matching. - **Resource limits** (memory ceiling, timeout, row/byte caps) log at warn. The query is valid, it just asked for more than it is allowed to spend. - **Invalid SQL** (`NOT_AN_AGGREGATE`, `UNKNOWN_IDENTIFIER`, `SYNTAX_ERROR`, the type and parse families) logs at warn **only when the caller wrote the SQL**. - **Everything else** keeps alerting. That gate matters. The client is shared, so the identical rejection on TRQL *we* generated is our bug and has to stay at error. Callers opt in with `userAuthoredQuery`: | caller | who wrote the SQL | opts in | | --- | --- | --- | | public query API | the customer | yes | | query editor | the customer | yes | | agent charts | the agent's model | yes | | built-in dashboard tiles | us, in code | no | | queue metric cards | us, in code | no | | health report | us, in code | no | The agent is the one judgement call. Its TRQL is not typed by a person, but it is also not something a code fix makes correct, so a query it gets wrong is not worth waking anyone for. The same endpoint serves built-in tiles whose TRQL we do write, so the opt-in lives with the caller rather than the route. Separately, when one of these queries did fail, the log recorded the generated ClickHouse SQL but not the query the caller actually wrote, which made the reports hard to act on. `queryWithStats` takes an optional `logFields` that `executeTSQL` uses to attach the original TSQL. ## Events were attributed to the wrong request Chasing the above turned up something broader: only a tenth of the events on that alert pointed at the query API. The rest were pinned to unrelated requests that happened to be in flight at the same time, so the alert looked like the trigger endpoint was failing. `Sentry.init` runs with `skipOpenTelemetrySetup: true`, because we register our own OTel pipeline. That skips `initOpenTelemetry`, and one of the things it does is: ```js api.context.setGlobalContextManager(new SentryContextManager()); ``` The async-context strategy is still installed, but `withIsolationScope` only marks the OTel context and delegates the actual fork to that context manager: ```js // "We depend on the otelContextManager to handle the context/hub" return api.context.with(ctx.setValue(SENTRY_FORK_ISOLATION_SCOPE_CONTEXT_KEY, true), ...) ``` `provider.register()` installed a plain `AsyncLocalStorageContextManager`, which does not know that key. The lookup found no scopes on the context and fell back to the process-global default isolation scope, so every request wrote its request data into the same object and the last writer won. The tracer now registers `SentryContextManager`, which subclasses `AsyncLocalStorageContextManager`, so OTel behaviour is unchanged. It is also registered on the path where tracing is disabled, which previously never called `register()` at all and so had no context manager of its own. Tenant tags were always correct, because those come from our own async local storage rather than the isolation scope. That is why the attribution being wrong was not obvious. This affects every error report the webapp sends, not just the query API. ## Verification `internal-packages/clickhouse`: 76 tests pass, including eight covering each level decision against a real ClickHouse container. Three pairs pin the gate open and shut at both layers: an invalid query, a compile failure, and a real limit breach driven with `max_rows_to_read` each log at warn with `userAuthoredQuery` and at error without it. The isolation fix has a test that reproduces the leak before asserting the fix. Two overlapping requests each tag their own isolation scope; with the plain context manager the slower one reads back the other's tag, and with `SentryContextManager` each reads back its own. Measured separately against a faithful reproduction of the server's wiring (own OTel pipeline, CommonJS entry) at 200 concurrent requests: per-request attribution goes from 0.5% to 100%, while span nesting, context propagation across awaits, and distinct trace IDs are identical before and after.
Configuration menu - View commit details
-
Copy full SHA for 6e5f0f0 - Browse repository at this point
Copy the full SHA 6e5f0f0View commit details -
Configuration menu - View commit details
-
Copy full SHA for 86b948b - Browse repository at this point
Copy the full SHA 86b948bView commit details
This comparison is taking too long to generate.
Unfortunately it looks like we can’t render this comparison for you right now. It might be too big, or there might be something weird with your repository.
You can try running this command locally to see the comparison on your machine:
git diff v4.5.8...v4.5.9