Skip to content
Permalink

Comparing changes

Choose two branches to see what’s changed or to start a new pull request. If you need to, you can also or learn more about diff comparisons.

Open a pull request

Create a new pull request by comparing changes across two branches. If you need to, you can also . Learn more about diff comparisons here.
base repository: triggerdotdev/trigger.dev
Failed to load repositories. Confirm that selected base ref is valid, then try again.
Loading
base: v4.5.8
Choose a base ref
...
head repository: triggerdotdev/trigger.dev
Failed to load repositories. Confirm that selected head ref is valid, then try again.
Loading
compare: v4.5.9
Choose a head ref
  • 20 commits
  • 245 files changed
  • 11 contributors

Commits on Jul 27, 2026

  1. perf(run-ops-database): index BatchTaskRun for the batches list on th…

    …e dedicated schema (#4396)
    
    ## Summary
    
    The batches list page orders by `(createdAt DESC, id DESC)`, which is
    why [#4361](#4361)
    added a matching index on `BatchTaskRun`. That index only landed in
    `@trigger.dev/database`.
    
    The dedicated run-ops database has its own migration history, so it
    never received the index. `BatchListPresenter` reads both databases and
    merges, so for environments whose batches live in the dedicated database
    the page kept falling back to a scan and in-memory sort, which is the
    exact behaviour #4361 set out to fix.
    
    ## Fix
    
    Adds the index to the run-ops schema with its own migration. `CREATE
    INDEX CONCURRENTLY IF NOT EXISTS`, so it is a no-op where the index
    already exists and still records its ledger row.
    
    The second half is the interesting part. Because the two packages own
    separate migration histories, a run-graph schema change has to be
    authored twice, and nothing made the miss visible: the run-ops status
    check truthfully reports "up to date" against its own history, so the
    apply step just skips.
    
    `schemaParity.test.ts` compares the physical shape of every model the
    run-ops schema declares against its counterpart in
    `@trigger.dev/database`: scalar fields with their attributes, plus
    `@@index`, `@@unique`, `@@id` and `@@map`. Relation navigation fields
    are excluded, since the run-ops schema deliberately drops relations that
    would cross a database boundary while keeping the scalar FK column. A
    field counts as a relation when its type resolves to a model name, which
    keeps enum-typed columns in scope.
    
    Two models are listed as run-ops-only: `CompletedWaitpoint` and
    `WaitpointRunConnection`, both explicit FK-free replacements for a
    control-plane implicit many-to-many, since an implicit m2m carries a
    foreign key that cannot resolve across databases. The test also asserts
    that exception list is exhaustive, so a new unpaired model fails rather
    than being silently skipped.
    
    Confirmed the guard actually fails: reverting the index turns
    `BatchTaskRun` red with the missing `@@index` named in the diff.
    ericallam authored Jul 27, 2026
    Configuration menu
    Copy the full SHA
    fc57643 View commit details
    Browse the repository at this point in the history
  2. ci: let the claude bot trigger the PR audit workflows (#4392)

    ## Summary
    
    PRs opened by the claude GitHub app fail both the agent instructions
    audit and the REVIEW.md drift audit before Claude gets a chance to run.
    `claude-code-action` refuses any actor whose account type is not `User`
    unless the actor is listed in `allowed_bots`:
    
    ```
    Workflow initiated by non-human actor: claude (type: Bot).
    Add bot to allowed_bots list or use '*' to allow all bots.
    ```
    
    So those PRs land with two permanently red checks and no audit coverage
    at all. Both workflows already allowlist Devin; this adds the claude app
    alongside it.
    
    ## Why this does not open the workflows up to outside contributors
    
    `allowed_bots` is only consulted for non-`User` actors. Humans,
    contributor or maintainer, take the separate write-permission path and
    are unaffected by what is in the list.
    
    Beyond that, both jobs are guarded by
    `github.event.pull_request.head.repo.full_name == github.repository`, so
    a fork PR skips the job entirely, and they trigger on `pull_request`
    rather than `pull_request_target`, so a fork-triggered run would get no
    API key and a read-only token anyway.
    
    The bot is named explicitly instead of using `"*"`, which would let
    every bot trigger these audits, dependabot's PR stream included.
    matt-aitken authored Jul 27, 2026
    Configuration menu
    Copy the full SHA
    3ed4851 View commit details
    Browse the repository at this point in the history

Commits on Jul 28, 2026

  1. fix(webapp): remove unused Electric sync trace routes (#4400)

    <!-- ccr-slack-attribution -->
    _Requested by **Eric Allam** · [Slack
    thread](https://triggerdotdev.slack.com/archives/C0AU83M3136/p1785222101937829?thread_ts=1785207509.304669&cid=C0AU83M3136)_
    
    Removes two dead Remix routes and the helpers only they used.
    
    `app/routes/sync.traces.runs.$traceId.ts` (`/sync/traces/runs/:traceId`)
    and `app/routes/sync.traces.$traceId.ts` (`/sync/traces/:traceId`) were
    added with the original ElectricSQL run page and lost their only
    consumers when the dashboard hooks that called them were deleted.
    Nothing in the repo references either route today.
    
    Also removed, because the deleted routes were their only callers:
    
    - `OtelTraceIdSchema`, `RESERVED_ELECTRIC_SHAPE_PARAMS`, `TraceScope`,
    `buildElectricTraceWhereClause` from `app/v3/electricShape.server.ts`
    (the file stays — `UNSAFE_REALTIME_TAG_CHARS` /
    `sanitizeRealtimeTagForSql` / `sanitizeRealtimeTagsForSql` are still
    used by `realtime.v1.runs.ts` and `realtimeClient.server.ts`)
    - the loader-specific cases in
    `apps/webapp/test/spanTraceRoutes.replicaLag.test.ts` and
    `internal-packages/run-store/src/runOpsStore.routesSpanTraceReadView.replicaLag.test.ts`
    
    `app/utils/longPollingFetch.ts` is untouched —
    `realtimeClient.server.ts` still uses it. `runOpsStore.ts` /
    `PostgresRunStore.ts` are untouched too; the unrouted-lookup mechanism
    there is generic and stays.
    
    As a plain code fact: the run lookup these loaders performed keyed on
    `TaskRun.traceId` alone, which is not an index-backed query shape. That
    is noted only as context for why the code is not worth keeping around
    unused.
    
    ### Judgement call worth a maintainer's opinion
    
    The request was specifically about `/sync/traces/runs/:traceId`, the
    route that looks up a run by `traceId`. This PR **also** deletes its
    sibling `/sync/traces/:traceId`. The reasoning:
    
    - both routes came in with the same ElectricSQL run-page work
    - both lost their only consumers in the same later commit
    - neither has any caller anywhere in the repo
    - they share the same helper module, so keeping one means keeping the
    helpers half-used
    
    If you would rather keep the sibling, reverting just that one file
    deletion is easy and does not affect the rest of this PR — say the word
    and I will restore it along with the helpers it needs.
    
    ## ✅ Checklist
    
    - [x] I have followed every step in the [contributing
    guide](https://github.com/triggerdotdev/trigger.dev/blob/main/CONTRIBUTING.md)
    - [x] The PR title follows the convention.
    - [x] I ran and tested the code works
    
    ---
    
    ## Testing
    
    Verification run locally from the repo root:
    
    | Command | Result |
    | --- | --- |
    | `pnpm run format` | clean, no changes produced |
    | `pnpm run lint:fix` | clean |
    | `pnpm run lint` | pass (exit 0, no findings) |
    | `pnpm run typecheck --filter webapp` | pass |
    | `pnpm run typecheck --filter @internal/run-store` | pass |
    
    A ripgrep sweep for `sync.traces`, `sync/traces`, `syncTraceRunsLoader`,
    `buildElectricTraceWhereClause`, `OtelTraceIdSchema` and
    `RESERVED_ELECTRIC_SHAPE_PARAMS` (excluding `node_modules`) returns zero
    hits.
    
    **Not fully verified:** both edited test files are testcontainers suites
    and need a Docker runtime, which was not available in my environment. I
    confirmed each file *collects* correctly with exactly the three intended
    remaining tests and no import errors — notably, dropping the
    `session.server` / `controlPlaneResolver.server` / `longPollingFetch` /
    `env.server` mocks does not break module loading for the surviving
    loaders. The assertions themselves then failed only on `Could not find a
    working container runtime strategy`. CI should be the real signal here.
    
    Per `apps/webapp/CLAUDE.md`, `pnpm run build --filter webapp` was
    deliberately not run.
    
    ---
    
    ## Changelog
    
    Removed two unused sync routes left over from the original ElectricSQL
    run page, along with the helpers and tests that existed only to serve
    them. No behaviour change — neither route had any caller.
    
    ---
    
    ## Screenshots
    
    _n/a — no user-visible surface changes._
    
    💯
    
    ---------
    
    Co-authored-by: Claude <noreply@anthropic.com>
    claude[bot] and claude authored Jul 28, 2026
    Configuration menu
    Copy the full SHA
    ec562c0 View commit details
    Browse the repository at this point in the history
  2. feat(webapp): org-gated internal API origin in run env vars (#4366)

    Adds an opt-in way for operators to route deployed runs' API traffic
    through a different origin than the public one, per organization. Set
    `INTERNAL_API_ORIGIN` on the webapp and enable the
    `internalApiOriginEnabled` feature flag (globally or per org, with the
    org override winning in both directions): deployed runs for enabled orgs
    then get `TRIGGER_API_URL` set to the internal origin instead of
    `API_ORIGIN`. Useful for gradually moving run traffic onto a private
    network path.
    
    ## Design
    
    The origin is resolved when an attempt starts, so flag changes take
    effect on the next attempt and roll back the same way, with no task
    redeploys. The org override is read fresh per attempt; the global
    default comes from the cached flags registry (a cold read fails safe to
    the public origin). When `INTERNAL_API_ORIGIN` is unset the flag is a
    no-op and no extra queries run, so existing deployments are unaffected.
    Dev runs always use the public origin, and `TRIGGER_STREAM_URL` remains
    unchanged.
    myftija authored Jul 28, 2026
    Configuration menu
    Copy the full SHA
    44eca4d View commit details
    Browse the repository at this point in the history
  3. Configuration menu
    Copy the full SHA
    38bf82a View commit details
    Browse the repository at this point in the history
  4. docs(wait): separate compute billing from concurrency release (#4405)

    The wait docs describe the 5 second compute-billing threshold as if it
    were also the suspension threshold. It isn't, and the gap is confusing
    when you're sizing a poll interval:
    
    - **Compute** stops being charged for any wait longer than 5 seconds.
    - **Concurrency** is only released once the machine has been snapshotted
    and shut down. For `wait.for` and `wait.until` that happens 60 seconds
    into the wait — a shorter wait stays `EXECUTING` and holds its
    concurrency slot for the whole wait, even though the compute is free.
    
    So `await wait.for({ seconds: 30 })` in a polling loop never releases
    its slot, which looks like a bug if the docs told you waits over 5
    seconds checkpoint.
    
    ## Changes
    
    **`docs/snippets/paused-execution-free.mdx`** — rendered on `/wait`,
    `/wait-for` and `/wait-until`. Drops "we checkpoint and" from the
    billing sentence so it's purely about compute, then adds one paragraph
    for the concurrency half.
    
    **`docs/queue-concurrency.mdx`** — the "Waits and concurrency" section
    states flatly that waiting runs don't consume slots. Adds a short
    subsection for the time-based exception.
    
    **`docs/how-to-reduce-your-spend.mdx`** — "Waits longer than 5 seconds
    automatically checkpoint your task, meaning you don't pay for compute" →
    the compute claim only. Code comments follow, plus a pointer that
    waiting doesn't always free concurrency.
    
    **`docs/how-it-works.mdx`** — the Checkpoint-Resume walkthrough used
    `wait.for({ seconds: 30 })` as *the* example of a wait that suspends.
    Bumped to 5 minutes and noted the sub-60s exception.
    
    No behaviour change — docs only.
    matt-aitken authored Jul 28, 2026
    Configuration menu
    Copy the full SHA
    205bdc3 View commit details
    Browse the repository at this point in the history
  5. fix(webapp): restyle the leave and remove team member dialogs (#4411)

    ## Summary
    
    The confirmation dialog for leaving a team or removing a teammate was
    still built on the old `Alert` primitive: the entire question sat in the
    title, there was no header divider or `Esc` affordance, and the footer
    used small buttons pinned to the right.
    
    It now uses the standard `Dialog` layout the rest of the dashboard uses.
    The title is static ("Remove team member" / "Leave team"), the question
    moves into the body with the person's name and the organization
    highlighted, and the footer is a bordered row with medium Cancel and
    confirm buttons. A member who has not set a name is now identified by
    their email instead of "them".
    
    Verified against a local dashboard on both dialogs. Confirming a removal
    posts the member id, deletes the membership and shows the success toast.
    Cancel, `Esc`, and Enter while Cancel is focused all close the dialog
    without issuing a request, leaving the member in place.
    
    No release note needed: this is a visual restyle of an existing dialog
    with no behaviour change.
    samejr authored Jul 28, 2026
    Configuration menu
    Copy the full SHA
    1e14e29 View commit details
    Browse the repository at this point in the history

Commits on Jul 29, 2026

  1. fix(webapp): fade overflowing side menu selector labels (#4412)

    Long organization, project, and environment names in the side menu were
    cut off mid-character. They now fade out at the right edge like the rest
    of the side menu items already did.
    
    ### Example of faded long names:
    <img width="246" height="200" alt="CleanShot 2026-07-28 at 22 59 24"
    src="https://github.com/user-attachments/assets/efa60b87-286f-4ab0-9d4e-490ef2de53e5"
    />
    samejr authored Jul 29, 2026
    Configuration menu
    Copy the full SHA
    a11e5ff View commit details
    Browse the repository at this point in the history
  2. Configuration menu
    Copy the full SHA
    15e160d View commit details
    Browse the repository at this point in the history
  3. docs: give docs pages unique title tags and redirect stale pages (#4416)

    ## Summary
    
    Several docs pages rendered identical `<title>` tags, which weakens
    search indexing and makes results ambiguous. Each affected page now has
    a unique, descriptive title while keeping its existing sidebar label
    unchanged.
    
    Alongside the retitles:
    
    - Removed two stale build-system upgrade pages that were no longer in
    the navigation, with redirects to the current package upgrade guide.
    - Redirected the build-extensions group index to its overview page so
    the two URLs stop sharing a title.
    - Dropped a leftover orphaned API reference page (its old URL already
    redirects to the management overview).
    
    No links break: nothing in the docs points at the removed pages, and
    every redirect target exists.
    D-K-P authored Jul 29, 2026
    Configuration menu
    Copy the full SHA
    8d321f8 View commit details
    Browse the repository at this point in the history
  4. fix(webapp): don't apply an invite's role to an existing org member (#…

    …4409)
    
    <!-- ccr-slack-attribution -->
    _Requested via [Slack
    thread](https://triggerdotdev.slack.com/archives/C097ZHVKZFA/p1785249693523749)_
    
    ## Summary
    
    Accepting an old invitation could change the role of someone who was
    already in the organization. A long-pending invite can carry a lower
    role than the member has since been promoted to, so accepting it was a
    silent demotion. When the accepting user was the organization's only
    Owner, the role layer refused that demotion, and the refusal (an
    expected, protective outcome) was logged as an error.
    
    An invitation now only sets a role on a membership the accept actually
    created, and people who are already in an organization are skipped when
    invitations are sent.
    
    ## How
    
    `acceptInvite` already skipped the `OrgMember` create when it found an
    existing membership, but the `rbac.setUserRole` call below it was gated
    only on `invite.rbacRoleId`. It now also tracks whether this accept
    created the membership. A create that loses the unique-constraint race
    counts as pre-existing, since whichever flow won it owns that
    membership's role.
    
    Skipping existing members outright would regress one case: a member with
    no RBAC role at all would never receive the invitation's role.
    `ensureOrgMember` handles that with `healMissingRoleAssignment`, which
    fills in a null role but never overwrites a real one, so
    `assignInviteRbacRole` takes the same gate. An established role is never
    touched; an absent one is filled in.
    
    `assignInviteRbacRole` branches on the result's machine-readable `code`
    instead of logging every refusal at `error`. `last_owner` goes to
    `logger.info`, matching the two directory-sync role paths; everything
    else, including a refusal that carries no code, goes to `logger.warn`.
    The helper is best-effort and never throws, so no outcome it produces
    warrants `error`. No string matching on the error text is involved.
    
    `inviteMembers` resolves the organization's members by email and skips
    those addresses before creating invites. The invite table's
    `@@unique([organizationId, email])` only dedupes *pending invites*, so
    it could never catch this.
    
    ## Invite surfaces
    
    Skipping addresses means a batch can now come back empty, and neither
    caller handled that:
    
    - The dashboard action built its redirect from
    `invites[0].organization`, so a batch where every address was skipped
    threw a `TypeError` that reached the admin as a raw error string. It
    also reported the submitted count rather than the created one. It now
    names what it skipped ("No invitations sent: 1 already a member of this
    organization") and counts what it actually created.
    - The invites API derived `alreadyInvited` as "everything not created",
    so an existing member was reported as though they had already been
    invited. `inviteMembers` now returns the two groups separately and the
    endpoint reports `alreadyMembers` alongside `alreadyInvited`.
    
    ## Testing
    
    `apps/webapp/test/member.server.test.ts` passes 16/16 locally, up from
    12.
    
    Getting there needed a harness fix. The `~/db.server` mock did not
    export `Prisma`, so any code reaching
    `PrismaNamespace.PrismaClientKnownRequestError` threw before it could
    branch, leaving every duplicate-key path in `member.server.ts`
    unreachable from tests. The mock now re-exports the real `Prisma`, and
    there is a case covering the pending-invite skip.
    
    New cases: the invite role is applied when the accept creates the
    membership; it is not applied when the member already has a role; it is
    applied when an existing member has no role assigned; the organization
    is still joined when the assignment is refused with `last_owner`; and
    `inviteMembers` reports members separately from pending invites. Forcing
    the gate off fails exactly the "already has a role" case, so the
    coverage is load-bearing.
    
    `pnpm run typecheck --filter webapp` and `oxfmt --check` both pass.
    
    ## Changelog
    
    Accepting an old invitation could change the role of someone who was
    already in the organization. An invitation now leaves an existing
    member's role untouched, people who are already in an organization are
    no longer sent invitations to it, and the invite form says which
    addresses it skipped instead of failing with an unhelpful error.
    
    ---
    
    ## ✅ Checklist
    
    - [x] I have followed every step in the [contributing
    guide](https://github.com/triggerdotdev/trigger.dev/blob/main/CONTRIBUTING.md)
    - [x] The PR title follows the convention.
    - [x] I ran and tested the code works.
    
    ## Screenshots
    
    No visual changes. The invite form's toast copy changes, as described
    above.
    
    ---------
    
    Co-authored-by: Claude <noreply@anthropic.com>
    Co-authored-by: Matt Aitken <matt@mattaitken.com>
    3 people authored Jul 29, 2026
    Configuration menu
    Copy the full SHA
    639eaf6 View commit details
    Browse the repository at this point in the history
  5. feat(webapp,run-engine): queue metrics and health dashboard (#4131)

    ## Summary
    
    Three related changes, each independently gated:
    
    **Queue metrics and health.** Per-queue depth, throughput (enqueued,
    started, completed), concurrency, whether a queue is throttled, and
    scheduling delay (how long a run waits between becoming eligible and
    actually starting), plus a per concurrency-key breakdown for keyed
    queues. Collected from inside the run queue itself, stored in
    ClickHouse, and surfaced on the Queues list, a new per-queue detail
    page, the task pages, and the run inspector. The question it answers is
    "does this queue have enough concurrency to keep up, and if not, which
    key or which limit is the constraint".
    
    **Percent-based queue concurrency limits.** A queue's concurrency
    override can now be expressed as a percentage of the environment limit,
    stored as the source of truth and re-materialized whenever the
    environment limit changes. Absolute overrides above the environment
    limit are now **rejected with a 400** instead of being silently capped,
    which is a behavior change on `POST
    /api/v1/queues/:queue/concurrency/override`.
    
    **The `health` report.** A server-computed verdict on whether work is
    flowing, whether the runs that do start are healthy, and whether
    telemetry is fresh, rendered as text with sparklines. Available as `GET
    /api/v1/reports/:key`, `trigger report`, and the `get_report` MCP tool
    (plus a `report` MCP prompt, which shows up as a slash command in hosts
    that support prompts).
    
    With the flags off, the Queues page renders the pre-metrics component
    verbatim, nothing is emitted, and nothing is written to ClickHouse.
    
    ## Configuration
    
    Two independent gates, on purpose. Emission is global so data accrues
    for everyone before anyone can look at it; the view is per organization
    so it can be turned on for one org at a time without a deploy.
    
    **Runtime flags (no restart)**
    
    | Flag | Store | Gates |
    | --- | --- | --- |
    | `queue_metrics:enabled` | run-queue Redis key (`"1"`/`"0"`, off by
    default) | All emission, gauges and counters. Cached in-process for 10s
    with stale-while-revalidate, warmed eagerly at boot so the first op
    after a deploy is not dropped. |
    | `queue_metrics:gauge_sample_rate` | run-queue Redis key, `0..1` |
    Fraction of queue ops that emit a gauge. Counters are never sampled, so
    throughput stays exact at any rate. |
    | `queueMetricsUiEnabled` | feature-flag catalog: global `FeatureFlag`
    row, per-org `Organization.featureFlags` override wins | Whether an org
    sees the metrics view at all: the Queues list variant, the queue detail
    route, the built-in Queues dashboard, the concurrency-keys endpoint, and
    the metrics blocks on task pages and the run inspector. Off by default;
    a gated org gets a 404 on the detail route rather than an empty page. |
    
    Both Redis keys are readable and writable from `/admin/queue-metrics`
    (super-admin UI, with a live per-shard stream-health table) and
    `GET`/`POST /admin/api/v1/queue-metrics` (admin PAT). The admin surface
    uses its own Redis client, so it works on any instance regardless of
    whether that instance runs the emitter or the consumer.
    
    **Environment variables (boot time)**
    
    | Variable | Default | Notes |
    | --- | --- | --- |
    | `QUEUE_METRICS_EMIT_ENABLED` | `0` | Constructs the emitter and
    injects it into the run engine. Without it the run queue has no emitter
    at all. |
    | `QUEUE_METRICS_CONSUMER_ENABLED` | `0` | Boots the stream consumer on
    this instance. Independent of emission, so consumers can be sized
    separately from the API. |
    | `QUEUE_METRICS_STREAM_SHARD_COUNT` | `4` | Stream shards, hashed per
    queue. |
    | `QUEUE_METRICS_CONSUMER_BATCH_SIZE` | `1000` | Poll batch equals
    insert batch, so an ack can never outrun a write. |
    | `QUEUE_METRICS_REDIS_{HOST,PORT,USERNAME,PASSWORD,TLS_DISABLED}` |
    falls back to the run-queue Redis | Set `HOST` to move the metrics
    stream onto a dedicated instance so a metrics backlog cannot compete
    with the run queue for memory. Self-hosters can leave it unset and get a
    single-Redis deployment. |
    | `QUEUE_METRICS_COUNTER_STREAM_MAXLEN` | `2000000` shared, `8000000`
    dedicated | Bound on how much a stalled consumer can hold. The default
    is deliberately lower when the stream shares the queue-critical Redis. |
    | `QUEUE_METRICS_COUNTER_ODOMETER_TTL_SECONDS` | `604800` | TTL on the
    per-queue cumulative counter key, refreshed on every write, so only
    queues idle for the whole window are purged. |
    | `QUEUE_METRICS_MAX_QUEUE_NAMES_PER_ENV` | `1000` | Distinct queue
    names tracked per environment; overflow collapses into `__overflow__`. |
    | `QUEUE_METRICS_MAX_CONCURRENCY_KEYS_PER_QUEUE` | `10000` | Same idea
    one level down, per queue. |
    | `QUEUE_METRICS_GAUGE_SAMPLE_RATE` | `1` | Default for the live
    sample-rate key above. |
    | `QUEUE_METRICS_QUERY_TABLES_VISIBLE` | `0` | Lists the queue-metrics
    tables in the Query page, its schema docs, the schema API and the AI
    query context. Off keeps them unlisted while the feature is dark; a
    query naming them still runs either way. |
    | `QUEUE_METRICS_CLICKHOUSE_URL` | falls back to the shared wiring |
    Runs queue metrics on their own ClickHouse service: the consumer's
    inserts and every queue-metrics read go through it, so a metrics-heavy
    chart refresh never competes with runs-list or trace reads. Unset
    reproduces the previous split exactly (inserts on `CLICKHOUSE_URL`,
    reads on the query pool). |
    | `QUEUE_METRICS_CLICKHOUSE_READER_URL` | the write URL | Reader split,
    so the consumer's inserts can never land on a read endpoint. |
    |
    `QUEUE_METRICS_CLICKHOUSE_{KEEP_ALIVE_ENABLED,KEEP_ALIVE_IDLE_SOCKET_TTL_MS,MAX_OPEN_CONNECTIONS,LOG_LEVEL,COMPRESSION_REQUEST}`
    | `1`, unset, `10`, `info`, `1` | Pool tuning, matching the other
    per-workload ClickHouse clients. |
    
    Migrations to apply: ClickHouse `036_create_queue_metrics_v1.sql`, and a
    Postgres migration adding the nullable
    `TaskQueue.concurrencyLimitOverridePercent`. Both are additive.
    
    ## How collection works
    
    Queue operations produce two kinds of signal, and they have opposite
    failure modes, so they are handled differently.
    
    **Gauges** (queued, running, queue limit, env queued, env running, env
    limit, throttled, plus keys-with-backlog and worst-key wait on keyed
    queues) are read *inside* the same Redis script that performs the
    enqueue or dequeue, so the reading is atomic with the operation it
    describes rather than a racy follow-up read. The script returns them on
    its reply and the app forwards them to the stream. Gauges are sampled
    and drop-tolerant: they are aggregated with `max`, so a lost reading
    costs resolution, never correctness.
    
    **Counters** (enqueued, started, completed, plus nack and dead-lettered)
    are cumulative odometers. Each event increments a per-queue key on the
    metrics Redis and emits the absolute total, and ClickHouse takes the
    difference across buckets at read time. This is the important property
    of the design: a summed-delta counter undercounts permanently on any
    lost event, while a cumulative one self-heals, because the next
    surviving reading restates the whole total. Only bucket granularity can
    be lost, never the total. A queue returning after its odometer TTL
    expired restarts at 1 and reset detection handles it, which is safe
    precisely because expiry only spans a window with no activity.
    
    Both land on one sharded Redis stream. A consumer reads it with a
    consumer group, reclaims stale pending entries on a 15s interval rather
    than on every poll, maps one entry to one or two ClickHouse rows
    (whole-queue and, for keyed queues, per-key), and acks only after the
    insert lands. Each batch carries a dedup token derived from its
    stream-entry ids, and the target tables set
    `non_replicated_deduplication_window`, so a retried batch cannot
    double-count either the raw rows or the aggregates that hang off them.
    Consumer and emitter both emit OTel metrics
    (`queue_metrics.emitter.emitted`,
    `queue_metrics.consumer.{entries,rows_inserted,insert_errors,insert_duration,stream_depth,group_lag,pending,lag_unknown}`);
    stream depth and group lag are the two worth alerting on, and
    `lag_unknown` exists because Redis can report a null lag after a trim,
    which must not be read as zero.
    
    ## Storage and read path
    
    `queue_metrics_raw_v1` is a short landing table with a 6 hour TTL. Four
    aggregate tiers are materialized straight from raw, never cascaded off
    each other, each with a 30 day TTL:
    
    - `queue_metrics_v1`, 10 second buckets per queue, the default read path
    - `queue_metrics_5m_v1`, 5 minute buckets per queue, for wide ranges and
    cross-queue ranking
    - `env_metrics_v1`, 10 second buckets per environment, queue-independent
    so it stays cheap at any range
    - `queue_metrics_ck_v1`, 10 second buckets per concurrency key
    
    Every tier is an MV from raw because the counter states do not survive a
    cascade: their merge is order sensitive, so a `-MergeState` chain off
    the 10s table inflates the result, and the same property means an
    aggregate state may only be merged inside one queue. That constraint is
    now enforced by the query engine rather than by reviewer discipline: a
    column can declare a `mergeGroupKey`, and any query that references it
    without grouping by, or pinning to a single value of, every named key
    fails to compile with an actionable message.
    
    On the read side, TRQL gains three tables (`queue_metrics`,
    `env_metrics`, and a `queue_metrics_by_key` that is hidden from the
    editor, schema docs and schema API but still queryable, so per-key rows
    can never silently merge into a plain per-queue query), plus
    `deltaSumTimestampMerge` and `quantilesTDigestMerge`. Two schema-level
    optimizations ride along: a table can declare coarser rollups, so a
    query whose bucket interval is 5 minutes or wider is routed to the 5m
    table with no change to the query itself, and it can opt into the
    ClickHouse query cache with time bounds floored to a fixed grid, so the
    auto-refreshing dashboards actually share cache entries instead of
    missing on every tick. Both are caller-side substitutions, so the
    printer stays unaware of physical layout.
    
    All of this can also live on its own ClickHouse service. A table
    declares the pool its reads run on, the three queue-metrics tables name
    the dedicated one, and the ingestion consumer writes through the same
    client, so both directions move together with one env var and nothing
    else routes differently.
    
    The other engine change is opt-in gap filling: charts can request rows
    for empty buckets, where counters zero-fill and gauges carry forward.
    Grouped gauge series are densified per group and carried inside a
    partition, so a quiet queue's line holds its last value without bleeding
    another queue's value into it.
    
    ## Queue concurrency limits
    
    `concurrencyLimitOverridePercent` on `TaskQueue` is the source of truth
    when an override is set as a percentage; the absolute `concurrencyLimit`
    is materialized from it (floored, clamped to at least 1 so a percentage
    can never act as a pause, and never above the environment limit). Every
    path that changes an environment limit now recalculates the
    environment's percent-based overrides afterwards, outside the
    transaction, and pushes changed limits to the engine. The push is
    attempted even when the stored value did not change, so a previously
    failed sync self-heals rather than leaving the database and the engine
    diverged; paused queues are skipped so a recalculation cannot
    effectively unpause one.
    
    The API accepts exactly one of `concurrencyLimit` or `percent`, and the
    reject-instead-of-clamp change above means a request asking for more
    than the environment allows now fails loudly. The percent bound (greater
    than 0, at most 100) is defined once and shared by the zod schema, the
    dashboard mutation handler and the service, so the three cannot drift.
    
    The concurrency-keys table on a queue is now paginated against the
    ClickHouse per-key tier, ranked by peak backlog with the total on every
    row from a single scan, and only the keys on the current page are
    enriched with live counts from Redis. That replaces a hard top-50 cap
    with something whose cost is a function of page size rather than key
    cardinality.
    
    ## The health report
    
    `GET /api/v1/reports/:key?period=&format=markdown|ansi|json`. The
    verdict is computed on the server and is deterministic, not
    model-generated. Three independent analyzers run over one input
    snapshot: flow (is work moving, and if not, is the cause a limit,
    throttling, one bad queue, or dead-lettering), execution (are the runs
    that start succeeding, and at what latency), and liveness (how fresh is
    the telemetry). When telemetry is genuinely stale, the first two are
    forced to unknown and every actionable field is stripped, so no surface
    ever advises action off stale data.
    
    Authorization is per query table rather than a blanket query grant: a
    JWT must be scoped to every table the report reads (`runs`,
    `env_metrics`, `queue_metrics`), so a narrowly scoped token cannot pull
    a report that reads more than it was granted. `period` is validated as a
    shorthand with a 90 day ceiling at the edge. The report catalog is a
    registry of `{ load, interpret }` entries, so the next report is a new
    entry and no change to the route, the view model, the renderers, the CLI
    or the MCP tool.
    
    `trigger mcp` no longer launches the install wizard when stdout is a
    TTY, which fixed a real failure: hosts spawn the server over a PTY, so
    the wizard would open and the client would time out waiting for a server
    that never started. The wizard now needs `trigger mcp --install`.
    
    ## The part that is live regardless of every flag
    
    The enqueue and dequeue scripts now return a 2-tuple so a gauge reading
    can ride back on the reply. Every return site in the eight affected
    scripts is wrapped, and a `nil` original is converted to `false` on the
    way out, because a raw `nil` in the first slot would make Lua truncate
    the multi-bulk reply and silently drop the gauge on the throttled and
    empty-queue paths. The reply shape and the destructuring on the app side
    are exercised on every queue operation whether or not metrics are
    enabled, so that is the part of `run-engine` worth the closest review.
    
    One behavior fix in the same area: the scheduling-delay anchor is set
    only on a run's first entry into the queue. Anchoring it to trigger time
    on re-enqueues made waitpoint and checkpoint resumes report the entire
    wait as scheduling delay. Queue ordering is untouched, so a re-enqueued
    run keeps its position, and nacks deliberately keep the original anchor
    because a rolled-back dequeue is the same continuous wait.
    
    A pending-version promotion still anchors to trigger time, on purpose:
    that promotion is the run's first real entry into the queue, since the
    trigger deliberately held it back waiting for a worker version, and the
    TTL is armed at the same point for the same reason. The consequence is
    worth naming, because it is a judgement call: a run that waits on a
    deployment reports that wait as scheduling delay on its queue, which is
    time unrelated to queue capacity.
    
    ## Verification
    
    Unit and integration suites across the new package, the run queue, the
    mapping layer, the query engine and ClickHouse (including a test that
    applies migration 036 through the same splitter CI uses, and a
    regression test that inserts the same batch three times to prove the
    aggregates do not inflate). Beyond that, the whole path was driven end
    to end against a live stack with real runs: emitter to Redis stream to
    consumer to ClickHouse to the dashboards, for both the local dev path
    and the deployed path where a supervisor drives the dequeue, with
    assertions on exact counter reconstruction per queue and per concurrency
    key, throttling, environment saturation, scheduling delay, and a
    deliberate mid-stream reading drop to confirm the cumulative counters
    still reconstruct the correct total. The gated-off state was checked on
    every touched surface.
    
    The dedicated ClickHouse service was verified against a second,
    separately-schema'd instance: with it configured, the driven counters
    reconstruct exactly on the dedicated instance, the shared instance gains
    no rows for that window, a read through the query API returns the value
    that exists only on the dedicated instance, and a `runs` query still
    succeeds (it would fail outright if it were mis-routed to a service
    without that table). With the variable unset, the full suite passes
    unchanged.
    
    ---------
    
    Co-authored-by: Katia Bulatova <katia@trigger.dev>
    Co-authored-by: Katia Bulatova <katherine.bulatova@gmail.com>
    Co-authored-by: James Ritchie <james@trigger.dev>
    4 people authored Jul 29, 2026
    Configuration menu
    Copy the full SHA
    4eb9292 View commit details
    Browse the repository at this point in the history
  6. feat(database,rbac): add multiple environment API key foundations (#4388

    )
    
    Adds the storage model and authorization contracts needed for multiple
    environment API keys. Credentials are represented by hashed values,
    revocation and expiration state, and persisted effective scopes.
    
    The built-in authorization fallback exposes full-access policy
    preparation, while optional authorization extensions can supply
    additional presets and task-aware scope generation. This change does not
    create, display, or authenticate additional keys.
    carderne authored Jul 29, 2026
    Configuration menu
    Copy the full SHA
    a81ad49 View commit details
    Browse the repository at this point in the history
  7. Configuration menu
    Copy the full SHA
    878c158 View commit details
    Browse the repository at this point in the history
  8. Configuration menu
    Copy the full SHA
    ed8f5e1 View commit details
    Browse the repository at this point in the history
  9. Configuration menu
    Copy the full SHA
    a098171 View commit details
    Browse the repository at this point in the history
  10. Configuration menu
    Copy the full SHA
    8ebc8a4 View commit details
    Browse the repository at this point in the history
  11. Configuration menu
    Copy the full SHA
    2f1734c View commit details
    Browse the repository at this point in the history

Commits on Jul 30, 2026

  1. fix(webapp,clickhouse): stop invalid customer queries alerting, and i…

    …solate Sentry scope per request (#4372)
    
    ## Summary
    
    A query sent to the query API with a typo in it, like a column name that
    does not exist, was being reported as a server error. That put customer
    SQL mistakes into our error alerting, where they made up almost all of
    the volume on one of our noisiest alerts, and it drowned out the
    failures that are actually ours to fix. This makes the level match who
    is at fault, and fixes two related problems found alongside it.
    
    ## Invalid queries are the caller's, not ours
    
    The query API route already got this right. It checks for `QueryError`,
    logs at warn, and returns a 400, with a comment saying the system
    handles it gracefully and no alert is needed.
    
    The layer underneath ignored that. `executeTSQL` logged every exception
    out of its catch block at error, including the compile failures the
    route was about to turn into a 400, and error-level logs are forwarded
    to error reporting.
    
    The TSQL package already draws the line we need:
    
    ```ts
    export class ExposedTSQLError extends BaseTSQLError {
      /** An exception that can be exposed to the user. */
    }
    
    export class InternalTSQLError extends BaseTSQLError {
      /** An internal exception in the TSQL engine. */
    }
    ```
    
    `SyntaxError` and `QueryError` extend the first. So the catch block now
    branches on `ExposedTSQLError` and logs those at warn, keeping error for
    `InternalTSQLError` and anything unanticipated, which is a genuine
    compiler bug.
    
    ## SQL the caller wrote is their mistake, not ours
    
    The same asymmetry showed up one level down. A query that compiles fine
    can still be rejected by ClickHouse at execution, and most of those
    rejections mean the caller's SQL is wrong rather than that we generated
    something bad.
    
    This is where the volume actually is. Checking production, one error
    group alone, a missing `GROUP BY` on the public query API
    (`NOT_AN_AGGREGATE`), accounts for over a million events across hundreds
    of users. It is by far the largest error group in the project, and
    classifying only by resource limit would have left every one of those at
    error level.
    
    So rejections are split three ways in `ClickhouseClient`, which is the
    only place holding the parsed `ClickHouseError` and its symbolic type.
    By the time the error reaches `executeTSQL` it has been wrapped and the
    type is gone, and the type never appears in the message text, so it
    cannot be recovered by string matching.
    
    - **Resource limits** (memory ceiling, timeout, row/byte caps) log at
    warn. The query is valid, it just asked for more than it is allowed to
    spend.
    - **Invalid SQL** (`NOT_AN_AGGREGATE`, `UNKNOWN_IDENTIFIER`,
    `SYNTAX_ERROR`, the type and parse families) logs at warn **only when
    the caller wrote the SQL**.
    - **Everything else** keeps alerting.
    
    That gate matters. The client is shared, so the identical rejection on
    TRQL *we* generated is our bug and has to stay at error. Callers opt in
    with `userAuthoredQuery`:
    
    | caller | who wrote the SQL | opts in |
    | --- | --- | --- |
    | public query API | the customer | yes |
    | query editor | the customer | yes |
    | agent charts | the agent's model | yes |
    | built-in dashboard tiles | us, in code | no |
    | queue metric cards | us, in code | no |
    | health report | us, in code | no |
    
    The agent is the one judgement call. Its TRQL is not typed by a person,
    but it is also not something a code fix makes correct, so a query it
    gets wrong is not worth waking anyone for. The same endpoint serves
    built-in tiles whose TRQL we do write, so the opt-in lives with the
    caller rather than the route.
    
    Separately, when one of these queries did fail, the log recorded the
    generated ClickHouse SQL but not the query the caller actually wrote,
    which made the reports hard to act on. `queryWithStats` takes an
    optional `logFields` that `executeTSQL` uses to attach the original
    TSQL.
    
    ## Events were attributed to the wrong request
    
    Chasing the above turned up something broader: only a tenth of the
    events on that alert pointed at the query API. The rest were pinned to
    unrelated requests that happened to be in flight at the same time, so
    the alert looked like the trigger endpoint was failing.
    
    `Sentry.init` runs with `skipOpenTelemetrySetup: true`, because we
    register our own OTel pipeline. That skips `initOpenTelemetry`, and one
    of the things it does is:
    
    ```js
    api.context.setGlobalContextManager(new SentryContextManager());
    ```
    
    The async-context strategy is still installed, but `withIsolationScope`
    only marks the OTel context and delegates the actual fork to that
    context manager:
    
    ```js
    // "We depend on the otelContextManager to handle the context/hub"
    return api.context.with(ctx.setValue(SENTRY_FORK_ISOLATION_SCOPE_CONTEXT_KEY, true), ...)
    ```
    
    `provider.register()` installed a plain
    `AsyncLocalStorageContextManager`, which does not know that key. The
    lookup found no scopes on the context and fell back to the
    process-global default isolation scope, so every request wrote its
    request data into the same object and the last writer won.
    
    The tracer now registers `SentryContextManager`, which subclasses
    `AsyncLocalStorageContextManager`, so OTel behaviour is unchanged. It is
    also registered on the path where tracing is disabled, which previously
    never called `register()` at all and so had no context manager of its
    own.
    
    Tenant tags were always correct, because those come from our own async
    local storage rather than the isolation scope. That is why the
    attribution being wrong was not obvious.
    
    This affects every error report the webapp sends, not just the query
    API.
    
    ## Verification
    
    `internal-packages/clickhouse`: 76 tests pass, including eight covering
    each level decision against a real ClickHouse container. Three pairs pin
    the gate open and shut at both layers: an invalid query, a compile
    failure, and a real limit breach driven with `max_rows_to_read` each log
    at warn with `userAuthoredQuery` and at error without it.
    
    The isolation fix has a test that reproduces the leak before asserting
    the fix. Two overlapping requests each tag their own isolation scope;
    with the plain context manager the slower one reads back the other's
    tag, and with `SentryContextManager` each reads back its own.
    
    Measured separately against a faithful reproduction of the server's
    wiring (own OTel pipeline, CommonJS entry) at 200 concurrent requests:
    per-request attribution goes from 0.5% to 100%, while span nesting,
    context propagation across awaits, and distinct trace IDs are identical
    before and after.
    ericallam authored Jul 30, 2026
    Configuration menu
    Copy the full SHA
    6e5f0f0 View commit details
    Browse the repository at this point in the history
  2. Configuration menu
    Copy the full SHA
    86b948b View commit details
    Browse the repository at this point in the history
Loading