Skip to content

[DE-XXXX] Add NucleusClient.merge_model_runs() - #474

Draft
luke-e-schaefer wants to merge 1 commit into
masterfrom
lukeschaefer/merge-model-runs
Draft

[DE-XXXX] Add NucleusClient.merge_model_runs()#474
luke-e-schaefer wants to merge 1 commit into
masterfrom
lukeschaefer/merge-model-runs

Conversation

@luke-e-schaefer

Copy link
Copy Markdown
Contributor

Adds NucleusClient.merge_model_runs(model_run_ids, name, *, model_id=None, metadata=None), wrapping POST /v1/nucleus/modelRun/merge.

Requires scaleapi PR https://github.com/scaleapi/scaleapi/pull/156919, which adds the endpoint.

Why

A benchmark evaluation names a single model run, and a benchmark's items may span several datasets. A model whose predictions were uploaded as separate runs — one per dataset, or one per inference batch — has no single run covering the benchmark, so every uncovered item scores as a false negative. Merge first, then pass the new run to create_benchmark_evaluation_v2().

merged = client.merge_model_runs(
    ["run_abc", "run_def", "run_ghi"],
    name="v3 — all benchmark datasets",
)
evaluation = client.create_benchmark_evaluation_v2(
    benchmark_id, merged["model_run_id"]
)

Semantics

Full union — predictions are copied, never deduplicated, and the source runs are left untouched. If two source runs predict on the same item with the same annotation_id, the colliding id is rewritten rather than dropped. The response reports predictions_copied, predictions_ignored and annotation_ids_rewritten so nothing is lost silently.

Verification

The test suite requires live API keys (conftest.py hard-asserts on NUCLEUS_PYTEST_API_KEY), so it could not be run locally, and there is no local backend to exercise the round trip against. Verified offline that the method builds the exact payload the server's Joi schema accepts (model_run_ids, name, optional model_id / metadata), posts to modelRun/merge, and raises on fewer than two distinct run ids before making a request.

Note on the version bump

Bumped to 0.21.0. PR #473 also bumps pyproject.toml / CHANGELOG.md; whichever lands second will need a trivial rebase on those two files.

🤖 Generated with Claude Code

A benchmark evaluation names a single model run, and a benchmark's items
may span several datasets. A model whose predictions were uploaded as
separate runs — one per dataset, or one per inference batch — therefore
had no single run covering the benchmark, and every uncovered item
scored as a false negative. Merging the runs produces one run that does
cover it, which can then be passed to create_benchmark_evaluation_v2().

The merge is a full union: predictions are copied, never deduplicated.
Colliding annotation_ids are rewritten rather than dropped, and the
response reports predictions_copied, predictions_ignored and
annotation_ids_rewritten so nothing is lost silently.

Wraps POST /v1/nucleus/modelRun/merge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant