Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,13 @@ All notable changes to the [Nucleus Python Client](https://github.com/scaleapi/n
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [0.21.1](https://github.com/scaleapi/nucleus-python-client/releases/tag/v0.21.0) - 2026-08-15

### Added
- **`NucleusClient.merge_model_runs()`.** Merges two or more model runs into one new run holding the union of their predictions, leaving the sources untouched. A benchmark evaluation names a single model run and a benchmark's items may span datasets, so a model uploaded as several runs previously had no single run covering the benchmark — every uncovered item scored as a false negative. Merge first, wait for the copy to finish, then pass the new run to `create_benchmark_evaluation_v2()`. All source runs must belong to the same model.

The copy runs asynchronously: the call returns `{"model_run_id", "dataset_ids", "job"}` immediately, but the new run is empty until the `job` completes — call `job.sleep_until_complete()` before evaluating. The merge is a full union: predictions are copied, never deduplicated, and colliding `annotation_id`s are rewritten rather than dropped. Copy counts (`predictions_copied`, `predictions_ignored`, `annotation_ids_rewritten`) are reported on the job.

## [0.21.0](https://github.com/scaleapi/nucleus-python-client/releases/tag/v0.21.0) - 2026-08-19

### Added
Expand Down
73 changes: 73 additions & 0 deletions nucleus/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -149,6 +149,7 @@
METRIC_TYPE_KEY,
MODEL_IDS_KEY,
MODEL_RUN_ID_KEY,
MODEL_RUN_IDS_KEY,
MODEL_TAGS_KEY,
MODEL_TRAINED_SLICE_IDS_KEY,
NAME_KEY,
Expand Down Expand Up @@ -539,6 +540,78 @@ def delete_model_run(self, model_run_id: str):
{}, f"modelRun/{model_run_id}", requests.delete
)

def merge_model_runs(
self,
model_run_ids: List[str],
*,
name: Optional[str] = None,
metadata: Optional[Dict[str, Any]] = None,
) -> Dict[str, Any]:
"""Merge several model runs into one new run holding all their predictions.

A benchmark evaluation names a single model run, and a benchmark's items may
span several datasets. A model whose predictions were uploaded as separate runs
— one per dataset, or one per inference batch — therefore has no single run
covering the benchmark, and every uncovered item scores as a false negative.
Merging the runs produces one run that does cover it, which you can then pass to
:meth:`create_benchmark_evaluation_v2`.

All source runs must belong to the same model.

The merge is a full union of all predictions. If two
source runs predict on the same item with the same ``annotation_id``, the
colliding id is rewritten rather than dropped. The source runs are left
untouched.

**Asynchronous.** The new run is created and returned immediately, but its
predictions are copied by a background job. The run is *empty until the job
completes*, so wait on the returned job before evaluating::

result = client.merge_model_runs(["run_abc", "run_def"])
result["job"].sleep_until_complete()
client.create_benchmark_evaluation_v2(
benchmark_id, result["model_run_id"]
)

Parameters:
model_run_ids: Two or more distinct model run ids (``run_*``) to merge.
name: Display name for the merged run. Defaults server-side to the model's
own name when omitted.
metadata: Optional metadata for the merged run. The merge always records
``merged_from_model_run_ids`` alongside whatever you pass.

Returns:
Dict describing the newly created (still-populating) run::

{
"model_run_id": str, # the new run, usable once the job finishes
"dataset_ids": List[str], # datasets the new run spans
"job": AsyncJob, # copy progress; poll or sleep_until_complete()
}

The copy's counts (``predictions_copied``, ``predictions_ignored``,
``annotation_ids_rewritten``, errors) are reported on the job, not here.
"""
unique_model_run_ids = list(dict.fromkeys(model_run_ids))
if len(unique_model_run_ids) < 2:
raise ValueError(
"merge_model_runs needs at least two distinct model run ids, got "
f"{sorted(unique_model_run_ids)}"
)
payload: Dict[str, Any] = {
MODEL_RUN_IDS_KEY: unique_model_run_ids,
}
if name is not None:
payload[NAME_KEY] = name
if metadata is not None:
payload[METADATA_KEY] = metadata
response = self.make_request(payload, "modelRun/merge")
return {
MODEL_RUN_ID_KEY: response[MODEL_RUN_ID_KEY],
DATASET_IDS_KEY: response[DATASET_IDS_KEY],
"job": AsyncJob.from_json(response, self),
}

def create_dataset_from_project(
self,
project_id: str,
Expand Down
1 change: 1 addition & 0 deletions nucleus/constants.py
Original file line number Diff line number Diff line change
Expand Up @@ -112,6 +112,7 @@
MODEL_TRAINED_SLICE_IDS_KEY = "trained_slice_ids"
MODEL_ID_KEY = "model_id"
MODEL_RUN_ID_KEY = "model_run_id"
MODEL_RUN_IDS_KEY = "model_run_ids"
MODEL_PREDICTION_ID_KEY = "model_prediction_id"
MODEL_PREDICTION_LABEL_KEY = "model_prediction_label"
NAME_KEY = "name"
Expand Down