Workflow metadata views

OMERO MapAnnotations are a searchable view of workflow history, not the source used for detached recovery. The event store remains authoritative. Full CSV provenance is exported independently and is not reduced by these view policies.

The core function biomero.provenance.render_workflow_metadata(tracker, workflow_id) renders namespace/value pairs without writing events or OMERO objects. Both result scripts use this function. Deploy matching scripts and core revisions together.

View policies

v0 is the only supported view, using the legacy-compatible layout. It retains scientific task parameters, including false and zero values, and the existing task/job fields. It excludes detached launcher and result-shallower coordination annotations, duplicate output_settings parameters, and unselected workflow parameters in orchestration tasks where workflow selection is known. These records remain available in the event store and full CSV.

Job Command, Env_* and Result_Message fields are retained, as are the Name, Input_Data and Created_On fields used by batching. Import-result tasks retain the SLURM_Get_Results.py namespace used for result discovery. Canonical and shallow-Zarr metadata use separate namespaces and are not changed by this policy.

The view adds Metadata_View_Version and Aggregate_Version. The existing Version field still describes software, not the metadata schema. A workflow and its tasks have independent aggregate versions. Metadata written during import can therefore legitimately retain IMPORTING even after the workflow finishes. Refreshing its view does not advance that snapshot.

Result storage provenance

The view includes additive storage fields on the import task when the event store contains observed provenance for the requested target_key (for example Plate:123). Storage_Format distinguishes shallow-zarr from full-zarr and Storage_Shallow is an explicit boolean string. These are per-result facts, not workflow feature flags.

The scripts collect importer outcome receipts and the result’s shallow manifest, then record a Task.StorageProvenanceRecorded event before rendering metadata. Core stores plain evidence and renders it without accessing files or connections. Recorded execution location, tool version, container reference, remote task/job IDs and report checksum are retained. Missing historical tool or location evidence is not inferred from current configuration.

Canonical source biocodes appear inline when small. Larger collections use a count and reference to Shallow_Manifest with its SHA-256 checksum. Full per-target evidence also remains in CSV provenance and the event store. render_workflow_metadata and plan_metadata_refresh accept an optional target_key; callers must supply it to select a result’s storage facts.

Existing snapshots without this event remain unchanged by a historical refresh; the updater does not invent missing storage history from current disk contents.

Planning a view refresh

Core creates data structures; the scripts layer owns connections, backups and annotation updates. Core does not import annotation client libraries.

from biomero.provenance import MetadataAnnotation, plan_metadata_refresh

existing = [
    MetadataAnnotation(namespace=namespace, values=values)
    for namespace, values in stored_annotations
]
changes = plan_metadata_refresh(
    tracker, workflow_uuid, existing, view_version="v0")

Each change contains a before and after view. An absent after view requests removal of the object’s link to that annotation, not deletion of the annotation. The caller is responsible for applying these changes safely.

Legacy snapshots are resolved by an exact, unique Modified_On match against aggregate history. New annotations use their explicit aggregate version. Ambiguous or incomplete snapshots and conflicting identities are refused rather than guessed. Unknown namespaces and additional custom keys are preserved. Existing CSV references for oversized values are retained.

The administrative refresh adapter is provided by biomero-scripts in admin/SLURM_Init_environment.py as an optional metadata refresh operation. The administrator guide describes dry runs, backups, shared-annotation checks and updating existing views in place. No automatic migration runs during initialization. New result scripts continue to write v0.

Versioning and compatibility

Metadata_View_Version identifies the rendering policy; Aggregate_Version identifies the exact event-sourced snapshot. Neither changes the existing software Version field. A refresh reapplies the selected policy to the original snapshots, not the latest workflow state. It does not rewrite events, rerun analysis or upgrade the contents of existing CSV attachments.

v0 restores the pre-detached task/job layout while adding revision markers and recorded storage provenance. It is not a byte-for-byte reproduction: internal launcher/helper annotations and excluded parameters are removed, and fields emitted by the renderer may be added to an existing annotation. Unknown custom fields and existing reduced CSV references are preserved. Missing annotations or unresolved historical snapshots cause the planner to refuse the update rather than synthesize a partial history.

Callers should use a dry run to inspect the proposed field changes before applying a refresh. Core returns MetadataChange objects; the scripts decide how to display differences, persist backups and update annotation links.

Scripts persistence adapter

This API belongs to the separate biomero-scripts repository, not core. Run it in the OMERO script runtime with the scripts’ admin directory on the Python import path:

from SLURM_Init_environment import refresh_workflow_metadata

# conn is the script's administrator gateway; tracker is WorkflowTracker.
plan = refresh_workflow_metadata(
    conn, tracker, "Plate", plate_id, workflow_uuid, view_version="v0")
# Inspect the plan before applying it.
result = refresh_workflow_metadata(
    conn, tracker, "Plate", plate_id, workflow_uuid,
    view_version="v0", dry_run=False,
    backup_path="/data/biomero-metadata-backups/plate-before-refresh.json")

The adapter reads existing annotations and asks core to plan plain-data changes. OMERO connections never enter the core API. The adapter checks administrative access, preflights the target against intervening changes, and applies the plan. Retained annotations keep their IDs, namespaces and creation events. Obsolete internal-task annotations are unlinked from the selected object, not deleted. Unknown namespaces and custom keys, repeated legacy Input_Data pairs and existing CSV references are preserved. Reapplying a view is idempotent.

Bulk execution groups each object’s workflow views into the same worker lane. Each lane owns its gateway and event-store reader, joins the administrative execution session with keepalive, and detaches without terminating the parent session. Its database sessions and connections are released when the lane ends. Missing history and refused plans are reported as skips. Write failures are reported separately and may leave partial updates because multiple OMERO writes do not form one transaction.

Optional backup snapshots contain original annotation IDs, values and links. There is no automated restore API. Manual recovery can restore retained values and relink original annotations after checking current state and permissions. Result import scripts only write new result metadata; existing annotations are maintained through the administrative adapter.

Keep the activity Message concise. Detailed field diffs for small dry runs belong in normal logger output, captured by the standard activity log. Bulk sweeps log progress and outcomes rather than complete metadata maps. Detached requests use the worker log and maintenance status for progress after handoff.

Detached administrative refresh

When detached execution is enabled, compatible scripts can queue an apply request through biomero.maintenance.queue_metadata_refresh. The returned UUID identifies the maintenance request, not a workflow being refreshed. One request may cover many workflow/result pairs. Dry runs remain inline.

MetadataRefresh records QUEUED, RUNNING and DONE/FAILED states with compact progress counters. metadata_refresh_statuses(tracker) returns active and recent terminal requests as plain data; the Check Setup script exposes this information to administrators. These requests do not appear as analysis workflows in the progress projection.

After a worker interruption, the supervisor can retry an unfinished sweep. This relies on idempotent annotation updates in the scripts, not a core checkpoint for every result. Completed and failed requests are not retried automatically. For execution policy, see the supervisor documentation.