Skip to content

fix: de-duplicate identical pipelines when combining Processing objects - #1875

Open
seanmcculloch wants to merge 1 commit into
devfrom
dedupe-pipelines-on-processing-add
Open

seanmcculloch wants to merge 1 commit into
devfrom
dedupe-pipelines-on-processing-add

Conversation

@seanmcculloch

@seanmcculloch seanmcculloch commented Aug 27, 2026 •

Copy link
Copy Markdown

Currently, the metadata manager merges Processing from multiple upstream sources, that usually share a pipeline, and it used to de-dup those itself. Moving that aggregation onto the schema's "+" operator and from_metadata lost the de-dup, since Processing.add concatenates pipelines without removing duplicates. This restores it in add (keyed on full identifiers.Code identity/__eq__, not name), so every + consumer gets it and the manager can drop its local shim.

What changed

Processing.__add__ concatenated pipelines without de-duplication, so combining N Processing objects that share a pipeline produced N identical Code entries. It now removes duplicates after merging.

remove_duplicates gains an equality-based fallback so it handles unhashable elements such as pydantic Code models, not just hashable scalars.

Review notes

  • Dedup preserves first-seen order and runs only when the merged list is non-empty.
  • Distinct pipelines that share a name are kept — identity is full model equality, not name.

seen = set()
output_list = [x for x in lst if not (x in seen or seen.add(x))]
except TypeError:
# Unhashable elements (e.g. pydantic models): fall back to equality

@seanmcculloch seanmcculloch Sep 16, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Eg, removing duplicate aind_data_schema.components.identifiers.Code objects like pipelines

@dbirman

dbirman commented Sep 16, 2026

Copy link
Copy Markdown
Member

How can this happen -- also isn't this a risk if someone runs the same code in multiple branches of a pipeline?

@seanmcculloch

seanmcculloch commented Sep 16, 2026 •

Copy link
Copy Markdown
Author

@dbirman We could end up with duplicate entries in the Processing.pipelines list easily during the step of aggregating processing metadata from the different capsules in the pipeline.

de-duping used to be handled my the metadata manager: https://github.com/AllenNeuralDynamics/aind-metadata-manager/blob/3d342db6a8aea4a5e157d0686de0ed913dd7a3ce/src/aind_metadata_manager/metadata_manager.py#L416-L440

I am currently working on this PR, which drops a lot of aggregation logic in the metadata manager, in favor or relying on the data schema itself (through __add__ and from_metadata), and we should de-duplicate this list.

BUT: it is risky in a way - since the description notes that Processing.pipelines is only matched via DataProcess.pipeline_name

pipelines: Optional[List[Code]] = Field(
default=None,
title="Pipelines",
description=(
"For processing done with pipelines, list the repositories here. Pipelines must use the name field "
",and be referenced in the pipeline_name field of a DataProcess."
),
)

However, Code.commit_hash and Code.version are optional, so I'm not sure we can make the linking stronger without enforcing that they exist.. So it is indeed risky for different commits of the same pipeline.

@dbirman

dbirman commented Sep 16, 2026

Copy link
Copy Markdown
Member

Individual capsules in a pipeline shouldn't be writing the pipeline though? How would they "know" that they are inside of a pipeline? It's the manager during aggregation that should assign them all the pipeline_name and add that object.

@dbirman dbirman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll go a bit farther than my comment above -- I don't think this is logic that the schema should be implementing. I think this kind of management has to happen downstream of the merge code, which we should try to keep as minimal as possible.

@seanmcculloch

Copy link
Copy Markdown
Author

@dbirman The intention is that this will be called by the metadata-manager, as it is right now. just that this PR is moving that code from the metadata-manager into the data-schema itself. Perhaps it should stay in the metadata-manager for now then?

The change here itself only removes exact duplicates within the Processing.pipelines: List[Code].

Can you confirm that this PR should be closed and that logic should stay in the metadata-manager?

@dbirman

dbirman commented Sep 16, 2026

Copy link
Copy Markdown
Member

I think the situation that this PR is solving for, which is roughly: "Two identical pipelines were run in parallel and are now being merged" is a situation that should never happen, which is why I'm pushing back.

I think that's actually not what you're trying to fix though, you're trying to fix: "A single pipeline has capsules whose processing needs to be merged, but each of the capsules wrote their own pipeline information". That's a bug in the capsules -- capsules (in my opinion) should not write pipeline information at all. It's only at the last step in the pipeline when all capsules are done that the metadata manager should capture all the individual capsule processing and then wrap them all with the pipeline info.

Does that make more sense? Hop on a call if it's not clear.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants