Skip to content

Accelerate row-level MERGE / UPDATE / DELETE plans (MergeRowsExec, ReplaceDataExec, WriteDeltaExec) #5122

Description

@andygrove

What is the problem the feature request solves?

Row-level MERGE / UPDATE / DELETE plans fall back to Spark entirely: MergeRowsExec (Spark 3.5+), ReplaceDataExec, WriteDeltaExec, and InsertOnlyMergeExec (Spark 4.2+).

CDC upsert into Iceberg is one of the most common modern ETL workloads. The expensive parts of a MERGE plan (the source/target join, filters, projections) are operators Comet already accelerates, but the merge tail forces the stage back to Spark, so in practice the whole statement runs without acceleration.

Describe the potential solution

MergeRowsExec itself is close to a projection: it routes each joined row through matched / not-matched instruction lists that are ordinary Catalyst expressions, emitting updated, inserted, or deleted rows plus the row-operation column. That makes it a candidate for a native operator independent of native writes, which would at least keep the join and merge logic in one native stage before handing off to the JVM writer.

Full acceleration of ReplaceDataExec / WriteDeltaExec also needs native DataSource V2 writes, tracked in #5121.

Additional context

Related: #5121 (DataSource V2 writes), #1625 (EPIC: native Parquet writes), #3756 (Iceberg feature matrix).

Activity

  1. peterxcli commented on Jul 31, 2026

    @peterxcli
    Member

    take, will start with MergeRowsExec first

  2. unikdahal commented on Aug 7, 2026

    @unikdahal
    Contributor

    Hi @peterxcli

    I didn't realise we had an issue for this. I was going through the native iceberg write implementation #4487 and found out that merge into was not supported.

    I have a working implementation for MergeRowsExec built currently on top of #4487 . which he's now splitting into parts landing against the merged split-operator work (#4658; #5298 is part 2, part 3 not yet started).

    If you're already working on MERGE INTO, that's totally fine, happy to step back. If not, I have a working implementation done and tested on my end for native CometMergeRowsExec, just waiting on part 3 to land before it can target main.

    Let me know either way!

  3. peterxcli commented on Aug 7, 2026

    @peterxcli
    Member

    I haven't started to work. Please go ahead. I think you could open draft PR first before the final split of iceberg native patch is landed, please also ping me to review when your PR is ready.

  4. unikdahal commented on Aug 7, 2026

    @unikdahal
    Contributor

    Thanks @peterxcli, appreciate it! Will polish it up and open a draft PR soon, will tag you for review once it's ready.

  5. andygrove commented on Sep 22, 2026

    @andygrove
    MemberAuthor

    Status update, since some of this has landed:

    Plan node State on main
    ReplaceDataExec (copy-on-write DELETE / UPDATE / MERGE) Done for Iceberg via the split-operator plan and native writer, behind the off-by-default spark.comet.write.iceberg.splitOperator.enabled and spark.comet.iceberg.write.enabled flags. Copy-on-write MERGE still falls back to the JVM write path because MergeRowsExec isn't native yet.
    MergeRowsExec In progress in #5318 (@unikdahal), awaiting review.
    WriteDeltaExec (merge-on-read) Not started. The split plan never intercepts Iceberg WriteDelta.
    InsertOnlyMergeExec (Spark 4.2+) Not started.

    The dependency on native DataSource V2 writes mentioned above now lives in #5649 (#5121 is closed in its favor), and this issue is tracked there under phase 3. Keeping this open for the merge-on-read and insert-only MERGE paths.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions