Skip to content

Enable the split-operator plan and native Iceberg writes by default #5644

Description

@andygrove

What is the problem the feature request solves?

Native Iceberg writes ship behind two flags that both default to false: spark.comet.write.iceberg.splitOperator.enabled (the writer/committer split from #4658) and spark.comet.write.iceberg.enabled (the iceberg-rust data-file writer from #5361, renamed from spark.comet.iceberg.write.enabled by #5306). This issue defines what has to be true before Comet's Iceberg write path becomes the default.

#6664 makes spark.comet.write.iceberg.enabled the one setting for both and defaults it to true in 1.2.0, on every Spark version; it plans the split operator by itself, and spark.comet.write.iceberg.splitOperator.enabled stays a testing setting that plans the split operator with the native writer off. 1.2.0 is targeted for late October or early November (#6550), so flipping now gives the write path nightly runs before the release. The split plan changes only the plan shape. The native writer writes the data files of each write its eligibility gate accepts; every other write falls back to iceberg-java's writer, and iceberg-java commits every write. Setting spark.comet.write.iceberg.enabled=false restores Spark's own write operator, as in Comet 1.1.0.

Describe the potential solution

Before #6664 merges

These are #6664's blockers. With the native writer on by default, each of them would reach every user who writes Iceberg tables.

Split plan:

Native writer correctness:

Failure handling:

Memory and performance:

Coverage and evidence:

In #6664

  • spark.comet.write.iceberg.enabled moves from CATEGORY_TESTING to CATEGORY_EXEC and switches both layers, and the docs that describe the flags as off by default are updated: user-guide/latest/iceberg-writes.md, iceberg.md, operators.md, contributor-guide/iceberg-writes.md and the Iceberg write review skill
  • Release note: a 1.2.0 upgrade-guide entry names the setting and how to turn it off. Iceberg write plans show IcebergCommit over IcebergWrite (or CometIcebergWrite) in place of Spark's AppendData, OverwriteByExpression, OverwritePartitionsDynamic and ReplaceData exec nodes. The split happens at physical planning, so analyzed and optimized logical plans are unchanged, and listeners that read logical plans (lineage tools) are unaffected. Code that matches Spark's physical write nodes will not find them.

Spark 3.4 gets the same defaults. No Iceberg Spark test job covers Iceberg 1.5.2, which the 3.4 profile pins, so on 3.4 the evidence is Comet's own suites.

Before 1.2.0 ships

Evidence gathered while the defaults run in CI and nightly:

Not blocking the defaults

Additional context

Part of #5649. Refreshed on 2026-10-06 to match #6664, which makes spark.comet.write.iceberg.enabled the one switch and defaults it to true in 1.2.0 on every Spark version. The 2026-10-05 version flipped the native writer one release after the split plan and left Spark 3.4 open; it had folded in the original list from 2026-09-02 and the issues the 2026-09-23 audit added (#6138 to #6148). Related: #4658, #5298, #5361.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions