You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Enable the split-operator plan and native Iceberg writes by default #5644
Native Iceberg writes ship behind two flags that both default to false: spark.comet.write.iceberg.splitOperator.enabled (the writer/committer split from #4658) and spark.comet.write.iceberg.enabled (the iceberg-rust data-file writer from #5361, renamed from spark.comet.iceberg.write.enabled by #5306). This issue defines what has to be true before Comet's Iceberg write path becomes the default.
#6664 makes spark.comet.write.iceberg.enabled the one setting for both and defaults it to true in 1.2.0, on every Spark version; it plans the split operator by itself, and spark.comet.write.iceberg.splitOperator.enabled stays a testing setting that plans the split operator with the native writer off. 1.2.0 is targeted for late October or early November (#6550), so flipping now gives the write path nightly runs before the release. The split plan changes only the plan shape. The native writer writes the data files of each write its eligibility gate accepts; every other write falls back to iceberg-java's writer, and iceberg-java commits every write. Setting spark.comet.write.iceberg.enabled=false restores Spark's own write operator, as in Comet 1.1.0.
spark.comet.write.iceberg.enabled moves from CATEGORY_TESTING to CATEGORY_EXEC and switches both layers, and the docs that describe the flags as off by default are updated: user-guide/latest/iceberg-writes.md, iceberg.md, operators.md, contributor-guide/iceberg-writes.md and the Iceberg write review skill
Release note: a 1.2.0 upgrade-guide entry names the setting and how to turn it off. Iceberg write plans show IcebergCommit over IcebergWrite (or CometIcebergWrite) in place of Spark's AppendData, OverwriteByExpression, OverwritePartitionsDynamic and ReplaceData exec nodes. The split happens at physical planning, so analyzed and optimized logical plans are unchanged, and listeners that read logical plans (lineage tools) are unaffected. Code that matches Spark's physical write nodes will not find them.
Spark 3.4 gets the same defaults. No Iceberg Spark test job covers Iceberg 1.5.2, which the 3.4 profile pins, so on 3.4 the evidence is Comet's own suites.
Before 1.2.0 ships
Evidence gathered while the defaults run in CI and nightly:
A stretch of nightly history with both defaults on and no failure that only the native writer causes (length to be agreed)
Benchmark results (CometIcebergWriteBenchmark, Add a native Iceberg write benchmark #5647) show no regression against iceberg-java for unpartitioned, clustered, fanout and copy-on-write delete writes
Part of #5649. Refreshed on 2026-10-06 to match #6664, which makes spark.comet.write.iceberg.enabled the one switch and defaults it to true in 1.2.0 on every Spark version. The 2026-10-05 version flipped the native writer one release after the split plan and left Spark 3.4 open; it had folded in the original list from 2026-09-02 and the issues the 2026-09-23 audit added (#6138 to #6148). Related: #4658, #5298, #5361.
What is the problem the feature request solves?
Native Iceberg writes ship behind two flags that both default to
false:spark.comet.write.iceberg.splitOperator.enabled(the writer/committer split from #4658) andspark.comet.write.iceberg.enabled(the iceberg-rust data-file writer from #5361, renamed fromspark.comet.iceberg.write.enabledby #5306). This issue defines what has to be true before Comet's Iceberg write path becomes the default.#6664 makes
spark.comet.write.iceberg.enabledthe one setting for both and defaults it totruein 1.2.0, on every Spark version; it plans the split operator by itself, andspark.comet.write.iceberg.splitOperator.enabledstays a testing setting that plans the split operator with the native writer off. 1.2.0 is targeted for late October or early November (#6550), so flipping now gives the write path nightly runs before the release. The split plan changes only the plan shape. The native writer writes the data files of each write its eligibility gate accepts; every other write falls back to iceberg-java's writer, and iceberg-java commits every write. Settingspark.comet.write.iceberg.enabled=falserestores Spark's own write operator, as in Comet 1.1.0.Describe the potential solution
Before #6664 merges
These are #6664's blockers. With the native writer on by default, each of them would reach every user who writes Iceberg tables.
Split plan:
IcebergWriteStrategyrespectsspark.comet.enabled, so disabling Comet restores Spark's planTransactionalExec)Native writer correctness:
voidfield whose source column was dropped falls back, so the write fails with iceberg-java's own error (PR fix: decline native Iceberg writes through a void partition field whose source column was dropped #6677)Failure handling:
Memory and performance:
Coverage and evidence:
run-all-spark-profiles)In #6664
spark.comet.write.iceberg.enabledmoves fromCATEGORY_TESTINGtoCATEGORY_EXECand switches both layers, and the docs that describe the flags as off by default are updated:user-guide/latest/iceberg-writes.md,iceberg.md,operators.md,contributor-guide/iceberg-writes.mdand the Iceberg write review skillIcebergCommitoverIcebergWrite(orCometIcebergWrite) in place of Spark'sAppendData,OverwriteByExpression,OverwritePartitionsDynamicandReplaceDataexec nodes. The split happens at physical planning, so analyzed and optimized logical plans are unchanged, and listeners that read logical plans (lineage tools) are unaffected. Code that matches Spark's physical write nodes will not find them.Spark 3.4 gets the same defaults. No Iceberg Spark test job covers Iceberg 1.5.2, which the 3.4 profile pins, so on 3.4 the evidence is Comet's own suites.
Before 1.2.0 ships
Evidence gathered while the defaults run in CI and nightly:
CometIcebergWriteBenchmark, Add a native Iceberg write benchmark #5647) show no regression against iceberg-java for unpartitioned, clustered, fanout and copy-on-write delete writesNot blocking the defaults
revertToSparkkeeps the write node (PR fix: restore Spark write execs when reverting transition-heavy stages #5957). It is reached only withspark.comet.exec.transitionRevert.enabled, which is off by default.Additional context
Part of #5649. Refreshed on 2026-10-06 to match #6664, which makes
spark.comet.write.iceberg.enabledthe one switch and defaults it totruein 1.2.0 on every Spark version. The 2026-10-05 version flipped the native writer one release after the split plan and left Spark 3.4 open; it had folded in the original list from 2026-09-02 and the issues the 2026-09-23 audit added (#6138 to #6148). Related: #4658, #5298, #5361.