Skip to content

Native Iceberg write fails on a V1 spec that mixes a live field with a void field whose source column was dropped #6141

Description

@andygrove

Describe the bug

On a format-version 1 table, a partition spec can mix a live field with a void field whose source column has since been dropped (iceberg-java keeps a dropped V1 partition field as a void transform). The native Iceberg writer fails every task on such a spec, and the query does not fall back.

CometLocationGenerator::try_new (native/core/src/execution/operators/iceberg_partition_path.rs:66-70) skips resolving the partition type only when the spec is entirely void (is_unpartitioned()). For a mixed spec it calls PartitionSpec::partition_type(schema), which errors with "No column with source column id" because the void field's source column is gone. The partition-value computation hits the same resolution.

The fixes for #5691 and #5693 cover the all-void spec only.

This was found by reading the code and has not been reproduced yet.

Steps to reproduce

CREATE TABLE t (id INT, a STRING, b STRING) USING iceberg
  PARTITIONED BY (a, b) TBLPROPERTIES ('format-version'='1');
ALTER TABLE t DROP PARTITION FIELD b;
ALTER TABLE t DROP COLUMN b;
-- with spark.comet.iceberg.write.enabled=true
INSERT INTO t VALUES (1, 'x');

Expected behavior

The insert succeeds, as it does with iceberg-java. Either the native writer resolves the partition type without the void field's source column (a void field never contributes a value), or the gate declines the native write when a void field's source column is missing from the schema.

Additional context

Found in an audit of the native Iceberg write path before enabling it by default. Part of #5649. Related: #5691, #5693.

Activity

  1. 0lai0 commented on Sep 24, 2026

    @0lai0
    Contributor

    take

  2. added
    priority:mediumFunctional bugs, performance regressions, broken features
    and removed on Sep 28, 2026
  3. andygrove commented on Oct 5, 2026

    @andygrove
    MemberAuthor

    @0lai0 a heads-up, since you picked this up. Running the statements above showed that iceberg-java cannot write this table either:

    Writer Iceberg 1.8.1 (Spark 3.5) Iceberg 1.11 (Spark 4.1)
    iceberg-java ValidationException: Cannot find source column for partition field IllegalArgumentException: Cannot build accessor for field: null
    native CometNativeException: ... Field not found the same

    So the outcome to match is iceberg-java's failure. Resolving the spec in the native writer would let Comet write a table that Spark without Comet cannot. #6677 adds a gate rule that declines the native write for a spec that mixes a live field with a void field whose source column was dropped, so the write fails with iceberg-java's own error, and it closes this issue. If you had work in progress on this, please take a look at #6677 and let us know.

  4. 0lai0 commented on Oct 5, 2026

    @0lai0
    Contributor

    Thanks for the heads-up and for running the repro against both Iceberg versions! I hadn't started on this yet, so no conflict on my side. Matching iceberg-java's failure via the gate makes sense to me. I'll take a look at #6677 and leave a review if I spot anything.
    Thanks @andygrove

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:Icebergarea:writerNative Parquet writerbugSomething isn't workingpriority:mediumFunctional bugs, performance regressions, broken features

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions