Skip to content

input_file_name() returns empty values over a native Iceberg scan #6707

Description

@andygrove

Describe the bug

With the native Iceberg scan, which is on by default (spark.comet.scan.icebergNative.enabled=true), input_file_name(), input_file_block_start() and input_file_block_length() return "", -1 and -1 for every row of an Iceberg table. Spark returns each row's data file and block offsets. The query succeeds, so the wrong values are silent.

These expressions read InputFileBlockHolder, which Iceberg's Spark reader sets as it opens each data file. The native scan never sets it. #3312 made CometScanRule fall back from the native Parquet scan when the plan uses these expressions, but transformV2Scan has no such check, and it is not handed the plan to make one.

No Comet operator is needed between the scan and the projection:

*(1) Project [input_file_name() AS input_file_name()#27, input_file_block_start() AS input_file_block_start()#28L, input_file_block_length() AS input_file_block_length()#29L, id#26L]
+- *(1) CometColumnarToRow
   +- CometIcebergNativeScan [id#26L], .../db/t/metadata/v4.metadata.json

Steps to reproduce

spark.sql("CREATE TABLE cat.db.t (id BIGINT) USING iceberg")
for (i <- 0 until 3) {
  spark.sql(s"INSERT INTO cat.db.t SELECT id FROM range(${i * 1000}, ${(i + 1) * 1000})")
}
spark
  .sql("SELECT input_file_name(), input_file_block_start(), input_file_block_length(), id FROM cat.db.t")
  .show(3, false)

Comet returns ["", -1, -1, <id>] for all 3000 rows. SELECT input_file_name(), id FROM cat.db.t WHERE id >= 0, which puts a CometFilter above the scan, returns "" for every row too.

Expected behavior

Each row reports the data file it was read from, as Spark does. Falling back to Spark's Iceberg reader when the plan uses these expressions, as the native Parquet scan does, would also be acceptable.

Additional context

  • Reproduced on main at b80bf4e08 with the default Spark 4.1 profile and a Hadoop catalog, comparing against the same query with spark.comet.enabled=false.
  • From reading branch-1.0 and branch-1.1, the native Iceberg scan is on by default in both and neither has a guard, so this most likely ships in 1.0 and 1.1.
  • fix: keep a Spark scan unconverted when the plan reads input_file_name #6703 fixes the same problem for the Spark-to-Arrow conversions (input_file_name() returns empty values above a converted Spark Parquet scan #6573) and notes this gap. Its CometScanRule.readsInputFileBlock(plan) could be reused here once transformV2Scan has the plan.
  • The "Current limitations" list in docs/source/user-guide/latest/iceberg.md should name the fallback once there is one.
  • The native CSV V2 scan (spark.comet.scan.csv.v2.enabled, off by default) goes through the same transformV2Scan. I have not tested it.

Activity

  1. zhangfengcdt commented on Oct 7, 2026

    @zhangfengcdt
    Member

    take

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions