Is your feature request related to a problem or challenge?
CometIcebergWriteExec rebuilds every natively written DataFile's metrics with iceberg-java's ParquetUtil.footerMetrics, so that metrics modes, truncation and bounds are iceberg-java's own decisions. It gets each footer by reading it back from storage (IcebergReflection.rebuildDataFilesWithJavaMetrics, through ParquetFileReader.open): two ranged reads per file, the 8-byte tail and then the footer, one file after another before the task returns its commit message. On S3 or GCS each read is a GET request, so a task that writes many small files, such as a fanout write over many partitions, pays two request round trips per file. iceberg-java takes the same metrics from the footer its writer still holds in memory.
Raised in review of #6664.
Describe the solution you'd like
Have the native writer send each file's serialized footer to the JVM with the manifest bytes, and parse it with parquet-mr's ParquetMetadataConverter instead of opening the file. iceberg-rust's ParquetWriter keeps the footer to itself, but every byte of the file passes through Comet's CountingFileWrite, which could keep the file's tail and cut the footer out of it at close.
Describe alternatives you've considered
Reading a task's footers concurrently would cut the wall-clock cost but not the request count.
Is your feature request related to a problem or challenge?
CometIcebergWriteExecrebuilds every natively writtenDataFile's metrics with iceberg-java'sParquetUtil.footerMetrics, so that metrics modes, truncation and bounds are iceberg-java's own decisions. It gets each footer by reading it back from storage (IcebergReflection.rebuildDataFilesWithJavaMetrics, throughParquetFileReader.open): two ranged reads per file, the 8-byte tail and then the footer, one file after another before the task returns its commit message. On S3 or GCS each read is a GET request, so a task that writes many small files, such as a fanout write over many partitions, pays two request round trips per file. iceberg-java takes the same metrics from the footer its writer still holds in memory.Raised in review of #6664.
Describe the solution you'd like
Have the native writer send each file's serialized footer to the JVM with the manifest bytes, and parse it with parquet-mr's
ParquetMetadataConverterinstead of opening the file. iceberg-rust'sParquetWriterkeeps the footer to itself, but every byte of the file passes through Comet'sCountingFileWrite, which could keep the file's tail and cut the footer out of it at close.Describe alternatives you've considered
Reading a task's footers concurrently would cut the wall-clock cost but not the request count.