Skip to content

Add an end-to-end HDFS write test for both native Parquet writers #6750

Description

@andygrove

What is the problem the feature request solves?

Both native Parquet writers open their file inside the committer's _temporary tree and rely on HDFS renames at task and job commit. CometWriteFilesExec has done this on Spark 4.0+ since #5763, and CometNativeWriteExec does it on 3.4/3.5 with #4746. No test runs either writer against HDFS. The native parquet_writer.rs HDFS tests are #[ignore] because they need a running cluster. The JVM suites write only to file:.

Things only an HDFS run would catch: the native HDFS writer creating the attempt directory under _temporary, the commit-time renames finding the file at the exact path the commit protocol chose, escapedHdfsDestination and checkNativeWriteDestination on real paths, and bytes_written (which comes from std::fs::metadata and reads 0 on HDFS).

Describe the potential solution

An end-to-end write test against a MiniDFSCluster, run on Spark 3.5 and on a 4.x profile so it covers both writers. WithHdfsCluster already starts one for the read benchmark. Check the plan is native, that the output has _SUCCESS and Spark's file names, that nothing is left under _temporary, and that the rows read back. Also check that a failed task leaves nothing behind.

Additional context

Asked for in the #4746 review on 2026-10-02.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions