You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied. Corrections should be made directly on the affected issue.
No type reclassifications. Every bug or enhancement label an author had already applied matched the issue's content. The 18 issues that arrived without a type label were classified from their bodies.
Area labels: area:scan, area:expressions; also carries EPIC, correctness
Rationale: An EPIC for timezone bugs, several of which returned silent wrong results. Confirms the author's label; see the escalations, since its only critical child is closed.
[EPIC] Match Spark's -0.0 and NaN semantics systematically instead of per expression (#6385)
Area labels: area:expressions; also carries EPIC, correctness
array_contains, arrays_overlap, array_distinct and array_union ignore string collation (#6470)
Area labels: area:expressions
Rationale: On Spark 4.x with default configs, the four native kernels compare collated strings by raw bytes, so array_contains returns false where Spark returns true and the set functions keep values that Spark merges. Step 1, the same class as CometSort sorts a multi-column key containing a collated string by raw bytes #6158.
Native sort orders null elements of array and struct keys by the key's null order, unlike Spark (#6476)
Area labels: none
Rationale: Under ASC NULLS LAST or DESC NULLS FIRST, a key holding a null element sorts at the opposite end from Spark, which also changes which rows a TopK keeps and how window order keys rank. That is silent wrong results at step 1, like the collated sort key in CometSort sorts a multi-column key containing a collated string by raw bytes #6158.
RANGE window frames over an array or struct key with a null element span the whole partition (#6477)
Area labels: none
Rationale: SUM(id) OVER (ORDER BY array(i)) returns the whole partition's total on the null-element row and every row after it, where Spark returns running sums. The plan is native and raises no error, so this is step 1.
Native corr, covariance, variance and stddev return wrong values for a constant fractional column merged from several partitions (#6481)
Area labels: area:aggregation; also carries correctness
Rationale: On default configs corr returns 0.878 where Spark returns NULL, or raises DIVIDE_BY_ZERO under ANSI, and the variance and covariance functions return tiny non-zero values instead of 0.0. That is silent wrong results at step 1, including the "Comet returns a value where Spark raises" shape. It has been present since 1.0.0.
percentile_approx returns different percentiles from Spark when its input holds a NaN with the sign bit set (#6519)
Area labels: area:aggregation
Rationale: CometApproxPercentile reports Compatible for floats and returns different percentiles at every percentage once a partition merges more than one head buffer. Arithmetic produces sign-bit NaNs on x86-64, so this is step 1, the same class as the other [EPIC] Match Spark's -0.0 and NaN semantics systematically instead of per expression #6385 children.
Native map construction doesn't match Spark 4.0+ float key normalization (-0.0 keys, missing DUPLICATED_MAP_KEY) (#6549)
Codegen dispatcher initializes kernels with an incorrect partition index under UNION ALL and coalesce (#6570)
Area labels: area:expressions
Rationale: The dispatcher initializes each kernel with TaskContext.partitionId(), so a dispatched spark_partition_id(), monotonically_increasing_id(), rand or uuid in a later UNION ALL branch, or below CometCoalesceExec, sees the task's partition index instead of the index of the partition being computed, which Spark uses. The dispatcher is on by default (spark.comet.exec.scalaUDF.codegen.enabled=true), so this is silent wrong results at step 1 (see escalations).
Comet shuffle does not report Spark 4.1's order-independent shuffle checksum (#6414)
Area labels: area:shuffle; also carries spark 4.1, spark 4.2
Rationale: With the opt-in orderIndependentChecksum configs, every Comet map output reports a checksum of 0, so Spark's protection against mixed retry output is silently off. The configs are off by default and nothing is wrong unless an indeterminate stage is retried, so step 3. Confirms the author's label.
Investigate TPC-DS q54 and q68 slowdowns in the 1.1.0 benchmark on Spark 4.2 (#6418)
Area labels: none; also carries performance, spark 4.2
AQE does not coalesce shuffle partitions under a CometUnion when a branch is not a shuffle (#6454)
Area labels: area:shuffle; also carries performance
Rationale: A small union runs hundreds of near-empty tasks, because CoalesceShufflePartitions matches Spark's UnionExec and not CometUnionExec. Results are correct, so this is performance degradation at step 3. Confirms the author's label.
Native CASE WHEN names its struct result's fields after the ELSE branch instead of the first THEN branch (#6482)
Area labels: area:expressions; also carries correctness
Rationale: Values are right, and only native code that reads field names sees the ELSE branch's names. The reproduction needs native to_json under StructsToJson.allowIncompatible=true; default to_json dispatches and is correct. So it follows the [Bug] Native make_interval overflows valid time components or loses seconds precision #5131 rule at step 3 (see escalations).
Reduce retained buffer allocation when map lookups select short nested values (#6525)
Native CASE WHEN and COALESCE reconcile struct fields by name instead of position (#6532)
Area labels: area:expressions; also carries correctness
Rationale: When branch structs have case-distinct field names in swapped positions, the native common type can widen the wrong field. The issue observes it only through native to_json, which needs StructsToJson.allowIncompatible=true, so it follows the [Bug] Native make_interval overflows valid time components or loses seconds precision #5131 rule for Incompatible opt-ins (see escalations).
Area labels: area:scan; also carries documentation, test, area:Iceberg
Rationale: Item 1 is a defect: a native Iceberg scan of a table with a hostless alias metadata location never calls the configured credential provider and signs with the default chain, so reads fail with 403 or use broader credentials than the provider vends. It needs an opted-in alias scheme and was found by reading the code, so step 3 (see escalations). The other items are docs and tests.
Native Iceberg writer counts NaNs under NULL structs and NULL list or map entries (#6562)
Area labels: area:writer; also carries area:Iceberg
Native scans resolve the AWS SDK ProfileCredentialsProvider names differently from Hadoop (#6575)
Area labels: area:scan
Rationale: For the AWS SDK profile provider names on Spark 4.x, a few profile shapes make native reads sign with different credentials, or call a different STS endpoint, than Hadoop. The reporter notes that these are uncommon profile shapes and that a plain profile works, so step 3 (see escalations).
Native Azure credential lookup still differs from Hadoop ABFS in a few configurations (#6605)
Area labels: area:scan
Rationale: In the configurations listed (container-scoped OAuth keys on Hadoop 3.4.2+, Fabric hosts, stale account-scoped key forms, credential provider stores), the native scan either fails or authenticates with a different credential than Hadoop's ABFS driver would. Each needs an uncommon setup, so step 3 (see escalations).
Array lookup functions evaluate their second argument on rows where the array is NULL, raising ANSI errors that Spark skips (#6613)
Area labels: area:expressions
Rationale: Under ANSI, Comet fails a query that Spark completes. The failure is visible, and ANSI off or try_cast avoids it, so step 3. This is the opposite direction from the "Comet returns a value where Spark raises" corrections. Confirms the author's label.
Rationale: Since perf: read the shuffle natively in operators AQE reuses from the initial plan #6547, a re-plan on Spark 3.4 flips the join's build side, so the DPP subquery no longer reuses the join's broadcast. The test is right and the plan changed, so this is a product regression on default configs rather than a test-only failure. It has a workaround (spark.comet.shuffle.directRead.enabled=false), so step 3 (see escalations).
Native Iceberg writer's value and null counts for a float under a NULL struct differ from iceberg-java 1.9+ (#6659)
Area labels: area:writer; also carries area:Iceberg
days transform is evaluated in the session timezone, while hours and Iceberg use UTC (#6333)
Area labels: area:expressions
Rationale: Spark never evaluates these partition transforms, so there is no Spark answer to match, and the author rated it low. Confirms the author's label (see escalations).
Broadcast/hash join fallback reasons are lost from the AQE-final plan (#6442)
Area labels: none
Rationale: Only the EXPLAIN annotation is lost under AQE. The join still falls back, and the reason still reaches the fallback log. Diagnostics only, step 4.
test: query tolerance= in Comet SQL tests passes when either side is NaN (#6616)
Flaky test: CometIcebergWriteActionSuite "a failed write job deletes the data files of tasks that completed" (#6643)
Area labels: area:writer; also carries test, area:Iceberg
Rationale: An intermittent test failure, step 4. The likely cause is in the test: its UDF blocks a native runtime thread while it waits for the other tasks.
Enhancements
[EPIC] DataFusion 56 upgrade: regressions found by tracking DataFusion main (#6410)
Investigate nested TPC-H q21 slowdown: Comet 37% slower than Spark at SF1000 (#6467)
Area labels: none; also carries performance
Rationale: A performance gap against Spark, measured on a downstream build and not reproduced on main, with no regression identified. If the rising iteration times turn out to be a leak, that part is a bug.
Native existence join enumerates every duplicate build match (N*M candidates for M markers) (#6484)
Rationale: New operator support. The fallback is intentional today.
Run array_contains on float elements natively with Spark's equality instead of through the codegen dispatcher (#6520)
Area labels: area:expressions
Rationale: Moves an Incompatible path, which dispatches today, to a native kernel.
Clarify how to run native Comet -> Celeborn shuffle end-to-end (#6523)
Area labels: area:shuffle
Rationale: A request to document what a Celeborn client must provide for the native path. The fallback with released clients is intended and documented.
perf: Comet hash joins use more task time than Spark (skewed SHJ with nested payload, BHJ) (#6528)
[EPIC] Adaptive runtime pruning for native Iceberg scans (#6641)
Area labels: area:scan
Rationale: New runtime pruning for native Iceberg scans.
perf: explore reducing copies in native Celeborn shuffle writes (#6654)
Area labels: area:shuffle
Rationale: A performance exploration of the native Celeborn write path. Nothing is broken.
Escalations to consider
Codegen dispatcher initializes kernels with an incorrect partition index under UNION ALL and coalesce (#6570)
Labeled critical from the code: CometScalaUDFCodegen initializes kernels with TaskContext.partitionId() on main, and the dispatcher is on by default. The report comes from the feat: route map lookups through codegen dispatcher #5875 review and has no standalone reproducer yet, so the reviewer may want one before relying on the label.
Native aggregate fails after spilling when groups are few and large (collect_list, collect_set) (#6363)
Matches the trigger "A priority:medium bug is reported by multiple users or affects a common workload → consider escalating to priority:high". Any collect_list or collect_set over a low-cardinality key that spills can hit it. The last pass made the same note for A native final aggregate that has spilled can fail the task during its replay #6254.
Native CASE WHEN and COALESCE reconcile struct fields by name instead of position (#6532)
Held at medium because the issue observes the widened field only through native to_json, which needs allowIncompatible. The wrong field type could also reach native consumers that run by default, such as hash or a cast to string. If one of them returns a different answer from Spark, step 1 applies.
Native CASE WHEN names its struct result's fields after the ELSE branch instead of the first THEN branch (#6482)
Held at medium on the same basis as Native Azure credential lookup still differs from Hadoop ABFS in a few configurations #6605. Here the configured location-scoped provider is bypassed rather than mismatched, so a read can succeed with broader credentials than the provider would vend. That is closer to the security case in the guide's priority:critical row, but it needs an opted-in alias scheme with hostless locations, and it hasn't been reproduced.
Labeled medium rather than low, because the failing Spark test reflects a plan change on Spark 3.4 with default configs. It doesn't block the merge queue, since the Spark 3.4 SQL tests run only with run-spark-3.4-tests. If the reviewer reads it as a test failure, priority:low fits.
days transform is evaluated in the session timezone, while hours and Iceberg use UTC (#6333)
#6576 has no reproduction on a supported path. The others are tracking or record issues rather than bug reports or feature requests. requires-triage was left in place on all five, so they reappear in the next pass until they are closed or classified.
Arrow struct writer should respect projected schema width (#6576)
The summary of the 2026-09-28 pass. Since then, its issues have had no priority changes, only the correctness and regression additions described above, so it can be closed once reviewed.
Triage pass over the open
requires-triagequeue, per the project Bug Triage Guide.priority:critical10,priority:high0,priority:medium17,priority:low4Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied. Corrections should be made directly on the affected issue.
Notes on this pass:
correctnessto five issues that pass had labeled critical (array_append: ANSI item error is swallowed when the array is NULL #6086, Field id matching misses a container id after INT96 coercion drops container metadata #6131, Nested duplicate names in a metadata-free Parquet file bypass the native resolver when the file schema equals the requested schema #6136, AQE reuses one exchange for scans with different dynamic pruning filters and drops rows #6264 and Storage-partitioned self-join of an Iceberg table returns duplicate rows with partially clustered distribution #6278), andregressionwas added to A native final aggregate that has spilled can fail the task during its replay #6254. Neither label is in the guide, so this pass adds neither. The critical bugs below that don't carrycorrectnessare array_contains, arrays_overlap, array_distinct and array_union ignore string collation #6470, Native sort orders null elements of array and struct keys by the key's null order, unlike Spark #6476, RANGE window frames over an array or struct key with a null element span the whole partition #6477, percentile_approx returns different percentiles from Spark when its input holds a NaN with the sign bit set #6519, Native map construction doesn't match Spark 4.0+ float key normalization (-0.0 keys, missing DUPLICATED_MAP_KEY) #6549, Codegen dispatcher initializes kernels with an incorrect partition index under UNION ALL and coalesce #6570 and input_file_name() returns empty values above a converted Spark Parquet scan #6573.bugorenhancementlabel an author had already applied matched the issue's content. The 18 issues that arrived without a type label were classified from their bodies.spark.comet.convert.parquet.enabled, which is off by default, and is labeled critical, as bug: df.count() returns 0 when Comet native scan is disabled and when Comet parquet conversion is enabled #4793 was under the same setting. The native Iceberg writer metadata bugs Native Iceberg writer counts NaNs under NULL structs and NULL list or map entries #6562 and Native Iceberg writer's value and null counts for a float under a NULL struct differ from iceberg-java 1.9+ #6659 stay at medium with Native Iceberg write over-counts NaNs for nested float fields when the input batch is already sliced #6146, since query results are unaffected. Native CASE WHEN names its struct result's fields after the ELSE branch instead of the first THEN branch #6482 and Native CASE WHEN and COALESCE reconcile struct fields by name instead of position #6532 are seen only through nativeto_json, anIncompatibleopt-in, so they follow [Bug] Native make_interval overflows valid time components or loses seconds precision #5131 at medium. See the escalations.spark 4as an area indicator, but the repository only hasspark 4.0/spark 4.1/spark 4.2/spark 3.x, so nothing was applied for it to the Spark 4-only issues array_contains, arrays_overlap, array_distinct and array_union ignore string collation #6470, Native map construction doesn't match Spark 4.0+ float key normalization (-0.0 keys, missing DUPLICATED_MAP_KEY) #6549, Support Spark 4.2 insert-only MERGE execution #6612 and test: move the Spark 4 collation suites to Comet SQL tests #6633.regressionis not in the guide, so this pass didn't add it. It fits Native CASE WHEN names its struct result's fields after the ELSE branch instead of the first THEN branch #6482 and Spark 3.4 SQL test SPARK-34637 fails since #6547 because AQE re-plans flip the DPP join's build side #6645, which come from changes onmainafter 1.1.0 (perf: optimize nativeCASE WHENandIF(up to 12x faster) #6350 and perf: read the shuffle natively in operators AQE reuses from the initial plan #6547).CometNativeCastSuiteTODO: nativetry_castof a float or double equal to 2^63 toBIGINTreturns NULL where Spark returnsLong.MaxValue. test: move expression-level codegen dispatcher tests to Comet SQL tests #6637 says EXPLAIN on Spark 4.x may misreport dispatchedjson_array_lengthandto_json. test: move hash and bitwise expression coverage from Scala suites to Comet SQL tests #6634 saysbit_get's out-of-range error text matches neither Spark 3.x nor 4.x. The first would be silent wrong results if confirmed.Bugs
priority:critical
area:scan,area:expressions; also carriesEPIC,correctnessarea:expressions; also carriesEPIC,correctnessmainwith default configs (max/min,hash, comparisons in aggregates and joins,greatest/least, array functions), and children such as Comparison operators on arrays and structs with floating-point leaves do not match Spark for signed zero #6157 and percentile_approx returns different percentiles from Spark when its input holds a NaN with the sign bit set #6519 are still open. Confirms the author's label.area:expressionsarray_containsreturnsfalsewhere Spark returnstrueand the set functions keep values that Spark merges. Step 1, the same class as CometSort sorts a multi-column key containing a collated string by raw bytes #6158.ASC NULLS LASTorDESC NULLS FIRST, a key holding a null element sorts at the opposite end from Spark, which also changes which rows a TopK keeps and how window order keys rank. That is silent wrong results at step 1, like the collated sort key in CometSort sorts a multi-column key containing a collated string by raw bytes #6158.SUM(id) OVER (ORDER BY array(i))returns the whole partition's total on the null-element row and every row after it, where Spark returns running sums. The plan is native and raises no error, so this is step 1.area:aggregation; also carriescorrectnesscorrreturns 0.878 where Spark returns NULL, or raisesDIVIDE_BY_ZEROunder ANSI, and the variance and covariance functions return tiny non-zero values instead of 0.0. That is silent wrong results at step 1, including the "Comet returns a value where Spark raises" shape. It has been present since 1.0.0.area:aggregationCometApproxPercentilereportsCompatiblefor floats and returns different percentiles at every percentage once a partition merges more than one head buffer. Arithmetic produces sign-bit NaNs on x86-64, so this is step 1, the same class as the other [EPIC] Match Spark's -0.0 and NaN semantics systematically instead of per expression #6385 children.area:expressionsmap_from_entrieskeeps a-0.0key where Spark returns0.0, and bothmap_from_entriesandmap_from_arraysreturn a map where Spark raisesDUPLICATED_MAP_KEY. Both serdes reportCompatible, so this is step 1, and it is the "Comet returns a value where Spark raises" shape that reviewers raised to critical in Duplicate field ids inside a struct are not validated when the file schema equals the requested schema and no predicate is pushed #5801 and Field id gating differs from Spark: root-only check and no dependence on fieldId.read.enabled #5936.area:expressionsTaskContext.partitionId(), so a dispatchedspark_partition_id(),monotonically_increasing_id(),randoruuidin a laterUNION ALLbranch, or belowCometCoalesceExec, sees the task's partition index instead of the index of the partition being computed, which Spark uses. The dispatcher is on by default (spark.comet.exec.scalaUDF.codegen.enabled=true), so this is silent wrong results at step 1 (see escalations).area:scanspark.comet.convert.parquet.enabled=true,input_file_name()and the two block columns return"",-1and-1for every row and the query succeeds, which is step 1. The conversion is off by default; bug: df.count() returns 0 when Comet native scan is disabled and when Comet parquet conversion is enabled #4793, a silent wrong count under the same setting, was labeled critical (see escalations).priority:medium
area:aggregationAdditional allocation failedwhere Spark completes, and more off-heap memory or turning off the native aggregate avoids it. A visible failure with workarounds is step 3, as with A native final aggregate that has spilled can fail the task during its replay #6254 (see escalations).EPICbranch-1.1, and A native final aggregate that has spilled can fail the task during its replay #6254 is a visible failure under memory pressure with workarounds, so step 3.area:shuffle; also carriesspark 4.1,spark 4.2orderIndependentChecksumconfigs, every Comet map output reports a checksum of 0, so Spark's protection against mixed retry output is silently off. The configs are off by default and nothing is wrong unless an indeterminate stage is retried, so step 3. Confirms the author's label.performance,spark 4.2OneRowRelationin Union branches forces Union and downstream aggregates off Comet (TPC-DS q77a) #4949 and Struct-typed scalar subquery result takes the consuming projection off Comet (widened by Spark 4.2 MergeSubplans) #5834. The cause isn't pinned on Comet yet, and q54 is a long-standing gap.area:shuffle; also carriesperformanceCoalesceShufflePartitionsmatches Spark'sUnionExecand notCometUnionExec. Results are correct, so this is performance degradation at step 3. Confirms the author's label.area:expressions; also carriescorrectnessto_jsonunderStructsToJson.allowIncompatible=true; defaultto_jsondispatches and is correct. So it follows the [Bug] Native make_interval overflows valid time components or loses seconds precision #5131 rule at step 3 (see escalations).area:expressionsarea:shuffle; also carriesperformanceOneRowRelationin Union branches forces Union and downstream aggregates off Comet (TPC-DS q77a) #4949 and Struct-typed scalar subquery result takes the consuming projection off Comet (widened by Spark 4.2 MergeSubplans) #5834.area:expressions; also carriescorrectnessto_json, which needsStructsToJson.allowIncompatible=true, so it follows the [Bug] Native make_interval overflows valid time components or loses seconds precision #5131 rule forIncompatibleopt-ins (see escalations).area:scan; also carriesdocumentation,test,area:Icebergarea:writer; also carriesarea:Icebergnan_value_countsis wrong for struct fields on every profile, and for list elements before Iceberg 1.10, while the data and query results match. Wrong metadata from the off-by-default native writer is step 3, as with Native Iceberg write over-counts NaNs for nested float fields when the input batch is already sliced #6146. Confirms the author's label.area:scanarea:scanarea:expressionstry_castavoids it, so step 3. This is the opposite direction from the "Comet returns a value where Spark raises" corrections. Confirms the author's label.area:shuffle,spark sql testsspark.comet.shuffle.directRead.enabled=false), so step 3 (see escalations).area:writer; also carriesarea:Icebergarea:aggregationinteger overflowand notry_addsuggestion instead of Spark'slong overflowparameters. Only the error parameters differ, which is the tier of ANSI integral overflow errors: wrong error class for Byte/Short arithmetic and missing try_ suggestion #6217 and its predecessor ANSI arithmetic overflow errors: wrong error class for Byte/Short, wrong type names, missing try_ suggestions #5071, so step 3.priority:low
area:expressionsEXPLAINannotation is lost under AQE. The join still falls back, and the reason still reaches the fallback log. Diagnostics only, step 4.query tolerance=in Comet SQL tests passes when either side is NaN (#6616)testtolerance=are never compared, so those fixtures can't catch a wrong NaN. Test-only, step 4, as with In-memory cache tests that use checkSparkAnswer compare the cache with itself #6203. Any real mismatch that the fix exposes should be filed separately.area:writer; also carriestest,area:IcebergEnhancements
EPICarea:ffiarea:expressionsto_csvisIncompatible, so its divergences are filed as enhancements. Confirms the author's label.area:ciperformancemain, with no regression identified. If the rising iteration times turn out to be a leak, that part is a bug.area:expressions; also carriesperformancearea:scan_metadatavalues it caused already fall back (Native Parquet scan returns wrong _metadata.file_block_start when Spark splits a file #6505, fix: fall back to Spark for _metadata.file_block_start and file_block_length #6510). Confirms the author's label.area:expressionsIncompatiblepath, which dispatches today, to a native kernel.area:shuffleperformanceOneRowRelation.area:scan; also carriesperformance,EPICCometSparkToColumnarExec.area:aggregation; also carriesarea:memoryarea:expressions; also carriesperformancearea:shuffle; also carriesperformancearea:shuffle; also carriesperformancearea:shuffle; also carriesperformancearea:writerarea:shufflearea:writerarea:writerarea:expressionsarea:expressions; also carriestest,EPICtesttestConfigparser fix it includes affects only fixture values that contain=.documentation,testarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:aggregation; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:expressions; also carriestestarea:scanarea:shuffleEscalations to consider
CometScalaUDFCodegeninitializes kernels withTaskContext.partitionId()onmain, and the dispatcher is on by default. The report comes from the feat: route map lookups through codegen dispatcher #5875 review and has no standalone reproducer yet, so the reviewer may want one before relying on the label.spark.comet.convert.parquet.enabled=true, which is off by default. bug: df.count() returns 0 when Comet native scan is disabled and when Comet parquet conversion is enabled #4793 had the same precondition and was labeled critical, as revertToSpark erases CometIcebergWriteExec / CometNativeWriteExec because originalPlan is the node's own child #5719 was for another off-by-default feature. The reviewer may preferpriority:highunder the "core path over experimental" principle.priority:criticalis confirmed, but its only critical child, Native timestamp_seconds returns a TIMESTAMP_NTZ-typed array, so downstream expressions ignore the session timezone #6328, is closed. The open children are timestamp_trunc panics on DST-transition timestamps in a DST timezone #5633 (priority:high) and days transform is evaluated in the session timezone, while hours and Iceberg use UTC #6333 (priority:low), sopriority:highfits if an EPIC's priority should follow its open children.priority:mediumbug is reported by multiple users or affects a common workload → consider escalating topriority:high". Anycollect_listorcollect_setover a low-cardinality key that spills can hit it. The last pass made the same note for A native final aggregate that has spilled can fail the task during its replay #6254.to_json, which needsallowIncompatible. The wrong field type could also reach native consumers that run by default, such ashashor a cast to string. If one of them returns a different answer from Spark, step 1 applies.to_jsonunderallowIncompatibleexposes the ELSE branch's field names. If a default-config consumer that reads field names turns up, step 1 applies.priority:high: the native Azure store authenticates with a credential other than the one Hadoop's ABFS driver would use. Native Azure store lets ambient AZURE_* environment variables override or corrupt explicit Hadoop auth config #5542 hit the common AKS workload identity setup, while each case here needs an uncommon configuration. If one of them turns out to be common,priority:highfits.HadoopS3ACredentialProviderAdapter, which adds a classpath requirement to configurations that work today.priority:criticalrow, but it needs an opted-in alias scheme with hostless locations, and it hasn't been reproduced.run-spark-3.4-tests. If the reviewer reads it as a test failure,priority:lowfits.priority:low. Comet returns a value where Spark raisesPARTITION_TRANSFORM_EXPRESSION_NOT_IN_PARTITIONED_BY, the shape that reviewers raised to critical in Duplicate field ids inside a struct are not validated when the file schema equals the requested schema and no predicate is pushed #5801 and Field id gating differs from Spark: root-only check and no dependence on fieldId.read.enabled #5936. It is held at low because Spark's error rejects calling a partition transform as an expression rather than reporting something about the data, and it only matters when something evaluatesdays()directly.Skipped (needs more info)
#6576 has no reproduction on a supported path. The others are tracking or record issues rather than bug reports or feature requests.
requires-triagewas left in place on all five, so they reappear in the next pass until they are closed or classified.correctnessandregressionadditions described above, so it can be closed once reviewed.