Skip to content

Bug triage results: 2026-10-05 #6665

Description

@andygrove

Triage pass over the open requires-triage queue, per the project Bug Triage Guide.

  • Date: 2026-10-05
  • Total issues processed: 87 (82 triaged, 5 skipped, 0 failed)
  • Type counts: 31 bugs, 51 enhancements
  • Priority counts applied: priority:critical 10, priority:high 0, priority:medium 17, priority:low 4
  • Guide: docs/source/contributor-guide/bug_triage.md

Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied. Corrections should be made directly on the affected issue.

Notes on this pass:

Bugs

priority:critical

  • [EPIC] Timezone handling bugs (#6335)
    • Area labels: area:scan, area:expressions; also carries EPIC, correctness
    • Rationale: An EPIC for timezone bugs, several of which returned silent wrong results. Confirms the author's label; see the escalations, since its only critical child is closed.
  • [EPIC] Match Spark's -0.0 and NaN semantics systematically instead of per expression (#6385)
  • array_contains, arrays_overlap, array_distinct and array_union ignore string collation (#6470)
  • Native sort orders null elements of array and struct keys by the key's null order, unlike Spark (#6476)
  • RANGE window frames over an array or struct key with a null element span the whole partition (#6477)
    • Area labels: none
    • Rationale: SUM(id) OVER (ORDER BY array(i)) returns the whole partition's total on the null-element row and every row after it, where Spark returns running sums. The plan is native and raises no error, so this is step 1.
  • Native corr, covariance, variance and stddev return wrong values for a constant fractional column merged from several partitions (#6481)
    • Area labels: area:aggregation; also carries correctness
    • Rationale: On default configs corr returns 0.878 where Spark returns NULL, or raises DIVIDE_BY_ZERO under ANSI, and the variance and covariance functions return tiny non-zero values instead of 0.0. That is silent wrong results at step 1, including the "Comet returns a value where Spark raises" shape. It has been present since 1.0.0.
  • percentile_approx returns different percentiles from Spark when its input holds a NaN with the sign bit set (#6519)
  • Native map construction doesn't match Spark 4.0+ float key normalization (-0.0 keys, missing DUPLICATED_MAP_KEY) (#6549)
  • Codegen dispatcher initializes kernels with an incorrect partition index under UNION ALL and coalesce (#6570)
    • Area labels: area:expressions
    • Rationale: The dispatcher initializes each kernel with TaskContext.partitionId(), so a dispatched spark_partition_id(), monotonically_increasing_id(), rand or uuid in a later UNION ALL branch, or below CometCoalesceExec, sees the task's partition index instead of the index of the partition being computed, which Spark uses. The dispatcher is on by default (spark.comet.exec.scalaUDF.codegen.enabled=true), so this is silent wrong results at step 1 (see escalations).
  • input_file_name() returns empty values above a converted Spark Parquet scan (#6573)

priority:medium

priority:low

  • days transform is evaluated in the session timezone, while hours and Iceberg use UTC (#6333)
    • Area labels: area:expressions
    • Rationale: Spark never evaluates these partition transforms, so there is no Spark answer to match, and the author rated it low. Confirms the author's label (see escalations).
  • Broadcast/hash join fallback reasons are lost from the AQE-final plan (#6442)
    • Area labels: none
    • Rationale: Only the EXPLAIN annotation is lost under AQE. The join still falls back, and the reason still reaches the fallback log. Diagnostics only, step 4.
  • test: query tolerance= in Comet SQL tests passes when either side is NaN (#6616)
  • Flaky test: CometIcebergWriteActionSuite "a failed write job deletes the data files of tasks that completed" (#6643)
    • Area labels: area:writer; also carries test, area:Iceberg
    • Rationale: An intermittent test failure, step 4. The likely cause is in the test: its UDF blocks a native runtime thread while it waits for the other tasks.

Enhancements

Escalations to consider

Skipped (needs more info)

#6576 has no reproduction on a supported path. The others are tracking or record issues rather than bug reports or feature requests. requires-triage was left in place on all five, so they reappear in the next pass until they are closed or classified.

  • Arrow struct writer should respect projected schema width (#6576)
  • Audit the PRs in 1.1.0 for regressions since 1.0.0 (#6399)
  • [EPIC] Bug fixes to consider backporting to branch-1.0 (#6201)
    • A release-management tracker, skipped on the same grounds as in the 2026-09-28 pass.
  • Bug triage results: 2026-09-28 (#6321)
    • The summary of the 2026-09-28 pass. Since then, its issues have had no priority changes, only the correctness and regression additions described above, so it can be closed once reviewed.
  • Bug triage results: 2026-08-24 (#5454)
    • The summary of the 2026-08-24 pass. The 2026-09-28 pass found nothing outstanding in it, so it can be closed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions