Skip to content

[EPIC] Timezone handling bugs #6335

Description

@andygrove

What / Why

I audited how Comet handles timezones, starting from #2730. The model is simple and mostly sound. Spark's TimestampType is a UTC instant, so nothing is converted at the JVM/native boundary. Comet passes the raw microseconds in both directions and labels them Timestamp(Microsecond, "UTC"). TimestampNTZType is Timestamp(Microsecond, None). The session timezone never becomes part of a value. Each timezone-aware expression carries it, and it's applied inside the native kernel, or the expression runs through the codegen dispatcher with Spark's own timeZoneId. #6337 adds a contributor guide page that describes the model in more detail.

The bugs cluster where that model breaks down:

  • native expressions that emit a TimestampType value with some other label
  • session timezone IDs that the native parser can't read
  • timezone rules that come from a different database than the JVM's

Label drift is easy to miss. The scan and shuffle boundaries cast every column back to its declared type, so a test that only projects the result passes. It shows up when the result is compared, goes through a CASE, or feeds another native expression.

All of the new bugs below reproduce on main at 764936187, on Spark 3.5 and 4.1.

Bugs

Status (2026-10-07)

Seven of the eight bugs are fixed on main. #6351 also replaced every getOrElse("UTC") fallback with CometTimeZone.nativeId, and #6347 turned the asserts in array_with_timezone into errors, so the section on #2730 below describes the code before those changes.

For #5633, #5956 replaced the literal-format kernel. It truncates in local time the way Spark does, follows LocalDate.atStartOfDay for the date levels, and handles the whole timestamp range in UTC sessions. When the format comes from a column, date_trunc still goes through the old helpers and can panic, and #6354 is being reworked on top of #5956 to cover that path.

#6686 brought the contributor guide page up to date, and #6340 points the PR review skills at it. Both are merged.

Follow-ups:

The UTC fallbacks in #2730

I instrumented every timeZoneId.getOrElse("UTC") site. Then I ran the datetime, cast, SQL-file, JSON, CSV, fuzz and expression suites. About 11,700 serde calls happened across 1,020 tests, and about 900 of them arrived without a timezone. Every one of those was a cast that Spark doesn't consider timezone-sensitive: numeric casts, Comet's own nullability-widening casts, and the cast inside IntegralDivide. None was a timezone-aware expression, which fits Spark refusing to resolve one without a timezone. So the fallback isn't a correctness bug today. The helper proposed in #6329 would replace it. The fallback can't simply be removed, though, because array_with_timezone asserts a non-empty timezone even for casts that don't use one.

Related

Already documented: Python Arrow UDFs see timestamps labelled UTC rather than the session timezone, and chrono-tz's DST rules end around 2100. Not in the user guide yet: spark.sql.parquet.int96TimestampConversion=true disables Comet for the session.

Fixed earlier, same class: #2720 (SparkToColumnar labelled timestamps with the session timezone), #2649 via #4761 (the date_trunc schema mismatch, whose fix introduced the label in #6330), and #5556 (the Python runner accepts Etc/UTC for UTC).

Test gaps

The SQL-file tests use UTC, America/Los_Angeles, America/New_York, Asia/Kolkata and +05:30. None of them use Etc/UTC, or the offset and short-ID forms from #6329. Most expression tests only project their result. Adding GMT+8 to the datetime files' ConfigMatrix would have caught #6329. #6328, #6330 and #6327 need more than a timezone setting: a test that compares each native timestamp-returning expression with another timestamp, or uses it in a CASE, in both a non-UTC session and Etc/UTC.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

EPICarea:expressionsExpression evaluationarea:scanParquet scan / data readingbugSomething isn't workingcorrectnesspriority:criticalData corruption, silent wrong results, security issues

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions