Describe the bug
Since #4761, TimestampTruncExpr stamps its output with the session timezone and declares Timestamp(Microsecond, <session timezone>) as its type (native/spark-expr/src/datetime_funcs/timestamp_trunc.rs:93, native/spark-expr/src/kernels/temporal.rs:132 and :834). Everything else in a native plan represents TimestampType as Timestamp(Microsecond, "UTC"). Arrow's comparison kernels require identical types, so comparing a truncated timestamp with any other timestamp fails.
On its own, that would only matter with spark.comet.expression.TruncTimestamp.allowIncompatible=true. But CometTruncTimestamp treats Etc/UTC as UTC (spark/src/main/scala/org/apache/comet/serde/datetime.scala:610) and runs the native path by default there, and the output is then labelled Etc/UTC. Etc/UTC is the JVM's default timezone on Linux hosts and containers whose /etc/localtime points at it, which is the default on Ubuntu and Debian images. That makes it Spark's default spark.sql.session.timeZone there too.
In an Etc/UTC session, comparing date_trunc(...) with another timestamp fails. That covers =, >=, <=, BETWEEN, <=>, nullif and join conditions. CASE and coalesce panic instead (#6327). IF, IN lists, greatest, equi-join keys, UNION, aggregates, sorts and windows all work.
Steps to reproduce
On main at 764936187, with the default config, on Spark 3.5 and 4.1:
CREATE TABLE events USING parquet AS
SELECT * FROM VALUES (TIMESTAMP'2024-01-15T18:30:45Z'), (TIMESTAMP'2024-06-30T23:30:00Z') AS v(ts);
SET spark.sql.session.timeZone=Etc/UTC;
SELECT ts FROM events WHERE date_trunc('DAY', ts) >= TIMESTAMP'2024-06-01 00:00:00';
Spark returns 2024-06-30 23:30:00. Comet fails with Invalid argument error: Invalid comparison operation: Timestamp(µs, "Etc/UTC") >= Timestamp(µs, "UTC"). The same query works with spark.sql.session.timeZone=UTC. With allowIncompatible=true it fails the same way in every other zone, for example Timestamp(µs, "Asia/Tokyo") >= Timestamp(µs, "UTC").
Expected behavior
The same result as Spark.
Additional context
This has been in place since 1.0.0. #5556 already had to accept Etc/UTC and UTC as equivalent in the Python runner because of this label. Rather than teaching each consumer, could TimestampTruncExpr do the truncation in the session timezone but keep the input's label on the output, and return the child's type from data_type()? Declared and actual types would still agree, so the RowConverter mismatch #4761 fixed stays fixed. #5956 changes the same kernel. Normalizing the session timezone (#6329) would hide this for Etc/UTC, but the label would still leak in every other zone under allowIncompatible.
The incompatibility reason for non-UTC date_trunc still cites #2649, which #4761 closed. The remaining non-UTC gaps are #5633 and the chrono-tz horizon described in the datetime compatibility guide.
Describe the bug
Since #4761,
TimestampTruncExprstamps its output with the session timezone and declaresTimestamp(Microsecond, <session timezone>)as its type (native/spark-expr/src/datetime_funcs/timestamp_trunc.rs:93,native/spark-expr/src/kernels/temporal.rs:132and:834). Everything else in a native plan representsTimestampTypeasTimestamp(Microsecond, "UTC"). Arrow's comparison kernels require identical types, so comparing a truncated timestamp with any other timestamp fails.On its own, that would only matter with
spark.comet.expression.TruncTimestamp.allowIncompatible=true. ButCometTruncTimestamptreatsEtc/UTCas UTC (spark/src/main/scala/org/apache/comet/serde/datetime.scala:610) and runs the native path by default there, and the output is then labelledEtc/UTC.Etc/UTCis the JVM's default timezone on Linux hosts and containers whose/etc/localtimepoints at it, which is the default on Ubuntu and Debian images. That makes it Spark's defaultspark.sql.session.timeZonethere too.In an
Etc/UTCsession, comparingdate_trunc(...)with another timestamp fails. That covers=,>=,<=,BETWEEN,<=>,nullifand join conditions.CASEandcoalescepanic instead (#6327).IF,INlists,greatest, equi-join keys,UNION, aggregates, sorts and windows all work.Steps to reproduce
On
mainat764936187, with the default config, on Spark 3.5 and 4.1:Spark returns
2024-06-30 23:30:00. Comet fails withInvalid argument error: Invalid comparison operation: Timestamp(µs, "Etc/UTC") >= Timestamp(µs, "UTC"). The same query works withspark.sql.session.timeZone=UTC. WithallowIncompatible=trueit fails the same way in every other zone, for exampleTimestamp(µs, "Asia/Tokyo") >= Timestamp(µs, "UTC").Expected behavior
The same result as Spark.
Additional context
This has been in place since 1.0.0. #5556 already had to accept
Etc/UTCandUTCas equivalent in the Python runner because of this label. Rather than teaching each consumer, couldTimestampTruncExprdo the truncation in the session timezone but keep the input's label on the output, and return the child's type fromdata_type()? Declared and actual types would still agree, so theRowConvertermismatch #4761 fixed stays fixed. #5956 changes the same kernel. Normalizing the session timezone (#6329) would hide this forEtc/UTC, but the label would still leak in every other zone underallowIncompatible.The incompatibility reason for non-UTC
date_truncstill cites #2649, which #4761 closed. The remaining non-UTC gaps are #5633 and the chrono-tz horizon described in the datetime compatibility guide.