Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -97,4 +97,34 @@ The following cases are not supported by Comet and always fall back to Spark, re
- Descending order in `WITHIN GROUP (ORDER BY ... DESC)` is not supported.
- Only numeric input types are supported.

## RegrIntercept

The following incompatibilities cause `RegrIntercept` to fall back to Spark by default. Set `spark.comet.expression.RegrIntercept.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_intercept` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrR2

The following incompatibilities cause `RegrR2` to fall back to Spark by default. Set `spark.comet.expression.RegrR2.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_r2` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrReplacement

The following incompatibilities cause `RegrReplacement` to fall back to Spark by default. Set `spark.comet.expression.RegrReplacement.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_sxx` and `regr_syy` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrSXY

The following incompatibilities cause `RegrSXY` to fall back to Spark by default. Set `spark.comet.expression.RegrSXY.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_sxy` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrSlope

The following incompatibilities cause `RegrSlope` to fall back to Spark by default. Set `spark.comet.expression.RegrSlope.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_slope` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

<!--END:EXPR_COMPAT-->
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,12 @@ By default, `ArrayContains` is evaluated in the JVM using Spark's own code-gener

- Spark compares array elements with ordering.equiv, so -0.0 matches +0.0 and all NaNs match each other; Comet's native array_contains compares the raw Arrow values bitwise

## ArrayDistinct

The following incompatibilities cause `ArrayDistinct` to fall back to Spark by default. Set `spark.comet.expression.ArrayDistinct.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Floating-point elements match Spark's signed-zero and NaN semantics natively only on Spark 4.2.0, whose optimizer normalizes the arguments (SPARK-54918)

## ArrayExcept

By default, `ArrayExcept` is evaluated in the JVM using Spark's own code-generated implementation (run inside the Comet pipeline), which matches Spark exactly. Set `spark.comet.expression.ArrayExcept.allowIncompatible=true` to opt into Comet's native implementation instead, which has the following differences from Spark:
Expand All @@ -47,6 +53,12 @@ By default, `ArrayJoin` is evaluated in the JVM using Spark's own code-generated
- array_join does not propagate non-UTF8_BINARY collations to the output string (https://github.com/apache/datafusion-comet/issues/2190)
- array_join evaluates its delimiter and null replacement eagerly, while Spark short-circuits past them (https://github.com/apache/datafusion-comet/issues/3178)

## ArrayUnion

The following incompatibilities cause `ArrayUnion` to fall back to Spark by default. Set `spark.comet.expression.ArrayUnion.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Floating-point elements match Spark's signed-zero and NaN semantics natively only on Spark 4.2.0, whose optimizer normalizes the arguments (SPARK-54918)

## ArraysZip

The following cases are not supported by Comet and always fall back to Spark, regardless of any `allowIncompatible` setting:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -97,4 +97,34 @@ The following cases are not supported by Comet and always fall back to Spark, re
- Descending order in `WITHIN GROUP (ORDER BY ... DESC)` is not supported.
- Only numeric input types are supported.

## RegrIntercept

The following incompatibilities cause `RegrIntercept` to fall back to Spark by default. Set `spark.comet.expression.RegrIntercept.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_intercept` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrR2

The following incompatibilities cause `RegrR2` to fall back to Spark by default. Set `spark.comet.expression.RegrR2.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_r2` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrReplacement

The following incompatibilities cause `RegrReplacement` to fall back to Spark by default. Set `spark.comet.expression.RegrReplacement.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_sxx` and `regr_syy` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrSXY

The following incompatibilities cause `RegrSXY` to fall back to Spark by default. Set `spark.comet.expression.RegrSXY.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_sxy` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrSlope

The following incompatibilities cause `RegrSlope` to fall back to Spark by default. Set `spark.comet.expression.RegrSlope.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_slope` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

<!--END:EXPR_COMPAT-->
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,12 @@ By default, `ArrayContains` is evaluated in the JVM using Spark's own code-gener

- Spark compares array elements with ordering.equiv, so -0.0 matches +0.0 and all NaNs match each other; Comet's native array_contains compares the raw Arrow values bitwise

## ArrayDistinct

The following incompatibilities cause `ArrayDistinct` to fall back to Spark by default. Set `spark.comet.expression.ArrayDistinct.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Floating-point elements match Spark's signed-zero and NaN semantics natively only on Spark 4.2.0, whose optimizer normalizes the arguments (SPARK-54918)

## ArrayExcept

By default, `ArrayExcept` is evaluated in the JVM using Spark's own code-generated implementation (run inside the Comet pipeline), which matches Spark exactly. Set `spark.comet.expression.ArrayExcept.allowIncompatible=true` to opt into Comet's native implementation instead, which has the following differences from Spark:
Expand All @@ -47,6 +53,12 @@ By default, `ArrayJoin` is evaluated in the JVM using Spark's own code-generated
- array_join does not propagate non-UTF8_BINARY collations to the output string (https://github.com/apache/datafusion-comet/issues/2190)
- array_join evaluates its delimiter and null replacement eagerly, while Spark short-circuits past them (https://github.com/apache/datafusion-comet/issues/3178)

## ArrayUnion

The following incompatibilities cause `ArrayUnion` to fall back to Spark by default. Set `spark.comet.expression.ArrayUnion.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Floating-point elements match Spark's signed-zero and NaN semantics natively only on Spark 4.2.0, whose optimizer normalizes the arguments (SPARK-54918)

## ArraysZip

The following cases are not supported by Comet and always fall back to Spark, regardless of any `allowIncompatible` setting:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -106,4 +106,34 @@ The following cases are not supported by Comet and always fall back to Spark, re
- Descending order in `WITHIN GROUP (ORDER BY ... DESC)` is not supported.
- Only numeric input types are supported.

## RegrIntercept

The following incompatibilities cause `RegrIntercept` to fall back to Spark by default. Set `spark.comet.expression.RegrIntercept.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_intercept` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrR2

The following incompatibilities cause `RegrR2` to fall back to Spark by default. Set `spark.comet.expression.RegrR2.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_r2` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrReplacement

The following incompatibilities cause `RegrReplacement` to fall back to Spark by default. Set `spark.comet.expression.RegrReplacement.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_sxx` and `regr_syy` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrSXY

The following incompatibilities cause `RegrSXY` to fall back to Spark by default. Set `spark.comet.expression.RegrSXY.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_sxy` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrSlope

The following incompatibilities cause `RegrSlope` to fall back to Spark by default. Set `spark.comet.expression.RegrSlope.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_slope` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

<!--END:EXPR_COMPAT-->
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,12 @@ By default, `ArrayContains` is evaluated in the JVM using Spark's own code-gener

- Spark compares array elements with ordering.equiv, so -0.0 matches +0.0 and all NaNs match each other; Comet's native array_contains compares the raw Arrow values bitwise

## ArrayDistinct

The following incompatibilities cause `ArrayDistinct` to fall back to Spark by default. Set `spark.comet.expression.ArrayDistinct.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Floating-point elements match Spark's signed-zero and NaN semantics natively only on Spark 4.2.0, whose optimizer normalizes the arguments (SPARK-54918)

## ArrayExcept

By default, `ArrayExcept` is evaluated in the JVM using Spark's own code-generated implementation (run inside the Comet pipeline), which matches Spark exactly. Set `spark.comet.expression.ArrayExcept.allowIncompatible=true` to opt into Comet's native implementation instead, which has the following differences from Spark:
Expand All @@ -47,6 +53,12 @@ By default, `ArrayJoin` is evaluated in the JVM using Spark's own code-generated
- array_join does not propagate non-UTF8_BINARY collations to the output string (https://github.com/apache/datafusion-comet/issues/2190)
- array_join evaluates its delimiter and null replacement eagerly, while Spark short-circuits past them (https://github.com/apache/datafusion-comet/issues/3178)

## ArrayUnion

The following incompatibilities cause `ArrayUnion` to fall back to Spark by default. Set `spark.comet.expression.ArrayUnion.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Floating-point elements match Spark's signed-zero and NaN semantics natively only on Spark 4.2.0, whose optimizer normalizes the arguments (SPARK-54918)

## ArraysZip

The following cases are not supported by Comet and always fall back to Spark, regardless of any `allowIncompatible` setting:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -106,4 +106,34 @@ The following cases are not supported by Comet and always fall back to Spark, re
- Descending order in `WITHIN GROUP (ORDER BY ... DESC)` is not supported.
- Only numeric input types are supported.

## RegrIntercept

The following incompatibilities cause `RegrIntercept` to fall back to Spark by default. Set `spark.comet.expression.RegrIntercept.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_intercept` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrR2

The following incompatibilities cause `RegrR2` to fall back to Spark by default. Set `spark.comet.expression.RegrR2.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_r2` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrReplacement

The following incompatibilities cause `RegrReplacement` to fall back to Spark by default. Set `spark.comet.expression.RegrReplacement.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_sxx` and `regr_syy` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrSXY

The following incompatibilities cause `RegrSXY` to fall back to Spark by default. Set `spark.comet.expression.RegrSXY.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_sxy` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

## RegrSlope

The following incompatibilities cause `RegrSlope` to fall back to Spark by default. Set `spark.comet.expression.RegrSlope.allowIncompatible=true` to enable Comet acceleration despite these differences.

- Comet merges the partial aggregates of `regr_slope` in a different floating-point operation order from Spark. When a group's rows come from more than one partial aggregate and a variable is constant at a value that binary floating point cannot represent exactly, such as 0.1, Comet returns a wrong value where Spark returns NULL, 0.0 or 1.0 (https://github.com/apache/datafusion-comet/issues/6423)

<!--END:EXPR_COMPAT-->
Loading