Skip to content

Include residual predicates in join estimates #25616

Description

@gabotechs

Describe the bug

The join estimator receives equality keys but no residual filter expression.
With exact scan counts, adding a cross-table inequality roughly halves actual
output without changing the estimate.

To Reproduce

From the repository root, with the CLI fix from
PR #25570 applied:

cargo build --profile ci --locked -p datafusion-benchmarks --bin dfbench
cargo install tpchgen-cli --version 1.1.1 --locked # if not already installed
repro_dir=$(mktemp -d)
tpchgen-cli --scale-factor 1 --format parquet \
  --parquet-compression 'ZSTD(1)' --parts 1 --output-dir "$repro_dir/data"
cat > "$repro_dir/repro.sql" <<'SQL'
SET datafusion.execution.target_partitions = 1;
SET datafusion.optimizer.enable_dynamic_filter_pushdown = false;
SELECT l_orderkey FROM lineitem JOIN part ON l_partkey = p_partkey;

SELECT l_orderkey FROM lineitem JOIN part ON l_partkey = p_partkey AND l_quantity < p_size;
SQL
target/ci/dfbench statistics \
  --path "$repro_dir/data" --query_path "$repro_dir/repro.sql"

Observed with tpchgen-cli 1.1.1 at
6c320561b5.
Inspect the SELECT reports; ignore the empty SET reports.

SELECT HashJoinExec node Estimated rows Actual rows
Equality only 0 6,001,215 6,001,215
Add l_quantity < p_size 0 6,001,215 2,932,286

Expected behavior

Estimate residual predicates from mapped input statistics. Treat inner join pair
counts and semi/anti matching probabilities separately.

Additional context

The first join estimate is accurate, so the second query isolates new error
rather than inherited filter or NDV error.

Part of #25610.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions