Skip to content

Improve NDV estimates for sparse join keys #25618

Description

@gabotechs

Describe the bug

An orders–lineitem join underestimates output fourfold despite exact input row
counts. Without explicit NDVs, the estimator uses the key range on the repeated
foreign-key side, although the keys are sparse.

To Reproduce

From the repository root, with the CLI fix from
PR #25570 applied:

cargo build --profile ci --locked -p datafusion-benchmarks --bin dfbench
cargo install tpchgen-cli --version 1.1.1 --locked # if not already installed
repro_dir=$(mktemp -d)
tpchgen-cli --scale-factor 1 --format parquet \
  --parquet-compression 'ZSTD(1)' --parts 1 --output-dir "$repro_dir/data"
cat > "$repro_dir/repro.sql" <<'SQL'
SET datafusion.execution.target_partitions = 1;
SET datafusion.optimizer.enable_dynamic_filter_pushdown = false;
SELECT l_orderkey FROM orders JOIN lineitem ON o_orderkey = l_orderkey;
SQL
target/ci/dfbench statistics \
  --path "$repro_dir/data" --query_path "$repro_dir/repro.sql"

Observed with tpchgen-cli 1.1.1 at
6c320561b5.
Inspect the SELECT reports; ignore the empty SET reports.

Operator Node Estimated rows Actual rows
HashJoinExec 0 1,500,303 6,001,215

Expected behavior

Use reliable NDVs or declared key relationships where available, and distinguish
sparse-domain estimates from range bounds. Preserve sensible behavior for dense
and unmatched keys.

Additional context

Inputs contain 1,500,000 orders and 6,001,215 lineitems. Keys span roughly six
million values but contain only 1.5 million distinct order keys. Related:
#20766.

Part of #25610.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions