Skip to content

[6.x] Fix quadratic required-word matching in Comb search - #15578

Open
edalzell wants to merge 1 commit into
statamic:6.xfrom
edalzell:fix/comb-required-words-performance
Open

edalzell wants to merge 1 commit into
statamic:6.xfrom
edalzell:fix/comb-required-words-performance

Conversation

@edalzell

@edalzell edalzell commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Comb's boolean mode (the default) checks +word with an unanchored lookahead, #(?=.*zakat)#iu. On a record without the word, it retries from every position and scans to the end each time: quadratic in record length.

On a real site (2.4k–3.7k documents, 5–16 MB per index), +zakat -ramadan took 9.7s, 18.5s and 26.9s on three indexes, versus ~0.5s with non-matching documents filtered out first.

Fix: one preg_match('#'.$word.'#iu', $record) per required word, a linear scan each. Anchoring the old pattern instead hits pcre.backtrack_limit on records over ~1 MB and silently drops them.

Words are still not preg_quoted, same as before. One intentional change: required words on different lines of a multi-line field now match (. didn't cross newlines).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant