Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -200,6 +200,14 @@ Known limitations of the paging loop:
- **Concurrent writes.** Offset paging is not keyset-stable: a row inserted or deleted by another client while a multi-page read is in flight shifts the remaining rows by one offset, so a row can be missed or duplicated until the next refetch.
- **Count cost.** Every collection read issues one `Prefer: count=exact` on its first request so the loop knows when to stop. On very large tables, or under heavy RLS, that is a full `count(*)` per read.

### Oversized `IN` lists

Every read is a PostgREST `GET`, so every filter — including a large `inArray(...)`, the shape a lazy join's on-demand collection produces for its join keys — lives in the query string. Supabase's API gateway rejects request lines over about 8 KB with `414 URI Too Long`, and postgrest-js's own `urlLengthLimit` (default 8000) is purely diagnostic: it only adds a hint to the error, it never splits or reroutes the request.

This adapter does the splitting itself. When a `loadSubset`'s rendered URL would exceed a fixed 8000-character budget (matching postgrest-js's own `urlLengthLimit` default), it slices the request's largest top-level `inArray(...)` filter into several disjoint requests that each fit, runs them with a concurrency cap, and concatenates the results. This is transparent to the rest of the collection: the query key stays the full, unchunked subset, so one query still owns every row it returns, and `useLiveQuery` code needs no changes.

Only a positive, top-level AND-ed `inArray(...)` is ever split — never a `not(inArray(...))`, and never one nested inside an `or(...)`, since chunking either would change which rows the union of requests matches. When no such filter exists, or the subset still cannot be made to fit, the request is sent as-is and the server's response (success or error) surfaces exactly as it would without this adapter.

**Evaluated Client-Side**

These operations fetch the required rows and process them in memory:
Expand Down
132 changes: 118 additions & 14 deletions src/functions.ts
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ import {
cursorWhereFromToSearch,
keyColumnsToSearch,
loadSubsetOptionsToSearch,
splitLoadSubsetOptions,
subsetParamsToSearch,
} from "./postgrest-filters"
import { postgrestRequest } from "./postgrest-request"
Expand Down Expand Up @@ -118,19 +119,18 @@ async function fetchAllPages(
return rows
}

export const supabaseQueryFn = async (
/**
* Read one already-in-budget subset's full window: the keyset tie request (if
* any) plus the main paged request. This is the entire body `supabaseQueryFn`
* used to run directly; it is now also what runs once per chunk when
* `splitLoadSubsetOptions` had to split an oversized subset into several.
*/
async function loadWindow(
supabase: SupabaseClient,
tableName: string,
ctx: {
client: QueryClient
queryKey: readonly unknown[]
signal: AbortSignal
meta: QueryMeta | undefined
pageParam?: unknown
direction?: unknown
}
) => {
const options = ctx.meta?.loadSubsetOptions ?? {}
options: LoadSubsetOptions,
signal: AbortSignal
) {
const search = loadSubsetOptionsToSearch(options)
// Cursor reads pin their own window start, so they never carry an offset (see
// loadSubsetOptionsToSearch); only a cursor-less read advances from `offset`.
Expand All @@ -152,7 +152,7 @@ export const supabaseQueryFn = async (
return await fetchAllPages(supabase, tableName, search, {
limit: options.limit,
offset: startOffset,
signal: ctx.signal,
signal,
})
}

Expand All @@ -162,19 +162,123 @@ export const supabaseQueryFn = async (
// larger than the cap is not truncated.
fetchAllPages(supabase, tableName, tiesSearch, {
offset: 0,
signal: ctx.signal,
signal,
}),
fetchAllPages(supabase, tableName, search, {
limit: options.limit,
offset: startOffset,
signal: ctx.signal,
signal,
}),
])
// `whereCurrent` (== boundary) and `whereFrom` (> / < boundary) are disjoint,
// so concatenation never duplicates; the collection re-sorts locally.
return [...ties, ...rows]
}

// PostgREST has no body-carried filters for reads and the Supabase API gateway
// rejects request lines over about 8 KB with 414, so a big `inArray(...)`
// predicate (the common shape a lazy join's on-demand collection produces)
// can make a subset's URL too long to send. This matches postgrest-js's own
// `urlLengthLimit` default so the two agree on what counts as "too long".
export const MAX_URL_LENGTH = 8000

// At most this many chunk requests run at once; splitting a large IN list can
// produce far more chunks than are worth having in flight simultaneously.
const MAX_CONCURRENT_CHUNKS = 6

/**
* The length of the longest URL a subset would produce: the base table URL
* plus the main request's query string, and — when the subset has a cursor —
* the unlimited `whereCurrent` tie request's query string too, since
* `loadWindow` sends both for a cursor read. Measuring rendered URLs (rather
* than estimating) accounts for quoting and percent-encoding exactly.
*/
function measureSubset(
supabase: SupabaseClient,
tableName: string,
options: LoadSubsetOptions
): number {
const base = supabase.from(tableName).url.toString().length
const mainLength = base + loadSubsetOptionsToSearch(options).toString().length
const tiesSearch = cursorCurrentToSearch(options)
return tiesSearch
? Math.max(mainLength, base + tiesSearch.toString().length)
: mainLength
}

/**
* Runs one `loadWindow` per chunk with at most `MAX_CONCURRENT_CHUNKS` in
* flight at a time, and concatenates their rows. Chunks are disjoint on the
* split column (see `splitLoadSubsetOptions`), so concatenation order does not
* matter and needs no dedupe.
*
* On the first failure, every other chunk's request is aborted through a
* local `AbortController` combined with the caller's own signal
* (`AbortSignal.any`) — this keeps the same all-or-nothing semantics an
* unsplit request has ("one failed page fails the whole load") instead of
* surfacing a partial result, and the first error encountered is what
* `Promise.all` rejects with.
*/
async function loadChunks(
supabase: SupabaseClient,
tableName: string,
chunks: LoadSubsetOptions[],
signal: AbortSignal
): Promise<any[]> {
const controller = new AbortController()
const combinedSignal = AbortSignal.any([signal, controller.signal])
const results: any[][] = new Array(chunks.length)
let nextIndex = 0

const runWorker = async (): Promise<void> => {
while (nextIndex < chunks.length) {
const index = nextIndex++
try {
results[index] = await loadWindow(
supabase,
tableName,
chunks[index],
combinedSignal
)
} catch (error) {
controller.abort()
throw error
}
}
}

const workerCount = Math.min(MAX_CONCURRENT_CHUNKS, chunks.length)
await Promise.all(Array.from({ length: workerCount }, runWorker))
return results.flat()
}

export const supabaseQueryFn = async (
supabase: SupabaseClient,
tableName: string,
ctx: {
client: QueryClient
queryKey: readonly unknown[]
signal: AbortSignal
meta: QueryMeta | undefined
pageParam?: unknown
direction?: unknown
}
) => {
const options = ctx.meta?.loadSubsetOptions ?? {}
// The query key stays the full, unchunked subset (subsetOptionsToQueryKey
// never sees these chunks), so this fan-out is invisible to
// query-db-collection: one query still owns every row it returns.
const chunks = splitLoadSubsetOptions(options, {
maxUrlLength: MAX_URL_LENGTH,
measure: (chunkOptions) => measureSubset(supabase, tableName, chunkOptions),
})

if (chunks.length === 1) {
return await loadWindow(supabase, tableName, chunks[0], ctx.signal)
}
return await loadChunks(supabase, tableName, chunks, ctx.signal)
}

export const supabaseOnInsert = async (
supabase: SupabaseClient,
tableName: string,
Expand Down
4 changes: 2 additions & 2 deletions src/postgrest-filters/common.ts
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,7 @@ export function paramsToSearch(params: PostgrestParam[]): URLSearchParams {
return search
}

function flattenAnd(expr: Expression): Expression[] {
export function flattenAnd(expr: Expression): Expression[] {
return expr.type === "func" && expr.name === "and"
? expr.args.flatMap(flattenAnd)
: [expr]
Expand Down Expand Up @@ -168,7 +168,7 @@ function toFilterString(
// Preserve the live adapter's union of duplicate IN demands. This deliberately
// fetches a superset, which TanStack re-filters. Never do this inside OR/NOT or
// for server aggregates, where client-side filtering cannot repair the result.
function mergeInFilters(filters: Expression[]): Expression[] {
export function mergeInFilters(filters: Expression[]): Expression[] {
const merged: Expression[] = []
const valuesByField = new Map<string, unknown[]>()
for (const filter of filters) {
Expand Down
4 changes: 4 additions & 0 deletions src/postgrest-filters/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,4 +6,8 @@ export {
} from "./load-subset-options-to-search"
export { queryIrToSearch } from "./query-ir-to-search"
export { realtimeFiltersToSearch } from "./realtime-filters-to-search"
export {
type SplitLoadSubsetOptionsConfig,
splitLoadSubsetOptions,
} from "./split-load-subset-options"
export { subsetParamsToSearch } from "./subset-params-to-search"
Loading
Loading