Skip to content

feat(clustering): add SuperClusterAlgorithm for mega-scale marker clustering and configurable badge formatting - #1799

Open
dkhawk wants to merge 1 commit into
mainfrom
feat/supercluster-algorithm
Open

dkhawk wants to merge 1 commit into
mainfrom
feat/supercluster-algorithm

Conversation

@dkhawk

@dkhawk dkhawk commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR introduces SuperClusterAlgorithm, a high-performance hierarchical greedy clustering algorithm designed to handle 100,000 to 1,000,000+ markers on Android with sub-millisecond query performance and minimal garbage collection pressure. It also adds configurable cluster badge precision and compact formatting to both DefaultClusterRenderer and DefaultAdvancedMarkersClusterRenderer, alongside an interactive sample demo.

Key Additions & Features

  1. SuperClusterAlgorithm & FlatKdTree:

    • Implements a bottom-up hierarchical zoom pyramid using flat contiguous primitive arrays (DoubleArray, IntArray).
    • Eliminates object allocations (QuadItem, HashSet, HashMap) during runtime pan/zoom, preventing ART GC pauses.
    • Zero-allocation internal ZoomLevel structures and direct flat KD-tree spatial partitioning for coincident and nearby points.
    • Direct raw coordinate buffer ingestion (setCoordinates(coords, itemFactory)) enabling instant loading and lazy item instantiation for 1,000,000 markers.
    • Screen-based viewport range queries execute in $O(\log N + K)$ time (< 1 ms across 100k points).
    • Supports incremental location updates, dynamic item additions/removals, and custom cluster radius configuration.
    • OnClusteringProgressListener callback interface reporting granular progress (0%..100%) during pyramid construction.
  2. Cluster Renderer Formatting Enhancements:

    • Added maxNonZeroDigits: Int (default: 1) for configurable significant digit precision (e.g., 5, 10+, 50+, 100+, 1k+).
    • Added showExactCount: Boolean (default: false) to display exact item counts only when explicitly requested.
    • Added useCompactNumberFormatting: Boolean and compactUnitUppercase: Boolean for SI notation (k/m vs K/M).
    • Implemented reactive clearIconCache() and forceRecluster invalidation on property mutation.
  3. SuperCluster100kDemoActivity:

    • Demonstrates 100,000 markers clustered around the San Francisco Bay Area and 1,000,000 markers distributed across 26 major US metropolitan hubs.
    • Quick dataset switcher card with solid opaque styling and embedded LinearProgressIndicator displaying live indexing progress.
    • Custom unclustered marker renderer displaying playful gremlins at high zoom levels.
    • Smooth logarithmic color spectrum across cluster sizes (2 to 1,000,000+ items).
    • Interactive Material cluster settings dialog allowing real-time adjustment of dataset size, radius, minClusterSize, non-zero digits, and formatting toggles.
  4. Tests & Documentation:

    • Comprehensive unit test suite across 5 test classes (FlatKdTree, SuperClusterAlgorithmTest, SuperClusterInvariantProofTest, SuperClusterRobolectricTest, DefaultClusterRendererTest) using standard MockK and Robolectric with zero external dependencies.
    • Mathematical proof test verifying conservation of total items and cluster size invariants across all zoom levels.
    • Performance benchmarks and architectural comparisons in clustering/README.md.

@dkhawk
dkhawk force-pushed the feat/supercluster-algorithm branch from 7441365 to 6d9edd7 Compare September 30, 2026 03:33
@dkhawk
dkhawk requested a review from kikoso September 30, 2026 16:49
@dkhawk
dkhawk force-pushed the feat/supercluster-algorithm branch from 6d9edd7 to 68f0f73 Compare September 30, 2026 17:54
@dkhawk
dkhawk removed the request for review from kikoso September 30, 2026 17:55
@dkhawk
dkhawk force-pushed the feat/supercluster-algorithm branch 2 times, most recently from ece6c6a to cb6fecb Compare September 30, 2026 20:38
@dkhawk
dkhawk marked this pull request as ready for review September 30, 2026 20:38
@dkhawk
dkhawk requested a review from LoyalAbbas September 30, 2026 20:38
@snippet-bot

snippet-bot Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Here is the summary of changes.

You are about to add 9 region tags.

This comment is generated by snippet-bot.
If you find problems with this result, please file an issue at:
https://github.com/googleapis/repo-automation-bots/issues.
To update this comment, add snippet-bot:force-run label or use the checkbox below:

  • Refresh this comment

@dkhawk
dkhawk requested review from kikoso and removed request for LoyalAbbas September 30, 2026 20:39
Comment thread demo/src/main/res/values/strings.xml Fixed
@dkhawk
dkhawk force-pushed the feat/supercluster-algorithm branch from cb6fecb to 3847032 Compare September 30, 2026 20:53
@googlemaps-bot

googlemaps-bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Code Coverage

Overall Project 58.69% -2.48% 🍏
Files changed 72.85% 🍏

Module Coverage
Kover Gradle Plugin XML report for :clustering 45.16% -7.82% 🍏
Files
Module File Coverage
Kover Gradle Plugin XML report for :clustering FlatKdTree.kt 99.51% -0.49% 🍏
SuperClusterAlgorithm.kt 85.63% -14.37% 🍏
ClusterManager.kt 72.36% -2.64% ❌
SuperCluster.kt 43.33% -56.67% ❌
DefaultClusterRenderer.kt 29.25% -4.9% 🍏
DefaultAdvancedMarkersClusterRenderer.kt 0% -20.59% ❌

@kikoso kikoso left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@dkhawk thanks for this, the flat KD-tree approach looks really nice and the invariant test is a good idea.

Main things before merging: CI is red because :clustering:apiCheck fails, so it needs an ./gradlew apiDump with the new public API (and we should look at the dump, since a lot of new public surface lands here). And the renderer change, which I think affects all existing users and not only SuperCluster ones, see inline.

Also ClusterRendererMultipleItems still has the old bucket logic, is it intended to leave it behind?

protected open fun getDescriptorForCluster(cluster: Cluster<T>): BitmapDescriptor {
val bucket = getBucket(cluster)
var descriptor = mIcons[bucket]
val key = if (showExactCount) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking I think. getDescriptorForCluster doesn't call getBucket() anymore, so nothing in the class reads buckets (the new property) or BUCKETS, and anyone who overrides getBucket() in a subclass gets silently ignored after upgrading.

It also changes the default badges for every existing user, not just SuperCluster ones. With maxNonZeroDigits = 1 a cluster of 350 now shows "300+" where it used to show "200+", and 5,000 shows "5000+" instead of "1000+". Should we keep the old bucket path as the default and only use the significant-digit rounding when one of the new flags is set? Same thing in DefaultAdvancedMarkersClusterRenderer.

val formatted = if (thousands % 1.0 == 0.0) {
"${thousands.toInt()}"
} else {
String.format(Locale.US, "%.1f", thousands)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This can overstate the count. With maxNonZeroDigits = 3, 1,450 rounds down to 1450 but then %.1f rounds half-up to "1.5k+", which reads as "at least 1,500". The + only works if we always truncate, so should we floor here instead of formatting?

Also the compact formatting is basically written three times in this file (here, the millions branch and formatCompactNumber), and then again in DefaultAdvancedMarkersClusterRenderer. Can we pull it into one internal helper with a small unit test, so the two renderers can't drift?

if (useCompactNumberFormatting) {
return formatCompactNumber(bucketOrSize)
}
return String.format(Locale.US, "%,d", bucketOrSize)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Locale.US here means a German user sees "1,234" instead of "1.234". Since this is user visible text on the map, shouldn't we use the default locale (or NumberFormat.getInstance())?


public companion object {
private val BUCKETS = intArrayOf(10, 20, 50, 100, 200, 500, 1000)
public val DEFAULT_BUCKETS: IntArray = intArrayOf(10, 20, 50, 100, 200, 500, 1000)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be private, or a List<Int>? A public IntArray in the companion means anyone can do DEFAULT_BUCKETS[0] = 5 and change it for every renderer in the process, and with BCV now it becomes part of the API we have to keep.

*/
public fun setCoordinates(
coords: DoubleArray,
itemFactory: ((index: Int, position: LatLng) -> T)? = null,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If itemFactory is null we hand out DefaultLeafItem(pos) as T, which is unchecked. For SuperClusterAlgorithm<MyItem> that compiles fine and then blows up with a ClassCastException somewhere in the renderer or in a click listener, far away from where the mistake was made.

Should the factory be required? Or we could have a separate non-generic entry point for the raw-coordinates case.

* Maps to [radius] in pixels to satisfy the [Algorithm] interface contract.
*/
public override var maxDistanceBetweenClusteredItems: Int
get() = mRadius.toInt()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to double check the units: with extent = 512 the radius ends up as radius / 2 in dp (world is 256 * 2^z dp wide), so maxDistanceBetweenClusteredItems = 100 gives 50dp here, while for NonHierarchicalDistanceBasedAlgorithm it is 100dp. People switching algorithms via ClusterManager would see clusters twice as tight.

Should the getter/setter convert, or at least say it in the KDoc? (The demo comment says "140px" too, which is really 70dp.)

leafCoords[2 * leafCount + 1] = y
leafSizes[leafCount] = neighborBuffer.size
leafIndices[leafCount] = -1
leafClusterIds[leafCount] = -(idCounter++)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Small one: cluster ids can collide. The first coincident group gets -(0) = 0 if it's item 0, and coincident groups get -1, -2, ... here while nextClusterId below also starts at -1. equals also compares position and size so the sets are fine, but clusterId is public, so anyone keying by it will get clashes. Maybe one shared counter for both?

val step = maxZoom - z + 1
val progress = 0.15f + 0.82f * (step.toFloat() / totalZoomSteps.toFloat())
val pct = (progress * 100).toInt()
onProgressListener?.onClusteringProgress(progress, "Building zoom level $z/$maxZoom ($pct%)")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These status strings end up in a public listener, but they're hard-coded English and not localizable, so apps can't really show them to users. Should we drop status from OnClusteringProgressListener (or make it a stage enum) and keep just the progress? Also the KDoc mentions -1.0f for indeterminate but we never send it.

Related, ClusterManager now does is SuperClusterAlgorithm<*> in two places to wire this up. Would be cleaner with a small interface the algorithm implements, so the manager doesn't depend on a concrete algorithm.

}
// [END maps_android_utils_benchmark_comparison_runner]

@Test

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This one only prints, there's no assertion, and it runs the old algorithms at 100k on every CI run. Can we @Ignore it (or move it to a benchmark task) and keep the numbers in the README? I suspect it's a good part of why the test job takes almost 5 minutes now.

MyItem(pos.latitude, pos.longitude, "Gremlin #$id", "1M US Supercluster Point")
}
}
val elapsedMs = System.currentTimeMillis() - startBuild

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

setCoordinates / addItems just store data, the pyramid only gets built later inside cluster() on the background thread. So the "Indexed in Xms" toast basically measures the array assignment, not the indexing. Maybe take the time from the progress listener reaching 1.0 instead?

…stering and configurable badge formatting

- Implement SuperClusterAlgorithm with a bottom-up hierarchical zoom pyramid powered by FlatKdTree flat contiguous arrays, enabling sub-millisecond viewport queries on 100k+ markers with minimal GC allocation.
- Optimize SuperClusterAlgorithm internals with zero-allocation flat primitive ZoomLevels, eliminating LinkedHashSet and LinkedHashMap overhead.
- Add raw contiguous buffer ingestion (setCoordinates) for zero-allocation initialization of 1,000,000+ points with lazy item instantiation.
- Add OnClusteringProgressListener to ClusterManager and SuperClusterAlgorithm for granular progress reporting during intensive spatial indexing passes.
- Support location updates, dynamic item addition/removal, and custom cluster radius configuration.
- Add configurable non-zero digit precision (maxNonZeroDigits), compact SI unit formatting (k/m and K/M), and exact count toggling (showExactCount) to DefaultClusterRenderer and DefaultAdvancedMarkersClusterRenderer with automatic icon cache invalidation.
- Add SuperCluster100kDemoActivity showcasing 100,000 markers in California and 1,000,000 markers across the United States, unclustered gremlin markers, logarithmic color stops, and an interactive cluster settings dialog.
- Add quick dataset switcher with opaque card styling, live LinearProgressIndicator progress feedback, and background coroutine loading.
- Include comprehensive performance benchmarks and architecture documentation in README files.
- Add full unit test coverage validating FlatKdTree, SuperCluster spatial partitioning, location updates, progress reporting, mathematical invariants, and renderer label formatting.
@dkhawk
dkhawk force-pushed the feat/supercluster-algorithm branch from 7905b79 to d22e4df Compare October 1, 2026 21:30

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants