Repository navigation
Make CI run on the contributor forks #4289
Description
Activity
- addedpriority:lowMinor issues, test failures, tooling, cosmeticMinor issues, test failures, tooling, cosmeticarea:ciCI/CD, GitHub Actions, build toolingCI/CD, GitHub Actions, build toolingand removed
on May 18, 2026 @comphead could you add some notes on what you are tried already and what you have learned about this? Is this still a good direction to pursue now that we have the merge queue?
@comphead could you add some notes on what you are tried already and what you have learned about this? Is this still a good direction to pursue now that we have the merge queue?
Thanks @andygrove makes sense to me.
Contributor forks is a promising direction following Apache Spark strategy.
I checked all CI ops we do usually expect and it worked well, the thing to keep in mind though: the contributor machines are smaller than machines from ASF shared pool and CI on contributor fork was observed 30-40% slower.However it still might be considered as last resort if local CI doesn't work
Could we start with PR-tier CI on pushes to feature branches in contributor forks, similar to Ozone? I currently open draft PRs in my fork just to run CI. This would let contributors validate changes before opening an upstream PR, while keeping the upstream merge queue unchanged.
Fork pushes would also need to select the test jobs, since the current push policy mainly refreshes caches.
@andygrove , @sunchao
could you take a look? thanks!Thanks @comphead. That looks useful for the Spark SQL and Iceberg suites. I was hoping to run PR-tier checks on fork pushes without opening a draft PR each time. Could we support that alongside the local script?
We keep this on hold for now, its considered as a last resort. Reason being the contributor GH machines are smaller than ASF leading to longer CI
What is the problem the feature request solves?
Follow up on #4281
Currently Comet pipeline runs heavyweight Apache Spark test pipelines and with increased number of PRs/commits the upstream repo resources cannot handle CI runs in terms of stability or performance.
The idea is to follow Apache Spark best practices and shift CI to contributor forks
Some of resources to consider
https://spark.apache.org/contributing.html
https://github.com/apache/spark/blob/master/.github/workflows/notify_test_workflow.yml
Describe the potential solution
No response
Additional context
No response