Skip to content

Make CI run on the contributor forks #4289

Description

@comphead

What is the problem the feature request solves?

Follow up on #4281

Currently Comet pipeline runs heavyweight Apache Spark test pipelines and with increased number of PRs/commits the upstream repo resources cannot handle CI runs in terms of stability or performance.

The idea is to follow Apache Spark best practices and shift CI to contributor forks
Some of resources to consider

https://spark.apache.org/contributing.html
https://github.com/apache/spark/blob/master/.github/workflows/notify_test_workflow.yml

Describe the potential solution

No response

Additional context

No response

Activity

  1. self-assigned this
    on May 11, 2026
  2. added
    priority:lowMinor issues, test failures, tooling, cosmetic
    area:ciCI/CD, GitHub Actions, build tooling
    and removed on May 18, 2026
  3. andygrove commented on Sep 15, 2026

    @andygrove
    Member

    @comphead could you add some notes on what you are tried already and what you have learned about this? Is this still a good direction to pursue now that we have the merge queue?

  4. comphead commented on Sep 15, 2026

    @comphead
    ContributorAuthor

    @comphead could you add some notes on what you are tried already and what you have learned about this? Is this still a good direction to pursue now that we have the merge queue?

    Thanks @andygrove makes sense to me.
    Contributor forks is a promising direction following Apache Spark strategy.
    I checked all CI ops we do usually expect and it worked well, the thing to keep in mind though: the contributor machines are smaller than machines from ASF shared pool and CI on contributor fork was observed 30-40% slower.

    However it still might be considered as last resort if local CI doesn't work

  5. rich7420 commented on Sep 21, 2026

    @rich7420
    Contributor

    Could we start with PR-tier CI on pushes to feature branches in contributor forks, similar to Ozone? I currently open draft PRs in my fork just to run CI. This would let contributors validate changes before opening an upstream PR, while keeping the upstream merge queue unchanged.

    Fork pushes would also need to select the test jobs, since the current push policy mainly refreshes caches.

  6. rich7420 commented on Sep 22, 2026

    @rich7420
    Contributor

    @andygrove , @sunchao
    could you take a look? thanks!

  7. comphead commented on Sep 22, 2026

    @comphead
    ContributorAuthor

    Thanks @rich7420 we currently have a script to run CI locally #5974 as an alternative, would that work?

  8. rich7420 commented on Sep 23, 2026

    @rich7420
    Contributor

    Thanks @comphead. That looks useful for the Spark SQL and Iceberg suites. I was hoping to run PR-tier checks on fork pushes without opening a draft PR each time. Could we support that alongside the local script?

  9. comphead commented on Sep 23, 2026

    @comphead
    ContributorAuthor

    We keep this on hold for now, its considered as a last resort. Reason being the contributor GH machines are smaller than ASF leading to longer CI

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:ciCI/CD, GitHub Actions, build toolingenhancementNew feature or requestpriority:lowMinor issues, test failures, tooling, cosmetic

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions