Skip to content

ci: govulncheck step should retry transient vuln-DB fetch failures #74

Description

@cristim

The Security Scanning job's govulncheck step died mid-scan on a transient network error (fetching vulnerabilities: Get https://vuln.go.dev/index/modules.json.gz: connection reset by peer) on run 29516272147, cascading into the SARIF upload steps and a red job despite 0 actual vulnerabilities (verified locally across all 6 modules at the pinned v1.1.4). Add a small bounded retry (e.g. 3 attempts, backoff) around the per-module govulncheck invocation in .github/workflows/ci.yml so a dropped connection to the live vuln DB does not fail CI. Keep the pin; a genuine vulnerability finding must still fail the job. Same pattern could guard the SARIF upload steps ('Path does not exist' cascades).

Activity

  1. cristim commented on Sep 11, 2026

    @cristim
    MemberAuthor

    This also occurs during Go module loading, not only vulnerability-database fetches. PR LeanerCloud/cloud-commitments-cli#2090 at unchanged head 3f74875049e191c7c27e829b50fc78c346807ba0 hit two consecutive govulncheck transport failures:

    • Attempt 1, job103333089302: proxy.golang.org returned HTTP/2 stream437 INTERNAL_ERROR while downloading Azure armbillingbenefits v1.0.0.
    • Attempt 2, job103337240718: the same error on stream407 while downloading AWS SDK config v1.29.12. Subsequent invalid-import errors were consequences of the missing module, not independent source defects.
    • Attempt 3: govulncheck passed without code changes; other security steps are still running at this checkpoint. A passed step is not yet a passed job.

    The first two attempts never produced a vulnerability verdict. npm audit passed on the same head. Other security scanners continued as designed, so do not change their independence or weaken the aggregate gate.

    Please include proxy module-download failures when investigating bounded transport recovery in .github/workflows/ci.yml:632-654. Retain the pinned scanner and all six modules. Prove actual vulnerability findings and exhausted retries still fail, with deterministic tests that distinguish transport failure from a security verdict. Do not disable HTTP/2 or suppress scanner failures speculatively.

    Re-triaged to P2, medium severity, this-sprint urgency, internal impact, small effort: these failures delayed an actual dependency-security merge across multiple full security-job attempts. Same-head reruns remain a workaround; no production outage or newly discovered CVE is established. This comment extends the existing issue rather than creating duplicate follow-up work. Related originating dependency issue: LeanerCloud/cloud-commitments-cli#2089.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions