Skip to content

ubuntu-slim runners time out after 15 mins while it usually takes <3 mins #64972

Description

@trivikr

Last successful run which landed a PR

git node land was successful

...
+ MULTIPLE_COMMIT_POLICY=--oneCommitMax
+ git node land --autorebase --yes --oneCommitMax 63949
+ cat output
- Loading data for nodejs/node/pull/63949
...

Details: https://github.com/nodejs/node/actions/runs/30765170760/job/91542478131

First timeout

...
+ MULTIPLE_COMMIT_POLICY=--oneCommitMax
+ git node land --autorebase --yes --oneCommitMax 64772
Error: The operation was canceled.

Details: https://github.com/nodejs/node/actions/runs/30783328197/job/91592125884

Activity

  1. added
    metaIssues and PRs related to the general management of the project.
    on Aug 3, 2026
  2. trivikr commented on Aug 3, 2026

    @trivikr
    MemberAuthor

    The retry attempt took close to 15 mins, but was successful
    https://github.com/nodejs/node/actions/runs/30783328197/job/91601263185

  3. trivikr commented on Aug 3, 2026

    @trivikr
    MemberAuthor

    This can be closed if issue doesn't persist.

  4. Renegade334 commented on Aug 3, 2026

    @Renegade334
    Member

    This happens to GHA jobs on the slim runners from time to time.

  5. aduh95 commented on Aug 4, 2026

    @aduh95
    Contributor

    There's clearly been an influx of those, I suspect a change on GH side. The issue is that the CQ can time out after removing the commit-queue PRs queued for automated landing through the Commit Queue. label, but before being able to actually merge, so the PR is effectively kicked out of the CQ without notifying anyone. We could add another GHA job to try clean up, but if the problem is GHA reliability, it might be hopeless to try to solve it with GHA 🫣

  6. panva commented on Aug 7, 2026

    @panva
    Member

    This also happens to linters.yml, format-cpp, lint-cpp, and lint-js-and-md i've re-ran more than enough over the span of the last 2 weeks

  7. changed the title [-]commit-queue times out after 15 mins while it usually takes <3 mins[/-] [+]ubuntu-slim runners time out after 15 mins while it usually takes <3 mins[/+] on Aug 8, 2026
  8. aduh95 commented on Aug 8, 2026

    @aduh95
    Contributor

    I wonder if it's a memory issue, in which case maybe adding some ESLint cache might help for lint-js-and-md. Switching back to ubuntu-latest is also an option, but a costly one (we prepare security releases from a private fork, where we have to pay for GHA usage)

  9. 9 remaining items

  10. MikeMcC399 commented on Sep 21, 2026

    @MikeMcC399
    Contributor

    I'm also continuing to see this issue occurring. It causes delays and additional actions for team members who need to re-run failed workflows.

    I have lost count of the number of times that I've restarted the lint-js-and-md job for contributor PR submissions after it has hit the 15 minute limit, running on 1 CPU.

    What is the financial impact of returning to ubuntu-latest?

  11. MikeMcC399 commented on Sep 29, 2026

    @MikeMcC399
    Contributor

    What about migrating the job lint-js-and-md in the GitHub Actions workflow .github/workflows/linters.yml from ubuntu-slim to ubuntu-26.04-arm? Although other jobs may be affected, I'm seeing this one fail a lot, just in the PRs that I've been in involved in. Even clean commits to the main branch are being hit (example https://github.com/nodejs/node/actions/runs/36535043076/job/109297044751).

    1 CPU and 5 GB that ubuntu-slim offers is clearly not enough, and is causing the job lint-js-and-md to regularly time out after 15 minutes. This wasn't happening before the job was migrated from ubuntu-latestto ubuntu-slim.

    4 CPUs and 16 GB that ubuntu-26.04-arm includes would speed things up and stop the timeouts occurring.

  12. inoway46 commented on Sep 30, 2026

    @inoway46
    Contributor

    +1 to moving this job to ARM. I tested lint-js-and-md on ubuntu-26.04-arm and observed no timeouts in 256 runs. Separately, #66408 improves performance on both runners, but slim still times out.

    “after” includes the optimization in #66408.

    Runner Success rate Avg Min P90 P99 Max Cost / 256 runs*
    ubuntu-slim (before) 243/256 (94.9%) 9m 03s 6m 15s 12m 41s 16m 17s 17m 41s $4.88
    ubuntu-slim (after) 246/256 (96.1%) 7m 32s 5m 28s 10m 02s 16m 12s 17m 25s $4.12
    ubuntu-26.04-arm (before) 256/256 (100%) 3m 15s 2m 04s 3m 22s 3m 28s 3m 33s $5.10
    ubuntu-26.04-arm (after) 256/256 (100%) 2m 54s 1m 54s 3m 02s 3m 09s 3m 12s $4.00

    * Costs use private rates and per-job minute rounding. Private ARM has 2 CPUs versus 4 in these public runs, so actual private costs remain unverified.

    Representative batches: baseline slim/ARM (combined-slim / combined-arm), optimized ARM.

  13. MikeMcC399 commented on Sep 30, 2026

    @MikeMcC399
    Contributor

    @inoway46

    Would you want to submit a PR for this change, since you have done all the preparation? It looks like it would be a positive outcome.

  14. inoway46 commented on Sep 30, 2026

    @inoway46
    Contributor

    FYI @MikeMcC399, I closed #66409 after feedback that reducing the workload or using caching would be preferable to changing the runner.

  15. MikeMcC399 commented on Sep 30, 2026

    @MikeMcC399
    Contributor

    We do need a solution soon. Viewing the list https://github.com/nodejs/node/actions/workflows/linters.yml from yesterday (CEST timezone) Sep 29, 2026, there are 6 instances of cancelled workflows due to lint-js-and-md timeout:

    https://github.com/nodejs/node/actions/runs/36577916798/job/109438217698
    https://github.com/nodejs/node/actions/runs/36539193751/job/109310277791
    https://github.com/nodejs/node/actions/runs/36535043076/job/109297044751
    https://github.com/nodejs/node/actions/runs/36535043076/job/109297044751
    https://github.com/nodejs/node/actions/runs/36508967793/job/109216667037
    https://github.com/nodejs/node/actions/runs/36492431339/job/109163934793

    This does not show instances where lint-js-and-md timed out and a repo member re-ran the failed test.

    If there is another alternative way of solving this that can be implemented soon, then that would be fine. The runner could however be migrated to the more powerful one as an interim solution until the other proposed solution is available.

  16. trivikr commented on Sep 30, 2026

    @trivikr
    MemberAuthor

    after feedback that reducing the workload or using caching would be preferable to changing the runner.

    I think we can do both?

  17. inoway46 commented on Oct 1, 2026

    @inoway46
    Contributor

    Yes, we can do both. Sorry, my previous comment was misleading.

  18. MikeMcC399 commented on Oct 3, 2026

    @MikeMcC399
    Contributor

    #66412 adds npm caching for dependencies & ESLint caching, both for lint-js-and-md. In my own tests it did not solve the issue here. We should get this merged and see what the result is in practice.

  19. added a commit that references this issue on Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    metaIssues and PRs related to the general management of the project.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions