Repository navigation
Deadlock at process shutdown #54918
Description
Activity
As written in #52550 (comment), applying the patch from #47452 seems to fix the issue, but it is not a solution since that patch brings other issues.
- addedconfirmed-bugIssues and PRs for confirmed bugs.Issues and PRs for confirmed bugs.
on Sep 13, 2024 This can been seen in various CI runs, assuming confirmed-bug
The output of
lldbon macOS:$ lldb -p 5033 (lldb) process attach --pid 5033 Process 5033 stopped * thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP frame #0: 0x00007ff802363c3a libsystem_kernel.dylib`__psynch_cvwait + 10 libsystem_kernel.dylib`: -> 0x7ff802363c3a <+10>: jae 0x7ff802363c44 ; <+20> 0x7ff802363c3c <+12>: movq %rax, %rdi 0x7ff802363c3f <+15>: jmp 0x7ff8023617e0 ; cerror_nocancel 0x7ff802363c44 <+20>: retq Target 0: (node) stopped. Executable module set to "/Users/luigi/code/node/out/Release/node". Architecture set to: x86_64h-apple-macosx-. (lldb) bt * thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP * frame #0: 0x00007ff802363c3a libsystem_kernel.dylib`__psynch_cvwait + 10 frame #1: 0x00007ff8023a16f3 libsystem_pthread.dylib`_pthread_cond_wait + 1211 frame #2: 0x000000010df93b23 node`uv_cond_wait(cond=0x00007f9a4d404108, mutex=0x00007f9a4d404098) at thread.c:798:7 [opt] frame #3: 0x000000010d0be52b node`node::NodePlatform::DrainTasks(v8::Isolate*) [inlined] node::LibuvMutexTraits::cond_wait(cond=0x00007f9a4d404108, mutex=0x00007f9a4d404098) at node_mutex.h:175:5 [opt] frame #4: 0x000000010d0be511 node`node::NodePlatform::DrainTasks(v8::Isolate*) [inlined] node::ConditionVariableBase<node::LibuvMutexTraits>::Wait(this=0x00007f9a4d404108, scoped_lock=<unavailable>) at node_mutex.h:249:3 [opt] frame #5: 0x000000010d0be511 node`node::NodePlatform::DrainTasks(v8::Isolate*) [inlined] node::TaskQueue<v8::Task>::BlockingDrain(this=0x00007f9a4d404098) at node_platform.cc:640:20 [opt] frame #6: 0x000000010d0be4fb node`node::NodePlatform::DrainTasks(v8::Isolate*) [inlined] node::WorkerThreadsTaskRunner::BlockingDrain(this=0x00007f9a4d404098) at node_platform.cc:214:25 [opt] frame #7: 0x000000010d0be4fb node`node::NodePlatform::DrainTasks(this=0x0000600002718000, isolate=<unavailable>) at node_platform.cc:467:33 [opt] frame #8: 0x000000010cf18a50 node`node::SpinEventLoopInternal(env=0x00007f9a2c814e00) at embed_helpers.cc:44:17 [opt] frame #9: 0x000000010d084083 node`node::NodeMainInstance::Run() [inlined] node::NodeMainInstance::Run(this=<unavailable>, exit_code=<unavailable>, env=0x00007f9a2c814e00) at node_main_instance.cc:111:9 [opt] frame #10: 0x000000010d083fe9 node`node::NodeMainInstance::Run(this=<unavailable>) at node_main_instance.cc:100:3 [opt] frame #11: 0x000000010cfe4159 node`node::Start(int, char**) [inlined] node::StartInternal(argc=<unavailable>, argv=<unavailable>) at node.cc:1489:24 [opt] frame #12: 0x000000010cfe3f54 node`node::Start(argc=<unavailable>, argv=<unavailable>) at node.cc:1496:27 [opt] frame #13: 0x00007ff802015345 dyld`start + 1909 (lldb)@aduh95 see #54918 (comment). It seems the same patch.
No it's not, it's based on it but it's meant to fix the test failures.
Reacted by Luigi PincaThe issue persists 😞
gdb -p 18560 GNU gdb (Ubuntu 15.0.50.20240403-0ubuntu1) 15.0.50.20240403-git Copyright (C) 2024 Free Software Foundation, Inc. License GPLv3+: GNU GPL version 3 or later <http://gnu.org/licenses/gpl.html> This is free software: you are free to change and redistribute it. There is NO WARRANTY, to the extent permitted by law. Type "show copying" and "show warranty" for details. This GDB was configured as "x86_64-linux-gnu". Type "show configuration" for configuration details. For bug reporting instructions, please see: <https://www.gnu.org/software/gdb/bugs/>. Find the GDB manual and other documentation resources online at: <http://www.gnu.org/software/gdb/documentation/>. For help, type "help". Type "apropos word" to search for commands related to "word". Attaching to process 18560 [New LWP 18621] [New LWP 18620] [New LWP 18619] [New LWP 18618] [New LWP 18574] [New LWP 18565] [New LWP 18564] [New LWP 18563] [New LWP 18562] [New LWP 18561] Downloading separate debug info for /lib/x86_64-linux-gnu/libstdc++.so.6 Downloading separate debug info for system-supplied DSO at 0x7ffdf50f4000 [Thread debugging using libthread_db enabled] Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1". Download failed: Invalid argument. Continuing without source file ./nptl/./nptl/futex-internal.c. 0x00007fca5f88ad61 in __futex_abstimed_wait_common64 (private=21976, cancel=true, abstime=0x0, op=393, expected=0, futex_word=0x55d834ea10e0) at ./nptl/futex-internal.c:57 warning: 57 ./nptl/futex-internal.c: No such file or directory (gdb) bt #0 0x00007fca5f88ad61 in __futex_abstimed_wait_common64 (private=21976, cancel=true, abstime=0x0, op=393, expected=0, futex_word=0x55d834ea10e0) at ./nptl/futex-internal.c:57 #1 __futex_abstimed_wait_common (cancel=true, private=21976, abstime=0x0, clockid=0, expected=0, futex_word=0x55d834ea10e0) at ./nptl/futex-internal.c:87 #2 __GI___futex_abstimed_wait_cancelable64 (futex_word=futex_word@entry=0x55d834ea10e0, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0, private=private@entry=0) at ./nptl/futex-internal.c:139 #3 0x00007fca5f88d7dd in __pthread_cond_wait_common (abstime=0x0, clockid=0, mutex=0x55d834ea1060, cond=0x55d834ea10b8) at ./nptl/pthread_cond_wait.c:503 #4 ___pthread_cond_wait (cond=0x55d834ea10b8, mutex=0x55d834ea1060) at ./nptl/pthread_cond_wait.c:627 #5 0x000055d82f131fed in uv_cond_wait (cond=<optimized out>, mutex=<optimized out>) at ../deps/uv/src/unix/thread.c:814 #6 0x000055d82e08c853 in node::NodePlatform::DrainTasks(v8::Isolate*) () #7 0x000055d82dec528d in node::SpinEventLoopInternal(node::Environment*) () #8 0x000055d82e04b364 in node::NodeMainInstance::Run() () #9 0x000055d82df9276a in node::Start(int, char**) () #10 0x00007fca5f81c1ca in __libc_start_call_main (main=main@entry=0x55d82debac50 <main>, argc=argc@entry=2, argv=argv@entry=0x7ffdf50a9a48) at ../sysdeps/nptl/libc_start_call_main.h:58 #11 0x00007fca5f81c28b in __libc_start_main_impl (main=0x55d82debac50 <main>, argc=2, argv=0x7ffdf50a9a48, init=<optimized out>, fini=<optimized out>, rtld_fini=<optimized out>, stack_end=0x7ffdf50a9a38) at ../csu/libc-start.c:360 #12 0x000055d82dec0785 in _start () (gdb)Also,
test/wasi/test-wasi-initialize-validation.jsandtest/wasi/test-wasi-start-validation.jsfail.$ ./node test/wasi/test-wasi-initialize-validation.js (node:13188) ExperimentalWarning: WASI is an experimental feature and might change at any time (Use `node --trace-warnings ...` to show where the warning was created) Mismatched noop function calls. Expected exactly 1, actual 0. at Proxy.mustCall (/home/luigi/node/test/common/index.js:451:10) at Object.<anonymous> (/home/luigi/node/test/wasi/test-wasi-initialize-validation.js:154:18) at Module._compile (node:internal/modules/cjs/loader:1557:14) at Module._extensions..js (node:internal/modules/cjs/loader:1702:10) at Module.load (node:internal/modules/cjs/loader:1328:32) at Module._load (node:internal/modules/cjs/loader:1138:12) at TracingChannel.traceSync (node:diagnostics_channel:315:14) at wrapModuleLoad (node:internal/modules/cjs/loader:217:24) at Function.executeUserEntryPoint [as runMain] (node:internal/modules/run_main:166:5) $ ./node test/wasi/test-wasi-start-validation.js (node:13199) ExperimentalWarning: WASI is an experimental feature and might change at any time (Use `node --trace-warnings ...` to show where the warning was created) Mismatched noop function calls. Expected exactly 1, actual 0. at Proxy.mustCall (/home/luigi/node/test/common/index.js:451:10) at Object.<anonymous> (/home/luigi/node/test/wasi/test-wasi-start-validation.js:151:18) at Module._compile (node:internal/modules/cjs/loader:1557:14) at Module._extensions..js (node:internal/modules/cjs/loader:1702:10) at Module.load (node:internal/modules/cjs/loader:1328:32) at Module._load (node:internal/modules/cjs/loader:1138:12) at TracingChannel.traceSync (node:diagnostics_channel:315:14) at wrapModuleLoad (node:internal/modules/cjs/loader:217:24) at Function.executeUserEntryPoint [as runMain] (node:internal/modules/run_main:166:5)- addedhelp wantedIssues that need assistance from volunteers or PRs that need help to proceed.Issues that need assistance from volunteers or PRs that need help to proceed.
on Oct 15, 2024 Hey, I've added the help wanted
Issues that need assistance from volunteers or PRs that need help to proceed. label to this issue. While all issues deserve TLC from volunteers, this one is causing a serious CI issue, and a little extra attention wouldn't hurt.From my testing, the following snippet hangs occassionally, which might be related to this?:
Line 640 in 2545b9e
tasks_drained_.Wait(scoped_lock); outstanding_tasks_is2when it hangs,0otherwise. It also looks like one of those tasks is completed, but the other one is not?Some task isn't completing (but that was probably already known), do we know what task, or how to find out?
Update 1
It seems a task is getting stuck during
BlockingPop, which might be connected to the issue.Update 2TheNodeMainInstanceis never properly destructed, which prevents theNodePlatformfrom shutting down. However, this situation appears to be a catch-22. TheNodeMainInstanceseems to finish only afterDrainTasksis completed, butDrainTaskscan’t finish untilNodeMainInstanceis shut down.Update 3Removingplatform->DrainTasks(isolate);fromSpinEventLoopInternalprevents the error from occurring (the draining still happens elsewhere without interruption, although other errors persist).Update 4
It's not a catch-22, because the hanging process appears to be hanging before the process destruction would even occur. I suspect that if you
btall the threads, you'll see thatDrainTasksisn't the only thing hanging,Indeed, our hanging worker thread appears to be:
Thread 8 (Thread 0x7fb7c74006c0 (LWP 1322244) "node"): #0 0x00007fb7cd49e22e in __futex_abstimed_wait_common64 (private=0, cancel=true, abstime=0x0, op=393, expected=0, futex_word=0x55a4a5eefdc8) at ./nptl/futex-internal.c:57 #1 __futex_abstimed_wait_common (futex_word=futex_word@entry=0x55a4a5eefdc8, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0, private=private@entry=0, cancel=cancel@entry=true) at ./nptl/futex-internal.c:87 #2 0x00007fb7cd49e2ab in __GI___futex_abstimed_wait_cancelable64 (futex_word=futex_word@entry=0x55a4a5eefdc8, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0, private=private@entry=0) at ./nptl/futex-internal.c:139 #3 0x00007fb7cd4a0990 in __pthread_cond_wait_common (abstime=0x0, clockid=0, mutex=0x55a4a5eefd78, cond=0x55a4a5eefda0) at ./nptl/pthread_cond_wait.c:503 #4 ___pthread_cond_wait (cond=0x55a4a5eefda0, mutex=0x55a4a5eefd78) at ./nptl/pthread_cond_wait.c:618 #5 0x000055a46a13ad8c in void heap::base::Stack::SetMarkerForBackgroundThreadAndCallbackImpl<v8::internal::LocalHeap::ExecuteWhileParked<v8::internal::CollectionBarrier::AwaitCollectionBackground(v8::internal::LocalHeap*)::{lambda()#1}>(v8::internal::CollectionBarrier::AwaitCollectionBackground(v8::internal::LocalHeap*)::{lambda()#1})::{lambda()#1}>(heap::base::Stack*, void*, void const*) () #6 0x000055a46acb27e3 in PushAllRegistersAndIterateStack () #7 0x000055a46a13b2df in v8::internal::CollectionBarrier::AwaitCollectionBackground(v8::internal::LocalHeap*) () #8 0x000055a46a1ad972 in v8::internal::HeapAllocator::AllocateRawWithLightRetrySlowPath(int, v8::internal::AllocationType, v8::internal::AllocationOrigin, v8::internal::AllocationAlignment) () #9 0x000055a46a1ae0a8 in v8::internal::HeapAllocator::AllocateRawWithRetryOrFailSlowPath(int, v8::internal::AllocationType, v8::internal::AllocationOrigin, v8::internal::AllocationAlignment) () #10 0x000055a46a1e4891 in v8::internal::LocalFactory::AllocateRaw(int, v8::internal::AllocationType, v8::internal::AllocationAlignment) () #11 0x000055a46a17a4fe in v8::internal::FactoryBase<v8::internal::LocalFactory>::NewProtectedFixedArray(int) () #12 0x000055a46a364a6d in v8::internal::DeoptimizationData::New(v8::internal::LocalIsolate*, int) () #13 0x000055a46a7b8d3f in v8::internal::maglev::MaglevCodeGenerator::GenerateDeoptimizationData(v8::internal::LocalIsolate*) () #14 0x000055a46a7b991b in v8::internal::maglev::MaglevCodeGenerator::BuildCodeObject(v8::internal::LocalIsolate*) [clone .part.0] () #15 0x000055a46a7d9948 in v8::internal::maglev::MaglevCodeGenerator::Assemble() () #16 0x000055a46a82f269 in v8::internal::maglev::MaglevCompiler::Compile(v8::internal::LocalIsolate*, v8::internal::maglev::MaglevCompilationInfo*) () #17 0x000055a46a830b49 in v8::internal::maglev::MaglevCompilationJob::ExecuteJobImpl(v8::internal::RuntimeCallStats*, v8::internal::LocalIsolate*) () #18 0x000055a469ffb6bb in v8::internal::OptimizedCompilationJob::ExecuteJob(v8::internal::RuntimeCallStats*, v8::internal::LocalIsolate*) () #19 0x000055a46a831047 in v8::internal::maglev::MaglevConcurrentDispatcher::JobTask::Run(v8::JobDelegate*) () #20 0x000055a46b21b45f in v8::platform::DefaultJobWorker::Run() () #21 0x000055a469cb125f in node::(anonymous namespace)::PlatformWorkerThread(void*) () #22 0x00007fb7cd4a1732 in start_thread (arg=<optimized out>) at ./nptl/pthread_create.c:447 #23 0x00007fb7cd51c2b8 in __GI___clone3 () at ../sysdeps/unix/sysv/linux/x86_64/clone3.S:78
Update 5
The hanging tasks is pushed via the following bt:
1: 0x5582238b587b node::TaskQueue<v8::Task>::Push(std::unique_ptr<v8::Task, std::default_delete<v8::Task> >) [out/Release/node] 2: 0x5582238b3d47 node::NodePlatform::PostTaskOnWorkerThreadImpl(v8::TaskPriority, std::unique_ptr<v8::Task, std::default_delete<v8::Task> >, v8::SourceLocation const&) [out/Release/node] 3: 0x558224e25304 v8::platform::DefaultJobState::CallOnWorkerThread(v8::TaskPriority, std::unique_ptr<v8::Task, std::default_delete<v8::Task> >) [out/Release/node] 4: 0x558224e2548f [out/Release/node] 5: 0x558223c0ca77 [out/Release/node] 6: 0x558223c0e337 v8::internal::Compiler::CompileOptimized(v8::internal::Isolate*, v8::internal::Handle<v8::internal::JSFunction>, v8::internal::ConcurrencyMode, v8::internal::CodeKind) [out/Release/node] 7: 0x55822423e64a v8::internal::Runtime_CompileOptimized(int, unsigned long*, v8::internal::Isolate*) [out/Release/node] 8: 0x5581c496f276So it looks like this is a V8 hanging-task issue? I'll reach out to them and see if they know anything. (https://issues.chromium.org/issues/374285493)
- addedv8 engineIssues and PRs related to the V8 dependency.Issues and PRs related to the V8 dependency.
on Oct 18, 2024 - added a commit that references this issue
on Oct 29, 2024 From the discussion in https://issues.chromium.org/issues/374285493 (closed as Won't Fix - Intended Behavior) the issue is not in V8 but in Node.js.
Reacted by jakecastelli94 remaining items
I think this
Lines 584 to 602 in df52c21
// FIXME(54918): we should not be blocking on the worker tasks on the // main thread in one go. Doing so leads to two problems: // 1. If any of the worker tasks post another foreground task and wait // for it to complete, and that foreground task is posted right after // we flush the foreground task queue and before the foreground thread // goes into sleep, we'll never be able to wake up to execute that // foreground task and in turn the worker task will never complete, and // we have a deadlock. // 2. Worker tasks can be posted from any thread, not necessarily associated // with the current isolate, and we can be blocking on a worker task that // is associated with a completely unrelated isolate in the event loop. // This is suboptimal. // // However, not blocking on the worker tasks at all can lead to loss of some // critical user-blocking worker tasks e.g. wasm async compilation tasks, // which should block the main thread until they are completed, as the // documentation suggets. As a compromise, we currently only block on // user-blocking tasks to reduce the chance of deadlocks while making sure // that criticl user-blocking tasks are not lost. is the reason why it is still open.
- added a commit that references this issue
on Aug 6, 2026 We're hitting what looks like the exit-time variant of this deadlock in CI, and we captured
per-thread kernel + userspace stacks that appear to show the complete wait cycle: the main
thread is inNodePlatform::Shutdownjoining the platform workers, while the remaining
V8Worker threads are inside concurrent Maglev/Sparkplug compile jobs, parked in
CollectionBarrier::AwaitCollectionBackgroundwaiting for a main-thread GC that can never
happen. (The OP's trace shows the main thread inDrainTasks; ours is the same shape one
stage later, inShutdown→uv_thread_join.)Environment: stock Node v24.18.1 x64 on Ubuntu 24.04, 4-vCPU CI runner shared with
MySQL + OpenSearch (i.e. CPU-contended).node --testwith--test-isolation=processand
--test-force-exit. The children run plain pre-compiled JS (esbuild output, no TS loaders);
noworker_threadsin the code under test; the only module customization is a synchronous
module.registerHooks()resolve hook (in-thread, no hook workers). Kernel log is clean — no
OOM kills.A test child finished its tests, then wedged on exit. Because it's blocked in a futex rather
than the event loop,--test-timeout/--test-force-exitcan't reach it, and the parent
runner (healthy, 10 threads, main thread idle inepoll_pwait) waits on it forever. State
below was captured ~3.5 minutes into the wedge; every thread at 0% CPU, all in
futex_do_waitper/proc/<pid>/task/*/stack.Thread table of the wedged child:
TID S WCHAN %CPU ELAPSED COMMAND 7732 S futex_do_wait 0.1 03:42 MainThread 7734 S futex_do_wait 0.0 03:42 V8Worker 7735 S futex_do_wait 0.0 03:42 V8Worker 7737 S futex_do_wait 0.0 03:42 V8Worker 7738 S futex_do_wait 0.0 03:42 SignalInspector(The DelayedTaskScheduler and one of the four V8Workers are already gone — consistent with
WorkerThreadsTaskRunner::Shutdownhaving stopped the scheduler and joined one worker
before wedging on the next.)Main thread (
eu-stack, trimmed):#2 uv_thread_join #3 node::WorkerThreadsTaskRunner::Shutdown() #4 node::NodePlatform::Shutdown() #5 node::DefaultProcessExitHandlerInternal(node::Environment*, node::ExitCode) #6 node::Environment::Exit(node::ExitCode) ... JIT frames ... #15 v8::internal::Execution::TryRunMicrotasks #17 v8::internal::MicrotaskQueue::PerformCheckpoint ... JIT frames ... #28 node::Environment::CheckImmediate(uv_check_s*) #30 uv_run #31 node::SpinEventLoopInternali.e. the exit was initiated from JS via a microtask checkpoint — consistent with the test
runner's--test-force-exitprocess.exit()after the file's tests completed.Two of the three surviving V8Workers (TIDs 7734, 7735) are identical, mid-Maglev-compile:
#1 absl::synchronization_internal::FutexWaiter::WaitUntil #4 absl::CondVar::WaitCommon #7 v8::internal::CollectionBarrier::AwaitCollectionBackground(v8::internal::LocalHeap*) #8 v8::internal::HeapAllocator::AllocateRawWithRetryOrFailSlowPath #11 v8::internal::FactoryBase<v8::internal::LocalFactory>::NewProtectedFixedArray #12 v8::internal::DeoptimizationData::New #13 v8::internal::maglev::MaglevCodeGenerator::GenerateDeoptimizationData #16 v8::internal::maglev::MaglevCompiler::Compile #19 v8::internal::maglev::MaglevConcurrentDispatcher::JobTask::Run #20 v8::platform::DefaultJobWorker::Run() #21 node::(anonymous namespace)::PlatformWorkerThread(void*)The third (TID 7737) is the same wait via the baseline compiler instead
(ConcurrentBaselineCompiler::JobDispatcher::Run→
Factory::CodeBuilder::BuildInternal→ same
AllocateRawWithRetryOrFailSlowPath→AwaitCollectionBackground).So the cycle is:
- main:
Environment::Exit→NodePlatform::Shutdown→
WorkerThreadsTaskRunner::Shutdown→uv_thread_join, waiting for the workers to drain; - workers: concurrent compile
JobTask::Run→ allocation failure →
CollectionBarrier::AwaitCollectionBackground, waiting for the main thread to perform a
collection — which it never will, because it's already inShutdown.
The previous day we captured the same process shape on a different suite at an earlier
teardown stage (kernel wait channels only, no userspace stacks yet): main thread
futex-parked with the DelayedTaskScheduler and two of four V8Workers already gone, the rest
futex-parked — so the race window appears to span the whole worker-join loop, and it isn't
specific to one test suite.We can't reproduce on demand — we've captured it twice on consecutive days across a CI that
runs ~28node:testsuites per commit, always on contended 4-vCPU runners. Happy to attach
the full dumps (per-thread/proc/<pid>/task/*/{stack,syscall}+eu-stackfor every
process in the job) or test candidate patches in our CI if that's useful.- main:
- added a commit that references this issue
on Aug 30, 2026 - added a commit that references this issue
on Sep 28, 2026 - added a commit that references this issue
on Sep 28, 2026 - added a commit that references this issue
on Sep 29, 2026 - added a commit that references this issue
on Sep 29, 2026 - added a commit that references this issue
on Oct 2, 2026
Version
v23.0.0-pre
Platform
Subsystem
No response
What steps will reproduce the bug?
There is a deadlock that prevents the Node.js process from exiting. The issue is causing a lot (all?) of timeout failures in our CI. It can be reproduced by running a test in parallel with our
test.pytool, for exampleSee also
test-stream-readable-unpipe-resume#54133 (comment)test-net-write-fully-async-bufferas flaky #52959 (comment)How often does it reproduce? Is there a required condition?
Rarely, but often enough to be a serious issue for CI.
What is the expected behavior? Why is that the expected behavior?
The process exits.
What do you see instead?
The process does not exit.
Additional information
Attaching
gdbto two of the hanging processes obtained from the command above, produces the following outputs: