Skip to content

Unhandled TLSSocket error - internal Node error #44751

Description

@prettydiff

Version

18.9.0, 17.1.0

Platform

Bodhi Linux (Ubuntu like) - VM guest on Windows 10 host

Subsystem

No response

What steps will reproduce the bug?

I am running my Node based peer-to-peer application in test automation mode that creates persistent sockets as TLS websockets. One the first test automation run, initiated from the windows host machine, everything is fine. The application instance on the VMs continues to run so that I can frequently run test automation instances without ever touching the VMs. When I attempt to run test automation a second time only 1 of my 4 VM application instances fails.

This appears to be a node defect. The error message mentions an unhandled error event but I have extensive error handling on just about everything in my application, most especially my socket management. The stack trace also does not indicate any code from my application.

How often does it reproduce? Is there a required condition?

100% reproducible. I am running 4 virtual machines each with a nearly identical install. This problem only occurs on one of those 4 VMs.

What is the expected behavior?

Socket not crashing.

What do you see instead?

node:events:491
      throw er; // Unhandled 'error' event

Error: read ECONNRESET
    at TLSWrap.onStreamRead (node:internal/stream_base_commons:217:20)
Emitted 'error' event on TLSSocket instance at:
    at emitErrorNT (node:internal/streams/destroy:151:8)
    at emitErrorCloseNT (node:internal/streams/destroy:116:3)
    at process.processTicksAndRejections (node:internal/process/task_queues:82:21) {
  errno: -104,
  code: 'ECONNRESET',
  syscall: 'read'
}

Node.js v18.9.0

Additional information

No response

Activity

  1. mscdex commented on Sep 23, 2022

    @mscdex
    Contributor

    Do you have a minimal test case that reproduces the issue?

  2. prettydiff commented on Sep 23, 2022

    @prettydiff
    ContributorAuthor

    @mscdex here is what I am doing to produce that scenario.

    1. Have 5 computers of Windows or any flavor of Linux. It doesn't matter if they are VMs, physical, or containers so long as they can all ping each other in a mesh configuration and have a modern web browser available.
    2. git clone -b remotes https://github.com/prettydiff/share-file-systems.git share
    3. node share/install
    4. On the last 4 computers execute share test_browser remote. This will turn on the application in a standby listening capacity awaiting test automation instructions.
    5. On the first computer execute networked test automation share test_browser device.
    6. For me the test automation fails around step 17 or 24 due to a regression. This is fine as its my application regression and not a node defect (yet). The last 4 boxes will continue to run the application awaiting further test automation instructions.
    7. On the first computer press CTRL+C and close the browser to kill the application, because we need to run it again.
    8. On the first computer run the application again share test_browser device.
    9. The Node problem occurs as the first computer attempts to open a connection to the 4 other computers which crashes the application on one of those 4 computers.

    In summary once everything is set up I am just running my test script 2 times consecutively.


    In my case I am able to reproduce this issue 100% of the time. I don't know how reproducible it will be for anybody else because it only occurs on one (the same one) of 4 VMs that were all clones 6 months ago. It might be a better use of everybody's time if I try to dig into this myself, but I don't know anything about Node's code so I will need to getting started guidance.

  3. prettydiff commented on Sep 23, 2022

    @prettydiff
    ContributorAuthor

    I am pinpointed that the error occurs when I close the application on the primary computer, and thus close/destroy the socket on the primary computer. Such an action expectedly sends a close message to the remote end and the remote end is generating an error on either its TLSSocket close, end, or destroy events and terminating the application with an unhandled error message.

    https://github.com/nodejs/node/blob/main/lib/internal/stream_base_commons.js indicates the problematic error occurs at the destroy event.

  4. added
    tlsIssues and PRs related to the tls subsystem.
    netIssues and PRs related to the net subsystem.
    on Sep 23, 2022
  5. prettydiff commented on Sep 24, 2022

    @prettydiff
    ContributorAuthor

    I have attempted to reproduce the code changes that resulted in this problem in a new branch one step at a time. I have it narrowed down to a very tiny refactor.

    prettydiff/share-file-systems@fa8568e#diff-f4938bd883bd4bfdf58503f680667c1b2662c0ff4f66544879c3ded5a2a5f6d0

    Before defect exposed

    Listener code

    listener: function terminal_server_transmission_transmitWs_listener(socket:websocket_client):void {
        const processor = function terminal_server_transmission_transmitWs_listener_processor(buf:Buffer):void {
            // ...
        };
        socket.on("data", processor);
    }

    After defect exposed

    Listener code, the wrapping function is removed

    listener: function terminal_server_transmission_transmitWs_listener(buf:Buffer):void {
        // ...
        const socket:websocket_client = this
        // ...
    }
    ```.socket.on("data", transmit_ws.listener);

    For some reason the defect is limited to how I associate a socket to its "data" event listener.

    Without guidance on diving into the internals of Node I do not think I can be more helpful. Please let me know how I can continue to help.

  6. afanasy commented on Aug 17, 2023

    @afanasy
    Contributor

    @prettydiff Have you tried

    https.createServer(...).
      // to catch Error: read ECONNRESET at TLSWrap.onStreamRead
      on('secureConnection', socket => {
        socket.on('error', err => {
          console.log(err)
        })
      }).
      listen(443)
  7. prettydiff commented on Aug 19, 2023

    @prettydiff
    ContributorAuthor

    I believe the problem was in my own code. I am struggling to remember from last year. I now have the same error handler for all sockets whether client or server side. I cannot remember if that was sufficient to solve for this.

    I believe the problem was Node writing an unhandled error messaging to stderr even when I had the proper error handling in the place in my application, because the unhandled error occurred only deep in a Node library. I no longer see this behavior now, so I changed something in my own code to prevent the error state from occurring within Node's deeper library, but I cannot remember what it is.

    I will review my commit history later today to see if I can identify the change.

  8. added a commit that references this issue on Jan 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    netIssues and PRs related to the net subsystem.tlsIssues and PRs related to the tls subsystem.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions