Hi there! I'm trying to build a production system leveraging River and I'm liking it so far.
My biggest issue has been that our jobs can take several hours to run, and we're often running into states where a job is "stuck" in a running state because of a Kubernetes job or something (even though we're attempting a graceful shutdown). I don't want to set RescueStuckJobsAfter to be too high, because that might unnecessarily slow down the system.
I don't fully understand why RescueStuckJobsAfter has to be an effective hard max on the runtime of a job. Would it be possible to expose some sort of "job heartbeat" that lets River know that a job is still alive and well, and use that to check whether a job needs to be rescued?
Hi there! I'm trying to build a production system leveraging River and I'm liking it so far.
My biggest issue has been that our jobs can take several hours to run, and we're often running into states where a job is "stuck" in a running state because of a Kubernetes job or something (even though we're attempting a graceful shutdown). I don't want to set RescueStuckJobsAfter to be too high, because that might unnecessarily slow down the system.
I don't fully understand why RescueStuckJobsAfter has to be an effective hard max on the runtime of a job. Would it be possible to expose some sort of "job heartbeat" that lets River know that a job is still alive and well, and use that to check whether a job needs to be rescued?