Skip to content

fix(healthcheck): recover unhealthy agents automatically - #466

Open
levilentz wants to merge 1 commit into
mainfrom
fix/health-state-recovery
Open

levilentz wants to merge 1 commit into
mainfrom
fix/health-state-recovery

Conversation

@levilentz

@levilentz levilentz commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

This is a fix to the agentex way of handling an agent going into "unhealthy" status. This is encountered when a kubernetes node is evicted, preventing the agent from coming back up cleanly.

RetriggerConfidence Score: 2/5

This PR is not ready to merge: recovery can leave agents unhealthy, schedule errors can stop the worker, and disabling health checks does not stop an existing sweep.

Fix All in CursorFindings

  1. P1 Recovered agent stays unhealthy ▶
  2. P1 Schedule error stops worker ▶
  3. P1 Disabled checks keep restarting ▶
  4. P2 Large sweeps miss deadline ▶
  5. P2 Status string breaks enum rule ▶
Fix with agent prompt
### Issue 1
agentex/src/temporal/workflows/healthcheck_workflow.py:57-60
When registration starts a new monitor for an agent already marked `Unhealthy`, it passes no `initial_status`. The workflow starts with `is_unhealthy=False`, so successful probes never request `Ready`. Pass the stored status when starting the monitor so a recovered agent does not remain marked unavailable.

### Issue 2
agentex/src/temporal/run_worker.py:302-304
If Temporal rejects schedule creation or briefly loses its connection, `ensure_reconciliation_schedule` raises before the shared worker starts. The startup error handler only covers the later sweep. Health checks, retention cleanup, and scheduled agent runs then stop too. Keep a schedule-creation failure from stopping the shared worker.

### Issue 3
agentex/src/temporal/activities/healthcheck_reconciliation_activities.py:17-21
Turning off `ENABLE_HEALTH_CHECK_WORKFLOW` skips startup setup but leaves a schedule already saved in Temporal. When that schedule runs, this activity starts monitors without checking the flag. An operator can disable health checks and still have them restart and change agent statuses. Check the flag when the scheduled sweep runs.

### Issue 4
agentex/src/temporal/workflows/healthcheck_reconciliation_workflow.py:16-20
The five-minute limit covers a scan of every monitored agent, with one awaited start call per agent. If the fleet is large or Temporal calls are slow, the activity times out and retries from page one. A long retry can also cause the schedule to skip its next run, delaying missing monitors. Bound or split the sweep so it can finish within its limit.

### Issue 5
agentex/src/temporal/workflows/healthcheck_workflow.py:59
This new comparison hardcodes `"Unhealthy"` even though `AgentStatus.UNHEALTHY` is available. The repository requires enum values instead of hardcoded status strings when an enum exists. Use the enum here; this requirement must be met before merging.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

Agent health checks keep running after agents become unhealthy and restore Ready status after two successful probes. Startup and recurring Temporal sweeps restart missing monitors, while deletion stops the agent’s monitor.

  • Keeps checking unhealthy agents and carries health state between workflow runs.
  • Reconciles monitors for ready and unhealthy agents at startup and every 300 seconds.
  • Stops monitors on deletion and leaves non-monitored agent statuses unchanged.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Boot[Worker startup] --> Schedule[Create sweep schedule]
  Schedule --> Sweep[Scan Ready and Unhealthy agents]
  Sweep --> Monitor[Start missing monitors]
  Monitor --> Probe[Probe ACP endpoint]
  Probe --> Status[Update agent status]
  Schedule -->|Every five minutes| Sweep
Loading

Reviews (1) · Last reviewed commit: "fix(healthcheck): recover unhealthy agen..."

@levilentz
levilentz requested a review from a team as a code owner October 2, 2026 01:51
Comment on lines +57 to +60
is_unhealthy = workflow_args.get(
"is_unhealthy",
workflow_args.get("initial_status") == "Unhealthy",
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Recovered agent stays unhealthy
When registration starts a new monitor for an agent already marked Unhealthy, it passes no initial_status. The workflow starts with is_unhealthy=False, so successful probes never request Ready. Pass the stored status when starting the monitor so a recovered agent does not remain marked unavailable.

Knowledge Base Used: Recover Unhealthy Agents After Successful Health Probes

Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/workflows/healthcheck_workflow.py
Line: 57-60

Comment:
**Recovered agent stays unhealthy**
When registration starts a new monitor for an agent already marked `Unhealthy`, it passes no `initial_status`. The workflow starts with `is_unhealthy=False`, so successful probes never request `Ready`. Pass the stored status when starting the monitor so a recovered agent does not remain marked unavailable.

**Knowledge Base Used:** [Recover Unhealthy Agents After Successful Health Probes](https://app.greptile.com/scale-ai/-/custom-context/knowledge-base/scaleapi/scale-agentex/-/reverts/incident-mitigation_417-20260903-unhealthy-agent-recovery-5b4d188.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

Comment on lines +302 to +304
adapter = TemporalAdapter(global_dependencies.temporal_client)
await ensure_reconciliation_schedule(adapter, task_queue)
try:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Schedule error stops worker
If Temporal rejects schedule creation or briefly loses its connection, ensure_reconciliation_schedule raises before the shared worker starts. The startup error handler only covers the later sweep. Health checks, retention cleanup, and scheduled agent runs then stop too. Keep a schedule-creation failure from stopping the shared worker.

Knowledge Base Used: Temporal health checks and retention cleanup

Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/run_worker.py
Line: 302-304

Comment:
**Schedule error stops worker**
If Temporal rejects schedule creation or briefly loses its connection, `ensure_reconciliation_schedule` raises before the shared worker starts. The startup error handler only covers the later sweep. Health checks, retention cleanup, and scheduled agent runs then stop too. Keep a schedule-creation failure from stopping the shared worker.

**Knowledge Base Used:** [Temporal health checks and retention cleanup](https://app.greptile.com/scale-ai/-/custom-context/knowledge-base/scaleapi/scale-agentex/-/docs/temporal-health-and-retention.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

Comment on lines +17 to +21
return await reconcile_healthcheck_workflows(
self.agent_repo,
self.temporal_adapter,
self.task_queue,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Disabled checks keep restarting
Turning off ENABLE_HEALTH_CHECK_WORKFLOW skips startup setup but leaves a schedule already saved in Temporal. When that schedule runs, this activity starts monitors without checking the flag. An operator can disable health checks and still have them restart and change agent statuses. Check the flag when the scheduled sweep runs.

Knowledge Base Used: Temporal health checks and retention cleanup

Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/activities/healthcheck_reconciliation_activities.py
Line: 17-21

Comment:
**Disabled checks keep restarting**
Turning off `ENABLE_HEALTH_CHECK_WORKFLOW` skips startup setup but leaves a schedule already saved in Temporal. When that schedule runs, this activity starts monitors without checking the flag. An operator can disable health checks and still have them restart and change agent statuses. Check the flag when the scheduled sweep runs.

**Knowledge Base Used:** [Temporal health checks and retention cleanup](https://app.greptile.com/scale-ai/-/custom-context/knowledge-base/scaleapi/scale-agentex/-/docs/temporal-health-and-retention.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

Comment on lines +16 to +20
return await workflow.execute_activity(
RECONCILE_HEALTHCHECK_WORKFLOWS_ACTIVITY,
start_to_close_timeout=timedelta(minutes=5),
retry_policy=RetryPolicy(
maximum_attempts=3,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Large sweeps miss deadline
The five-minute limit covers a scan of every monitored agent, with one awaited start call per agent. If the fleet is large or Temporal calls are slow, the activity times out and retries from page one. A long retry can also cause the schedule to skip its next run, delaying missing monitors. Bound or split the sweep so it can finish within its limit.

Knowledge Base Used: Temporal health checks and retention cleanup

Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/workflows/healthcheck_reconciliation_workflow.py
Line: 16-20

Comment:
**Large sweeps miss deadline**
The five-minute limit covers a scan of every monitored agent, with one awaited start call per agent. If the fleet is large or Temporal calls are slow, the activity times out and retries from page one. A long retry can also cause the schedule to skip its next run, delaying missing monitors. Bound or split the sweep so it can finish within its limit.

**Knowledge Base Used:** [Temporal health checks and retention cleanup](https://app.greptile.com/scale-ai/-/custom-context/knowledge-base/scaleapi/scale-agentex/-/docs/temporal-health-and-retention.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

recovery_counter = workflow_args.get("recovery_counter", 0)
is_unhealthy = workflow_args.get(
"is_unhealthy",
workflow_args.get("initial_status") == "Unhealthy",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Status string breaks enum rule
This new comparison hardcodes "Unhealthy" even though AgentStatus.UNHEALTHY is available. The repository requires enum values instead of hardcoded status strings when an enum exists. Use the enum here; this requirement must be met before merging.

Rule Used: Use enum values instead of hardcoded strings when available. Reference existing enums from shared-types package rather than using string literals. (source)

Learned From
scaleapi/scaleapi#126557

Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/workflows/healthcheck_workflow.py
Line: 59

Comment:
**Status string breaks enum rule**
This new comparison hardcodes `"Unhealthy"` even though `AgentStatus.UNHEALTHY` is available. The repository requires enum values instead of hardcoded status strings when an enum exists. Use the enum here; this requirement must be met before merging.

**Rule Used:** Use enum values instead of hardcoded strings when available. Reference existing enums from shared-types package rather than using string literals. ([source](https://app.greptile.com/scale-ai/-/custom-context?memory=c0c58ddb-09dc-4e8a-9837-ec1bb02f9579))

**Learned From**
[scaleapi/scaleapi#126557](https://github.com/scaleapi/scaleapi/pull/126557)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Cursor Fix in Claude Code Fix in Codex

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant