Repository navigation
Conversation
| is_unhealthy = workflow_args.get( | ||
| "is_unhealthy", | ||
| workflow_args.get("initial_status") == "Unhealthy", | ||
| ) |
There was a problem hiding this comment.
Recovered agent stays unhealthy
When registration starts a new monitor for an agent already marked Unhealthy, it passes no initial_status. The workflow starts with is_unhealthy=False, so successful probes never request Ready. Pass the stored status when starting the monitor so a recovered agent does not remain marked unavailable.
Knowledge Base Used: Recover Unhealthy Agents After Successful Health Probes
Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/workflows/healthcheck_workflow.py
Line: 57-60
Comment:
**Recovered agent stays unhealthy**
When registration starts a new monitor for an agent already marked `Unhealthy`, it passes no `initial_status`. The workflow starts with `is_unhealthy=False`, so successful probes never request `Ready`. Pass the stored status when starting the monitor so a recovered agent does not remain marked unavailable.
**Knowledge Base Used:** [Recover Unhealthy Agents After Successful Health Probes](https://app.greptile.com/scale-ai/-/custom-context/knowledge-base/scaleapi/scale-agentex/-/reverts/incident-mitigation_417-20260903-unhealthy-agent-recovery-5b4d188.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| adapter = TemporalAdapter(global_dependencies.temporal_client) | ||
| await ensure_reconciliation_schedule(adapter, task_queue) | ||
| try: |
There was a problem hiding this comment.
Schedule error stops worker
If Temporal rejects schedule creation or briefly loses its connection, ensure_reconciliation_schedule raises before the shared worker starts. The startup error handler only covers the later sweep. Health checks, retention cleanup, and scheduled agent runs then stop too. Keep a schedule-creation failure from stopping the shared worker.
Knowledge Base Used: Temporal health checks and retention cleanup
Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/run_worker.py
Line: 302-304
Comment:
**Schedule error stops worker**
If Temporal rejects schedule creation or briefly loses its connection, `ensure_reconciliation_schedule` raises before the shared worker starts. The startup error handler only covers the later sweep. Health checks, retention cleanup, and scheduled agent runs then stop too. Keep a schedule-creation failure from stopping the shared worker.
**Knowledge Base Used:** [Temporal health checks and retention cleanup](https://app.greptile.com/scale-ai/-/custom-context/knowledge-base/scaleapi/scale-agentex/-/docs/temporal-health-and-retention.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| return await reconcile_healthcheck_workflows( | ||
| self.agent_repo, | ||
| self.temporal_adapter, | ||
| self.task_queue, | ||
| ) |
There was a problem hiding this comment.
Disabled checks keep restarting
Turning off ENABLE_HEALTH_CHECK_WORKFLOW skips startup setup but leaves a schedule already saved in Temporal. When that schedule runs, this activity starts monitors without checking the flag. An operator can disable health checks and still have them restart and change agent statuses. Check the flag when the scheduled sweep runs.
Knowledge Base Used: Temporal health checks and retention cleanup
Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/activities/healthcheck_reconciliation_activities.py
Line: 17-21
Comment:
**Disabled checks keep restarting**
Turning off `ENABLE_HEALTH_CHECK_WORKFLOW` skips startup setup but leaves a schedule already saved in Temporal. When that schedule runs, this activity starts monitors without checking the flag. An operator can disable health checks and still have them restart and change agent statuses. Check the flag when the scheduled sweep runs.
**Knowledge Base Used:** [Temporal health checks and retention cleanup](https://app.greptile.com/scale-ai/-/custom-context/knowledge-base/scaleapi/scale-agentex/-/docs/temporal-health-and-retention.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| return await workflow.execute_activity( | ||
| RECONCILE_HEALTHCHECK_WORKFLOWS_ACTIVITY, | ||
| start_to_close_timeout=timedelta(minutes=5), | ||
| retry_policy=RetryPolicy( | ||
| maximum_attempts=3, |
There was a problem hiding this comment.
Large sweeps miss deadline
The five-minute limit covers a scan of every monitored agent, with one awaited start call per agent. If the fleet is large or Temporal calls are slow, the activity times out and retries from page one. A long retry can also cause the schedule to skip its next run, delaying missing monitors. Bound or split the sweep so it can finish within its limit.
Knowledge Base Used: Temporal health checks and retention cleanup
Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/workflows/healthcheck_reconciliation_workflow.py
Line: 16-20
Comment:
**Large sweeps miss deadline**
The five-minute limit covers a scan of every monitored agent, with one awaited start call per agent. If the fleet is large or Temporal calls are slow, the activity times out and retries from page one. A long retry can also cause the schedule to skip its next run, delaying missing monitors. Bound or split the sweep so it can finish within its limit.
**Knowledge Base Used:** [Temporal health checks and retention cleanup](https://app.greptile.com/scale-ai/-/custom-context/knowledge-base/scaleapi/scale-agentex/-/docs/temporal-health-and-retention.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| recovery_counter = workflow_args.get("recovery_counter", 0) | ||
| is_unhealthy = workflow_args.get( | ||
| "is_unhealthy", | ||
| workflow_args.get("initial_status") == "Unhealthy", |
There was a problem hiding this comment.
Status string breaks enum rule
This new comparison hardcodes "Unhealthy" even though AgentStatus.UNHEALTHY is available. The repository requires enum values instead of hardcoded status strings when an enum exists. Use the enum here; this requirement must be met before merging.
Rule Used: Use enum values instead of hardcoded strings when available. Reference existing enums from shared-types package rather than using string literals. (source)
Learned From
scaleapi/scaleapi#126557
Prompt To Fix With AI
This is a comment left during a code review.
Path: agentex/src/temporal/workflows/healthcheck_workflow.py
Line: 59
Comment:
**Status string breaks enum rule**
This new comparison hardcodes `"Unhealthy"` even though `AgentStatus.UNHEALTHY` is available. The repository requires enum values instead of hardcoded status strings when an enum exists. Use the enum here; this requirement must be met before merging.
**Rule Used:** Use enum values instead of hardcoded strings when available. Reference existing enums from shared-types package rather than using string literals. ([source](https://app.greptile.com/scale-ai/-/custom-context?memory=c0c58ddb-09dc-4e8a-9837-ec1bb02f9579))
**Learned From**
[scaleapi/scaleapi#126557](https://github.com/scaleapi/scaleapi/pull/126557)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
This is a fix to the agentex way of handling an agent going into "unhealthy" status. This is encountered when a kubernetes node is evicted, preventing the agent from coming back up cleanly.
This PR is not ready to merge: recovery can leave agents unhealthy, schedule errors can stop the worker, and disabling health checks does not stop an existing sweep.
Fix with agent prompt
Summary
Agent health checks keep running after agents become unhealthy and restore
Readystatus after two successful probes. Startup and recurring Temporal sweeps restart missing monitors, while deletion stops the agent’s monitor.Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR Boot[Worker startup] --> Schedule[Create sweep schedule] Schedule --> Sweep[Scan Ready and Unhealthy agents] Sweep --> Monitor[Start missing monitors] Monitor --> Probe[Probe ACP endpoint] Probe --> Status[Update agent status] Schedule -->|Every five minutes| SweepReviews (1) · Last reviewed commit: "fix(healthcheck): recover unhealthy agen..."