Project hosts
Understand project host reliability
Read host reliability, availability, outage exposure, planned downtime, and day-grid signals.
What the Reliability view measures
The host Reliability tab summarizes recent host availability. It is not a generic cloud SLA and it is not a project success metric. It answers: when this host was intended to be online, how often was it actually reporting online?
The modal and tab show:
- current state, such as online, planned downtime, or recovering
- current uptime
- window availability over the selected lookback period
- reliability over intended-online periods
- unplanned outage count
- unplanned exposure time
- planned downtime, when present
Reliability versus availability
Reliability measures uptime only during periods when the host was intended to be online. Planned downtime is excluded from the reliability denominator.
Availability divides online time by the selected window after subtracting periods recorded as Unobserved. Planned downtime remains in that denominator. A host intentionally stopped for most of the month can therefore have low availability but good reliability. Check the displayed unobserved duration too: missing observations are not evidence that a host was healthy.
Reading the day grid
The small day squares summarize the recent window. Green days were reporting online. Yellow or red indicates unplanned exposure. Gray can indicate planned downtime or unobserved time. Hovering a day shows the distinction and durations.
If the host is currently unavailable, the top alert distinguishes planned unavailability from unplanned or recovering state.
Admin annotations
Admins can annotate recent non-online events. Use this to distinguish planned maintenance, provider incidents, testing, billing holds, or known user-driven stops. Public notes should be written carefully because they can be shown to users.
Agent notes
Use reliability when deciding whether a host is suitable for long-running workloads. If a user reports intermittent failures, compare reliability, current state, host logs, active operations, spot recovery state, and project events before blaming a notebook or terminal.