Project hosts
Spot recovery strategy for project hosts
Explain spot retry windows, standard fallback, probes, and returning from fallback to spot.
Why spot recovery exists
Spot hosts can be much cheaper than standard on-demand hosts, but the cloud provider can reclaim them at any time. CoCalc's spot recovery strategy controls what happens after that interruption: retry spot, optionally fall back to a standard VM, and later probe whether spot capacity is available again.
Spot recovery is active only when the host uses spot pricing and Interruption restore is set to Restore immediately. The Spot Recovery Strategy modal shows the recovery states as a diagram, but the diagram is read-only; the settings below it control behavior.
Retry spot first
After a spot interruption, CoCalc first tries to restore the same kind of spot capacity. The key settings are:
- Spot retry window (minutes): how long CoCalc keeps retrying spot before moving on.
- Retry backoff (seconds): the base delay between spot restore attempts. The worker adds exponential backoff up to a cap.
- Max restore attempts before fallback: a count-based limit. Use a positive
value. Entering
0currently uses the default count instead of disabling the attempt limit.
Use a short window when user-facing uptime matters. Use a longer window when cost matters more than immediate recovery.
Standard fallback
When Allow standard fallback is enabled, CoCalc can temporarily switch the host to a standard on-demand VM if spot recovery fails. The host remains configured as a spot host, but it is running as a standard fallback. The UI shows this as standard fallback and explains the current standard rate and the spot rate when restored.
The fallback settings are:
- Minimum standard runtime (minutes): how long the standard fallback should run before CoCalc starts trying to return to spot.
- Spot probe interval (minutes): how often to check the same zone and machine type for spot availability.
- Require successful probe before returning to spot: when enabled, CoCalc only switches back after a matching probe VM starts successfully.
Returning to spot
While a host is on standard fallback, CoCalc probes for spot availability. After a successful probe and the minimum runtime window, it can move back to spot. Returning to spot is itself disruptive because the underlying VM changes, so schedule sensitive workloads accordingly.
Agent notes
When explaining spot recovery, distinguish three states: desired pricing (spot), effective pricing (possibly standard fallback), and recovery phase. Use spot for cost-sensitive workloads that tolerate interruption. Use standard hosts for workloads that must not be interrupted by cloud spot reclamation.