Project hosts
Inspect host metrics and logs
Interpret host metrics and logs to investigate resource pressure, provisioning, and runtime failures.
What host logs are for
The host Logs tab shows operational history for the project host itself: provisioning, bootstrap, lifecycle actions, software reconcile, daemon state, provider errors, and recent host-controller activity. Use it when the host is not starting, projects cannot be placed, software is drifting, or a lifecycle action looks stuck.
Host logs are different from project logs. A notebook kernel crash, terminal process failure, or web app error may be a project problem. A failed provision, daemon rollout, provider API error, or unavailable host service is a host problem.
First things to check
Start with the drawer Overview and the relevant tab:
- Overview for current state and active operations.
- Reliability for recent online/offline history.
- Runtime for software lifecycle, drift, and daemon health.
- Logs for the event stream behind those summaries.
For CLI inspection, use a small recent tail first:
cocalc host logs HOST_ID --tail 200
If the recent tail is not enough, narrow by the time of the failed action instead of dumping unrelated history.
Read resource measurements before changing capacity
Use the selected host's Overview tab and Current metrics card to check whether a slow or interrupted job coincides with host pressure. These are host-wide observations, including other projects and host services; they do not identify your job's bottleneck on their own.
Check the sample time first. metrics pending, metrics stale, an absent field, or an empty history is missing or outdated evidence, not zero usage. A recent host action can make an earlier sample irrelevant. Compare samples from the period when the problem happened with the relevant project activity.
A computation is slow
Inspect host CPU percentage and load averages; the project's process/activity view.
Host CPU is aggregated across cores and projects. Check whether the affected program uses multiple cores and whether other work is competing. One host percentage does not establish that adding cores will speed up this job.
A notebook kernel is killed or memory grows
Inspect the kernel's memory display, project RAM limit, and host available memory.
The kernel display includes its child processes but excludes unrelated project processes. All of those processes share the project limit. Compare both scopes using memory troubleshooting and host RAM policy.
Writes fail or a project cannot start because of storage
Inspect project disk quota, host disk and filesystem-metadata space, root-disk space, and shared scratch where applicable.
These are separate limits and locations. Match the error to the affected storage before deleting files or changing disk size. See host storage; do not assume host free space means the project has free quota.
File-heavy work stalls while CPU use is modest
Inspect available I/O containment status, sampled per-project read/write rates, and the workload's own timings.
Check capability and sampling errors first. The top-project list is sampled and can be truncated; it is not a complete per-process profiler. Transfer rates alone do not prove that disk hardware is the bottleneck.
A GPU allocation fails or GPU work is slow
Inspect device and framework measurements on the machine that actually runs the code.
Current host metrics do not report GPU utilization or GPU memory. Host RAM is not GPU memory; missing GPU measurements do not mean the device is idle or has space.
The notebook usage display does not currently report CPU or memory use for a remote Jupyter kernel. Inspect that remote machine instead of interpreting the CoCalc host's measurements as remote usage.
Inspect current and recent metrics from the CLI
Use the CLI authentication guide to select the correct site and account. Metrics history requires the host owner, a host manager, or a site administrator. Membership in a project alone does not grant that host-level access. Replace HOST_ID with an existing host you are authorized to inspect:
cocalc --json host metrics HOST_ID --window 1h --points 60
The JSON success response wraps the host identity and metric results in data. Check data.current.collected_at and the timestamps in data.history.points. data.current can be null, history points can be empty, and returned points can be compacted to the requested maximum. Requesting more points does not create measurements that were never collected. data.derived summarizes sampled storage risk and can be null. It is not a benchmark or a promise about future capacity.
When data.current.io_containment is present, check capability, capability_reason, sampling_error, sampled_project_count, total_project_count, and truncated before interpreting top_projects. An unsupported collector or a partial sample is not evidence of no I/O load.
Command syntax was checked with CoCalc CLI 1.0.3 on 2026-09-11. This reference does not include an observed host measurement or a before/after performance comparison.
Record the input size, command or notebook, machine, concurrent work, sample timestamps, elapsed time, and output check for a representative run. Use those observations to choose one change, then compare the same input and verified result. Keep raw host output private when it includes identifiers for other projects. If the data does not identify a constraint, retain that uncertainty and inspect the application before changing capacity.
Reading log patterns
Provider errors usually point to credentials, quotas, unavailable machine types, pricing mode, spot interruptions, region or zone capacity, or network setup. Bootstrap errors usually point to package installation, image setup, SSH/connector availability, or first-start configuration. Runtime errors point to daemon health, version drift, reconcile failures, or project-host service rollout.
When a host is recovering from spot interruption or fallback, logs are most useful when read together with the spot recovery state and current effective pricing.
Sharing logs safely
Logs may include host ids, project ids, paths, provider names, and operational context. Avoid pasting large raw logs into public channels. Prefer a short tail around the failure time, plus the host id, action attempted, current state, and any active operation id.
Agent notes
When helping with host debugging:
- Ask for or select the host id.
- Open Logs, Runtime, and Reliability rather than using logs alone.
- Capture the action attempted, approximate time, current host state, active operation, and whether the host is spot or standard.
- Use
cocalc host logs HOST_ID --tail 200for a focused first pass. - Route host inspection to the host-owning bay; do not assume the browser's current project bay owns the host.