What Happens When My Training Crashes¶
TraceML preserves training failures and saves stdout and stderr from the supervised training command. Start with those logs to find the failing rank and the original error.
Training status and telemetry health are separate: training can fail while TraceML successfully saves the measurements collected before the failure.
Where to look first¶
- Read the final training message and the command's exit code.
- Open
logs/<run-name>/nodes/node_<node-rank>/training.stderr.log. - Find your exception, a
Fatal Python errorframe dump, or torchrun'sRoot Causeblock identifying the failing worker. - Check
statusandtelemetry_statusinlogs/<run-name>/manifest.json. - If telemetry failed, inspect
aggregator/process.stderr.logon node 0.
A traceback for torchrun's ChildFailedError describes a worker failure. It
is not necessarily the original exception from your training code.
What is saved¶
Paths below are relative to logs/<run-name>/:
| Artifact | What to look for |
|---|---|
nodes/node_<n>/training.stderr.log |
Python traceback, fatal-signal frame dump when available, and torchrun's failing-rank details. |
nodes/node_<n>/training.stdout.log |
Training output before the failure. |
manifest.json |
Training status, telemetry health, and saved artifact paths. |
aggregator/process.stderr.log |
Aggregator diagnostics; owned by node 0 in multi-node runs. |
final_summary.json, final_summary.txt |
Measurements saved if history was enabled and summary finalization succeeded. |
aggregator/finalization_warning.json |
Missing ranks or other finalization warnings, when present. |
aggregator/finalization_error.json |
Finalization failure details, when present. |
Training stream saving is enabled by default. --no-save-training-output
lets training inherit the terminal and creates no training stdout/stderr
files. It does not disable aggregator diagnostics.
Direct python or torchrun launches with traceml serve leave training
output capture to your terminal, scheduler, or container runtime.
Python exceptions¶
An uncaught exception propagates through the worker. TraceML stops its runtime in cleanup without suppressing the training error.
On the torchrun path, TraceML uses PyTorch's error recorder. The saved stderr
includes the traceback and torchrun's Root Cause block. PyTorch may also
write a temporary error.json and report its path as error_file; TraceML
does not copy that file into the run folder.
Normal cleanup sends remaining telemetry and a rank-finished marker. If the
aggregator also finalizes successfully, the report is saved and telemetry
can be complete even though training failed. In a distributed failure,
other workers may be terminated before completing their cleanup.
Native crashes and signal deaths¶
A segmentation fault or a killed process can bypass Python cleanup. Telemetry still queued in that worker is lost.
On the torchrun error-recorder path, Python's fault handler can print a frame
dump for fatal signals such as SIGSEGV:
Fatal Python error: Segmentation fault
Current thread ...:
File "/path/to/train.py", line 23 in main
This shows Python frames, not the native C, C++, or CUDA stack. Some deaths,
such as SIGKILL, provide no fault-handler dump. In those cases, inspect
scheduler or system logs as well as torchrun's worker details.
A torchrun Root Cause block can show the worker's signal:
rank : 0 (local_rank: 0)
exitcode : -11 (pid: ...) (SIGSEGV)
error_file: <N/A>
traceback : Signal 11 (SIGSEGV) received by PID ...
The CLI returns the supervised process's exit code. On the torchrun path,
this is torchrun's nonzero exit code; the worker's signal is in its
Root Cause block. If the supervised process itself dies from a signal,
TraceML returns 128 + signal number.
What the terminal prints¶
Summary and dashboard modes show training output live. CLI mode keeps it out of the live display and, after a failure, prints a saved stderr excerpt of up to 40 lines and 8 KiB.
With stream saving enabled, the final output includes the log paths:
[TraceML] Stderr: /path/to/logs/<run-name>/nodes/node_0/training.stderr.log
[TraceML] Stdout: /path/to/logs/<run-name>/nodes/node_0/training.stdout.log
The aggregator-owning launcher also reports telemetry health before the final training outcome. Other nodes do not report authoritative aggregator health. A telemetry failure does not replace the training exit code after training has started.
Telemetry and the final summary after a crash¶
A report after a crash contains only telemetry that reached the aggregator. An early failure can leave too few completed steps for a diagnosis. A final summary is not guaranteed if the aggregator fails, history is disabled, or summary generation cannot finish.
When a dead worker leaves a missing rank-finished marker, the aggregator can
finalize after all rank connections close and incoming telemetry settles.
Successful finalization with missing ranks records a warning and normally
reports telemetry_status: degraded with
telemetry_reason: finalization_warning.
If connections remain open, finalization waits within its configured budget:
--finalize-timeout-sec, 300 seconds by default. This is a telemetry deadline;
it does not terminate a hung training job.
See Reading the Output before interpreting an incomplete run as evidence of a performance problem.
Manifest fields¶
| Field | Meaning |
|---|---|
status |
completed for exit code 0, failed for a nonzero training exit, or interrupted when the launcher receives an interruption. |
lifecycle.training_ended_at |
When the supervised training process exited, if recorded. |
artifacts |
Paths to saved training streams, aggregator diagnostics, and reports that exist. |
telemetry_status, telemetry_reason |
Whether telemetry completed, degraded, failed, or was unavailable, and why. |
aggregator_exit_code |
Aggregator process outcome, when known to its owner. |
Use the command's exit code and torchrun's worker details for the training exit code or signal; the manifest does not have a general training-exit-code field.