Skip to content

Catch Training Regressions in CI

Compare a reference training run with a candidate and fail CI when Step Time increases beyond your threshold.

Regression Guard is experimental. It checks that both runs declared the same workload and topology, completed successfully, and have comparable Step Time measurements.

1. Declare your workload

Add a guard section to traceml.yaml in your training project:

mode: summary
history_enabled: true

guard:
  schema_version: 1
  workload:
    name: resnet50-imagenet-training
    parameters:
      model: resnet50
      data_version: imagenet-1k-v1
      precision: bf16
      per_rank_batch_size: 32
  measurement:
    start_step: 10
    completed_steps: 50

Use the same declaration for both runs. Record parameters that materially change the workload. TraceML searches the launch directory and its parents for traceml.yaml.

The measurement fields declare the intended range; they do not select the steps used for comparison. The check uses aggregate Step Time from each saved summary. Different analyzed ranges and step counts are allowed.

2. Run the reference and candidate

Run your reference code:

traceml run train.py --run-name reference

Then run your candidate code with the same workload declaration:

traceml run train.py --run-name candidate

Use comparable hardware and the same process topology. Choose a fresh run name for every execution; guarded runs do not overwrite an existing manifest.

These commands assume your script uses a supported automatic trainer or the required explicit integration. Regression Guard requires traceml run, summary mode, and history recording.

For DDP, use the same node count and processes per node in both runs. See Multi-node requirements for shared storage and per-node launch configuration.

3. Compare the results

traceml compare \
  logs/reference/final_summary.json \
  logs/candidate/final_summary.json \
  --max-step-time-regression-pct 5 \
  --output compare/reference-vs-candidate

The reference comes first. A candidate more than 5% slower fails the check. TraceML saves JSON and text comparison reports.

Both runs need matching normalized declarations and expected topology, completed training and launcher outcomes, and positive Step Time on a common CPU or GPU clock. Only Step Time determines the CI result; other measurements remain comparison context.

Result Exit code Meaning
WITHIN_THRESHOLD_IN_THIS_PAIR 0 Difference is within the threshold.
FASTER_IN_THIS_PAIR 0 Candidate is faster beyond the threshold.
SLOWER_IN_THIS_PAIR 4 Candidate is slower beyond the threshold.
INCONCLUSIVE 3 Required evidence is missing or incompatible.

Invalid input or output failures return 1; invalid command-line usage returns 2. Without --max-step-time-regression-pct, comparison is exploratory and does not enforce a CI threshold. See Compare Runs for decision details.

4. Add the check to CI

The experimental pilot deliberately leaves reference selection to the user. Make the chosen reference final_summary.json available to the job, run the candidate, and pass both explicit paths to traceml compare. For example:

- name: Check Step Time
  run: |
    traceml compare \
      artifacts/reference/final_summary.json \
      logs/candidate/final_summary.json \
      --max-step-time-regression-pct 5 \
      --output compare/reference-vs-candidate

- name: Preserve comparison evidence
  if: always()
  uses: actions/upload-artifact@v4
  with:
    name: traceml-performance-comparison
    path: compare/

The comparison step succeeds for a faster candidate or a result within the threshold. It fails for a slower candidate, invalid input, or inconclusive evidence. See the exit-code table when a CI system needs to distinguish those outcomes.

5. Choose a useful threshold

Repeat the reference workload on comparable hardware before choosing a threshold. Allow for normal variation, and keep the model, data, precision, batch size, and process topology consistent.

This check describes one pair of runs. It does not establish statistical significance or equivalent model quality. TraceML checks your declarations; it does not verify that the training program actually used those parameters.

Try the complete CPU DDP workflow

Run a small reference-and-candidate trial

The checked-in minimal DDP example provides a small end-to-end trial on a CPU-only machine. From the repository root, create this traceml.yaml:

mode: summary
history_enabled: true

guard:
  schema_version: 1
  workload:
    name: ddp-minimal-guard-trial
    parameters:
      model: tiny-mlp
      data_version: synthetic-v1
  measurement:
    start_step: 1
    completed_steps: 20

Run the same declared workload twice with fresh run names:

OMP_NUM_THREADS=1 traceml run examples/distributed/ddp_minimal.py \
  --run-name guard-reference \
  --nproc-per-node 2 \
  --args --steps 20

OMP_NUM_THREADS=1 traceml run examples/distributed/ddp_minimal.py \
  --run-name guard-candidate \
  --nproc-per-node 2 \
  --args --steps 20

Then evaluate the pair:

traceml compare \
  logs/guard-reference/final_summary.json \
  logs/guard-candidate/final_summary.json \
  --max-step-time-regression-pct 1000 \
  --output compare/guard-reference-vs-candidate

The first summary is always the reference and the second is the candidate. TraceML writes both JSON and text comparison artifacts before returning the CI exit code. This tiny CPU workload can vary by tens of percent between identical runs, so the wide threshold demonstrates the complete workflow rather than a meaningful performance verdict.

Advanced reference

Configuration fields

schema_version: 1 identifies the contract format, not a stable 1.0 release.

Field Required Type Rules
guard.schema_version Yes Integer Must be 1.
guard.workload.name Yes String Nonempty, no surrounding whitespace, at most 128 characters.
guard.workload.parameters No Mapping Defaults to {} and may contain at most 32 entries.
Parameter key Yes per entry String Nonempty, no surrounding whitespace, at most 64 characters.
Parameter value Yes per entry Scalar A nonempty string, signed integer, finite float, or Boolean. Strings may contain at most 256 characters.
guard.measurement.start_step Yes Integer Declared first completed step; must be at least 1.
guard.measurement.completed_steps Yes Integer Declared number of completed steps; must be at least 1.

Unknown fields, lists, nested parameter mappings, and null parameter values are rejected. Integers and the inclusive measurement end must fit in a signed 64-bit value. Workload names, parameter keys, and string values cannot contain Unicode category C characters, including control and format characters.

Workload identity

workload.name is the only required workload field. Choose a stable name that identifies the work being compared. TraceML stores it exactly as written, so surrounding whitespace is rejected instead of being removed silently.

workload.parameters is optional. Record values that materially change the work, such as model, data revision, precision, batch size, sequence length, or image dimensions. Model and data fields are recommended for training but are not required because TraceML also supports synthetic and pipeline-focused workloads.

Parameter values may be strings, integers, finite floating-point values, or Booleans. Lists, nested mappings, and null values are not supported in schema 1. Two guarded runs are eligible for a CI decision only when their normalized workload declarations match exactly.

YAML scalar spelling affects the stored type. PyYAML reads 1e-3 as a string, while 0.001 and 1.0e-3 are floating-point values; yes, no, on, and off are Booleans. It also follows YAML 1.1 integer forms: unquoted 010 is the integer 8, and 1:30 is the integer 90. TraceML accepts those parsed integers; for clarity, write step fields as ordinary decimal integers such as 10. Quote a value such as "010" when it is an identifier that must remain a string. Use an explicit decimal form for floating-point parameters.

These values are user declarations. TraceML records them but does not infer or verify that the training program used them. Do not include secrets, credentials, or machine-local paths because the values are persisted in run artifacts.

Measurement and history retention

TraceML completed steps are numbered from 1. start_step is inclusive and completed_steps is the exact requested count. This example:

measurement:
  start_step: 10
  completed_steps: 50

declares an intended range of steps 10 through 59. If --trace-max-steps is used, it must include the complete requested range.

The complete requested window should remain inside history_retention until the run is finalized. In the experimental pilot, this declaration is a compatibility fingerprint: two runs must declare the same range, but TraceML does not compare individual step IDs or use the declaration to re-aggregate telemetry. The CI decision uses the aggregate Step Time already stored for each final summary's step_time.global.window. That analyzed range can differ from the requested range or from the other run, and its analyzed-step count remains visible in the comparison.

Multi-node requirements

One-process and fixed-size DDP runs on one or more nodes are allowed. traceml watch, traceml serve, and direct traceml.init() launches ignore the guard declaration.

For multi-node DDP, use the ordinary launch shape in Distributed Training. Every launcher must discover the same guard declaration and use the same explicit --run-name, --nnodes, and --nproc-per-node; only --node-rank differs. Resolve --logs-dir to the same shared directory on every node. Node 0 records the declaration, and each launcher writes its outcome under that directory.

Saved artifacts and completion checks

Manifest, node outcomes, and final-summary context

The launcher validates and normalizes the declaration before starting the aggregator or training workers. After writing the manifest, it prints a short Measurement contract captured confirmation and stores the result in manifest.json:

{
  "guard": {
    "contract": {
      "schema_version": 1,
      "workload": {
        "name": "resnet50-imagenet-training",
        "parameters": {
          "data_version": "imagenet-1k-v1",
          "model": "resnet50",
          "per_rank_batch_size": 32,
          "precision": "bf16"
        }
      },
      "measurement": {
        "start_step": 10,
        "completed_steps": 50
      }
    }
  }
}

The source configuration path is not part of the portable contract. Changing traceml.yaml after launch cannot change the declaration captured for that run.

After its local training process exits, every guarded launcher also writes:

logs/<run-name>/nodes/node_<node-rank>/guard_outcome.json

The atomic, versioned record contains the public run name, node topology, normalized-contract digest, training status and exit code, and completion time. It excludes command arguments, paths, hostnames, environment contents, device identifiers, and credentials. Exit code 0 records completed training; any other observed exit code records failed training.

The fresh run name, topology, and contract digest bind the record to its run.

Node 0 gives outcome collection up to the configured finalization timeout, then writes a bounded result beside the contract in manifest.json. Aggregator shutdown is a separate bounded phase.

{
  "guard": {
    "training": {
      "status": "completed",
      "nodes_expected": 2,
      "nodes_observed": 2,
      "reasons": [],
      "nodes": [
        {"node_rank": 0, "exit_code": 0},
        {"node_rank": 1, "exit_code": 0}
      ]
    }
  }
}

completed means every expected launcher reported exit code 0. Missing, malformed, conflicting, or nonzero outcomes produce incomplete with stable reason codes. This establishes training-command completion only; it does not prove telemetry or measurement-window completeness. Collection and recording failures emit warnings but never replace the supervised training command's exit code.

Reason Meaning
node_outcome_missing An expected launcher did not report before the timeout.
node_outcome_invalid An expected record was malformed, unreadable, or too large.
node_outcome_conflict A record's run name, topology, or contract did not match.
node_training_failed At least one valid record contained a nonzero exit code.
node_outcome_collection_failed Node 0 could not complete collection because of an internal error.

guard.training is independent of the manifest's top-level run and telemetry statuses. Read all three fields when diagnosing an incomplete run.

The final report copies the portable declaration, expected topology, and aggregate launcher result into final_summary.json under run_context. Internal coordination fields and individual node outcomes remain only in the manifest and node artifacts. This lets local CI comparison consume one summary per run without reading SQLite or joining a second artifact. An ordinary run has no declaration or launcher_completion; when no readable launcher manifest exists, run_context is an empty object. An interrupted run can omit launcher_completion when its summary finishes before node 0 consolidates the launcher outcomes; run.status still records the interruption.

To inspect the captured declaration, format the manifest and look under guard.contract:

python -m json.tool logs/reference/manifest.json

Tested environments

The pilot keeps implementation support separate from environments exercised end to end in automated tests.

Environment Current qualification
Linux, CPU, two-rank Gloo DDP A built wheel runs two guarded jobs and compares their summaries in CI.
Single-process and ordinary single-node DDP Covered by launcher, reporting, and comparison tests.
Multi-node outcome coordination Covered on one machine with two launchers and a shared run directory. A real multi-machine environment is not yet qualified.
CUDA Supported by ordinary TraceML paths, but the complete guarded comparison workflow has no dedicated GPU CI qualification yet.
Windows The wheel and CLI surface are smoke tested; the complete guarded DDP workflow is not qualified there yet.

The comparison is one observation about one explicit pair. It does not claim repeatability, statistical significance, model-quality equivalence, complete telemetry delivery, or support for changing process membership.

Troubleshooting

Symptom What to check
A run refuses to start because its manifest exists Choose a fresh --run-name; guarded runs never overwrite an earlier run.
final_summary.json is missing Inspect the training and aggregator output. Comparison requires a finalized summary from each run.
The result is INCONCLUSIVE because run context is missing Produce both summaries with guarded traceml run executions rather than an ordinary or direct-SDK run.
The declarations do not match Use the same normalized guard section for both runs, including parameter types and values.
The topology does not match Use the same node count and processes per node for the reference and candidate.
Launcher completion is incomplete Inspect the per-node launcher output and the root manifest to find the missing, invalid, conflicting, or failed node outcome.
Step Time is unavailable on a common clock Confirm that both finalized summaries contain measured positive CPU Step Time, or measured positive GPU Step Time.

Preserve both final_summary.json files and the generated comparison JSON and text report when investigating a CI failure. Manifests and SQLite databases are not inputs to the comparison.