A forecasting pipeline can finish successfully while answering the wrong question. It can write valid Parquet, train XGBoost, and rank rows convincingly even though a feature contains a future event or a negative label means only that collection stopped early. Those errors survive ordinary checks for missing files and failed training runs.

In Cloud Orchestrator, I repair that meaning before comparing models. The Borg track declares a causal schema, complete label horizons, and chronological partitions. The active controller uses a separate pair of next-window predictors. The distinction follows through to their artifacts and interpretation: sharing a model library does not make two prediction targets interchangeable.

Repair the contract before the model

The joined dataset records a schema version, the configured failure-event definition, and an explicit complete-event boundary. The advanced feature builder rejects missing schema fields, mixed versions, or a different failure definition. It does not silently interpret an older artifact using today’s rules.

The contract connects each failure mode to a specific check:

Failure modeRequired behavior
A task’s final state leaks into an earlier rowOnly backward as-of task and machine state enters features
Incomplete future observation appears healthyLeave the target unknown until the whole horizon is covered
Temporal features depend on arbitrary row orderSort by cluster, time, and task identity before grouped history
A derived request/capacity ratio has a zero or missing denominatorPreserve missingness rather than inventing zero utilization
A model sees all available columnsSelect the declared feature list; retain future fields only as metadata
A schema repair leaves old files in placeReject incompatible artifacts and rebuild into an explicitly chosen location

Resource ratios and rolling features retain missingness flags because absence carries meaning. “No CPU request was recorded” gives me different information from “CPU use was zero.” If both become zero, the model loses that distinction and a later feature-importance report cannot recover it.

Historical terminal fields also remain available for inspection without becoming training inputs. That separation helps investigate why a row was labeled without letting its eventual outcome influence its prediction. Feature construction and input allowlist

Multi-horizon labels need multi-horizon discipline

A five-minute failure target and a longer target ask different questions. A row can have a mature short-horizon label while its longer-horizon label remains unknown. The builder evaluates each horizon against the declared coverage boundary independently.

The essential label calculation is:

pl.when(
    pl.col("end_time") + horizon_us
    <= pl.col("event_observation_end_time")
).then(
    (
        (pl.col("next_failure_time") > pl.col("end_time"))
        & (pl.col("next_failure_time") <= pl.col("end_time") + horizon_us)
    ).fill_null(False)
).otherwise(None)

The False case is valid only inside the complete-coverage branch. Moving that fallback outside the branch would turn missing evidence into apparently healthy examples.

The target horizon therefore belongs to the model’s identity. A probability needs an event definition and an interval: a five-minute failure score and a one-hour failure score can differ without either model being inconsistent.

Split time before searching parameters

I reserve training, development validation, and final test periods before tuning. The default split uses 20% validation and 20% final test, with the earlier interval available for training. Boundaries come from observation timestamps; the target horizon then determines which rows must be purged.

For a 15-minute target, a training observation just before validation begins can include events from validation in its label. The implementation therefore requires the training horizon to end strictly before the validation boundary. Validation uses the same rule against the final test boundary.

An illustrative partition looks like this:

training observations  | purged boundary | validation observations
                                                | purged boundary | final test

Each retained label's horizon must remain inside its own interval.

Purging reduces the usable sample, but it preserves the purpose of the held-out period. With random row splitting, nearby observations of the same task can place almost identical history—and the same eventual failure—on both sides of the evaluation. Chronological partitions alone do not solve that problem when label horizons cross their boundaries. Temporal split and sampling implementation

Do not make validation artificially easy

Failure data is imbalanced, so the training path can sample negatives to keep a large experiment manageable. Validation needs to reveal how the predictor behaves under the population it will encounter. Its bounded sample is therefore independent of the label, avoiding deliberate class rebalancing. A finite sample can still differ from the full population; it is not selected to make failure easier to find.

Average precision is especially sensitive to this distinction. A model evaluated on an enriched positive sample can look useful while producing too many alerts on the natural population. The metric implementation also handles tied scores without depending on row order and returns an undefined result when validation has no positive evidence.

An undefined result is useful here. It tells me that the window cannot support this comparison, whereas a convenient numerical default could be mistaken for a measured difference between models.

After development selects a model, the advanced evaluator reads the complete reserved final holdout once. It checks the frozen model and source identities and refuses reuse or a changed artifact. If the final result motivates another design choice, that choice needs new reserved evidence. The held-out interval cannot become a permanent tuning dashboard.

The controller predicts a different target

The active stack does not simply load the multi-horizon Borg classifier and call it a cloud controller. It has its own causal dataset contract and next-window targets:

PredictorDeclared target
Stack risk predictorNext-window saturation or new-failure proxy
Stack demand predictorNext-window maximum CPU or memory fraction
Advanced Borg classifierQualifying trace failure within a named future horizon

For the stack predictors, partitions are established before constructing next-window pairs. The last observation of a partition has no following label in that partition; the code does not fabricate one or borrow the first row of the next partition.

This distinction changes how I describe a dashboard number. A high next-window saturation proxy is evidence of modeled pressure. It is not a calibrated probability that a tenant will miss its latency SLO, and it is not the same target as a Borg terminal event. Stack dataset contract, predictor implementation

A model score does not own an action

The next boundary is operational. A risk score can inform recovery reasoning, but it cannot authorize arbitrary replication, migration, or throttling. The coordinated controller assigns service sizing to D, placement to E, and corroborated restart to A. Recovery requests go to their owners, and every proposal must satisfy current feasibility and operator scope.

That design avoids turning a few probability thresholds into hidden mutation authority. A model can be wrong, stale, unsupported, or out of distribution. The controller still needs valid state, a permissible action, and an adapter capable of carrying it out.

The repaired pipeline makes a prediction traceable to its data, target, and evaluation interval. Feature importance and development metrics can then help explain it. Moving from that explanation to a cloud-control claim requires another experiment: the large external Borg artifacts still need a schema-2 rebuild, and workload transfer and action-response calibration still need measurements. That is the boundary the controller architecture takes up next.

Previous: Part 1 — A Causal Foundation. Next: Part 3 — Eight Roles, One Mutation.

Implementation references link to a pinned snapshot of the private research repository and require repository access.