When I investigated a Kubernetes incident, the information I needed was usually spread across several clocks. Resource metrics described a sampling window, the Horizontal Pod Autoscaler evaluated its own signals, and new capacity took time to become usable. Pending pods made the pressure visible, but by the time I connected the charts to the events, I was often reconstructing something that had already happened.

I started with failure forecasting because it offered a smaller, testable question: could recent resource history provide a useful signal before the controller chose an action? That work became the data foundation for Cloud Orchestrator, a research prototype that also coordinates eight specialist roles and evaluates customer and provider outcomes. Borg traces provide forecasting data and the project’s starting point; the controller is not an implementation of Google’s Borg scheduler.

Before choosing a model, I needed to settle two definitions: what was observable when a prediction was made, and which later event would make that prediction correct?

Three workspaces with different responsibilities

The repository contains three related tracks. Each has a different artifact contract, so I keep their data and model outputs separate. A baseline model is useful for checking the trace pipeline; a controller predictor must answer the question its runtime actually asks.

TrackLocationResponsibility
Baseline forecastingroot scripts, ~/Documents/borg_data, ~/Documents/borg_processedJoin trace evidence and build a simple failure-risk baseline
Advanced XGBoostsrc/advanced_xgboost/, ~/Documents/borg_xgboost_workspaceMulti-horizon features, model selection, and frozen holdout evaluation
Active controllerorchestrator_stack/Observations, predictions, specialist proposals, execution, and research evidence

Large raw and processed files stay outside Git. The advanced workspace has its own raw, processed, models, reports, runtime, and config directories. Schema explanations also live beside generated data. A repository README is useful, but a Parquet file copied into another workspace needs its meaning to travel with it.

This becomes important after a data-contract change. A model can still load even when its label definition is no longer valid for the experiment. Compatibility checks connect an artifact to the schema and target that produced it; the directory structure makes that relationship easier to inspect.

The observation has a clock

Each joined row represents a completed usage window. Its prediction time is end_time, measured in microseconds. CPU, memory, task state, and machine context must describe information available at or before that point.

A plain join on task identity is insufficient. The task can have many events, including later termination. A machine can also change after the observation. Attaching a final task record or the latest machine record without a time condition lets the model see an answer that the live controller could not have known.

The current joined dataset declares causal schema version 2 and uses backward as-of joins for task and machine features. The relevant shape in the dataset builder is:

usage.sort("end_time").join_asof(
    events,
    left_on="end_time",
    right_on="event_time",
    by=TASK_KEYS,
    strategy="backward",
    suffix="_event",
    check_sortedness=False,
).join_asof(
    machines,
    left_on="end_time",
    right_on="machine_event_time",
    by="machine_id",
    strategy="backward",
    check_sortedness=False,
)

The join direction gives the row its causal meaning. Task and machine features describe the latest known state at the observation boundary, while future failure information is joined separately for labels and excluded from the declared model feature list. next_failure_time, last_event_time, and final_event_type can help explain outcomes; their presence in the dataset does not make them legitimate predictor inputs. Dataset construction source

Unknown is a necessary label

For a prediction made at time t with horizon h, the label asks whether a qualifying failure occurs in (t, t + h]. A failure at t is already observed; it is not the future event being predicted. The builder finds the next strictly future qualifying failure rather than assuming a task’s final event answers every earlier window.

Even that is insufficient without coverage. Suppose a 15-minute label starts at 10:00, but the event extract ends at 10:08. Seeing no failure by 10:08 does not establish survival through 10:15.

The operator therefore declares the complete event boundary through BORG_EVENTS_COMPLETE_THROUGH_US. The maximum timestamp in a file is not a substitute: a late event says that one event was collected, not that every earlier event was collected.

Available evidenceLabel
Entire horizon covered; qualifying failure inside itPositive
Entire horizon covered; no qualifying failure inside itNegative
Horizon not fully coveredUnknown (null)

Only mature labels enter supervised training. The coverage requirement applies consistently, including to rows with an observed failure. That costs some examples, but it ensures that positive and negative labels come from the same observation contract. Feature and label contract

A baseline should expose the data

The baseline produces validation predictions and high-risk alert rows alongside metrics and feature statistics. Those row-level outputs let me ask concrete questions: which observation received a high score, which features existed at that time, and was the target mature?

The same discipline applies to evaluation. Before selecting a model, I reserve separate training, development-validation, and final-test intervals. A training row is removed near the validation boundary if its forecast horizon reaches into validation; the same rule separates validation from final test. Otherwise adjacent rows can share future outcome information even when their observation timestamps differ.

The baseline trainer reports development validation and explicitly records that it has not evaluated the final test. The advanced path owns the later frozen holdout procedure. Keeping that distinction visible is more useful than attaching the word “test” to every metric. Temporal partition implementation, baseline trainer

What this foundation establishes

The forecasting code can construct causal inputs, refuse incompatible schemas, represent missing future coverage, and evaluate development predictions without crossing the declared time boundaries. Small regression runs exercise real XGBoost training and artifact restoration.

That does not mean every historical data file has been rebuilt. The research-readiness record leaves the large external Borg rebuild outstanding; preserved metrics from older artifacts cannot be carried forward as results for schema 2. It also does not establish that a Borg-trained score is calibrated for an EKS workload. Transfer requires its own measurements. Readiness evidence and remaining work

A useful failure score starts with a traceable observation. I need to recover the information available at prediction time, the future interval being predicted, and the evidence that makes its label complete. With those boundaries in place, a surprising score can lead to a data or model investigation instead of an argument about what the dataset meant.

Next: Part 2 — Data Repair, Temporal Validation, and Two Different Prediction Tasks. The current learning architecture is covered in Part 5.

Implementation references link to a pinned snapshot of the private research repository and require repository access.