On this page
Training a model and giving it control are different engineering decisions. Lower prediction error can identify a promising candidate without establishing that it will help tenants. A candidate can also outperform an earlier baseline while being incompatible with the workload, permissions, or incumbent policy in the next experiment.
I designed the learning path in Cloud Orchestrator around that handoff. The project began with Borg trace forecasting and grew into a six-layer control and evaluation prototype: eight specialist roles propose bounded operations, a Referee checks and selects them, and an independent evaluator assesses customer and provider outcomes. The source repository is Cloud Orchestrator Architecture; the shorter public name reflects the broader system.
Learning adds a second loop. It fits candidates from eligible development feedback, freezes a candidate for comparison with the current policy, and activates it only after new measured evidence passes the promotion gate. The loop can repeat, but retaining the incumbent is a valid result. Continuous learning describes the workflow, not a promise of continuous improvement.
Two loops, two kinds of authority
The control loop has immediate operational responsibility:
current state and forecasts
→ A–H proposals
→ ownership and feasibility checks
→ guarded proposal selection
→ one operation
→ verified postcondition and next observation
The learning loop has a different responsibility:
independent development outcomes
→ exact decision/measurement join
→ bounded fitting and temporal validation
→ inactive candidate
→ fresh frozen candidate-versus-incumbent experiment
→ independent outcome gate
→ atomic promotion or keep the incumbent
The watcher can continue processing eligible feedback until stopped. Experiment collection and matched resets remain external responsibilities. During a frozen trial, the controller explicitly selects the candidate or incumbent without changing the active policy pointer, so testing a candidate cannot silently turn into deploying it. Continuous-learning architecture
What actually learns
The eight specialists remain deterministic constrained policies:
- A coordinates recovery and owns corroborated restart.
- B owns provider allocation; C owns admission and fairness.
- D owns service sizing; E owns placement.
- F advises on interference; G advises on incident diagnosis.
- H dispatches dependency-ready work inside accepted DAGs.
The optional learned component is a central ridge-regression ranker. It estimates the outcome associated with an eligible action in its observed context. Its feature representation has eight action-kind blocks and ten bounded context/action features per block, including pressure, demand, headroom, budget, waiting time, and requested change.
I chose this limited learning problem because it preserves the specialists’ ownership and proposal contracts. The historical A/B/C PPO implementation remains a separate legacy_abc path with different observations and actions. Its checkpoints do not supply policies for the eight-role controller.
The ranker does not generate arbitrary operations. It can order proposals that the specialists and feasibility checks have already made eligible. Recovery precedence, admission fairness, long-wait protection, cooldowns, contraction stabilization, and uncertain-operation guards remain fixed. There is still at most one mutation per observation.
The diagnosis baseline forms bounded structured hypotheses, and the evaluator computes its result from measured records. Neither requires an LLM. That keeps the learning question focused on ranking supported choices rather than introducing another source of authority over execution or acceptance. Adaptive ranker implementation
Keep the policy identity visible
The Learning & artifacts view places decision-learning metadata beside the separate legacy training signals. Its fields identify the decider mode, policy generation, policy used for the recorded decision, frozen evaluation protocol, and watcher status. Those identities are useful when an operator needs to explain why a decision used one policy rather than another.

Current interface captured on 29 September 2026 from the repository’s synthetic saved fixture. The fixture supplies no adaptive-policy or watcher record, so the corresponding fields remain unavailable or not reported. This shows missing metadata explicitly; it does not demonstrate fitting or promotion.
A dashboard file path or training status can help locate an artifact, but it does not activate it. The decision record must identify the policy actually used, and the registry must retain the evidence behind any promotion. Keeping missing values visible makes that distinction easier to audit. Learning-view implementation
Small training windows do not justify broad confidence
The first learning cycle requires 20 accepted development windows. Later cycles require at least four new windows, and a fit uses at most the most recent 2,048. The temporal development split is 80/20, with overlapping windows purged; at least 12 training and four validation windows must remain.
Each scored action kind also needs at least four training examples. Input support must remain inside the recorded feature range with a bounded margin. If any eligible choice lacks support, the entire selection returns to deterministic arbitration. If observed tenant contracts differ from the registry’s frozen contract identity, selection falls back to the deterministic path.
These thresholds establish a minimum supported fitting problem. They do not show that 20 windows are enough for a reliable deployment estimate. The controller also checks whether the present choice resembles the recorded support, because a numerical prediction alone says little about the reliability of an extrapolation.
Unsupported input and invalid registry state lead to different outcomes. A valid model can abstain when its input is unfamiliar and return selection to the deterministic policy. Corrupted or incompatible learning state blocks selection so that the underlying failure remains visible.
Only the executed decision earns a feedback row
Suppose D proposes more memory, E proposes moving a task, and the Referee selects D. A later successful outcome cannot become a positive training example for both proposals. E’s alternative was not executed.
The feedback join binds the full observation, selected original proposal, active policy, controller identity, and verified native operation to a following independent measurement window. Incomplete, duplicate, changed, overlapping, future, or locally confounded feedback is rejected. Accepted historical development windows can still participate in later bounded fits; duplicate feedback IDs or modified consumed records cannot enter as new evidence.
The development target combines normalized goodput and provider cost per good request, with penalties for tenant SLO or budget failure. Objective weights and scales are part of registry identity. They cannot drift between fits while the result continues to claim the same experiment.
The fitted relationship is a selected-outcome association. Its examples describe what the controller executed and what followed, while the outcomes of rejected proposals remain unknown. Temporal prediction error can help assess the fit, but promotion and per-agent causal credit require comparisons that these observational records do not provide.
Frozen evaluation decisions are excluded from training even if their feedback is relabeled as development. Sealed final tests also cannot select or promote a candidate. Once evaluation outcomes feed model choice, they have become development data and need a different evidentiary role. Learning cycle and feedback checks
Promotion is a fresh experiment
The registry binds code, operator scope, adapters, workload, capacity, tenant contracts, and objective. A candidate’s baseline must still be both its parent and the current incumbent when the evidence is evaluated. This prevents a delayed result from promoting a model against a policy that has already changed.
Promotion uses the existing independent CSC/CSP gate. The experiment freezes its protocol after fitting and pins the protocol file’s exact digest before collecting new paired measurements from matched initial state. The promotion CLI separately checks the collected evidence file against a supplied digest. The offered workload, capacity, tenant contracts, and measurement definitions remain fixed across candidate and baseline.
Customer evidence retains every offered request, including errors and unsent work. Provider evidence includes declared accounting, while optional energy requires measured joules. Every tenant’s SLO and budget must pass. The selected primary endpoint must improve without violating the other stakeholder’s protection.
The learner recomputes acceptance from raw evidence rather than inheriting a saved passed flag. Synthetic registries can exercise fitting and rejection, but they cannot activate a native policy. The promotion path needs the declared measured comparison.
Its result remains scoped to that experiment. Artifact hashes preserve identity and integrity; collector authentication, storage permissions, and trial independence must come from deployment controls and an auditable collection process outside the controller’s write authority. Outcome protocol and trust boundary
Repeated attempts consume evidence and error budget
One nominal 95% comparison is not an unlimited license to keep testing candidates until something passes. Repeated opportunities to promote also create repeated opportunities to accept noise.
The development promotion loop assigns attempt n an error allocation:
alpha_n = 0.05 / (n × (n + 1))
These allocations sum to at most 0.05 over the registry’s lifetime. Promotion requires a correspondingly stricter paired Student-t lower bound to exceed the predeclared meaningful effect. This budget depends on valid fixed-sample tests and their statistical assumptions; it cannot rescue dependent or fabricated measurements.
Failed evaluated attempts consume their allocation and data too. New attempts need fresh, non-overlapping measurement windows and partitions. The exact pair count is declared before collection; the learner cannot append more pairs after inspecting a disappointing result.
Promotion history is therefore part of the experiment’s state. Deleting a registry would discard consumed evidence and reset the recorded error budget. Storage and archival procedures need to preserve those identities along with the models; retaining only the newest weights would lose part of the justification for using them.
Activation and rollback keep the execution history
Passed evidence is preserved first, then promotion history and the active pointer are updated atomically under a registry lock. The controller reads the active version on each decision. Candidate fitting leaves the pointer untouched.
An operator can withdraw the learned version and return to the deterministic baseline. That rollback preserves the native execution journal. If a resize request is still uncertain, reverting the ranker must not release another action against the same resource.
Rollback is an operator command. The implementation does not automatically detect post-promotion drift or initiate rollback from monitoring, so a deployment still needs an owner and intervention criteria for that decision. It also needs permissions that enforce the intended separation among learner, collector, and controller.
Assess a role without inventing credit
The A–H assessment path separates decision quality from contribution. It can inspect whether a proposal respected ownership, available evidence, and the role’s constraints. That does not show that the role caused a better service outcome.
A contribution experiment substitutes one preregistered safe baseline role while fixing the other policies, Referee, scope, capacity, workload, and fault schedule. F and G are assessed as advisory roles; H is meaningful only for DAG workloads. Selection frequency and common reward are not per-agent causal scores.
The implemented gate accepts one preregistered role hypothesis per report. Eight simultaneous significance claims would need a separately designed multiplicity protocol. More agents increase the importance of experimental discipline; they do not create more independent evidence automatically. Role assessment implementation
What is working, and what the next experiment must establish
This article describes source snapshot 9513ba717292b351895e26a8d7eddcdea5c0a96e. The recorded local verification on 29 September 2026 passed 452 Python tests and the dashboard behavior checks under Node 22. Those checks exercise the data, coordination, learning, execution, and evidence contracts; the current dashboard captures illustrate their interface with synthetic data.
Those results do not establish a newly trained full PPO curriculum, deployed application/provider gateways, externally enforced cloud permissions, or an independently measured cloud advantage. Native action tests include mocked API paths. Small real model and artifact tests are valuable, but their outcome is software verification.
The next empirical milestone needs a workload with explicit tenant contracts, collector authority outside the controller, raw provider accounting, reproducible matched resets, and a powered fixed-sample comparison. Action-response and predictor calibration must also be measured before broader transfer claims. The large historical Borg artifacts still need their separate causal-schema rebuild.
The engineering result is a controlled handoff from a fitted model to an active policy. A candidate has an identity, an incumbent to compare against, a defined evidence requirement, and a history that survives activation or rollback. When the evidence is inconclusive, the incumbent stays in place. That gives the next cloud experiment a clear question to answer and preserves the operational context needed to interpret its result.
Previous: Part 4 — Tuning a Controller Without Tuning Its Evidence. Start the series with Part 1 — A Causal Foundation.
Implementation references link to a pinned snapshot of the private research repository and require repository access.