TRAINING-OPS / ROBOT LEARNING
Delegate training.
Ground the decision.
A Training-Ops agent for long-running robot-learning work, from simulation to an evidence-backed evaluation decision.
Starts training, manages goal-driven iterations and compares policies. Leaves a traceable run record behind each decision.
The same starting point. Two policies. A separately recorded, illustrative simulation replay.
This is not footage from, or additional qualification evidence for, the original 120 episodes.
01 / THE DELEGATED WORK
Not just a run. A traceable loop.
Isaac Lab handles physics and RL training. Training Factory manages the goal, progress and evaluation conditions of the work around it.
- 01
Define
Set the job, goal, budget and evaluation conditions.
- 02
Run
Launch simulation and training processes and track their progress.
- 03
Evaluate
Compare policies under fixed conditions across environments and seeds.
- 04
Decide
Advance the goal loop and report policy qualification with its evidence.
The goal and iteration budget bound the loop. Qualification can also run without training. The incumbent policy is never replaced automatically.
02 / RECORDED DEMO RESULT
One better metric is not enough.
Some measurements improved in the warehouse demo. But the candidate did not meet every fixed acceptance condition. Reference and repeat passes reached the same rejection.
CANDIDATE POLICY
RejectedThe incumbent was retained.
DEMO WORKFLOW
CompletedMeasurements and decision were revalidated.
The same 9 of 33 rules failed in each pass.
Three examples of unmet conditions
Local verification: 9 September 2026| Condition | Required | Candidate | Outcome |
|---|---|---|---|
| Settled arrival · proxy | ≥ 20% | 0% | Not met |
| Final target error · held-out | ≤ 1.000 m | 1.006 m | Not met |
| Mean reward · proxy | ≥ 21.165 (baseline) | 20.678 | Not met |
Rounded for readability. Decisions used the unrounded measurements.
Two policies × two environments × three seeds × five episodes × two passes. 24 fresh seed processes, three distinct seeds (210/211/212). These are not 120 independent statistical samples.
03 / CURRENT BOUNDARIES
What it demonstrates. What it does not.
Working today
- Training loop, run tracking and reporting
- Multi-seed, dual-environment policy comparison
- Repeatable rejection under fixed rules
- A browser-based demo; no installation or download required
Explicit limits
- The product is local and in development, not a ready-made SaaS.
- This demo shows no new training or qualified drone policy.
- The warehouse has been used before; this is not unseen-environment generalization proof.
- Settled arrival is a same-sample distance/speed condition, not sustained hover or real-world safety.
This introduction does not distribute Training Factory application source, model weights, NVIDIA assets or private run archives.