TRAINING-OPS / ROBOT LEARNING

Delegate training.
Ground the decision.

A Training-Ops agent for long-running robot-learning work, from simulation to an evidence-backed evaluation decision.

Starts training, manages goal-driven iterations and compares policies. Leaves a traceable run record behind each decision.

WAREHOUSE / ILLUSTRATIVE REPLAY21 SECONDS

The same starting point. Two policies. A separately recorded, illustrative simulation replay.

This is not footage from, or additional qualification evidence for, the original 120 episodes.

GOLDEN DEMO / V1Rejected
120evaluation episodes
2evaluation environments
4/4repeat checks passed
0automatic policy changes

01 / THE DELEGATED WORK

Not just a run. A traceable loop.

Isaac Lab handles physics and RL training. Training Factory manages the goal, progress and evaluation conditions of the work around it.

  1. 01

    Define

    Set the job, goal, budget and evaluation conditions.

  2. 02

    Run

    Launch simulation and training processes and track their progress.

  3. 03

    Evaluate

    Compare policies under fixed conditions across environments and seeds.

  4. 04

    Decide

    Advance the goal loop and report policy qualification with its evidence.

The goal and iteration budget bound the loop. Qualification can also run without training. The incumbent policy is never replaced automatically.

02 / RECORDED DEMO RESULT

One better metric is not enough.

Some measurements improved in the warehouse demo. But the candidate did not meet every fixed acceptance condition. Reference and repeat passes reached the same rejection.

CANDIDATE POLICY

Rejected

The incumbent was retained.

DEMO WORKFLOW

Completed

Measurements and decision were revalidated.

The same 9 of 33 rules failed in each pass.

Three examples of unmet conditions

Local verification: 9 September 2026
ConditionRequiredCandidateOutcome
Settled arrival · proxy≥ 20%0%Not met
Final target error · held-out≤ 1.000 m1.006 mNot met
Mean reward · proxy≥ 21.165 (baseline)20.678Not met

Rounded for readability. Decisions used the unrounded measurements.

Two policies × two environments × three seeds × five episodes × two passes. 24 fresh seed processes, three distinct seeds (210/211/212). These are not 120 independent statistical samples.

03 / CURRENT BOUNDARIES

What it demonstrates. What it does not.

Working today

  • Training loop, run tracking and reporting
  • Multi-seed, dual-environment policy comparison
  • Repeatable rejection under fixed rules
  • A browser-based demo; no installation or download required

Explicit limits

  • The product is local and in development, not a ready-made SaaS.
  • This demo shows no new training or qualified drone policy.
  • The warehouse has been used before; this is not unseen-environment generalization proof.
  • Settled arrival is a same-sample distance/speed condition, not sustained hover or real-world safety.

This introduction does not distribute Training Factory application source, model weights, NVIDIA assets or private run archives.