Lab 2 · The Twin

Where training happens

In Hello Racer you drove the car. In First Training, the next lab, a computer learns to drive it. Not the real car: a car crashes, a battery runs out, and learning takes millions of tries. So training happens on a digital twin, a copy of the car that lives in a simulator.

Two NVIDIA tools do this:

  • Isaac Sim is the simulator. A 3D world with physics: gravity, friction, wheels that roll and slip, a motor that pushes. Think of a video game engine built for robots.
  • Isaac Lab sits on top of Isaac Sim and is built for learning. It can run thousands of copies of a robot at once, on one GPU, and it talks to the learning code. Every lab in the Train band is an Isaac Lab "task".
thousands of copies of the twin on one plane, each with its own waypoints
Isaac Lab runs thousands of copies of the twin on one GPU. Every copy is one car with its own course. That is why learning is fast.

The twin is simplified, on purpose

The twin is a file: the Fiesta chassis, four wheels, four struts, a motor, a steering joint, with the mass and the sizes of the real car. The physics engine moves it 60 times a second.

It is a simplified twin. The tires do not heat up. The ground is flat. The battery never sags. The radio has no lag. That is on purpose: a twin is as simple as it can be and still teach something that holds on the real car. Everything it leaves out is the sim-to-real gap, and later labs are about that gap.

The simulator on this page is simpler still: the twin's steering and motor chain on a bicycle model, in your browser. Good enough to see what the learner sees.

Drive it, see what the learner sees

loading the simulator
obs [ 0.00, 0.00, 0.00 ] racer.drive(steer=+0.00, throttle=+0.00) waypoints 0/6

// obs = [ distance, cos(heading error), sin(heading error) ]. Turn the car away from the waypoint and watch the sin go to +1 or -1. Drive to it and watch the distance go to 0. Those three numbers are all the policy gets.

In and out: the shape

The thing that learns is called a policy. It is a small function. Numbers go in, numbers come out, every tick.

  • In: the observation. What the policy is allowed to know. For the first task, GoatRacer-Fiesta-3obs-v1, that is three numbers: the distance to the waypoint, and the cos and sin of the heading error (the angle between where the car points and where the waypoint is). You saw them in the readout above.
  • Out: the action. Steer and throttle, each in [-1, 1]. The same two numbers your keys sent in Hello Racer.

The policy does not see the field, the car, or the waypoint. It sees three numbers. That is a design choice, and the Train labs show what it costs and what it buys.

The shape of a task is frozen. A policy trained on three numbers can only ever be fed three numbers. New observations mean a new task version and a new policy. When a later lab says "17 observations", that is a different task, not this one grown.

The network

The policy is a neural network: layers of numbers, each layer a weighted sum of the one before, with a bend (the activation) between them. Training changes the weights. Here is the one you train in First Training, from the recipe the GPU runs:

obs … 32 dense 32 · ELU … 32 dense 32 · ELU action 2 1,250 parameters

// 3 obs -> dense 32 -> ELU -> dense 32 -> ELU -> dense 2 · 1,250 params - the sizes come from sim/isaaclab/GoatRacer/GoatRacer/tasks/agents.py: actor 32, 32, critic 32, 32, elu.

Left to right: the three observations go in, each dense layer multiplies by its weights and adds a bias, the bend between layers is the activation, and two numbers come out: steer and throttle.

There are two networks like this. The actor is the policy: it drives. The critic is its coach: it guesses how good the current state is, and training uses that guess to tell the actor what to change. Only the actor goes to the car.

Count the actor's weights: 3 x 32, 32 x 32, 32 x 2, plus a bias per unit. About 1,250 numbers. That is the whole driver. It fits in a text file, runs in your browser, and runs on the car's Orin a thousand times a second.

// The Code card shows the lines of the env that build the observation. Not a copy: the page reads the file the GPU runs.

Follow the White Rabbit

  • Isaac Lab - the framework. The "Direct workflow" pages are the shape of our env files.
  • OpenUSD - the file format the twin is written in. A scene is a tree of prims.
  • 3Blue1Brown: neural networks - the clearest pictures of layers, weights and training there are.

Code

// sim/isaaclab/GoatRacer/GoatRacer/tasks/goatracer_fiesta_3obs_env.py - the exact lines the GPU runs to build the observation. policy is the (envs, 3) tensor: one row per copy of the twin.

    def _get_observations(self) -> dict:
        target = self._target_positions[self._all_ids, self._target_index]
        err_vec = target - self.robot.data.root_pos_w[:, :2]
        self._previous_position_error = self._position_error.clone()
        self._position_error = err_vec.norm(dim=-1)
        yaw = self.robot.data.heading_w
        t_yaw = torch.atan2(err_vec[:, 1], err_vec[:, 0])
        self._target_heading_error = torch.atan2(
            torch.sin(t_yaw - yaw), torch.cos(t_yaw - yaw))
        cols = [
            self._position_error,
            torch.cos(self._target_heading_error), torch.sin(self._target_heading_error),
        ]
        if self.cfg.observation_space != 3:                           # -v0: the zero padding
            z = torch.zeros_like(self._position_error)
            cols += [z, z, z, z, z]
        policy = torch.stack(cols, dim=-1)                            # (E,3) -v1 / (E,8) -v0
        out = {"policy": policy}
        self._nan_envs = torch.isnan(policy).any(dim=-1) | torch.isnan(
            self.robot.data.root_pos_w).any(dim=-1)

Read it top to bottom: the vector to the waypoint, its length (the distance), the angle to it minus the car's own heading (the heading error), then cos and sin of that. Three columns, stacked. That is the whole input.