Skip to content
nanoai

Stanford lab lets GPT Astra run a humanoid through five skills

HomeBody gives a frontier chat model a Unitree G1 body, a digital twin of the room and a short list of motor skills, with no robot action model trained in between. It has not published success rates.

By Siva Charakani3 min read

Stanford AI robotics research graphic showing GPT Astra controlling a humanoid robot through five household tasks, including picking up objects, moving, opening a cabinet, placing objects, and cleaning, with a callout noting that success rates were not published
Image: AI Generated

Key takeaways

  • HomeBody lets GPT Astra direct a Unitree G1 humanoid through five reusable motor skills.
  • The robot explores first and uses a digital twin in Isaac Sim as its memory of the room.
  • The project page shows demos but no success rates; the code is listed as coming soon.

Researchers at Stanford and Caltech have shown a Unitree G1 humanoid doing kitchen tasks under the direction of OpenAI’s GPT Astra, with no robot-specific action model trained in between. The project, called HomeBody, gives the frontier model a body through a stored model of the room and a small set of motor skills it can call by name.

The claim is about architecture, and the evidence so far is a set of demonstrations. The project page reports no success rates, and the code repository says the code is coming soon.

The layer HomeBody removes

Most humanoid systems stack three layers, as the HomeBody page describes them. A large vision-language model does the high-level reasoning. A learned vision-language-action model, or VLA, turns that into commands; it is trained on robot data to map images and instructions to movement. A low-level controller then executes the motion. Collecting enough of that data is slow and costly.

HomeBody asks whether the reasoning model can skip that middle layer and “directly orchestrate a library of reusable motor skills”. Its vocabulary is small: navigation, picking, placing, drawer opening and picking from drawers. GPT Astra picks each skill and its target from what the robot currently sees, the map, the state of its gripper, recalled observations and the result of the previous step.

Memory before action

Before it acts, the robot explores. Astra then works as what the team calls a Real2Sim agent, building a digital twin of the space in Isaac Sim, Nvidia’s robot simulator, from the humanoid’s own data: camera observations, measured SLAM geometry (a map built as the robot tracks its own position) and other signals. That twin serves as the robot’s spatial memory, so it can act on things it cannot currently see.

The tasks on the project page are ordinary on purpose: “Clean up all of the coffee bags and put them in the middle”, throw away milk and orange juice cartons that have gone bad, and fetch a medicine container from a drawer after a person says “I forgot my medicine, can you get it for me?”. The team says the kitchen was one the system had not seen before.

What the demos do not tell you

The authors are candid about limits. Real2Sim reconstruction adds setup time and API costs. Astra’s reasoning latency introduces pauses between skills. Long sessions run into finger-servo overheating, and heavier perception models may need more compute. The local stack runs on a single Razer Blade laptop with an RTX 4090 GPU, while Astra runs remotely over the network.

That last point deserves more attention than it will get. The robot’s decision-maker is a metered API call. OpenAI lists GPT-6 Astra at $10 per million input tokens and $50 per million output tokens, taking text and images in and returning text. A household robot built this way inherits a vendor’s pricing, latency and availability. Without success rates, you also cannot compare HomeBody with VLA-based systems on how often each finishes a task.

Why a short skill list may be the point

Five skills sounds like a limitation, and for general-purpose robots it is: every new kind of task needs a new skill written by people, which is the scaling problem VLAs are meant to solve. But a short list is also a boundary. On September 26, Axios reported that OpenAI, Anthropic and researchers are investigating tens of thousands of incidents in which frontier models bypassed guardrails or escaped sandboxes. A model that can only navigate, pick, place and open drawers has far fewer ways to surprise its owners than an agent with open-ended tools. That is our reading, not a claim the HomeBody team makes.

What to watch

A paper with task success rates, the code release, and whether the team swaps in other frontier models to see how much of the result depends on Astra. The open question is whether skill libraries can grow fast enough to matter outside a single kitchen.

  • humanoid robots
  • robotics
  • embodied AI
  • Stanford
  • GPT-6 Astra

Sources

  1. HomeBody: A Humanoid That Explores, Remembers, and Acts on Its Own — Stanford University (The Movement Lab)
  2. GitHub - Stanford-TML/homebody — GitHub (Stanford-TML)
  3. GPT-6 Astra Model | OpenAI API — OpenAI
  4. Scoop: Top AI companies probing tens of thousands of security incidents — Axios via Yahoo Tech, Sep 26, 2026

Follow The Nano AI: Instagram · X · LinkedIn · YouTube

Was this article helpful?

Comments

No comments yet. Start the conversation.

Be respectful. Comments are moderated.

Related stories

The AI briefing, without the noise.

The stories that matter in AI, sourced and explained. Free, and you can unsubscribe at any time.