Part XI: Tactile, Embodied, and Robotic Sensing
Chapter 58  [R]

Vision-Language-Action and Embodied Foundation Models

Placeholder chapter page. Sections are produced by the book-skills pipeline.

Sections

  1. 58.1 From perception to action: the VLA idea
  2. 58.2 RT-1/RT-2 and web-scale knowledge transfer
  3. 58.3 Open generalist policies (Octo, OpenVLA)
  4. 58.4 Flow-matching and open-world VLAs (π0/π0.5)
  5. 58.5 Humanoid and general robot foundation models (GR00T, Gemini Robotics)
  6. 58.6 Sim-to-real and Open X-Embodiment data
  7. 58.7 Evaluation, safety, and generalization

Lab 58

fine-tune or evaluate an open VLA policy on a manipulation benchmark subset (LIBERO/CALVIN).