Google AgentHands gives AI agents spatial hand gestures in XR
Google Research has introduced AgentHands, an LLM-powered XR prototype that gives AI assistants synchronized virtual hands able to point, demonstrate actions and interact with objects in physical space.

AI assistants are beginning to move beyond the screen
Conversational AI has become increasingly capable of understanding images, video and physical environments. But explaining a spatial task through speech still presents a fundamental limitation.
An assistant can tell someone to turn a knob, inspect a specific component or move an object in a particular direction. The user must then translate those verbal instructions into actions in physical space. Google Research describes this as a “mental mapping gap”.
On August 25, Google Research presented AgentHands, an experimental XR system designed to reduce that gap by giving an AI agent something humans naturally use when explaining physical tasks: hands. Instead of only saying what to do, the AI can point to an object, trace a shape, demonstrate a movement or accompany a warning with a physical gesture.
From conversational AI to embodied AI
AgentHands is an LLM-powered research prototype developed for spatially grounded conversations in XR and published at CHI 2026. The system generates virtual hands whose movements are synchronized with the agent's speech and positioned relative to objects in the user's environment.
The concept builds on something fundamental in human communication. When people explain physical tasks, they rarely rely exclusively on words: they point, indicate dimensions, reproduce movements and use gestures to emphasize warnings or instructions. AgentHands attempts to introduce that same communication layer into an AI assistant.
How AgentHands understands physical space
Before the agent can gesture toward something, it needs to understand where that object exists. AgentHands combines eye gaze and scene reconstruction to register physical objects within a three-dimensional spatial map. A user can look at an object such as a plant or a laptop, allowing the system to create a corresponding 3D bounding box that the AI can reference when generating its response.
The process effectively connects three elements: language, spatial understanding and physical gesture. When the user asks a question, the LLM generates its verbal response together with embedded GestureEvents that describe which gesture should occur and when.
A local parser running on the XR device then synchronizes text-to-speech with the animation system using word-level timestamps. As the AI speaks, its virtual hands perform the corresponding movements in the user's physical environment.
The AI can point, demonstrate and warn
Google organized the gesture library around three semantic categories.
- Deictic gestures reference something in the environment, such as pointing toward a particular component
- Iconic gestures represent an action, dimension or shape — for example, demonstrating how to rotate a control
- Expressive gestures communicate social or emotional information
The system can also combine movement with XR visual effects. A warning about a hot surface, for example, can include a hand gesture together with a visual glow highlighting the danger. That turns the AI response from a purely linguistic output into a spatial performance.
Google tested AgentHands on physical tasks
The research team evaluated the system with 12 participants, comparing AgentHands against a speech-only AI assistant. Participants performed two procedural activities: caring for an orchid, including identifying plant components and performing watering and fertilizer tasks; and operating a 3D printer, identifying hardware components and following steps involving the printer controls.
The study found statistically significant improvements in several areas. Participants found it easier to identify locations and directions when the AI used spatial gestures, complex actions were easier to follow, and gestural and visual warnings made safety information more noticeable. Google additionally reports reduced cognitive load when instructions combined speech with spatial movement.
The sample is small and AgentHands remains a research prototype, so these results should not be interpreted as evidence of production-scale effectiveness. They do, however, provide an early indication of how embodied AI interfaces could change interaction inside XR environments.
Spatial computing changes what an AI interface can be
Most current AI assistants remain fundamentally screen-based. Even multimodal systems capable of interpreting a camera feed generally communicate through text, voice or graphical overlays. XR introduces another possibility: because an XR system understands the geometry of the environment, an AI agent can communicate inside that geometry.
Instead of saying “turn the knob on the left clockwise”, the interface can point directly toward the correct knob and demonstrate the rotation. Instead of “be careful with the hot nozzle”, it can place a warning exactly where the hazard exists. That difference becomes particularly relevant for procedural environments such as training, maintenance, manufacturing, healthcare, education and technical support.
Android XR could become a platform for embodied AI
AgentHands is not currently announced as a commercial Android XR feature; Google explicitly describes it as research. However, the company states that it is continuing to explore the concept within the Android XR ecosystem, including personalized gestures based on dominant hand and user-specific spatial routines.
That direction is significant as Android XR expands across headsets and spatial glasses. AI agents operating on these devices may eventually have access not only to cameras and microphones but also to gaze, hand tracking, scene reconstruction and persistent spatial information.
"The interface between a user and an AI assistant could become less about opening an application and more about the agent understanding and communicating directly within the surrounding environment."
Sources
- Google Research — AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR (August 25, 2026): technical description, demos, architecture, user study and results
- Google Research / ACM CHI 2026 — AgentHands research paper: gesture taxonomy, system architecture and experimental evaluation
Editorial coverage of industry news and developments across spatial computing and immersive technologies.
Related reading

XREAL AURA passes 10,000 reservations before launch: Android XR gains traction in spatial computing
XREAL AURA has passed 10,000 global reservations before reaching the market. The spatial computing glasses pair Android XR, Google Gemini, a 70° field of view and a separate compute puck, signalling early demand for lighter XR devices.

Valve moves Steam Frame closer to launch: SteamOS now ships support for its wireless adapter
Valve added preliminary support for the Steam Frame Wireless Adapter in SteamOS 3.8.25 Beta. Together with the headset's recent FCC authorization and a growing "Great on Frame" catalog, the signals point to a launch drawing closer.

VITURE Pro 2 debuts at $299: XR glasses cut weight and price while improving image quality
VITURE launched its Pro 2 XR Glasses at $299.99. At 63 grams, with 1080p Sony Micro-OLED panels, 120 Hz and up to 1,600 nits, the device reflects a wider XR trend: bringing spatial displays into lighter, more affordable formats.