AI Labs · Spatial intelligence · Benchmark

High-End Spatial And Human Intelligence For Physical AI

VLGE transforms how 3D worlds are built, played, and explored into synchronized environments, telemetry, event streams, physics-grounded scenarios, and behavioral micro-signals, all anchored in real human interaction. Together, they form the human behavioral data layer for AI systems learning to perceive, reason, and operate in 3D space.

world / …
samples / 0

The benchmark

Labs Build Models. We Own The Benchmark

A benchmark is the reference set a lab measures its own agent and world models against. Ours is a real person's full trajectory through a captured space, route, dwell, interaction, hesitation and attention, and the human-grounded agents that reproduce that behavior at scale. The benchmark itself is published free, the way DROID, LIBERO and Open X-Embodiment are. Scoring against it requires the human reference data, which is ours.

Lab Xtheir agent · their world model
runsAgent A inside their own world model.
resultThe agent's simulated actions in the simulated world.
VLGEhuman-grounded agent · same world
runsThe same task, reproduced in the same world by an agent grounded in real human sessions.
resultHuman-believable actions, carrying the real behavioral distribution.
Benchmark score The gap

How far the lab's agent sits from the real human distribution, per task, per world, per metric.

Stage 1

Human Ground Truth

Real people build and play inside the world while every channel records: position, kinematics, hand and finger articulation, grasp episodes, gaze, dwell, hesitation and decision points. This is the reference distribution, not a scripted demonstration.

Stage 2 · 3

Human-Grounded Agents

LLM-based agents are conditioned on that observed behavior rather than on hand-written policies, then run the same worlds and tasks. Because they inherit the human distribution, they generate interactions, decisions and trajectories at a scale human capture alone cannot reach, without drifting from the people they model.

Score

Quantitative, Per Task

A lab's agent is scored against both: the human reference and the grounded agents reproducing it. Every metric is a number, and every number is traceable back to the synchronized streams it was computed from.

What gets scored

First public release in preparation. Reference worlds, tasks and the metric definitions above are fixed, baselines are being measured now. Request early access →

The session record

What A Session Looks Like

Every session is both a navigation trace and a manipulation demonstration. It becomes a continuous event feed of pose, velocity, state changes and dwell, time-aligned with source timestamps, carrying the layers below: per-finger articulation, the full arm and shoulder kinematic chain, complete grasp episodes with their failures, labelled interaction scenarios, physics-driven motion, a LiDAR channel, and a per-frame behavioral layer. Exports preserve the raw data material and can be adapted to the buyer's training purpose. Below are four real captures, replayed and interleaved as they're stored.

4.1

Hand & Finger Tracking

Per-finger flexion and abduction every frame, three joints per finger, fifteen per hand, at the same frequency as full-body movement. Fingertip position and velocity through grasp and release, and confirmed contact point, normal, distance and target, recorded in the object's own frame so a contact site is reusable across sessions as a grasp affordance.

Closes the gap between full-body motion capture and the fine motor control object handling needs.

4.2

Joint, Arm & Reach Kinematics

Shoulder, elbow and wrist angles and angular velocity; requested versus resolved wrist pose with the overrun in metres when a target sits beyond comfortable range; and the whole-body compensation, torso lowering, forward lean, shoulder assist, when the arm alone cannot reach. Local rotations retarget cleanly to robot arms with different link lengths.

Separates a reachable target from a strained one, for workspace design and reach planning.

4.3

Grasp & Manipulation Telemetry

Every grasp is a complete episode, not an event flag: grip taxonomy (power, pinch, palm) with the pre-contact hand pose, object mass, dimensions, compliance, break force and torque, closure progress, grip-volume occupancy, opposing-finger contact count, object pose in the hand frame every frame, and release linear and angular velocity. Failed attempts get their own record.

The engine's per-frame validation fields are a ready-made grasp-success classifier, with no annotation, and the failures give a policy its negative examples.

4.4

Interaction Scenarios & Events

A labelled library of pick-up, place, push, pull, open and close. Each scenario pairs a pre- and post-interaction snapshot to isolate the effect on the object; push and pull carry applied force direction, displacement and object velocity; open and close carry articulated state such as hinge angle or slide position. Every discrete interaction is also a semantic event with sub-second precision.

Trainable manipulation and object-affordance data straight out of the log, without annotation.

4.5

Physics-Based Motion

When the engine takes over a participant's movement, the full physical event is recorded at locomotion fidelity: fall height, impact velocity, contact point and angle, resulting body trajectory, object collision response, and the frame-by-frame recovery path back to controlled movement including residual instability.

A continuous, physically accurate record of body dynamics under uncontrolled conditions.

4.6

Spatial Sensor Data (LiDAR)

A dedicated channel from virtual sensor rigs mounted on participants, with the full sensor configuration and a raw scan payload up to 360°. Each frame adds per-ray 3D hit point, surface normal, distance, material tag, reflectivity and object identity, plus aggregated semantic and material counts.

Every frame is at once a geometric measurement, a semantic scene description and a material classification, at a scale that would be prohibitive to acquire physically.

4.7

Behavioral Layer (Edit & Play)

The BrainRecord pipeline adds 48 per-frame features and session-level analytics across world creation and world interaction: velocity decomposition, turning, acceleration and curvature; frame-level action labels; gaze, attention and spatial coverage; and hesitation, decision points, micro-corrections, commitment and control effort.

Outputs are analytics-, visualization- and training-ready, including engagement, hesitation, novelty and goal-directed-versus-wandering heatmaps.

Full field-by-field dictionaries available under MNDA.

The corpus

Every State, In One View

Each point is one timestamped state. Depicted are 61,736 points across 30 sessions, a browser-sized sample of a much larger captured corpus. Enough to read the structure, synchronization and signal quality for yourself. Drag to orbit; hover any point to read its raw JSON.

corpus / …

Process layer · decision-making

The Moment Of Decision

Where a person hesitated, showed intent, weighed it, and committed to a direction, read straight from the deliberation, hesitation and movement channels of one real session. This is the cognitive layer that no robotics or driving dataset captures.

Process layer · interaction

Not Just Where They Go, What They Do

Movement is half the story. The other half is the build itself: Every time the user places, moves, rotates, scales, or lights an object, we capture it. Below is one real Edit-Mode session replayed from its event log. Scrub the timeline, or switch builders. Each tick is a real recorded action.

Two People, One Space,
Two Completely Different Paths.

Most datasets capture what a space is. VLGE captures how it's used: one person wanders and lingers while the next walks straight through. The same world produces opposite behavior, and every number here is read straight from the two real captures below.
behavioral ground tracks ● 2 sessions · real capture
session Asession B straightness dwell pauses ≥0.5s radius of gyration

What data is available

One Session. Every Data Stream, Aligned

The console plays one real captured session end to end, with every panel reading the same playhead. As you scroll the streams, it reiterates that session through each one.

01 / 07

3D Environments

Production-grade interactive 3D scenes, from interiors and streets to towers and stores, authored to be walked and played inside online.

watertight geometry · materials · navmesh · floor plans
02 / 07

3DGS

Gaussian-splat captures of worlds built and recorded from reality: photoreal radiance fields for training perception and novel-view synthesis at the fidelity the physical world actually has, with derived perception labels for registered geometry.

splats · radiance fields · multi-view · intrinsics · depth · masks · 2D/3D boxes
03 / 07

Spatial Metadata

The structured layer beneath the geometry: labelled zones, object semantics, adjacencies, sightlines and coordinates, the grammar a model needs to reason about a space. Buyer exports define metrics in fixed time windows so fingerprints compare cleanly across sessions.

coordinates · semantics · zones · sightlines · time-window metrics · fingerprint
04 / 07

Continuous Event Streams

Timestamped behavior at inferred source cadence, with buyer-cadence resampled exports: pose, gaze, velocity, dwell, build and interaction events. This is the record that turns a static scene into how it's actually used, and it's the JSON streaming on the right.

pose · gaze · dwell · interactions · build-log · JSON · resampled exports
05 / 07

Physics-Applied Interactions

Realistic physics drive every interaction, enabling object manipulation, accident scenarios, creative puzzles, and emergent behaviors. Every action is paired with full-body motion capture, from body pose and hand articulation down to individual finger movements, creating a complete behavioral record stored alongside the event stream.

physics · object manipulation · full-body motion · hand tracking · finger tracking · collision events · interaction log · JSON
06 / 07

Benchmark Datasets

The same streams, packaged as a scoring set: a fixed reference world and task, one real human trajectory as ground truth, the human-grounded agent runs that reproduce it, and the metric definitions above. A lab runs its own agent on the world and gets a number per metric. How the benchmark works ↑

reference trajectory · grounded agent runs · metric definitions · per-task scores
07 / 07

Custom Data Collection & Delivery

Need a specific environment, task, population or modality? We build the world, run the collection, and deliver the dataset to your spec, ready to scrub, sample and export.

bespoke worlds · task instructions · audio/transcripts · co-present agents · scheduled delivery

Delivery package

Built For Technical Diligence, Not Just Demos

The live instruments prove synchronization; the benchmark proves performance. VLGE turns the same human-anchored streams into benchmark-ready datasets with structured exports, enrichment layers, and standardized splits, supporting quantitative evaluation through benchmark metrics and qualitative evaluation through synchronized behavioral, event, telemetry, and 3D streams.

Open any component for what it contains, the fields it ships, and where to see it running on this page.

01

Benchmark

Quantitative scoring against the human reference. A per-task number for how far an agent's behavior sits from the real human distribution.

Scored on the attributes the data actually carries, which is what makes it behavioral rather than task-completion. Route and dwell say where the agent went, decision and attention say how it chose, manipulation and reach say how it acted. Each is a distribution comparison against real human sessions in the same world on the same task, so runs are comparable across labs.

path_length · straightness_index · radius_of_gyration · revisit_rate · dwell_count · dwell_duration · spatial_entropy · hesitation_events · decision_points · micro_corrections · commitment_latency · gaze_targets · attention_shifts · coverage · grasp_success_rate · grip_taxonomy_match · object_pose_in_hand · release_dynamics · joint_angles · ik_residual · reach_overrun
See how the benchmark runs ↑
02

Physics-Based Worlds

Objects react through collision, gravity, constraints and force, never scripted animation. What the data records is a physical consequence, not a keyframe.

The same action produces different outcomes under different conditions, and both are captured at full fidelity. When the engine takes over a participant's body, the whole event is recorded: fall height, impact velocity, contact point and angle, the resulting trajectory, and the frame-by-frame recovery back to controlled movement.

mass · dimensions · friction · restitution · joint_compliance · break_force · break_torque · damping · collision_mode · impact_force · contact_point · contact_normal · fall_height · impact_velocity · recovery_frames
See the physics-based motion layer ↑
03

Structured Data

Source cadence runs roughly 90 to 280 Hz depending on session and channel, resampled on export to whatever fixed rate you train at. One clock, one coordinate convention, field dictionary included.

Pose, video, events and sensor frames arrive already aligned, so there is no reconciliation pass on your side. We match your axis convention, unit system and frame rate rather than asking you to match ours, and every export carries its own source cadence per channel so nothing is inferred.

session_id · frame_idx · t_source · t_aligned · source_hz · resampled_hz · pos[x,y,z] · rot[quat] · frame_convention · units · checksum
See the aligned streams on one playhead ↑
04

Perception Exports

What a camera would have seen. Calibrated views with labels that are exact rather than annotated, pixel-aligned to known 3D truth.

Because the geometry is authored rather than reconstructed, depth is the true depth buffer, masks are true object identity per pixel, and boxes are projections of actual object bounds. Intrinsics and extrinsics ship per frame, so views are reprojectable for novel-view synthesis, depth, segmentation and detection from one capture.

rgb · depth_m · camera_intrinsics(K) · extrinsics(R,t) · semantic_mask · instance_mask · bbox_2d · bbox_3d · surface_normals · splat/radiance-field source
See the 3DGS and environment streams ↑
05

Behavior Enrichment

The BrainRecord layer: 48 per-frame features plus session analytics, across both world creation and world interaction.

This is what separates a trajectory from a decision. Velocity, turning, acceleration and curvature are decomposed per frame; hesitation, decision points, micro-corrections, commitment and control effort are measured; every frame carries an action label. Outputs are heatmap-ready, including engagement hotspots, hesitation zones and goal-directed versus wandering behavior.

velocity_decomp · turn_rate · curvature · action_label · gaze_target · attention_shift · coverage · hesitation · decision_point · micro_correction · commitment · control_effort · surprise · novelty · info_value · learning_potential
See a real decision moment read from these channels ↑
06

VLA Task Layer

Instruction, execution and outcome for vision-language-action training. What the person was asked to do, what they did, and whether it worked.

A session run against a written instruction has intent attached to it rather than being untargeted exploration. Audio and transcripts sit alongside the action stream, and each attempt carries an outcome label including the repair, the second try after a first attempt fails, which is the part most instruction-following datasets are missing.

instruction_text · task_id · transcript · audio · sub_goal · attempt_idx · outcome (success · fail · repair) · failure_reason · t_start · t_end
See how a session is recorded ↑
07

Interaction Histories

Every manipulation, movement and environmental change as a replayable event stream, including the attempts that failed.

A grasp is stored as a complete episode, not a flag: the grip formed before contact, the object's pose in the hand frame every frame, closure progress and opposing-finger contact, and the velocity at release that separates a placement from a drop. Scenarios pair pre- and post-interaction snapshots, and failures get their own record, the negative example a recovery policy needs.

grip_taxonomy (power · pinch · palm) · closure_progress · grip_volume_occupancy · opposing_contacts · contact_point_object_frame · object_pose_in_hand · release_linear_vel · release_angular_vel · scenario (pick · place · push · pull · open · close) · pre_state · post_state · hinge_angle · applied_force · failure_record
See a real edit-mode event log replayed ↑
08

Human Motion Fidelity

Full-body pose, the complete arm and shoulder chain, and per-finger articulation, all at one frequency.

Conventional motion capture gives you the body and loses the hand. We record flexion and abduction for each digit, fifteen joints per hand, at body-movement rate, with fingertip position and velocity through grasp and release. The arm chain adds the IK residual between requested and resolved wrist pose, and the whole-body compensation when the arm alone cannot reach. Local rotations retarget cleanly to robot arms with different link lengths.

body_pose · joint_angles · angular_velocity · finger_flexion · finger_abduction · fingertip_pos · fingertip_vel · wrist_pose_requested · wrist_pose_resolved · ik_residual_m · torso_lower · forward_lean · shoulder_assist
See the hand, finger and reach layers ↑
09

Spatial Sensor Data

A dedicated LiDAR channel from virtual rigs mounted on participants, with full sensor configuration and a raw scan payload up to 360°.

Beyond the point cloud, every frame carries per-ray hit point, surface normal, distance, material tag, reflectivity and object identity, plus aggregated semantic and material counts. Each frame is a geometric measurement, a scene description and a material classification at once, at a scale that would be prohibitive to acquire physically.

sensor_config · scan_fov · rays_per_frame · hit_point_xyz · surface_normal · distance_m · material_tag · reflectivity · object_id · semantic_counts · material_counts
See the LiDAR layer ↑
10

Multi-Agent Capture

Several real people in one world on one clock, for social navigation, collision avoidance and crowd modeling.

Co-present sessions are one synchronized capture, not separate traces stitched afterwards, so inter-personal distance, yielding, path crossing and mutual gaze are measurable frame by frame. Humans and grounded agents can share the same world and are recorded identically, so a mixed population is available where pure human capture would not scale.

participant_id · shared_clock · inter_agent_distance · yield_event · path_crossing · mutual_gaze · personal_space_violation · co_presence_window
See two people take opposite paths through one world ↑

Example applications

Built For The Systems Learning To Operate In 3D

01
build logs · geometry

World Models

Ground generative world models in how real spaces are built and changed.

02
human trajectories

Simulation

Populate simulators with real human trajectories, not scripted agents.

03
timestamped pose · navmesh

Robotics

Navigation & manipulation priors from human behavior in 3D.

04
9 behavioral channels

Embodied AI

Agents that learn to perceive, plan and act in physical space.

05
spatial fingerprints · coverage

Spatial Intelligence

How people read, cover and reason about space, measured per session.

In depth

The Four Layers, And Where To Read Each One

01 / Worlds

Explorable 3D and 3DGS environments of any real space, a museum, a street, a store, a tower, a game, authored to be walked, built and played inside online. Watertight geometry, materials, navmesh and floor plans, or a photoreal Gaussian-splat capture of the place itself.

The environment and 3DGS streams ↑
02 / Data

Every session becomes a synchronized record on one clock: pose and kinematics, hand and finger articulation, grasp episodes, semantic events, physics, LiDAR, and a per-frame behavioral layer covering dwell, hesitation, attention and decision points.

What a session looks like ↑
03 / Agents

LLM-based agents that create a world and play inside it, perceiving, reasoning, acting and changing the world as they go. Conditioned on observed human behavior rather than hand-written policies, they scale human-believable data far past what human capture alone can reach.

How grounded agents run ↑
04 / Benchmark

Agents grounded in real human ground truth close the sim-to-real gap for embodied AI. The benchmark built on that data is free for any lab to run, and scores how far their agent sits from the real human distribution, per task, per world, per metric.

The benchmark framework ↑

Benchmark Today. Build For Tomorrow. A Growing Library Of Spatial Datasets Designed For Rigorous Evaluation And Future Research

Request data access Talk to the data team →