3D Environments
Production-grade interactive 3D scenes, from interiors and streets to towers and stores, authored to be walked and played inside online.
AI Labs · Spatial intelligence
VLGE turns how worlds are built, played, and explored into synchronized 3D environments, telemetry, event streams, physics-applied scenarios and behavioral micro-signals: the human behavioral data layer for systems learning to operate in a 3D space.
Every session people build and play becomes a continuous event feed of pose, velocity, state changes and dwell, time-aligned with source timestamps. Exports preserve the raw data material and can be adapted to the buyer's training purpose. Below are four real captures, replayed and interleaved as they're stored.
The corpus
Each point is one timestamped state. Depicted are 61,736 points across 30 sessions, a browser-sized sample of a much larger captured corpus. Enough to read the structure, synchronization and signal quality for yourself. Drag to orbit; hover any point to read its raw JSON.
Process layer · decision-making
Where a person hesitated, showed intent, weighed it, and committed to a direction, read straight from the deliberation, hesitation and movement channels of one real session. This is the cognitive layer that no robotics or driving dataset captures.
Process layer · interaction
Movement is half the story. The other half is the build itself: Every time the user places, moves, rotates, scales, or lights an object, we capture it. Below is one real Edit-Mode session replayed from its event log. Scrub the timeline, or switch builders. Each tick is a real recorded action.
Two people, one space,
two completely different paths.
What data is available
The console plays one real captured session end to end, with every panel reading the same playhead. As you scroll the streams, it reiterates that session through each one.
Production-grade interactive 3D scenes, from interiors and streets to towers and stores, authored to be walked and played inside online.
Gaussian-splat captures of worlds built and recorded from reality: photoreal radiance fields for training perception and novel-view synthesis at the fidelity the physical world actually has, with derived perception labels for registered geometry.
The structured layer beneath the geometry: labelled zones, object semantics, adjacencies, sightlines and coordinates, the grammar a model needs to reason about a space. Buyer exports define metrics in fixed time windows so fingerprints compare cleanly across sessions.
Timestamped behavior at inferred source cadence, with buyer-cadence resampled exports: pose, gaze, velocity, dwell, build and interaction events. This is the record that turns a static scene into how it's actually used, and it's the JSON streaming on the right.
Realistic physics drive every interaction, enabling object manipulation, accident scenarios, creative puzzles, and emergent behaviors. Every action is paired with full-body motion capture, from body pose and hand articulation down to individual finger movements, creating a complete behavioral record stored alongside the event stream.
Need a specific environment, task, population or modality? We build the world, run the collection, and deliver the dataset to your spec, ready to scrub, sample and export.
Delivery package
The live instruments prove synchronization. For qualified evaluations, VLGE packages the same streams into structured exports, enrichment layers and benchmark-ready splits through direct engagement.
Field-level mapping, coordinate frames, units, timestamps and video/telemetry alignment handled with the data team.
Optional per-frame families for data quality, egocentric kinematics, action tokens, attention, spatial coverage, surprise and micro-signals.
For registered geometry: calibrated cameras, depth, semantic/instance masks and projected 2D/3D boxes.
Instruction prompts, task transcripts, audio and success/repair labels for vision-language-action training.
Synchronized co-present sessions for social navigation, collision avoidance and crowd-modeling studies.
Reference splits and starter tasks for world models, embodied agents, robotics and simulation pipelines.
Every environment is governed by real-world physics, allowing objects to react naturally through collisions, gravity, constraints, and force-based interactions instead of scripted animations.
Every user action is captured as high-resolution motion data, including full-body pose, hand articulation, and individual finger movements, preserving how tasks are physically performed.
Every manipulation, movement, and environmental change is recorded as structured event streams, producing replayable interaction histories for AI training, robotics, and simulation.
Example applications
Ground generative world models in how real spaces are built and changed.
Populate simulators with real human trajectories, not scripted agents.
Navigation & manipulation priors from human behavior in 3D.
Agents that learn to perceive, plan and act in physical space.
Benchmarks & datasets for the research defining the field.