Benchmark ยท Behavioral realism

VLGE Benchmark

At VLGE, we are developing benchmarks that push agentic AI beyond conventional measures of success and task completion toward a more demanding objective: behavioral realism.

The grounding

Grounded In Real Human Behavior

Our approach is grounded in real human behavior. We capture how people perceive, decide, move, hesitate, interact, and complete diverse tasks in real-world environments, then use these behavioral distributions to ground and evaluate agentic systems. Rather than asking only whether an agent succeeds, we ask a harder question:

Does It Behave Like A Human While Doing So?

The dimension

Human-Likeness, Measured

This enables a new evaluation dimension for agentic AI: human-likeness at the behavioral level. Agents are compared against empirical human reference distributions using standardized metrics and evaluation protocols that measure how closely their decisions, timing, motion, exploration, and interaction patterns align with real human behavior.

How it runs

One Task, Two Populations, One Distance

Human sessions become the reference distribution and agent sessions the candidate. Both are reduced to the same behavioral features, and the score is the distance between the two distributions.

How the VLGE benchmark measures behavioral realism Human play sessions form the reference distribution and agent play sessions the candidate. Both are reduced to the same behavioral features, decision, timing, kinematic, exploration, progression and recovery, then scored as the distance between the agent and human distributions, which feeds back into agent fine-tuning. Human Play Sessions Reference Agent Play Sessions Candidate Behavioral Features Decision Timing Kinematic Exploration Progression Recovery Behavior Realism Distance agent human D( P(agent) ‖ P(human) ) Agent Fine-Tuning

Solid: the human reference. Dashed: the agent candidate and the loop back into it.

The release

Open Benchmarks For AI Labs

The resulting evaluations are packaged as open benchmarks for AI labs, providing metrics, protocols, and human-grounded reference distributions that can be used to measure, compare, and improve agentic systems. By establishing a common behavioral standard, VLGE aims to make human realism measurable, reproducible, and actionable across the next generation of embodied and agentic AI.

Make Human Realism Measurable. Score Your Agent Against Real People.

See the data behind it →