Benchmark ยท Behavioral realism
At VLGE, we are developing benchmarks that push agentic AI beyond conventional measures of success and task completion toward a more demanding objective: behavioral realism.
The grounding
Our approach is grounded in real human behavior. We capture how people perceive, decide, move, hesitate, interact, and complete diverse tasks in real-world environments, then use these behavioral distributions to ground and evaluate agentic systems. Rather than asking only whether an agent succeeds, we ask a harder question:
Does It Behave Like A Human While Doing So?
The dimension
This enables a new evaluation dimension for agentic AI: human-likeness at the behavioral level. Agents are compared against empirical human reference distributions using standardized metrics and evaluation protocols that measure how closely their decisions, timing, motion, exploration, and interaction patterns align with real human behavior.
How it runs
Human sessions become the reference distribution and agent sessions the candidate. Both are reduced to the same behavioral features, and the score is the distance between the two distributions.
Solid: the human reference. Dashed: the agent candidate and the loop back into it.
The release
The resulting evaluations are packaged as open benchmarks for AI labs, providing metrics, protocols, and human-grounded reference distributions that can be used to measure, compare, and improve agentic systems. By establishing a common behavioral standard, VLGE aims to make human realism measurable, reproducible, and actionable across the next generation of embodied and agentic AI.