STRIDE : Automated Evaluation of Text-to-Trajectory
Alignment across Diverse Contexts

Wanchun Ni1,*, Tao Qi2, Leonel Aguilar1, Jiugeng Sun1,
Marlene Wagner1, Verena Zimmermann1, Mennatallah El-Assady1
1ETH Zurich  2Beijing University of Posts and Telecommunications
*Corresponding author
NeurIPS 2026 · Evaluations & Datasets Track

Evaluating Context Alignment Without Human Trajectory Data

Language-conditioned models now generate pedestrian trajectories from a scenario description, but their evaluation has not kept pace. Existing metrics compare generated trajectories with recorded human data. That does not scale: collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks do not transfer to this domain.

STRIDE is the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. Instead of asking an LLM to judge trajectories directly, it decomposes each scenario into behavioral questions and answers every question with deterministic trajectory measurements, so the evaluation is reproducible, interpretable and needs no human trajectory data.

Overview of STRIDE in two stages. Stage 1, STRIDE-Bench generation: a scenario (a firecracker thrown into a busy plaza causes a brief panic) is decomposed by an LLM, guided by the VRDST protocol and TrajFacts human trajectory knowledge, into behavioral questions such as 'Are people running in panic?'; each question is tied to measurement functions such as mean_speed, flow_alignment or speed_trend and to an expected answer range. Stage 2, STRIDE score calculation: a generated trajectory is measured with these functions, each question receives a Q-score, and their mean gives the scene score, 0.89 in this example. (full-size figure, opens in a new tab)

Overview of the STRIDE evaluation framework. Stage 1 constructs STRIDE-Bench by decomposing each scenario into behavioral questions, measurement functions, and expected answers. Stage 2 computes the STRIDE score by comparing function outputs against the expected answers.


How STRIDE Works

  • A protocol grounded in sociology. The VRDST protocol (Velocity, Realism, Direction, Spatial, Temporal) is derived from sociological theories and defines a complete evaluation space across individual, group and environment layers.
  • Scenario-adaptive behavioral questions. Guided by the protocol, each high-level scenario description is decomposed into behavioral questions that fit that scenario.
  • Deterministic measurements. Every question is resolved against a library of deterministic measurement tools, each computing an exact trajectory statistic, so the answers are reproducible.

STRIDE-Bench

We instantiate STRIDE in the crowd domain as STRIDE-Bench, with calibrated expected answers for every measurement.

  • 1Kscenarios
  • 6Kbehavioral questions
  • 11Kmeasurements
  • 30real-world maps
  • 80%human agreement

Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment. Evaluating several text-to-trajectory models with it shows limited context-alignment capability and persistent challenges in fine-grained context conditioning.

The benchmark is available on Hugging Face, and the code on GitHub.


Abstract

Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult.

We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data.

We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation.


Citation

@inproceedings{ni2026stride,
    title     = {{STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts}},
    author    = { Wanchun Ni              and
                  Tao Qi                  and
                  Leonel Aguilar          and
                  Jiugeng Sun             and
                  Marlene Wagner          and
                  Verena Zimmermann       and
                  Mennatallah El-Assady },
    year      = {2026},
    booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations \& Datasets Track},
    doi       = {10.48550/arXiv.2609.34799}
}