Neha Arora

I lead the MobilityAI team at Google Research. Mobility AI is focused on advancing AI to help cities tackle major transportation problems like traffic congestion, road safety, and emissions, enabling data-driven urban planning and operations. Read more about our work here - https://research.google/blog/introducing-mobility-ai-advancing-urban-transportation/ My prior experience has been in search ranking algorithms.
Authored Publications
Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
Mobility-Embedded POIs: Learning What a Place Is and How It’s Used from Human Movement
Shushman Choudhury
Shang-Ling Hsu
Cyrus Shahabi
Forty-third International Conference on Machine Learning (2026)
Preview abstract Recent progress in geospatial foundation models (GeoFMs) has highlighted the importance of learning general-purpose representations for real-world locations, particularly Points of Interest (POIs) where human activity concentrates. Yet, ex- isting POI representations remain largely static, drawing from textual metadata (e.g., category labels, descriptions) and spatial attributes (e.g., coordinates, neigh- borhood context), all of which describe what a place is, but not how it is actu- ally used. We argue that human mobility provides a complementary and dynamic signal, capturing real-world visitation patterns that reveal how places function in practice. To this end, we introduce Mobility Embedded POIs (ME-POIs), a pretraining framework that learns POI representations directly from sequences of human visits. Each visit is encoded as a contextualized embedding that captures the POI’s static attributes as well as its temporal and sequential context, includ- ing when the visit occurs and which visits surround it. These visit embeddings are aligned with learnable POI embeddings via a contrastive objective, grounding POI representations in their real-world usage patterns. To address the long tail of sparsely visited POIs, we transfer visitation distributions from data-rich anchors to sparse locations, leveraging multi-scale spatial proximity to capture local and regional patterns, and functional similarity to enable transfer across semantically related POIs. We demonstrate the utility of ME-POIs for a set of automated map enrichment tasks, critical in geospatial intelligence. We show empirically that by embedding visitation dynamics, ME-POIs outperform text- and location-only baselines, proving that mobility-informed embeddings provide a stronger founda- tion for modeling place function and change. View details
Preview abstract This paper introduces XMob, a novel differentiable traffic simulation framework built in JAX to advance traditional models like SUMO’s mesoscopic simulator. By leveraging JAX’s capabilities for vectorized, hardware-accelerated computation (GPU/TPU), XMob achieves orders-of-magnitude speedups, enabling large-scale urban network simulations and extensive counterfactual analyses. A key innovation is XMob’s inherent differentiability, facilitating direct integration with gradient-based optimization for tasks such as demand calibration and network parameter estimation, significantly outperforming black-box approaches. Furthermore, XMob can be used in Physics-Informed Machine Learning (PIML) pipelines to enhance data-driven augmentation, embedding domain principles like flow conservation and shockwave theory. This ensures physically plausible and robust predictions, even for unobserved scenarios such as lane modifications. The hybrid architecture, combining a deterministic JAX core with incremental machine learning, offers a scalable and efficient solution for modern traffic simulation and optimization challenges. View details
Preview abstract Estimating Origin-Destination (OD) travel demand is vital for effective urban planning and traffic management. Developing universally applicable OD estimation methodologies is significantly challenged by the pervasive scarcity of high-fidelity traffic data and the difficulty in obtaining city-specific prior OD estimates (or seed ODs), which are often prerequisite for traditional approaches. Our proposed method directly estimates OD travel demand by systematically leveraging aggregated, anonymized statistics from Google Maps Traffic Trends, obviating the need for conventional census or city-provided OD data. The OD demand is estimated by formulating a single-level, one-dimensional, continuous nonlinear optimization problem with nonlinear equality and bound constraints to replicate highway path travel times. The method achieves efficiency and scalability by employing a differentiable analytical macroscopic network model. This model by design is computationally lightweight, distinguished by its parsimonious parameterization that requires minimal calibration effort and its capacity for instantaneous evaluation. These attributes ensure the method's broad applicability and practical utility across diverse cities globally. Using segment sensor counts from Los Angeles and San Diego highway networks, we validate our proposed approach, demonstrating a two-thirds to three-quarters improvement in the fit to segment count data over a baseline. Beyond validation, we establish the method's scalability and robust performance in replicating path travel times across diverse highway networks, including Seattle, Orlando, Denver, Philadelphia, and Boston. In these expanded evaluations, our method not only aligns with simulation-based benchmarks but also achieves an average 13% improvement in it's ability to fit travel time data compared to the baseline during afternoon peak hours. View details
Toward Foundation Models for Mobility Enriched Geospatially Embedded Objects
Shang-Ling Hsu
Shushman Choudhury
Cyrus Shahabi
Proceedings of the 33rd ACM International Conference on Advances in Geographic Information Systems, ACM (2025)
Preview abstract Recent advances in large foundation models (FMs) have enabled learning general-purpose representations in natural language, vi- sion, and audio. Yet geospatial artificial intelligence (GeoAI) still lacks widely adopted foundation models that generalize across tasks. We argue that a key bottleneck is the absence of unified, general-purpose, and transferable representations for geospatially embedded objects (GEOs). Such objects include points, polylines, and polygons in geographic space, enriched with semantic context and critical for geospatial reasoning. Much current GeoAI research compares GEOs to tokens in language models, where patterns of human movement and spatiotemporal interactions yield contextual meaning similar to patterns of words in text. However, modeling GEOs introduces challenges fundamentally different from language, including spatial continuity, variable scale and resolution, temporal dynamics, and data sparsity. Moreover, privacy constraints and global variation in mobility further complicates modeling and gen- eralization. This paper formalizes these challenges, identifies key representational gaps, and outlines research directions for build- ing foundation models that learn behavior-informed, transferable representations of GEOs from large-scale human mobility data. View details
Toward Foundation Models for Mobility Enriched Geospatially Embedded Objects
Maria Despoina Siampou
Shang-Ling Hsu
Shushman Choudhury
Cyrus Shahabi
2025
Preview abstract Recent advances in large foundation models (FMs) have enabled learning general-purpose representations in natural language, vi- sion, and audio. Yet geospatial artificial intelligence (GeoAI) still lacks widely adopted foundation models that generalize across tasks. We argue that a key bottleneck is the absence of unified, general-purpose, and transferable representations for geospatially embedded objects (GEOs). Such objects include points, polylines, and polygons in geographic space, enriched with semantic context and critical for geospatial reasoning. Much current GeoAI research compares GEOs to tokens in language models, where patterns of human movement and spatiotemporal interactions yield contextual meaning similar to patterns of words in text. However, modeling GEOs introduces challenges fundamentally different from language, including spatial continuity, variable scale and resolution, temporal dynamics, and data sparsity. Moreover, privacy constraints and global variation in mobility further complicates modeling and gen- eralization. This paper formalizes these challenges, identifies key representational gaps, and outlines research directions for build- ing foundation models that learn behavior-informed, transferable representations of GEOs from large-scale human mobility data. View details
Improving simulation-based origin-destination demand calibration using sample segment counts data
Arwa Alanqary
Yechen Li
The 12th Triennial Symposium on Transportation Analysis conference (TRISTAN XII), Okinawa, Japan (2025)
Preview abstract This paper introduces a novel approach to demand estimation that utilizes partial observations of segment-level track counts. Building on established simulation-based demand estimation methods, we present a modified formulation that integrates sample track counts as a regularization term. This approach effectively addresses the underdetermination challenge in demand estimation, moving beyond the conventional reliance on a prior OD matrix. The proposed formulation aims to preserve the distribution of the observed track counts while optimizing the demand to align with observed path-level travel times. We tested this approach on Seattle's highway network with various congestion levels. Our findings reveal significant enhancements in the solution quality, particularly in accurately recovering ground truth demand patterns at both the OD and segment levels. View details
Preview abstract This talk explores the critical role of calibration in developing accurate and reliable digital twins of metropolitan transportation networks. We address the challenge of underdetermination in large-scale stochastic simulation calibration, where limited real-world data can hinder accurate counterfactual analysis. To overcome this, we introduce sample-efficient methodologies that leverage metamodels, machine learning, and diverse emerging data sources. Specially, we will discuss the use of abundant path travel times for demand calibration. The enhancement in simulation fidelity is demonstrated through case studies of multiple metropolitan areas in the US. View details
Towards a Trajectory-powered Foundation Model of Mobility
Ivan Kuznetsov
Shushman Choudhury
Aboudy Kreidieh
(2024)
Preview abstract This position paper advocates for the development of a geospatial foundation model based on human mobility trajectories in the built environment. Such a model would be widely applicable across many important societal domains currently addressed independently, including transportation networks, data-driven urban planning and management, tourism, and sustainability. Unlike existing large vision-language models, trained primarily on text and images, this foundation model should integrate the complex spatiotemporal and multimodal data inherent to human mobility. This paper motivates this challenging research agenda, outlining many downstream applications that would be differentially impacted or enabled by such a model. It then explains the critical spatial, temporal, and contextual factors that the model must capture to effectively analyze trajectories. Finally, it concludes with several research questions and directions, laying the foundations for future exploration in this exciting and emerging field. View details
Scalable Learning of Segment-Level Traffic Congestion Functions
Shushman Choudhury
Aboudy Kreidieh
Alexandre Bayen
IEEE Intelligent Transportation Systems Conference (2024)
Preview abstract We propose and study a data-driven framework for identifying traffic congestion functions (numerical relationships between observations of traffic variables) at global scale and segment-level granularity. In contrast to methods that estimate a separate set of parameters for each roadway, ours learns a single black-box function over all roadways in a metropolitan area. First, we pool traffic data from all segments into one dataset, combining static attributes with dynamic time-dependent features. Second, we train a feed-forward neural network on this dataset, which we can then use on any segment in the area. We evaluate how well our framework identifies congestion functions on observed segments and how it generalizes to unobserved segments and predicts segment attributes on a large dataset covering multiple cities worldwide. For identification error on observed segments, our single data-driven congestion function compares favorably to segment-specific model-based functions on highway roads, but has room to improve on arterial roads. For generalization, our approach shows strong performance across cities and road types: both on unobserved segments in the same city and on zero-shot transfer learning between cities. Finally, for predicting segment attributes, we find that our approach can approximate critical densities for individual segments using their static properties. View details
Traffic simulations: multi-city calibration of metropolitan highway networks
Yechen Li
Damien Pierce
The 27th IEEE International Conference on Intelligent Transportation Systems (ITSC), Edmonton, Canada (2024)
Preview abstract This paper proposes an approach to perform travel demand calibration for high-resolution stochastic traffic simulators. It employs abundant travel times at the path-level, departing from the standard practice of resorting to scarce segment-level sensor counts. The proposed approach is shown to tackle high-dimensional instances in a sample-efficient way. For the first time, case studies on 6 metropolitan highway networks are carried out, considering a total of 54 calibration scenarios. This is the first work to show the ability of a calibration algorithm to systematically scale across networks. Compared to the state-of-the-art simultaneous perturbation stochastic approximation (SPSA) algorithm, the proposed approach enhances fit to field data by an average 43.5% with a maximum improvement of 80.0%, and does so within fewer simulation calls. View details
×