Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11526 publications
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Geo experiment is a crucial, privacy-conscious option for measuring media effectiveness. Historically, two primary challenges have hindered the adoption of geo experiments: the high costs required to overcome large variance across geographical regions, and the unreliability of conventional analysis methods under real-world autocorrelation and non-stationary trends. To address these challenges, we introduce Meridian GeoX, Google’s open-source geo experiment framework. As a cornerstone of Google's modern measurement suite, Meridian GeoX is designed to standardize and optimize the end-to-end causal measurement lifecycle. The framework provides a unified suite of methodologies—including data-driven stratified sampling, advanced counterfactual models (Time-Based Regression, Synthetic Control, and Synthetic Difference-in-Differences), and a novel Design-Aware Placebo Inference engine. Extensive empirical benchmarking against industry alternatives demonstrates that Meridian GeoX delivers superior predictive accuracy, lowers Minimum Detectable Effects (MDE), and significantly reduces budget requirements. Furthermore, evaluations demonstrate that the framework's novel placebo inference maintains rigorous control over false positive rates while maximizing sensitivity under challenging conditions. By integrating seamlessly with the Meridian Marketing-Mix Model (MMM), this framework delivers a complete modern measurement solution. Ultimately, Meridian GeoX empowers advertisers with a powerful, cost-efficient, and methodologically robust solution for privacy-safe modern measurement, establishing a new industry standard for causal measurement and media effectiveness. View details
Looking to the brain to improve energy efficiency of AI
Taro Toyoizumi
Hakwan Lau
Michał Klincewicz
Seng Bum Michael Yoo
Megan Peters
taylor.w.webb@gmail.com
Current Biology (2026)
Preview abstract Modern artificial intelligence (AI) systems have achieved remarkable capabilities, but at an extraordinary energy cost. Training and running large-scale models can consume vast resources, posing environmental, economic, and social challenges. In contrast, biological brains perform lifelong learning, adaptive control, and flexible reasoning using orders of magnitude less energy for learning and adaptation over a lifetime. What accounts for this difference -- and how can it guide future AI development? In this article, we identify key biological principles that support energy-efficient capacities in biological brains, and consider how they might inform the design of more sustainable artificial systems. We organize our analysis around three domains: architectural constraints, signaling strategies, and learning algorithms. In each domain, we discuss concrete observations from biology -- from cell to circuit to cognitive level -- and describe how current and emerging AI systems mirror or diverge from these motifs. One striking feature of biological energy optimization is often overlooked: that brains are remarkably stable in their energy usage across heterogeneous modes, suggesting they may minimize energy needs during active environmental processing through maximizing the utility of “rest-like” background processes. Overall, rather than advocating for biomimicry for its own sake, we argue for biologically informed engineering. Understanding how natural systems minimize energetic cost while maximizing flexibility may help us build AI that is not only powerful, but also efficient, equitable, and environmentally responsible. View details
Preview abstract Source-to-source compilers may perform inefficiently by executing transpilation passes on scripts that do not contain the specific language features a pass is designed to transform, potentially leading to redundant processing. A compiler can analyze a script to generate a per-script feature map, for example, by identifying language features in its abstract syntax tree (AST). Before executing a transpilation pass, the compiler can check this map and may bypass the pass for that script if the specific feature targeted by the pass is not present. This feature map can also be dynamically updated throughout the compilation process as other passes transform the code. This method of conditional pass execution based on content-aware analysis may reduce redundant AST traversals, which could decrease overall compilation time and computational resource consumption. View details
Preview abstract Distributed federated SQL query engines frequently materialize query results into files in data lakes. Optimizing the sizes of these files (e.g., balancing their sizes) is crucial for the efficiency of not only the materialization queries themselves, but also subsequent queries that read these files. Existing techniques to manage file size include specifying output targets (e.g., number of partitions), using cardinality estimation for materialized data volume, performing full shuffles prior to materialization to obtain accurate statistics, and/or post (background) compactions to merge small files. All of these solutions have limitations in practice; they typically do not provide any strong guarantees or can be prohibitively expensive when they do (e.g., they require background compactions to merge small files, which requires additional I/O and CPU cost). Most of the existing techniques can produce many small files during materialization and are sensitive to data skew. In this paper, we propose a novel method that streamlines materialization within the same query execution and provides guarantees on the file sizes using statistics at runtime. Our approach does not require a full shuffle before the materialization operator in order to get accurate statistics, making it attractive and robust in practice. To show the effectiveness of our approach, we present production metrics from F1 Query at Google that has been running this functionality for the majority of its production workload over many quarters. Our approach reduces the number of files produced by a factor of 100 or more, and as a result, it avoids creating tens of billions of files per day, saving storage and computation cost. View details
Preview abstract Optimizing burst-heavy datacenter workloads necessitates fine-grained network control and visibility. We introduce CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header. The architecture captures 𝜇s-granularity switch metrics, such as available bandwidth, and signals them to end-hosts using in-band, line-rate operations. We propose Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production. Beyond transport-level performance, CSIG enables flow-aware observability by embedding 𝜇s-scale metrics into every packet, allowing individual application transfers to pinpoint their bottleneck location, such as the topology tier limiting their performance. CSIG thus transforms network telemetry from post-hoc correlation into a real time, contextaware capability. We demonstrate CSIG’s broad deployability by validating it across five generations of commodity switch hardware (up to 102.4 Tbps), four NIC generations, and five transport stacks. Our design proves that a streamlined Layer 2 approach, focusing exclusively on the principal path bottleneck, provides transport-agnostic gains without requiring forklift hardware upgrades. View details
Preview abstract A growing body of qualitative research has identified contextual risk factors that elevate people’s chances of experiencing digital-safety attacks. However, the lack of quantitative data on the population level distribution of these risk factors prevents policymakers and tech companies from developing targeted, evidence-based interventions to improve digital safety. To address this gap, we surveyed 5,001 adults in the United States to analyze: (1) the frequency of and relationship between digital-safety attacks (e.g., scams, harassment, account hacking), and (2) how these attacks align with 10 contextual risk factors. Nearly half of our respondents identify as resource constrained, which significantly correlates with higher likelihood of experiencing four common attacks. We also present qualitative insights to expand our understanding of the factors beyond the existing literature (e.g., “prominence” included high-visibility roles in local communities). This study provides the first large-scale quantitative analysis correlating digital-safety attacks with contextual risk factors and demographics. View details
Spectral amplification for ground-state energy estimation of electronic structure in first quantization
Alicja Dutkiewicz
Alec White
Guang Hao Low
Albert Eugene DePrince III
Marika Kieferova
Dominic Berry
arXiv:2607.15358 (2026)
Preview abstract We demonstrate an asymptotic gate complexity improvement in first-quantized ground-state energy estimation of electronic structure Hamiltonians in a plane wave basis by employing the sum-of-squares spectral gap amplification protocol. The improvement relies on identifying a sum-of-squares representation of the Hamiltonian which provides a lower bound certificate and low cost block encoding that leads to a provably lower quantum phase estimation gate cost. This is achieved by using a sum-of-squares operator generated by the total charge density operator resulting in a block encoding normalization improvement of $\lambda = \mathcal{O}\left(\eta\Delta^{-1.5}+\eta^{1.5}\Delta^{-1} \right)$ compared to prior work $\lambda = \mathcal{O}(\eta\Delta^{-2}+\eta^2\Delta^{-1})$ where $\eta$ is the number of electrons and $\Delta$ is the simulation grid spacing. The asymptotic reduction in block encoding normalization and similar block encoding costs to prior work is demonstrated to reduce resource estimates for materials and chemical systems by a factor of $2 - 44\times$ corresponding to the lowest cost estimates for \textit{ab initio} materials simulation. View details
Performance analysis of updated Sleep Tracking algorithms across Google and Fitbit wearable devices
Arno Charton
Linda Lei
Siddhant Swaroop
Marius Guerard
Michael Dixon
Logan Niehaus
Shao-Po Ma
Ross Wilkinson
Ryan Gillard
Conor Heneghan
Pramod Rudrapatna
Mark Malhotra
Shwetak Patel
Google, Google, 1600 Amphitheatre Parkway Mountain View, CA 94043 (2026) (to appear)
Preview abstract Background: The general public has increasingly adopted consumer wearables for sleep tracking over the past 15 years, but reports on performance versus gold standards such as polysomnogram (PSG), high quality sleep diaries and at-home portable EEG systems still show potential for improved performance. Two aspects in particular are worthy of consideration: (a) improved recognition of sleep sessions (times when a person is in bed and has attempted to sleep), and (b) improved accuracy on recognizing sleep stages relative to an accepted standard such as PSG. Aims: This study aimed to: 1) provide an update on the methodology and performance of a system for correctly recognizing valid sleep sessions, and 2) detail an updated description of how sleep stages are calculated using accelerometer and inter-beat intervals Methods: Novel machine learning algorithms were developed to recognize sleep sessions and sleep stages using accelerometer sensors and inter-beat intervals derived from the watch or tracker photoplethysmogram. Algorithms were developed on over 3000 nights of human-scored free-living sleep sessions from a representative population of 122 subjects, and then tested on an independent validation set of 47 users. Within sleep sessions, an algorithm was developed to recognize periods when the user was attempting to sleep (Time-Attempting-To-Sleep = TATS). For sleep stage estimation, an algorithm was trained on human expert-scored polysomnograms, and then tested on 50 withheld subject nights for its ability to recognize Wake, Light (N1/N2), Deep (N3) and REM sleep relative to expert scored labels. Results: For sleep session estimation, the algorithm had at least 95% overlap on TATS with human consensus scoring for 94% of nights from healthy sleepers. For sleep stage estimation, comparing with the current Fitbit algorithm, Cohen’s kappa for four-class determination of sleep stage increased from an average of 0.56 (std 0.13) to 0.63 (std 0.12), and average accuracy increased from 71% (std 0.10) to 77% (std 0.078) Conclusion: A set of new algorithms has been developed and tested on Fitbit and Pixel Watches and is capable of providing robust and accurate measurement of sleep in free-living environments. View details
Fair Allocation of Indivisible Goods with Variable Groups
Paul Golz
Warut Suksompong
Ayumi Igarashi
AAAI (2026)
Preview abstract We study the fair allocation of indivisible goods with variable groups. In this model, the goal is to partition the agents into groups of given sizes and allocate the goods to the groups in a fair manner. We show that for any number of groups and corresponding sizes, there always exists an envy-free up to one good (EF1) outcome, thereby generalizing an important result from the individual setting. Our result holds for arbitrary monotonic utilities and comes with an efficient algorithm. We also prove that the EF1 existence can be guaranteed even when the goods lie on a path and each group must receive a connected bundle. In addition, we consider a probabilistic model where the utilities are additive and drawn randomly from a distribution. We show that if there are n agents and the number of goods m is divisible by the number of groups k, then an envy-free outcome exists with high probability if m = ω(log n), and this bound is tight. On the other hand, if m is not divisible by k, then an envy-free outcome is unlikely to exist as long as m = o(√n). View details
Toward a test of medical AI superintelligence
Ethan Goh
David Wu
Chase Walton
Liam McCoy
Anastasia Perez
Laura Wegner
Fateme Nateghi Haredasht
Luyang Luo
Kathleen Lacar
Thomas Buckley
Austin Schoeffler
Peter Brodeur
Kameron C. Black
John Havlik
John Rumsfeld
Daniel Lopez-martinez
Paxton Maeder-York
Karan Singhal
David Gunning
Bon Ku
Haider Warraich
Shantanu Nundy
Vishnu Ravi
Arnold Milstein
Jason Hom
Kevin Schulman
Pranav Rajpurkar
Arjun Manrai
Robert Wachter, MD
Eric Topol
Eric horvitz
Adam Rodman
Jonathan Chen
Nature Medicine (2026)
Preview abstract Researchers urgently need a rigorous, task-based framework to define and measure medical AI ‘superintelligence’, because existing benchmarks are misleading and insufficient. View details
Fine-Grained Table Retrieval for Open-Domain Tabular Question Answering
Xingyu Ji
Wojciech Kosiuk
Madelon Hulsebos
Proceedings of the 11th Workshop on Automated Knowledge Base Construction (AKBC 2026), Association for Computational Linguistics
Preview abstract This work introduces a fine-grained table retrieval framework for grounding large language models in heterogeneous, open-domain relational data. Instead of encoding a query as a single vector, the approach decomposes natural language queries into semantic components and embeds each independently, enabling more precise matching of compositional query intent. These representations are used in a staged retrieval pipeline with component-level search, connectivity-aware grouping, and reranking. Experiments on three TARGET benchmark corpora show consistent improvements in capped recall@k and stronger alignment between query intent and tabular structure over dense retrieval baselines, particularly for longer and more complex queries and when using lightweight embedding models. View details
AIRS: Scaling Live Inference in Resource Constrained Environments
Xiaohao Yang
Tuan Do
Chelsea Chen
Harshvardhan GM
(2026)
Preview abstract Advancements in large language models (LLMs) have made them increasingly useful for complex reasoning tasks which previously required domain experts. One such task is quality evaluation of query responses produced by a search engine. Evaluation generates metrics necessary to study the quality, impact, and usefulness of product changes and features. Typically, to compute evaluation metrics, human experts are asked to rate various attributes of search responses. This process is generally quite expensive and requires several days to complete. As an alternative, LLMs are now being used to perform rating tasks with lower costs and latency. In addition, many new metrics are being developed to evaluate Google's new AI-based offerings, which require ratings too. As a result, there is much higher demand for LLM rating prediction tasks in comparison with the allocated TPU (Tensor Processing Unit) budget. A larger portion of the company's TPU resources are reserved for serving live user traffic. In this paper, we present the AI Rater Service (AIRS), an inference pipeline that employs several software engineering techniques to generate AI ratings with high reliability and low latency. AIRS maximizes LLM inference throughput by optimizing TPU resource utilization across various evaluation workflows, while minimizing latency for higher priority tasks. View details
Shuffles of Context-Free Languages along Regular Trajectories
Corentin Barloy
Michaël Cadilhac
Kyle Ockerlund
2026
Preview abstract In single-core processors, concurrency requires that multiple processes be interleaved into a single thread of execution by a scheduler. The language-theoretic operation that corresponds to this is the shuffle of two languages: the set of words obtained by interleaving a word from each language in an arbitrary, letter-wise fashion. It is well known that regular languages are closed under shuffles, while context-free languages (CFLs) are not. Following an established line of research, this paper considers shuffles according to regular "trajectories," that is, subject to scheduling constraints expressed by an automaton. Unsurprisingly, some trajectories allow for CFLs to be shuffled into CFLs (e.g., simple concatenation of the two words), while others do not. This paper provides a robust toolset to show that a given trajectory would always shuffle two nonregular CFLs into a nonCFL. In the case of deterministic CFLs (DCFLs), a salient trichotomy of trajectories depending on how they shuffle DCFLs is provided. These results are based on lemmata of independent interest regarding how pushdown automata (PDA) must invoke the stack when accepting a nonregular CFL or DCFL. The latter case relies on a recent result of Jančar and Šíma (MFCS'2021); answering an open question therein, it is demonstrated that said result cannot be generalized to arbitrary CFLs, leading to dedicated machinery for both cases. View details
×