Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 161 publications
Preview abstract Every abstraction layer in the modern software stack exists to solve legitimate problems—coordinating independent developers and enforcing trust across boundaries. However, each layer exacts a Cognitive Tax: overhead paid not for correctness, but for human coordination. Published measurements bound this non-computational overhead at ~59% across the ISA frontend, ABI, IEEE 754 logic, and per-die guard-band slack. We propose a shift from shipping static artifacts to distributing formal intent. Developer intent is translated into Z3 invariant bundles, which a local Neural-Symbolic Oracle then synthesizes into `Asemantic Code'—an artifact governed by load-time proof certificates and mathematically mutated to exploit the specific manufacturing physics of its execution die. View details
Preview abstract While MVP and CLEAN Architecture are popular Android patterns, they often introduce boilerplate or lack reactivity. This article introduces the Reactive Data Layer Architecture (RDLA), an offline-first, push-based data layer pattern designed for modern Android apps using Jetpack Compose and Room. Using a heart rate tracking example, we demonstrate how RDLA provides robust local-remote synchronization, clean separation of concerns without Use Case overhead, and simplified unit testing via the TestExtensions pattern. View details
Physical Design Aware Verification Methodology for Closing Coverage Gaps in SharedBus MBIST
Shivam Tulsyan
Vasudevan Pillai A
Maheedhar Jalasutram
Prachi Sinha
Mayank Parasrampuria
2026
Preview abstract The industry shift toward SharedBus MBIST architectures has successfully mitigated the Power, Performance, and Area (PPA) bottlenecks associated with traditional embedded memory testing. However, reusing functional paths for testing introduces severe verification challenges, as conventional MBIST algorithms often fail to detect intricate mapping errors like data-bus scrambling, tiedoff data bits, and irregular address bits decoding. If left undetected, these discrepancies in implementation result in silent coverage gaps and ineffective memory repair mechanisms. This paper proposes a robust assertion-based RTL verification methodology specifically designed to close these gaps in SharedBus MBIST implementations. By deploying a Walking-0 pattern and continuous monitors across SharedBus and physical memory interfaces, the methodology enforces a strict set of verification rules. Experimental results validate this approach, demonstrating the successful identification of critical implementation bugs across multiple vendor cores that escaped conventional verification. The paper concludes by proving that the overhead of this methodology is minimal and highly justified by the resulting improvements in silicon quality. View details
Preview abstract Lightweight execution environments like the Little Kernel (LK) are commonly deployed in post-silicon validation to assess software-hardware interactions. These bare-metal kernels, however, lack the sophisticated power management features present in full operating systems, such as the Generic Power Domain (GenPD) framework. Instead of building complex software abstractions that simulate production-grade power management drivers, this paper applies a Design Verification (DV) approach to post silicon. By discarding standard software paradigms, the introduced framework leverages fundamental bare-metal kernel primitives to intentionally engineer synthetic, non-deterministic stressors. Eliminating intermediate software layers allows us to utilize the bare-metal setup to drive chaotic, aggressive stimuli straight to the hardware execution layer. View details
DDRop: Generic Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes
Jesse Demeulemeester
Stefan Gloor
Patrick Jattke
David Oswald
Martin Thompson
Kaveh Razavi
Ingrid Verbauwhede
Jo Van Bulck
ACM Conference on Computer and Communications Security (CCS) (2026)
Preview abstract Trusted Execution Environments (TEEs) are increasingly deployed in the cloud to protect sensitive workloads through hardwareenforced isolation, remote attestation, and transparent memory encryption. However, to meet memory performance and size demands, modern TEEs omit cryptographic freshness guarantees, leaving them vulnerable to replay attacks by adversaries with physical memory access. Prior work demonstrated low-cost active interposition attacks on DDR4 without requiring expensive specialized equipment, but these techniques do not extend to DDR5, where existing approaches are limited to passive ciphertext side-channel analysis and rely on bus downclocking to accommodate legacy memory bus analyzers. We present the first low-cost (<200$) DDR5 RDIMM interposer capable of active fault injection at native speeds. By injecting targeted parity errors to silently discard cache line writebacks, we introduce DDRop, a new primitive that exploits the absence of cryptographic freshness to break the integrity of Intel TDX, Scalable SGX, and AMD SEV-SNP. Building on this primitive and targeting the APIs exposed by the TDX module and AMD Secure Processor, we show that adversaries can gain ciphertext access, copy arbitrary victim pages, and inject malicious secure page-table entries. We demonstrate end-to-end attacks on an up-to-date TDX platform, including forcing any TD into debug mode and forging attestation reports. While software-level mitigations, including timing-based interposer detection and API hardening, may reduce the attack surface, our results demonstrate that active DDR5 bus interposition is practical at low cost, highlighting the need for robust cryptographic memory integrity protections against physical adversaries. View details
Modeling multi-chiplet architectures for TPU co-design
Pritha Doddahosahally Narayanappa
Hung-Ming Hsu
Khai Tran
Hardie Cate
Narges Shahidi
Zhijie Deng
Avinash Lingamneni
Thejasvi Vijayaraj
Aditya Yanamandra
Sameer Kumar
Ming Liu
Lluis-Miquel Munguia
2026
Preview abstract We present an experience report on modeling multi-chiplet hardware accelerators for TPU co-design, and argue that detailed chiplet modeling is necessary to identify bottlenecks and optimizations for future hardware design. Our chiplet modeling tool (CMT) shows that at lower target serving latencies, serving queries/second/chip (QPS/chip) can be up to $2.2\times$ different with chiplet modeling compared to modeling with aggregated compute and memory resources. With this tool, we identify interdependencies between chiplet modeling and the rest of the system: we highlight an example where cost savings from prefetching weights from HBM to local SRAM need to be incorporated in the selection heuristic for chiplet sharding strategies. Using CMT, we explore performance sensitivity to chiplet properties in a design space exploration that highlights considerations for high-performance serving. We use an open-source MoE model (OSS-MoE) served on Ironwood-like TPU architectures as a case study. We conclude by discussing the need to establish best sharding practices to narrow this large design search space. View details
SMaCk: Efficient Instruction Cache Attacks via Self-Modifying Code Conflicts
Seonghun Son
Berk Gulmezoglu
ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) (2025)
Preview abstract Self-modifying code (SMC) allows programs to alter their own instructions, optimizing performance and functionality on x86 processors. Despite its benefits, SMC introduces unique microarchitectural behaviors that can be exploited for malicious purposes. In this paper, we explore the security implications of SMC by examining how specific x86 instructions affecting instruction cache lines lead to measurable timing discrepancies between cache hits and misses. These discrepancies facilitate refined cache attacks, making them less noisy and more effective. We introduce novel attack techniques that leverage these timing variations to enhance existing methods such as Prime+Probe and Flush+Reload. Our advanced techniques allow adversaries to more precisely attack cryptographic keys and create covert channels akin to Spectre across various x86 platforms. Finally, we propose a dynamic detection methodology utilizing hardware performance counters to mitigate these enhanced threats. View details
IM-DD vs. Coherent in Datacenters: A Revisit in 2025
Optical Fiber Communication (OFC) Conference 2025 (2025)
Preview abstract This tutorial examines the progress and scaling limitations of IM-DD based optical technologies and explores how datacenter use cases optimized coherent technology, including a newly proposed polarization-folding, time-diversity approach and a novel single-sideband coherent detection technology—can address some of these challenges View details
Necro-reaper: Pruning away Dead Memory Traffic in Warehouse-Scale Computers
Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Association for Computing Machinery (2025)
Preview abstract Memory bandwidth is emerging as a critical bottleneck in warehouse-scale computing (WSC). This work reveals that a significant portion of memory traffic in WSC is surprisingly unnecessary, consisting of unnecessary writebacks of deallocated data and fetches of uninitialized data. This issue is particularly acute in WSC, where short-lived heap allocations bigger than a cache line are prevalent. To address this problem, this work proposes a pragmatic approach tailored to WSC. Leveraging the existing WSC ecosystem of vertical integration, profile-guided compilation flows, and customized memory allocators, this work presents Necro-reaper, a novel software/hardware co-design that avoids dead memory traffic without requiring the hardware tracking of prior work. New ISA instructions enable the hardware to avoid unnecessary dead traffic, while extended software components, including a profile-guided compiler and memory allocator, optimize the utilization of these instructions. Evaluation across a diverse set of 10 WSC workloads demonstrates that Necro-reaper achieves a geomean memory traffic reduction of 26% and a geomean IPC increase of 6%. View details
ExfilState: Automated Discovery of Timer-Free Cache Side Channels on ARM CPUs
Fabian Thomas
Michael Torres
Michael Schwarz
ACM Conference on Computer and Communications Security (CCS) (2025) (to appear)
Preview abstract Summary: Silent Data Corruption by 10x Test Escapes Threatens Reliable Computing" highlights a critical issue: manufacturing defects, dubbed "test escapes," are evading current testing methods at an alarming rate, ten times higher than industry targets. These defects lead to Silent Data Corruption (SDC), where applications produce incorrect outputs without error indications, costing companies significantly in debugging, data recovery, and service disruptions. The paper proposes a three-pronged approach: quick diagnosis of defective chips directly from system-level behaviors, in-field detection using advanced testing and error detection techniques like CASP, and new, rigorous test experiments to validate these solutions and improve manufacturing testing practices. View details
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
Shiwei Liu
Guanchen Tao
Yifei Zou
Derek Chow
Zichen Fan
Kauna Lei
Bangfei Pan
Dennis Sylvester
Mehdi Saligane
Arxiv (2024)
Preview abstract The self-attention mechanism sets transformer-based large language model (LLM) apart from the convolutional and recurrent neural networks. Despite the performance improvement, achieving real-time LLM inference on silicon is challenging due to the extensively used Softmax in self-attention. Apart from the non-linearity, the low arithmetic intensity greatly reduces the processing parallelism, which becomes the bottleneck especially when dealing with a longer context. To address this challenge, we propose Constant Softmax (ConSmax), a software-hardware co-design as an efficient Softmax alternative. ConSmax employs differentiable normalization parameters to remove the maximum searching and denominator summation in Softmax. It allows for massive parallelization while performing the critical tasks of Softmax. In addition, a scalable ConSmax hardware utilizing a bitwidth-split look-up table (LUT) can produce lossless non-linear operation and support mix-precision computing. It further facilitates efficient LLM inference. Experimental results show that ConSmax achieves a minuscule power consumption of 0.2 mW and area of 0.0008 mm^2 at 1250-MHz working frequency and 16-nm CMOS technology. Compared to state-of-the-art Softmax hardware, ConSmax results in 3.35x power and 2.75x area savings with a comparable accuracy on a GPT-2 model and the WikiText103 dataset. View details
Limoncello: Prefetchers for Scale
Carlos Villavieja
Baris Kasikci
Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Association for Computing Machinery, New York, NY, United States (2024)
Preview abstract This paper presents Limoncello, a novel software system that dynamically configures data prefetching for high utilization systems. We demonstrate that in resource-constrained environments, such as large data centers, traditional methods of hardware prefetching can increase memory latency and decrease available memory bandwidth. To address this, Limoncello dynamically configures data prefetching, disabling hardware prefetchers when memory bandwidth utilization is high and leveraging targeted software prefetching to reduce cache misses when hardware prefetchers are disabled. Limoncello is software-centric and does not require any modifications to hardware. Our evaluation of the deployment on a real-world hyperscale system reveals that Limoncello unlocks significant performance gains for high-utilization systems: it improves application throughput by 10%, due to a 15% reduction in memory latency, while maintaining minimal change in cache miss rate for targeted library functions. View details
Pathfinder: High-Resolution Control-Flow Attacks with Conditional Branch Predictor
Hosein Yavarzadeh
Archit Agarwal
Max Christman
Christina Garman
Daniel Genkin
Andrew Kwong
Deian Stefan
Mohammadkazem Taram
Dean Tullsen
International Conference on Architectural Support for Programming Languages and Operating Systems, ACM (2024)
Preview abstract This paper presents novel attack primitives that provide adversaries with the ability to read and write the path history register (PHR) and the prediction history tables (PHTs) of the conditional branch predictor in modern Intel CPUs. These primitives enable us to recover the recent control flow (the last 194 taken branches) and, in most cases, a nearly unlimited control flow history of any victim program. Additionally, we present a tool that transforms the PHR into an unambiguous control flow graph, encompassing the complete history of every branch. This work provides case studies demonstrating the practical impact of novel reading and writing/poisoning primitives. It includes examples of poisoning AES to obtain intermediate values and consequently recover the secret AES key, as well as recovering a secret image by capturing the complete control flow of libjpeg routines. Furthermore, we demonstrate that these attack primitives are effective across virtually all protection boundaries and remain functional in the presence of all recent control-flow mitigations from Intel. View details
Hardware-Assisted Fault Isolation: Going Beyond the Limits of Software-Based Sandboxing
Shravan Narayan
Tal Garfinkel
Mohammadkazem Taram
Joey Rudek
Evan Johnson
Chris Fallin
Anjo Vahldiek-Oberwagner
Michael LeMay
Ravi Sahita
Dean Tullsen
Deian Stefan
IEEE Micro (2024)
Preview abstract Hardware-assisted Fault Isolation (HFI) is a minimal extension to current processors that supports secure, flexible, and efficient in-process isolation. HFI addresses the limitations of software-based isolation (SFI) systems including: runtime overheads, limited scalability, vulnerability to Spectre attacks, and limited compatibility with existing code. HFI can be seamlessly integrated into exisiting SFI systems (e.g. WebAssembly), or directly sandbox unmodified native binaries. To ease adoption, HFI proposes incremental changes to existing high-performance processors. View details
×