Long T. Le

Long T. Le

Long T. Le is a Staff Research Engineer Manager within Google Cloud AI Research, where he leads initiatives to bring advanced AI solutions to global applications. His current research and development focus on next-generation LLM solutions, including distillation, Retrieval-Augmented Generation (RAG), and AI agents—with a specialized emphasis on Agentic Security, Safety, and Cybersecurity. In his early years at Google, Long pioneered novel deep learning methodologies for tabular data, developed critical COVID-19 forecasting models, and advanced Recommendation AI systems. Before joining Google, he served as a Machine Learning Engineer at Capital One in New York City, where he developed high-impact models for loan optimization and first-party fraud detection. Long earned his Ph.D. in Computer Science from Rutgers University and holds a Bachelor of Computing from the National University of Singapore (NUS).
Authored Publications
Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
Preview abstract Large language model agents increasingly act in deployment environments where failures are contextual, user-specific, and costly. In such settings, a \emph{static general-purpose guardrail is often insufficient}: whether an action should be allowed may depend on local privacy norms, organizational rules, or evolving user expectations that are difficult to enumerate fully in advance. We study \emph{lifelong deployment-time guardrail adaptation}, where a fixed base guardrail improves over time from sparse, noisy user-reported failures without repeated fine-tuning. We propose a conservative policy induction framework organized as an online--offline loop. Online, the deployed guardrail uses structured policy memory to guide runtime decisions. Offline, newly accumulated reports are converted into reusable policy items and folded back into memory through periodic refresh. The method combines three ingredients: \emph{broad policy abstraction} for sparse failure generalization, \emph{conflict-aware local policies} for mixed-label regions where broad reuse becomes too coarse, and \emph{confidence-gated reuse} based on conservative posterior lower bounds so that weakly supported memory does not influence inference too early. Across PrivacyLens+, ConFaide+, and AgentHarm, the resulting system consistently improves over a lightweight base guardrail and strong memory-based baselines in sparse-feedback regimes, remains robust to noisy feedback, traces a better cost--performance frontier than scaling the base model alone, and jointly reduces over-refusal and over-acceptance without an explicit balance knob. View details
Preview abstract Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and dynamic environments, but this also introduces severe security risks. In particular, indirect prompt injection attacks can compromise agents through malicious instructions hidden in external sources such as web pages, emails, and retrieved documents. Existing defenses are largely reactive, while current automated red-teaming methods mainly optimize attack success rather than systematically uncovering hidden vulnerabilities within the agent pipeline. In this work, we propose PI-Hunter, an automated agentic red-teaming framework that shifts the focus from attack optimization to vulnerability exposure. By combining static attack-surface analysis, source-aware seeding, trajectory evaluation, and feedback-guided exploration, PI-Hunter proactively discovers vulnerable ingestion paths and localizes how malicious instructions propagate through agent reasoning. Extensive experiments across multiple benchmarks, agent architectures, attacks, and defenses show that \method~substantially improves vulnerability exposure and attack-surface coverage compared with existing automated red-teaming baselines, while remaining effective even under strong prompt injection defenses. View details
Preview abstract Large language models (LLMs) have shown promise in assisting cybersecurity tasks, yet existing approaches struggle with automatic vulnerability discovery and exploitation due to limited interaction, weak execution grounding, and a lack of experience reuse. We propose Code-RedTeam, a security-aware multi-agent framework designed to mirror real-world red-teaming workflows by integrating security-domain knowledge, code-aware analysis, execution-grounded iterative reasoning, and long-term memory. Code-RedTeam decomposes vulnerability analysis into coordinated discovery and exploitation stages, enabling agents to plan, execute, validate, and refine actions based on real execution feedback while learning from prior trajectories. Extensive evaluations on challenging security benchmarks demonstrate that Code-RedTeam consistently outperforms strong baselines across diverse backbone models, achieving over 60% attack success rate in vulnerability exploitation and up to 10% absolute improvement in vulnerability detection. Ablation and iteration studies further confirm the critical role of execution feedback, structured interaction, and memory for building robust and generalizable cybersecurity agents. View details
Preview abstract Artificial intelligence is rapidly evolving, marked by the emergence of Large Language Model (LLM) agents – systems capable of complex reasoning, planning, and interaction with digital and physical environments. These agents, powered by advancements in LLMs, demonstrate remarkable capabilities across diverse domains, including finance, healthcare, web navigation, software development, and daily task assistance. Unlike traditional AI systems, LLM agents can perceive their surroundings, formulate multi-step plans, utilize external tools and APIs, access memory or knowledge bases, and execute actions to achieve specified goals. This ability to act upon the world, however, introduces significant safety and security challenges. The safety paradigms developed for traditional LLMs, primarily focused on mitigating harmful textual outputs (e.g., toxicity, bias), are insufficient for safeguarding LLM agents. Agents interacting with dynamic environments and executing actions present a broader attack surface and new categories of risk. These include performing unsafe operations, violating privacy constraints through improper data handling or access control failures, deviating from user objectives (task misalignment), and susceptibility to novel manipulation techniques like indirect prompt injection and memory poisoning. Ensuring the trustworthy operation of these powerful agents is paramount, especially as they are integrated into high-stakes applications. To address this critical challenge, we introduce VeriGuard, a novel framework designed to enhance the safety and reliability of LLM agents by interactively verifying their policies and the actions. VeriGuard integrates a verification module that intercepts code-based actions proposed by the agent. In the first step, VeriGuard will generates and verifies the policies. The policies are rigorously checked against a set of predefined safety and security specifications Then each action will be verified to make sure it will align with the agent specification. This interactive verification loop ensures that the agent's behavior remains within safe operational bounds, effectively preventing the execution of harmful or unintended operations. By verifying each step, VeriGuard provides a robust safeguard, substantially improving the trustworthiness of LLM agents in complex, real-world environments. View details
Preview abstract Temporally consistent long video generation remains a fundamental challenge. Existing methods suffer from feature drift, where entities and environments gradually change unintentionally, or content collapse, where narratives fail to progress meaningfully. We introduce A${^2}$RD, an agentic autoregressive video generation architecture that decouples creative synthesis from consistency by modeling consistency as a test-time objective. A${^2}$RD features segment-by-segment generation augmented with a novel Multimodal Video Memory (\memory{}) that tracks segment contexts and dynamics and Test-Time Scaling algorithms that verify and refine generation. For each segment, it operates in a Retrieve--Synthesize--Refine--Update (RSRU) loop: the agent retrieves relevant contexts, determines the segment generation mode (extrapolation or interpolation) adaptively, synthesizes boundary frames then video segment with refinements applied at both frame and video levels, and updates \memory{} for subsequent generation. We further develop LVbench-C, a challenging benchmark measuring long-horizon entities and environments evolving in non-linear transitions. Extensive experiments on public and LVbench-C benchmarks across one-, three-, and five-minute video generation demonstrate that A${^2}$RD generates significantly more consistent, meaningful videos than existing baselines. Human evaluations confirm strong consistency in characters, objects, and environments, with smooth motion and meaningful narrative progression. View details
Preview abstract AI agents equipped with tool-calling capabilities are susceptible to \emph{Indirect Prompt Injection} (IPI) attacks. In this attack scenario, malicious commands hidden within \emph{untrusted} content trick the agent into performing unauthorized actions. Existing defenses can reduce attack success but often suffer from the \emph{over-defense dilemma}: they deploy expensive, \emph{always-on} sanitization that degrades utility and latency even in benign scenarios. We revisit IPI through an operational causal lens: a successful injection manifests as a \emph{grounding collapse} where the user request no longer provides decisive support for the agent's privileged action, while a particular untrusted segment provides disproportionate marginal support. Based on this signature, we propose \texttt{CausalArmor}, a selective defense framework that (i) computes lightweight, normalized leave-one-out attributions at privileged decision points, and (ii) triggers targeted sanitization only when an untrusted segment dominates the user intent. Additionally, CausalArmor employs \emph{retroactive Chain-of-Thought masking} to prevent the agent from acting on ``poisoned" reasoning traces. Experiments on AgentDojo and DoomArena demonstrate that CausalArmor matches the security of aggressive defenses with explainability while preserving utility and latency of AI agents. View details
Preview abstract Time-series forecasting has traditionally been evaluated solely on numerical accuracy, treating models as "black boxes'' that fail to capture the underlying reasoning. To address this gap, we introduce TFRBench, a novel benchmark designed to evaluate the reasoning capabilities of forecasting systems alongside their numerical accuracy. Unlike existing benchmarks, TFRBench requires models to generate verifiable natural language reasoning by analyzing cross-channel dependencies, identifying strategic trends, and justifying significant events using external context. To construct this benchmark, we propose a systematic multi-agent framework comprising Reasoning, Search, Verifier, Forecasting, and Summary agents. Our benchmark spans five diverse domains including Energy, Sales, Web/CloudOps, Transportation, and Finance, covering 10 distinct datasets. Qualitative evaluation confirms that our generated reasoning is highly faithful and effective; specifically, Large Language Models (LLMs) prompted with our generated reasoning demonstrate significantly improved forecasting accuracy compared to direct forecasting with LLMs. Conversely, benchmarking experiments reveal that off-the-shelf LLMs consistently struggle with both reasoning (shows lower LLM-as-Judge scores) and direct numerical forecasting (MAE and MASE), frequently failing to capture domain-specific dynamics. TFRBench thus establishes a new standard for interpretable, reasoning-based evaluation in time-series forecasting. View details
Preview abstract Recent knowledge distillation (KD) research made significant progress on improving smaller student models to match larger teachers' performances. Two noticeable methods, supervised KD and on-policy KD emerged as the state-of-the-art approaches. However, supervised KD for auto-regressive models suffers from distribution mismatch between training over fixed dataset and inference over student generated outputs. Conversely, on-policy KD, which uses student-generated samples for training, can suffer from low-quality training examples and the teacher's potential inaccuracies in assessing these samples. To address these limitations, we introduce Speculative Knowledge Distillation (SKD). Instead of solely training on teacher- or student-proposed samples, SKD leverages the student model to initially propose tokens following its own generation distribution. Subsequently, the teacher model is employed to replace tokens that are deemed out-of-distribution. Compared with supervised KD, the samples generated by SKD are more likely to align with the student's inference-time distribution, and 2) SKD can mitigate the generation of low-quality sequences by incorporating the teacher's feedback at each token. Furthermore, we demonstrate that SKD is a generic framework capable of implementing both supervised and on-policy knowledge distillation as specific instances. To validate SKD's effectiveness, we apply it to distill autoregressive large language models for various tasks, including translation, summarization, math, and instruction following. Our experiments consistently demonstrate SKD's superior performance compared to existing methods across different domains, tasks, data sizes, and model initialization strategies. View details
Preview abstract Artificial intelligence is rapidly evolving, marked by the emergence of Large Language Model (LLM) agents – systems capable of complex reasoning, planning, and interaction with digital and physical environments. These agents, powered by advancements in LLMs, demonstrate remarkable capabilities across diverse domains, including finance, healthcare, web navigation, software development, and daily task assistance. Unlike traditional AI systems, LLM agents can perceive their surroundings, formulate multi-step plans, utilize external tools and APIs, access memory or knowledge bases, and execute actions to achieve specified goals. This ability to act upon the world, however, introduces significant safety and security challenges. The safety paradigms developed for traditional LLMs, primarily focused on mitigating harmful textual outputs (e.g., toxicity, bias), are insufficient for safeguarding LLM agents. Agents interacting with dynamic environments and executing actions present a broader attack surface and new categories of risk. These include performing unsafe operations, violating privacy constraints through improper data handling or access control failures, deviating from user objectives (task misalignment), and susceptibility to novel manipulation techniques like indirect prompt injection and memory poisoning. Ensuring the trustworthy operation of these powerful agents is paramount, especially as they are integrated into high-stakes applications. To address this critical challenge, we introduce VeriGuard, a novel framework designed to enhance the safety and reliability of LLM agents by interactively verifying their policies and the actions. VeriGuard integrates a verification module that intercepts code-based actions proposed by the agent. In the first step, VeriGuard will generates and verifies the policies. The policies are rigorously checked against a set of predefined safety and security specifications Then each action will be verified to make sure it will align with the agent specification. This interactive verification loop ensures that the agent's behavior remains within safe operational bounds, effectively preventing the execution of harmful or unintended operations. By verifying each step, VeriGuard provides a robust safeguard, substantially improving the trustworthiness of LLM agents in complex, real-world environments. View details
Preview abstract Recent advances in knowledge distillation (KD) have enabled smaller student models to approach the performance of larger teacher models. However, popular methods such as supervised KD and on-policy KD, are adversely impacted by the knowledge gaps between teacher-student in practical scenarios. Supervised KD suffers from a distribution mismatch between training with a static dataset and inference over final student-generated outputs. Conversely, on-policy KD, which uses student-generated samples for training, can suffer from low-quality training examples with which teacher models are not familiar, resulting in inaccurate teacher feedback. To address these limitations, we introduce Speculative Knowledge Distillation (SKD), a novel approach that leverages cooperation between student and teacher models to generate high-quality training data on-the-fly while aligning with the student’s inference-time distribution. In SKD, the student proposes tokens, and the teacher replaces poorly ranked ones based on its own distribution, transferring high-quality knowledge adaptively. We evaluate SKD on various text generation tasks, including translation, summarization, math, and instruction following, and show that SKD consistently outperforms existing KD methods across different domains, data sizes, and model initialization strategies View details
×