<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://abraranwar.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://abraranwar.github.io/" rel="alternate" type="text/html" /><updated>2026-09-09T08:36:02+00:00</updated><id>https://abraranwar.github.io/feed.xml</id><title type="html">Abrar Anwar</title><subtitle>Abrar Anwar is a robotics researcher at NVIDIA on the Isaac Loco-Manipulation team, working on reinforcement learning post-training of robot foundation models. Previously a PhD student at USC with Jesse Thomason on language-guided robot learning and evaluation.</subtitle><entry><title type="html">Anticipating Unintended Robot Behaviors with Side Effect Critics</title><link href="https://abraranwar.github.io/side-effects.html" rel="alternate" type="text/html" title="Anticipating Unintended Robot Behaviors with Side Effect Critics" /><published>2026-09-09T00:00:00+00:00</published><updated>2026-09-09T00:00:00+00:00</updated><id>https://abraranwar.github.io/side-effects</id><content type="html" xml:base="https://abraranwar.github.io/side-effects.html"><![CDATA[<p>We introduce the concept of Side Effect Critics to anticipate and mitigate unintended robot behaviors when deploying learned policies in real-world environments.</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[We introduce the concept of Side Effect Critics to anticipate and mitigate unintended robot behaviors when deploying learned policies in real-world environments.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/side_effects.png" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/side_effects.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations</title><link href="https://abraranwar.github.io/steering.html" rel="alternate" type="text/html" title="Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations" /><published>2026-09-01T00:00:00+00:00</published><updated>2026-09-01T00:00:00+00:00</updated><id>https://abraranwar.github.io/steering</id><content type="html" xml:base="https://abraranwar.github.io/steering.html"><![CDATA[<p>Robotic Steering is a mechanistic, few-shot finetuning approach for vision-language-action models that selectively adapts task-specific attention heads, achieving more robust, efficient, and interpretable robot learning than LoRA across diverse tasks.</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[Robotic Steering is a mechanistic, few-shot finetuning approach for vision-language-action models that selectively adapts task-specific attention heads, achieving more robust, efficient, and interpretable robot learning than LoRA across diverse tasks.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/steering.png" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/steering.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration</title><link href="https://abraranwar.github.io/safecast.html" rel="alternate" type="text/html" title="SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration" /><published>2026-08-04T00:00:00+00:00</published><updated>2026-08-04T00:00:00+00:00</updated><id>https://abraranwar.github.io/safecast</id><content type="html" xml:base="https://abraranwar.github.io/safecast.html"><![CDATA[<p>VLA policies fail under deployment-time shifts like clutter, lighting changes, and reworded instructions. SAFECAST uses contrast set perturbations to train and calibrate hidden-state risk probes, improving failure detection ROC-AUC over state-of-the-art baselines on both real-world DROID and LIBERO simulation.</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[VLA policies fail under deployment-time shifts like clutter, lighting changes, and reworded instructions. SAFECAST uses contrast set perturbations to train and calibrate hidden-state risk probes, improving failure detection ROC-AUC over state-of-the-art baselines on both real-world DROID and LIBERO simulation.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/safecast.png" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/safecast.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents</title><link href="https://abraranwar.github.io/dips.html" rel="alternate" type="text/html" title="DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents" /><published>2026-07-02T00:00:00+00:00</published><updated>2026-07-02T00:00:00+00:00</updated><id>https://abraranwar.github.io/dips</id><content type="html" xml:base="https://abraranwar.github.io/dips.html"><![CDATA[<p>DiPS is a Q-learning framework that dynamically selects persuasion strategies based on the evolving conversational context. In a fire-rescue evacuation scenario, DiPS achieves higher persuasion success than zero-shot LLM and RAG-augmented approaches in both simulated and real human interactions.</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[DiPS is a Q-learning framework that dynamically selects persuasion strategies based on the evolving conversational context. In a fire-rescue evacuation scenario, DiPS achieves higher persuasion success than zero-shot LLM and RAG-augmented approaches in both simulated and real human interactions.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/dips.png" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/dips.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons</title><link href="https://abraranwar.github.io/robometer.html" rel="alternate" type="text/html" title="Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons" /><published>2026-06-01T00:00:00+00:00</published><updated>2026-06-01T00:00:00+00:00</updated><id>https://abraranwar.github.io/robometer</id><content type="html" xml:base="https://abraranwar.github.io/robometer.html"><![CDATA[<p>Robometer is a general-purpose, video-language-input dense reward model trained on RBM-1M, a dataset of over 1M trajectories spanning 21 robot embodiments. It improves robot learning across online RL, offline RL, model-based RL, failure detection, and data retrieval for imitation learning.</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[Robometer is a general-purpose, video-language-input dense reward model trained on RBM-1M, a dataset of over 1M trajectories spanning 21 robot embodiments. It improves robot learning across online RL, offline RL, model-based RL, failure detection, and data retrieval for imitation learning.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/robometer.png" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/robometer.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning</title><link href="https://abraranwar.github.io/isaaclab.html" rel="alternate" type="text/html" title="Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning" /><published>2025-11-06T00:00:00+00:00</published><updated>2025-11-06T00:00:00+00:00</updated><id>https://abraranwar.github.io/isaaclab</id><content type="html" xml:base="https://abraranwar.github.io/isaaclab.html"><![CDATA[<p>I helped introduce approaches for VLA post-training with IsaacLab!</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[I helped introduce approaches for VLA post-training with IsaacLab!]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/isaaclab.png" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/isaaclab.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">RobotFleet: An Open-Source Framework for Centralized Multi-Robot Task Planning</title><link href="https://abraranwar.github.io/robotfleet.html" rel="alternate" type="text/html" title="RobotFleet: An Open-Source Framework for Centralized Multi-Robot Task Planning" /><published>2025-10-12T00:00:00+00:00</published><updated>2025-10-12T00:00:00+00:00</updated><id>https://abraranwar.github.io/robotfleet</id><content type="html" xml:base="https://abraranwar.github.io/robotfleet.html"><![CDATA[<p>RobotFleet is an open-source, extensible framework that introduces a centralized, modular autonomy stack to simplify scalable planning, scheduling, and execution for heterogeneous multi-robot fleets in open-world tasks.</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[RobotFleet is an open-source, extensible framework that introduces a centralized, modular autonomy stack to simplify scalable planning, scheduling, and execution for heterogeneous multi-robot fleets in open-world tasks.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/robotfleet.png" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/robotfleet.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection</title><link href="https://abraranwar.github.io/seq-eval.html" rel="alternate" type="text/html" title="Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection" /><published>2025-08-28T00:00:00+00:00</published><updated>2025-08-28T00:00:00+00:00</updated><id>https://abraranwar.github.io/seq-eval</id><content type="html" xml:base="https://abraranwar.github.io/seq-eval.html"><![CDATA[<p>The space of language commands a robot can execute grows combinatorially with scene complexity. Evaluating a robot on this large domain is impractical + takes time, so we introduce contrast sets for robots to make small perturbations to test instances. This leads to good test set estimation and less experimenter effort.</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[The space of language commands a robot can execute grows combinatorially with scene complexity. Evaluating a robot on this large domain is impractical + takes time, so we introduce contrast sets for robots to make small perturbations to test instances. This leads to good test set estimation and less experimenter effort.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/seq_eval.jpg" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/seq_eval.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations</title><link href="https://abraranwar.github.io/rewind.html" rel="alternate" type="text/html" title="ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations" /><published>2025-08-28T00:00:00+00:00</published><updated>2025-08-28T00:00:00+00:00</updated><id>https://abraranwar.github.io/rewind</id><content type="html" xml:base="https://abraranwar.github.io/rewind.html"><![CDATA[<p>We design ReWiND rewards that use language-guided rewards to train bimanual arms on OOD tasks in 1 hour!
We use offline-to-online, lang-conditioned, visual RL on action-chunked transformers on a real robot and in simulation!</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[We design ReWiND rewards that use language-guided rewards to train bimanual arms on OOD tasks in 1 hour! We use offline-to-online, lang-conditioned, visual RL on action-chunked transformers on a real robot and in simulation!]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/rewind.png" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/rewind.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation</title><link href="https://abraranwar.github.io/remembr.html" rel="alternate" type="text/html" title="ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation" /><published>2025-05-20T00:00:00+00:00</published><updated>2025-05-20T00:00:00+00:00</updated><id>https://abraranwar.github.io/remembr</id><content type="html" xml:base="https://abraranwar.github.io/remembr.html"><![CDATA[<p>To tackle long-horizon spatio-temporal memory, we introduce a retrieval-based approach for building and reasoning over memory. 
We show improved performance on planning and embodied question answering given a long video history.</p>]]></content><author><name></name></author><category term="items" /><category term="research" /><summary type="html"><![CDATA[To tackle long-horizon spatio-temporal memory, we introduce a retrieval-based approach for building and reasoning over memory. We show improved performance on planning and embodied question answering given a long video history.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://abraranwar.github.io/assets/images/index/remembr.jpg" /><media:content medium="image" url="https://abraranwar.github.io/assets/images/index/remembr.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>