Joshua Terranova
Back to highlights

USC Viterbi

LiraLab · Undergraduate Researcher

2026–Present

LiraLab wordmark: Learning and Interactive Robot Autonomy Lab
Photo by LiraLab on Unsplash
Research

Undergraduate researcher in the Learning and Interactive Robot Autonomy Lab (Prof. Erdem Bıyık): robot learning, preference-based reward learning, and human-robot interaction.

I joined USC's Learning and Interactive Robot Autonomy Lab (LiraLab) under Prof. Erdem Bıyık in August 2026 as an undergraduate researcher.

The lab builds algorithms for robot learning, safe and efficient human-robot interaction, and multi-agent systems. My first public deliverable is Experiment 1: PEBBLE causal instrumentation on Walker-walk under Mistake teachers, run on USC CARC.

Journal

What I worked on during 2026–Present. Hover underlined terms for quick definitions.

Joining · Lab agenda

LiraLabUSC Learning and Interactive Robot Autonomy Lab (liralab.usc.edu), PI Erdem Bıyık develops robot learning methods that use explicit feedback (demonstrations, pairwise comparisons) and more implicit cues so agents can align with human goals and preferences.

I started on that agenda in August 2026. Agents that model other agents from demonstrations, comparisons, gaze, language, and related feedback, then adapt online without breaking alignment.

Prep · Preference learning and offline RL

Before joining, I built an independent preference-based reward learning pipeline on a simulated differential-drive point-goal task, with a ROS 2Robot Operating System 2 / Gazebo embodiment of the same layout: Bradley-Terry pairwise preference loss, neural reward ensembles, uncertainty via ensemble variance, then CQLConservative Q-Learning: offline RL that penalizes overestimation on out-of-distribution actions / behavior cloning on relabeled rewards.

That loop (preferences → learned reward → conservative policy) is the skill stack I brought into LiraLab, not a LiraLab paper. Field autonomy work at USC Field Robotics Lab (URC 2026) remains the closed-loop perception and control thread alongside this.

Experiment 1 · Accuracy decay vs Goodhart

When PEBBLEPreference-Based learning via Ensemble Bootstrapping and Learning (Lee, Smith & Abbeel, ICML 2021) degrades under noisy preference labels, standard B-Pref logs mostly final return, so reward-model accuracy decay and SAC overoptimization look the same. Experiment 1 adds measurements that separate those mechanisms.

PEBBLE updates the reward model only at discrete num_interactInteraction interval: preference queries and reward-model updates fire every N environment steps events, then relabels the replay buffer. Between updates the reward model is frozen. Hypothesis H1: collapse from reward-model quality failure under label noise. Hypothesis H2: SAC overoptimizes a roughly constant reward error (GoodhartWhen a proxy objective keeps improving while the true objective stagnates or falls).

Experiment 1 · Instrumentation and CARC protocol

I patched B-Pref PEBBLE with frozen-window tracking and four preference-accuracy sets (train, held-out 20%, on-policy, fixed oracle reference), then ran a Walker-walk Mistake sweep (ε ∈ {0, 0.05, 0.1, 0.2}) for 500k steps on USC CARC under biyik_1165 (max_feedback=1000, num_interact=20000). Inside each frozen window I track true return, proxy return from frozen r_hatLearned reward model used as the SAC training signal, and on-policy accuracy with clean oracle labels.

Training still uses the configured Mistake teachersB-Pref synthetic preference teachers that flip labels with probability ε; metric labels always use a clean oracle path. One seed per ε on this first sweep. Local Apple Silicon was recon only.

Experiment 1 · Results

End-of-training snapshot (last ~20 train episodes; last RM update near 190k when feedback hits 1000): at ε=0, true return ~880 with held-out accuracy ~0.87 and fixed-ref ~0.85. At ε=0.2, true return collapses to ~44 while held-out falls to ~0.60 and fixed-ref to ~0.64. Train-fit accuracy stays ~0.98 even at ε=0.2: the model fits the noisy labels.

Within-window Goodhart-like signatures (late-half proxy up by >5 while true down by >5) are rare: 0/10 windows at ε≤0.1 and 1/10 at ε=0.2. Primary reading: reward-model generalization failure under noise, not a clean Goodhart case. Caveat: single seed per ε, Walker only; multi-seed and a second domain are next.

Related highlights

All highlights