research AdSelect Fine-tuning multimodal LLMs for advertisement shot selection — a set-prediction approach to editing ads like professionals AdShot MLLM benchmark for advertisement video clipping — evaluating temporal reasoning at production scale Robin All in-house voice assistant for AI-CARING at Northeastern's PARCS Lab — runs on any tablet in the home, "Hey Robin" wake word, a Parakeet → LLM → Kokoro cascade on a self-hosted GPU server, joined by one WebSocket DepthDiT Diffusion Transformer for monocular depth estimation — adapting PixArt-Alpha and PixArt-Sigma for dense prediction GVHMR Multi-Person Extension Extending global-frame human motion recovery to multi-person scenes for IARPA MOVES industry Production CV Deployments at Collablens Built and deployed computer vision systems for ITC freight monitoring and Smartivity defect detection VenuMatch Agentic venue matching system using LangGraph, ChromaDB, FastAPI, and React/Vite. Production Diffusion Pipelines at Dashtoon Flux LoRA adapters for 20+ characters, Hidream-l1 MoE integration, Bollywood bias mitigation, and Google Veo 2 evaluation coursework LiDAR-IMU Fusion Odometry Point-to-plane ICP fused with IMU preintegration in a 15-state error-state Kalman filter, scored against GPS ground truth on KITTI