Robotics · AI/ML · Software

Jaiveer Bassi

I build and audit intelligent systems, from robot data pipelines and fault injection rigs to reproducible machine learning research.

I am an AI Robot Operator at Physical Intelligence and Founder & Executive Director of House of Seva. I hold an M.S. in Software Engineering and a B.S. in Computer Science.

Open to full time robotics software, AI/ML, and software engineering opportunities.

Portrait of Jaiveer Bassi

Current focus

Research

Reproducible work with the claim, evidence, controls, and limitations kept together.

Published HumanEval scores compared with the primary local pipeline and the no-newline diagnostic condition
Technical report v1.9Code models · Evaluation pipelines

One Byte, One Rank Reversal: Auditing HumanEval Pipeline Sensitivity for Base Code Models

A pinned five model reproduction on an RTX 4070 found StarCoder2 3B scoring 3 of 164 instead of its published 31.7 percent, reversing a published ranking. Removing one trailing newline from the evaluation prompt raised it to 49 of 164 and restored the order.

FP32 and INT8 latency distributions with backend specific speedups
Archived preprintEdge AI · Host CPU inference

When Does INT8 Actually Accelerate Host-CPU Inference? A Reproducible Audit of MobileNetV2 Conversion and Packaging

Shows that INT8 latency depends on backend delegation and quantization granularity, alongside paired accuracy and packaging controls. Version 1.2.0 is archived on Zenodo and has not been submitted for peer review.

Directional PGD transfer rates across budgets and model pairs
Under review at TMLRAdversarial ML · CIFAR 10

Conditioning and Directionality in Adversarial Transfer on CIFAR-10

A denominator aware full test study across three architectures and three training seeds. The manuscript is under review at Transactions on Machine Learning Research, with an archived reproducibility release.

Nearly identical MRI scans found across benchmark test and training folders
PreprintMedical imaging · Benchmark audit

Complete Patient-Level Leakage Among Traceable Images in a Widely Used Brain Tumor MRI Benchmark

Finds complete patient overlap among traceable tumor test images plus substantial duplicate contamination, and releases patient disjoint folds. Available as a preprint and Zenodo artifact; not yet submitted to a journal.

Closed loop success and open loop error across five demonstration failure modes
Under review at CoRL 2026 LfCRobot learning · Data quality

Not All Bad Demonstrations Are Equally Bad: Quantifying How Demonstration Failure Modes Degrade Closed-Loop Policy Performance

A controlled study of 660 policy fits across five failure modes, two policy families, ten seeds, and transition matched controls. Under review at the CoRL 2026 Learning from Corrections and Interventions workshop.

Trajectory smoothness metrics under episode length controls
Accepted at NeurIPS 2026 RoboPADRobot learning · Demonstration curation

Curation Metrics Are a Post-Training Decision: Auditing Trajectory Smoothness for Adapting Robot Foundation Models

Accepted as a poster at the NeurIPS 2026 RoboPAD workshop. It builds on the base preprint, “Smoothness Ranks Skill, Not Success: An Audit of Published Trajectory-Smoothness Curation Metrics Under Episode-Length Controls,” which shows that published smoothness curation metrics mainly rank operator skill rather than task success; the CoRL WEBP and CoRL Oops, I Erred versions remain under review.

Build, break, measure

Selected projects

Robotics infrastructure, fault characterization, ML tooling, and production software. More on my GitHub.

Live π0 FAST shadow inference dashboard for the four degree of freedom teleoperation arm

π0 FAST · LoRA · RTX 4070

π0 on a Budget

Fine-tuning π0-FAST on a consumer RTX 4070 using synchronized demonstrations from a custom four-degree-of-freedom teleoperation arm, followed by open-loop and closed-loop evaluation. The live camera, Arduino telemetry, RTX policy server, and real-time shadow-inference dashboard are operational; real demonstration collection and policy fine-tuning are next.

GitHub →
Closed-loop rollouts of policies trained at increasing demonstration contamination

Robot learning · Imitation · Simulation

Demonstration Quality Robustness

A controlled simulation study of 660 policy fits measuring how specific demonstration failure modes degrade closed-loop policy performance, and how poorly open-loop evaluation tracks that damage.

Simulated UR5e running a waypoint cycle with live joint angle telemetry

URSim · RTDE · Fault injection

UR5e RTDE Harness

Drives a simulated UR5e over RTDE, the same interface used by production UR5e cells, logs synchronized joint telemetry, and deliberately induces connection faults to characterize whether the client fails with a clean exception, a hang, or a native crash.

GitHub →
Same simulated scene, different language instructions, different robot behavior

openpi · π0.5 · LIBERO

openpi Language Steerability Demo

Runs the public π0.5 LIBERO checkpoint in a single simulated scene with different language instructions and scores success with the simulator's own goal checks. Trained and paraphrased prompts succeed 10 of 10 times, while novel object recombinations drop as low as 1 of 10.

GitHub →
Simulated two link arm grasp with bounding box, contact keypoint, and failure mode overlays

Computer vision · Labelbox · COCO

Grasp Annotation Pipeline

Samples public robot arm video into frames, provisions a Labelbox project with a grasp-focused ontology, and exports completed labels as COCO JSON. Each frame can carry a grasp_event bounding box, an object_contact keypoint, and a failure_mode label: drop, miss, collision, timeout, or success.

GitHub →
Low cost leader follower teleoperation rig

Article · Teleoperation · Robotics

A $30 Teleoperation Rig

Rebuilding the core idea behind GELLO style robot demonstration collection with direct joint space control.

Chess RL self-play games and training progress

C++ · PyTorch · Self play

Chess RL Engine

An AlphaZero inspired engine that learns only through self play: a C++ bitboard move generator exposed through pybind11, a PyTorch policy and value network on an 18 channel board encoding, and a PyGame interface showing move probabilities in real time.

GitHub →
Animated map of an MLP learning to classify shapes, showing activations, weights, and dropout

PyTorch · Matplotlib · Visualization

Neural Net Mapper

Trains an MLP on a synthetic shapes dataset and renders an animated map of its inner workings: neuron activations, weight signs and magnitudes, dropout, predictions, and live loss and accuracy curves.

GitHub →
CAN bus fault diagnostics rig

Article · CAN bus · Diagnostics

Simulating Real Robot Failures on a Breadboard

A three node CAN bus built to induce and diagnose failures that software logs alone cannot explain.

MedStract homepage with PubMed literature search

Flask · PubMed · NLP

MedStract

Retrieves peer-reviewed abstracts from PubMed and turns them into summaries calibrated to four reading levels, from the general public to the domain expert, with question answering over the results, publication-trend charts, and citation exports in six formats.

Live site →
Castline Studio order interface

Next.js · TypeScript · WebAssembly

Castline Studio Order System

An e commerce workflow with client side STL analysis, Etsy and shipping integrations, automated email, and deployment on Vercel.

Live site →

Timeline

Experience & education

Stanford University

Undergraduate Research Fellow

Stanford University · September 2026 to present

Physical Intelligence

AI Robot Operator

Physical Intelligence · May 2026 to present

House of Seva

Founder & Executive Director

House of Seva · June 2026 to present

Handshake AI

LLM Response Evaluator

Handshake AI · April to June 2026

Kaiser Permanente

Software Developer

Kaiser Permanente · October 2023 to June 2024

M.S., Software Engineering

Grand Canyon University · 2024 to 2025

University of Silicon Valley

B.S., Computer Science

University of Silicon Valley · 2021 to 2023

Credentials

Certifications

Selected technical and professional credentials.

Beyond work

Hobbies

Reading and philosophy

I read philosophy, theology, science, and the classics. Add me on Goodreads to see what I am reading, and follow my writing on Medium.

Building keyboards

I enjoy building custom keyboards and plan to release projects and articles about them whenever possible. Stay tuned.

3D printing

I design, print, and refine physical pieces through Castline Studio, combining digital fabrication with hands on experimentation.