Robotics · AI/ML · Software

Jaiveer Bassi

I build and audit intelligent systems, from robot data pipelines and fault injection rigs to reproducible machine learning research.

I am an AI Robot Operator at Physical Intelligence and Founder & Executive Director of House of Seva. I hold an M.S. in Software Engineering and a B.S. in Computer Science.

Open to full time robotics software, AI/ML, and software engineering opportunities.

Portrait of Jaiveer Bassi

Current focus

Research

Reproducible work with the claim, evidence, controls, and limitations kept together.

Published HumanEval scores compared with the primary local pipeline and the no-newline diagnostic condition
Technical report v1.9Code models · Evaluation pipelines

One Byte, One Rank Reversal: Auditing HumanEval Pipeline Sensitivity for Base Code Models

A pinned five model reproduction on an RTX 4070 found StarCoder2 3B scoring 3 of 164 instead of its published 31.7 percent, reversing a published ranking. Removing one trailing newline from the evaluation prompt raised it to 49 of 164 and restored the order.

FP32 and INT8 latency distributions with backend specific speedups
Archived preprintEdge AI · Host CPU inference

When Does INT8 Actually Accelerate Host-CPU Inference? A Reproducible Audit of MobileNetV2 Conversion and Packaging

Shows that INT8 latency depends on backend delegation and quantization granularity, alongside paired accuracy and packaging controls. Version 1.2.0 is archived on Zenodo and has not been submitted for peer review.

Directional PGD transfer rates across budgets and model pairs
Under review at TMLRAdversarial ML · CIFAR 10

Conditioning and Directionality in Adversarial Transfer on CIFAR-10

A denominator aware full test study across three architectures and three training seeds. The manuscript is under review at Transactions on Machine Learning Research, with an archived reproducibility release.

Nearly identical MRI scans found across benchmark test and training folders
PreprintMedical imaging · Benchmark audit

Complete Patient-Level Leakage Among Traceable Images in a Widely Used Brain Tumor MRI Benchmark

Finds complete patient overlap among traceable tumor test images plus substantial duplicate contamination, and releases patient disjoint folds. Available as a preprint and Zenodo artifact; not yet submitted to a journal.

Closed loop success and open loop error across five demonstration failure modes
Under review at CoRL 2026 LfCRobot learning · Data quality

Not All Bad Demonstrations Are Equally Bad: Quantifying How Demonstration Failure Modes Degrade Closed-Loop Policy Performance

A controlled study of 660 policy fits across five failure modes, two policy families, ten seeds, and transition matched controls. Under review at the CoRL 2026 Learning from Corrections and Interventions workshop.

Trajectory smoothness metrics under episode length controls
Three workshop submissionsRobot learning · Demonstration curation

Smoothness Ranks Skill, Not Success: An Audit of Published Trajectory-Smoothness Curation Metrics Under Episode-Length Controls

Base preprint plus three venue specific workshop submissions under review at CoRL WEBP, NeurIPS RoboPAD, and CoRL Oops, I Erred.

Credentials

Certifications

Selected technical and professional credentials.

Timeline

Experience & education

Stanford University

Undergraduate Research Fellow

Stanford University · September 2026 to present

Physical Intelligence

AI Robot Operator

Physical Intelligence · May 2026 to present

House of Seva

Founder & Executive Director

House of Seva · June 2026 to present

Handshake AI

LLM Response Evaluator

Handshake AI · April to June 2026

Kaiser Permanente

Software Developer

Kaiser Permanente · October 2023 to June 2024

M.S., Software Engineering

Grand Canyon University · 2024 to 2025

University of Silicon Valley

B.S., Computer Science

University of Silicon Valley · 2021 to 2023

Build, break, measure

Selected projects

Robotics infrastructure, fault characterization, ML tooling, and production software. More on my GitHub.

Live π0 FAST shadow inference dashboard for the four degree of freedom teleoperation arm

In progress · π0 FAST · LoRA · RTX 4070

π0 on a Budget

Fine-tuning π0-FAST on a consumer RTX 4070 using synchronized demonstrations from a custom four-degree-of-freedom teleoperation arm, followed by open-loop and closed-loop evaluation. The live camera, Arduino telemetry, RTX policy server, and real-time shadow-inference dashboard are operational; real demonstration collection and policy fine-tuning are next.

GitHub →
Robot telemetry pipeline dashboard

Kubernetes · Redis Streams · PostgreSQL

Fault Tolerant Robot Telemetry Pipeline

A multi stage pipeline designed for deliberate pod and network failures, with durable acknowledgements, replay safe writes, health probes, observability, and a chaos harness.

GitHub →
UR5e RTDE harness terminal demonstration

URSim · RTDE · Fault injection

UR5e RTDE Harness

Simulator backed control and synchronized telemetry with reproducible tests for connection failures, stale handles, hangs, and native crashes.

GitHub →
Low cost leader follower teleoperation rig

Article · Teleoperation · Robotics

A $30 Teleoperation Rig

Rebuilding the core idea behind GELLO style robot demonstration collection with direct joint space control.

Read on Medium →
Fleet Triage robot health dashboard

Python · Streamlit · Diagnostics

Fleet Triage

Robot fleet log analysis with fault classification, recurrence detection, anomaly detection, rule based root cause analysis, and a live health dashboard.

GitHub →
Robot grasp annotation workflow

Computer vision · Labelbox · COCO

Grasp Annotation Pipeline

A data operations workflow for robot arm grasp annotations with bounding boxes, keypoints, failure labels, and automated COCO JSON export.

GitHub →
AI Model Packager command line demonstration

Python · Docker · MLOps

AI Model Packager

A library and CLI for packaging machine learning models into portable, optimized Docker inference images.

GitHub →
Chess reinforcement learning interface

C++ · PyTorch · Self play

Chess RL Engine

An AlphaZero inspired system combining a C++ bitboard engine, PyTorch neural network, self play training, and a PyGame interface.

GitHub →
CAN bus fault diagnostics rig

Article · CAN bus · Diagnostics

Simulating Real Robot Failures on a Breadboard

A three node CAN bus built to induce and diagnose failures that software logs alone cannot explain.

Read on Medium →
Live MedStract biomedical literature analysis homepage

Python · PubMed · NLP

MedStract

A biomedical literature tool for searching PubMed, rewriting summaries for different audiences, exporting citations, and generating PDFs.

Live site →
Castline Studio order interface

Next.js · TypeScript · WebAssembly

Castline Studio Order System

An e commerce workflow with client side STL analysis, Etsy and shipping integrations, automated email, and deployment on Vercel.

Live site →
Live CareerTuner resume analysis homepage

OpenAI · React · Node.js · Supabase

CareerTuner

An AI resume analyzer with job specific feedback, ATS scoring, skill gap analysis, and PDF export.

Live site →

Beyond work

Hobbies

Reading and philosophy

I read philosophy, theology, science, and the classics. Add me on Goodreads to see what I am reading, and follow my writing on Medium.

Building keyboards

I enjoy building custom keyboards and plan to release projects and articles about them whenever possible. Stay tuned.

3D printing

I design, print, and refine physical pieces through Castline Studio, combining digital fabrication with hands on experimentation.