AI research engineer · Boston

I test how language models behave on real code and build the tools around them.

At Handshake AI, I write evaluation pipelines in Python, C++, and Go. I spend much of my time designing hard cases, tracing failures across a repository, and checking whether a fix holds up. Outside work, I build tools for code review, debugging, and browser automation.

Research interests
  • Model evaluation on code
  • Repository-scale debugging
  • Reliable developer tools

Projects.

A code-review benchmark, a debugging CLI, and a browser agent.

AI evaluation · Active prototype

Code Review Arena

A benchmark that checks the patch, not just the critique.

Code Review Arena scores bug detection and repair separately, then applies each patch and runs the required tests and validators.

10-case benchmark run10 cases
keyword_gamerFinds the defects
Detection
1.000
Validated repair
0.000
reference.patchRepairs the cases
Validated
10 / 10
View results ↗
02

Developer tool

DebugBrief

Published on PyPI

A debugging session becomes a clean engineering handoff.

DebugBrief records commands, exit codes, timestamps, and repository changes, then turns that event history into a pull-request, handoff, incident, or JSON report.

PyPIpublished package
03

Browser agent

Helm Browser Agent

Local prototype

A browser agent that checks whether the task is actually done.

Helm turns an instruction into a plan, runs one browser action at a time, and checks the page before it says the task is done.

E2Eoffline WebSocket path
View all projects

Education.

Machine learning and analytics at Northeastern, following undergraduate study in electronics and data computing.

Northeastern UniversityBoston, Massachusetts · May 2025

M.S. in Data Analytics Engineering

Machine Learning concentration

KL UniversityApril 2022

B.Tech. in Electronics and Communication Engineering

Data Computing specialization