Experience.

I have worked on model evaluation, analytics pipelines, production ML, and academic research.

3 languages

evaluation codebases

Python · C++ · Go

Docker

reproducible environments

Fail-to-pass test suites

Regression suites

validated fixes

Edge cases returned to evaluation

Mar 2026 to Present

Handshake AI

AI Research Fellow · Boston, Massachusetts

At Handshake AI, I write evaluation pipelines for Python, C++, and Go repositories. The work combines hard test design, multi-file debugging, and checking that fixes survive the full regression suite.

Evaluation pipelines
Develop LLM evaluation pipelines across Python, C++, and Go codebases, with Docker-based reproducible environments and fail-to-pass test suites that probe model behavior on production code.
Adversarial evaluation
Design prompts that surface weaknesses in long-context reasoning, multi-file edits, and tool use, then translate the findings into fine-tuning data and reusable prompt patterns.
Repository-scale debugging
Trace and resolve multi-file bugs in production-scale open-source repositories, validate fixes against regression suites, and return edge cases to evaluation tasks.
  • Python
  • C++
  • Go
  • Docker
  • LLM evaluation

Jun 2025 to Mar 2026

Amazon

Software Development Engineer · Boston, Massachusetts

At Amazon, I replaced weekly manual reporting work with a daily pipeline, added data-quality checks, and made common reporting questions self-service.

Analytics orchestration
Migrated manual analytics workflows to a Python- and Airflow-orchestrated pipeline on PostgreSQL, reducing weekly data preparation from 12 hours to under 2 and enabling daily reporting.
Data quality
Built modular Python validation with dbt-style tests to detect schema drift, anomalies, and outliers, eliminating recurring downstream report errors and improving training-data integrity.
Self-service reporting
Standardized organizational reporting through reusable SQL models and Tableau dashboards, replacing ad hoc analyst requests with a consistent self-service data layer.
  • Python
  • Airflow
  • PostgreSQL
  • SQL
  • Tableau
  • Data validation

Oct 2021 to Jun 2023

Accenture

Software Engineer · India

At Accenture, I built NLP, forecasting, churn, and recommendation models, then turned them into AWS services that could be versioned, tested, and monitored.

NLP systems
Designed end-to-end NLP pipelines with PyTorch, Hugging Face Transformers, and spaCy for text classification, named entity extraction, and intent recognition, shipping models into production applications.
Predictive ML
Built scikit-learn and XGBoost systems for forecasting, churn prediction, and recommendations, improving prediction accuracy by double digits over rule-based baselines.
Production ML
Productionized models on AWS with Docker, SageMaker, and FastAPI, adding versioning, A/B testing, and Prometheus monitoring to reduce iteration cycles from weeks to days.
  • PyTorch
  • Hugging Face
  • spaCy
  • scikit-learn
  • XGBoost
  • AWS
  • SageMaker
  • Docker
  • FastAPI
  • Prometheus

May 2020 to Nov 2021

KL University

Research Assistant · India

I supported academic research at KL University while completing my undergraduate studies.