About

I've spent 5+ years turning messy, real-world data into systems people rely on, across payments, healthcare, e-commerce, streaming, and logistics.

My work sits across three areas that usually live on different teams: experimentation, statistics, and causal inference that show what actually changed, production ML models with the monitoring to keep them healthy, and LLM and agentic systems that automate real workflows. I care most about the last step, getting a result into a decision someone makes.

Most of my days start in SQL. I write the queries and data models behind every analysis, then choose the statistical approach that fits the question: regression and Bayesian models, time series forecasts, a well-powered A/B test, or a causal design when randomization isn't possible.

60%less analysis time with an agentic copilot
55%fewer unhandled agent failures
45%fewer production regressions
30%lower measurement error in impact studies
50%faster clinical document processing
28%lower p95 inference latency

Selected work

Experiments and forecasts that teams actually use

Stripe
Problem
Product and operations teams wanted to test more ideas, but underpowered tests and slow, inconsistent readouts meant results often arrived too late or too vague to act on.
Approach
Standardized experiment design around a clear hypothesis, a primary metric with guardrails, and power and sample size set before launch. Built a GenAI-assisted tool that recommends test designs, interprets A/B results, and writes a plain-language summary. Alongside it, built Bayesian time series forecasts of key payments metrics in SQL, PyMC, and statsmodels.
Result
Product and operations teams adopted the experimentation tool, Finance and Strategy use the forecasts for planning and resource allocation, and readouts now go straight into roadmap and decision reviews.
Hypothesis Power + sample Run test Readout Decision GenAI assist: recommends designs, interprets results, drafts the summary
  • SQL
  • Experiment design
  • Power analysis
  • Bayesian forecasting
  • Tableau

An agentic copilot for financial operations

Stripe
Problem
Analysts spent hours reading filings, transaction logs, and dispute records to answer routine compliance questions, and a single-chain LLM setup broke midway too often to trust.
Approach
Designed a multi-agent system in LangGraph: a planner breaks the question down, an executor calls tools and retrieval, and a validator checks every answer before it returns, with deterministic routing and recovery when a step fails.
Result
Analysis time dropped 60% and unhandled agent failures fell 55% versus the single-chain version.
Question Planner Executor Validator Tools and RAG Retry or re-plan if a check fails
  • LangGraph
  • GPT-4
  • RAG
  • FastAPI
  • LangSmith

Measuring impact when you can't run an A/B test

Stripe
Problem
Some programs roll out to everyone at once, so there's no randomized control group, and simple before-and-after comparisons kept overstating impact.
Approach
Built a toolkit of causal methods, including Difference-in-Differences, Synthetic Control, and Double ML, that builds a credible counterfactual for what would have happened without the program.
Result
Measurement error fell 30% compared with the previous attribution approach, and leadership got estimates with honest uncertainty ranges that informed which programs to expand and which to rethink.
Program launch Estimated impact ObservedSynthetic control
  • Synthetic Control
  • Diff-in-Diff
  • Double ML
  • A/B testing
  • PyMC

Document intelligence for clinical records

CVS Health (Aetna)
Problem
Clinical notes, insurance documents, and medical reports were reviewed by hand, which was slow and hard to scale in a regulated environment.
Approach
Built a platform that extracts entities, classifies documents, and answers questions over records using RAG, with an evaluation framework tracking hallucination rate, accuracy, and latency before anything shipped.
Result
Processing time fell 50%, manual review dropped 35%, and reliability of AI-generated insights improved 30%.
Documents Extract Embed Retrieve + LLM Evaluate Hallucination, accuracy, and latency checked before release
  • LangChain
  • Hugging Face
  • FAISS
  • Pinecone
  • Azure AKS

Production fraud and risk models that stay healthy

Stripe
Problem
Fraud and risk models score high-volume payment data, and patterns shift constantly, so a model that looks great offline can quietly degrade in production.
Approach
Built classification and scoring models with XGBoost, LightGBM, and PyTorch, with feature engineering on transaction, account, and third-party vendor signals. Added calibrated scores, drift and data quality checks, automated evaluation before every release, and dashboards in Prometheus and Grafana.
Result
Production regressions fell 45% and incident detection got 50% faster, and risk and operations teams adopted the monitoring.
Features Train + tune Evaluate Deploy Monitor drift Retrain when drift or data quality checks fire
  • XGBoost
  • LightGBM
  • PyTorch
  • Feature engineering
  • MLOps

Making LLM inference faster and cheaper

Stripe
Problem
As usage grew, response times and token costs grew with it.
Approach
Added Redis caching for repeat requests, compressed prompts, and moved long-running work onto async task queues, with Prometheus and Grafana dashboards to watch the effect.
Result
p95 latency fell 28% and token usage fell 32% without losing output quality.
p95 latency BeforeAfter, -28% Token usage BeforeAfter, -32%
  • Redis
  • Async queues
  • Prompt compression
  • Prometheus
  • Grafana

Consulting across industries

Hexaware Technologies, client engagements

At Hexaware I worked with enterprise clients across e-commerce and marketplaces, streaming and consumer apps, fraud and trust and safety, logistics, and financial services. The work included recommendation and ranking models, pricing and promotion experiments, fraud and anomaly detection, and demand forecasting, usually presented straight to client leadership.

Improved model performance 25% through feature engineering and tuning, cut data processing latency 35% with Spark pipelines, and reduced manual reporting effort 70% with automated dashboards.

How I approach a problem

Whether the answer is an analysis, a model, or an AI system, I follow the same loop.

1. Frame

Start with the decision someone needs to make, then define the metric that reflects it, plus guardrail metrics that shouldn't get worse.

2. Get the data right

Write the SQL, check for missing data, leakage, and bias, and confirm the data can actually answer the question.

3. Choose the method

An A/B test with power and sample size set upfront when I can randomize, a causal design such as Difference-in-Differences or Synthetic Control when I can't, a forecast or ML model when the job is prediction, and an LLM system when the work is language-heavy.

4. Validate

Compare against a simple baseline, check calibration and robustness, and evaluate AI output against labeled examples before anything ships.

5. Drive the decision

Share a short recommendation with the uncertainty stated plainly, then keep monitoring after launch.

Experience

Jun 2025 to present

Data Scientist / ML Engineer, Stripe

  • Agentic LLM systems, RAG, and fine-tuned models for financial operations.
  • Experiment design, causal inference, and forecasting that Finance, Strategy, Product, and Operations teams use in planning and decision reviews.
  • Evaluation, monitoring, and inference optimization for production AI.
Jul 2024 to May 2025

AI/ML Engineer, CVS Health (Aetna)

  • Clinical document intelligence with NLP, RAG, and LLM evaluation.
  • A/B testing and MLOps in a HIPAA-regulated environment.
Oct 2020 to Aug 2023

Data Scientist, Hexaware Technologies

  • ML, experimentation, and analytics for enterprise clients across several industries.
Education

Master's in Computer Science, Saint Louis University

Skills

Data science

SQL (Snowflake, PostgreSQL, BigQuery, dbt), Python, R, statistics, hypothesis testing, regression, GLMs, Bayesian modeling (PyMC), statsmodels, A/B testing, experiment design, power analysis, causal inference, Difference-in-Differences, Synthetic Control, Double ML, propensity score matching, time series forecasting, product analytics, metrics and KPIs, Tableau, Looker, Power BI.

Machine learning

Scikit-learn, XGBoost, LightGBM, Random Forests, PyTorch, TensorFlow, feature engineering, classification, clustering, recommendation systems, ranking, anomaly detection, fraud detection, NLP, model evaluation, MLOps, MLflow, model monitoring, drift detection, Spark, PySpark, Airflow, Databricks, Docker, Kubernetes, CI/CD, AWS SageMaker, GCP Vertex AI, Azure ML.

AI systems

Generative AI, LLMs, GPT-4, Claude, RAG, Graph RAG, AI agents, multi-agent systems, LangChain, LangGraph, prompt engineering, fine-tuning (LoRA, QLoRA, PEFT), LLM evaluation, LangSmith, vector databases (Pinecone, FAISS, Qdrant, Weaviate), embeddings, vLLM, inference optimization, FastAPI.

Contact

I'm open to Data Scientist, Product Data Scientist, Applied Scientist, Machine Learning Engineer, and AI Engineer roles. The fastest way to reach me is email.

Email me Connect on LinkedIn Call +1 (314) 481-9645