AI/ML Engineer Data Scientist Applied Researcher

Amrutha Ravikumar

I got into data through business, which means I've always had a face in mind for who reads the output.

Amrutha Ravikumar

About

I came into analytics through business. That background shaped how I work: I start from the problem someone needs to solve, figure out whether the data can actually answer it, and make sure the person on the other end can use what I hand them.

MS in Applied Machine Intelligence at Northeastern University (4.0 GPA). Over the past year I've led first-author ML research with a six-person cross-institutional team, built evaluation pipelines for nonprofits across Maine, and designed databases reconciling messy federal data across 72,000+ census tracts.

I work across the full arc of a data problem: scoping, cleaning, modeling, validating, and communicating what it means to the people who need to act on it.

0GPA
0Penalty outcomes analyzed
0Census tracts integrated
1stAuthor, preprint on SportRxiv

Experience

ML Research Analyst

Jan 2026 · Jun 2026
Ahmedabad University · NCAA data partnership with Sacred Heart University, CT
  • Integrated 13,777 minor-penalty outcomes across 55 NCAA Division I programs with 5 seasons of play-by-play data into a unified analytic database.
  • Built Q-learning models (8-state reward structure, 10,000 iterations) and Markov Chain transition maps. Validated via leave-one-season-out cross-validation and 500-game bootstrap resampling (Wilcoxon p < 0.0001).
  • Delivered a decision-support framework giving coaching staff actionable personnel deployment recommendations.
  • First-author preprint on SportRxiv with a 6-author cross-institutional research team.
PythonQ-LearningMarkov ChainsScikit-learnBeautifulSoup

Data Analyst

Feb 2026 · Jun 2026
Northeastern University Roux Institute · Data for Social Good, LSHE Team
  • Competitively selected for a cohort evaluating AI-assisted program impact for nonprofits and service providers across Maine.
  • Co-designed a dual-phase evaluation framework combining Kirkpatrick Model survey methodology with a CIPP structured analysis pipeline.
  • Engineered a Python pipeline with Likert-scale normalization and NLP text mining (TF-IDF, cosine similarity) to audit qualitative stakeholder responses at scale.
  • Delivered evaluation report accepted by the LSHE team and briefed directly to program stakeholders.
PythonNLPTF-IDFPandasCosine Similarity

Projects

Food Intervention Priority Index (FIPI)

72,531 census tracts across 51 states

Designed a PostgreSQL database integrating four federal datasets (USDA, Census ACS, County Health Rankings, SNAP) across 72,531 U.S. census tracts. Built a composite vulnerability scoring model with statistical validation (p < 0.001). The SNAP gap analysis revealed that 45% of scored counties have safety nets that exist on paper but fail geographically. Interactive Tableau dashboards at state and county granularity.

PostgreSQLSQLTableauPythonData Engineering

NCAA Hockey Intelligence System

13,777 penalty outcomes, 55 D1 programs, first-author preprint

End-to-end research pipeline from web scraping to a preprint on SportRxiv. No public dataset existed for NCAA penalty-kill analysis, so I built one from scratch: scraped 5 seasons of play-by-play data, integrated macro and micro datasets, and ran Q-learning and Markov Chain models to produce personnel deployment recommendations for coaching staff. Validated via LOSO cross-validation and 500-game bootstrap resampling.

PythonScikit-learnBeautifulSoupSeleniumQ-LearningMarkov Chains

AI-Powered Financial Chatbot (10-K Analysis)

3 companies, 3 fiscal years, 12 financial metrics, deployed web app

Built a rule-based financial chatbot that analyzes Apple, Microsoft, and Tesla 10-K filings (FY2023 to FY2025). Extracted data directly from SEC EDGAR, engineered 12 derived metrics (margins, growth rates, cash conversion, leverage), and deployed a Flask web app with per-user session state. Caught and fixed a real normalization bug during testing where hyphenated queries silently returned wrong answers, and designed context-aware ratio explanations so the chatbot explains why a number is high or low, beyond simply stating it.

PythonFlaskPandasSEC EDGARFinancial Analysis

Completed as part of the BCG X GenAI Job Simulation (Forage), Aug 2026

Consumer Segmentation and Motivation Analysis

3,000+ survey responses, client-facing deliverable

Built segmentation models on 3,000+ survey responses for Wyman and Son (ME). PCA for dimensionality reduction and K-Means clustering with statistical validation (p < 0.001). Identified high-intent teen/tween households as the primary growth segment, directly informing client marketing strategy.

RPythonPCAK-MeansTableau

AI-Assisted Program Evaluation (Data for Social Good)

Accepted by Northeastern LSHE team

Designed and delivered an evaluation framework for healthcare and nonprofit programs across Maine. Combined Kirkpatrick survey methodology with CIPP analysis. Built a Python pipeline with TF-IDF text mining to audit qualitative stakeholder responses at scale. Report accepted by the LSHE team and briefed to program stakeholders.

PythonNLPTF-IDFCosine SimilarityProgram Evaluation

Writing

Research

First Author · Preprint · SportRxiv · 2026

Penalty-kill personnel deployment and offensive-value exposure in NCAA ice hockey: a box-score decision-support framework

Ravikumar, A., Kaya, T., Artan, N.S., Taber, C., Morris, J.R., Raval, M.S.

Analyzed penalty-kill decision-making across 55 NCAA Division I programs using reinforcement learning and probabilistic modeling. Built the dataset from scratch (no public source existed), validated via leave-one-season-out cross-validation and 500-game bootstrap resampling. The framework gives coaching staff data-grounded personnel recommendations for situations that were previously guided entirely by instinct and film review.

Certifications

BCG X · Forage

GenAI Job Simulation

August 2026 · Junior Data Scientist, GenAI Consulting Team

Completed a simulation involving AI-powered financial chatbot development for BCG's GenAI Consulting team. Integrated and interpreted complex financial data from 10-K and 10-Q reports, employed rule-based logic and Python (pandas) to create a chatbot providing user-friendly financial insights and analysis. Deployed as a live Flask web app.

Quantium · Forage

Data Analytics Job Simulation

May 2026 · Data Science Team

Completed a simulation focused on data analytics and commercial insights. Developed expertise in data preparation and customer analytics, utilizing transaction datasets to extract insights and deliver data-driven commercial recommendations. Identified benchmark stores for uplift testing on trial store layouts, enabling evidence-based decision-making. Created comprehensive reports for the Category Manager to support strategic decisions.

Northeastern · Roux Institute

Innovation Challenge Digital Badge

March 2026

Worked 1:1 with a nonprofit partner to scope a pressing challenge, develop a solution, and present it at the Roux Entrepreneurship showcase. View badge →

Contact

I'm looking for full-time roles where data drives real decisions. If you're building something where the model needs to work and the person using it needs to understand why, I'd like to hear about it.

ravikumar.amr@northeastern.edu