Full-stack data scientist working across ETL orchestration, machine learning, and applied AI — from cleaning 34M+ raw records to shipping RAG-powered chatbots in production.
I'm an Honours Data Science & Analytics student at Seneca Polytechnic (GPA 3.9/4.0), currently a Data Analyst Intern at Manulife building AI agents and PySpark workflows on Databricks. My work spans the full stack: orchestrating ETL pipelines in Airflow, training and evaluating ML models at scale, and building the APIs and interfaces that put them in front of real users.
I'm especially drawn to applied AI — retrieval-augmented generation, agentic workflows, and NLP — and to the engineering discipline that makes those systems reliable in production. Past work includes a co-authored IEEE paper on large-scale vehicle recognition and a top-10 national sales performance while balancing full-time study.
01
An Airflow-orchestrated backend extracts financial news, generates embeddings into a vector database, and serves a chatbot UI that answers natural-language questions grounded in retrieved context — deployed behind a FastAPI service. My most complete full-stack build: orchestration, retrieval, and a conversational frontend working together.
02
An end-to-end ETL platform ingesting live TTC GTFS transit data, enriching it with static schedule data, and loading curated records into PostgreSQL for downstream analytics — orchestrated with Apache Airflow and containerized with Docker for reproducible deploys.
03
Cleaned, validated, and transformed 34M+ NYC Yellow Taxi trip records with Python, Polars, and SQL, then engineered features and trained XGBoost and TensorFlow models to predict fares (RMSE ≈ 1.06, MAPE ≈ 4.2%). Results shipped as Tableau dashboards for non-technical stakeholders.
Programming
Machine Learning & AI
Data Engineering
Cloud & Tools
May 2026 — Present
Power BI dashboards for KPI tracking; an AI agent in Copilot Studio that turns natural-language questions into data queries; PySpark/Databricks workflows over large-scale conversational datasets.
Mar 2026 — May 2026
Built an AI pipeline that extracts structured data from RFQ emails and FastAPI endpoints to orchestrate extraction, pricing, and workflow automation; agent-based workflows with Pydantic AI.
May 2025 — Dec 2025
Deployed FastAPI services on Google Cloud Run; designed SQL schemas and relational models in Cloud SQL; delivered backend features in an Agile team.
Jul 2024 — Present
Ranked top 10 nationally for sales performance for three consecutive quarters; used Power BI and Excel to track trends and support decisions.
Dec 2024 — May 2025
Managed registration data in SQL databases and built weekly Power BI reports supporting planning and logistics under tight timelines.
Education
Seneca Polytechnic · Jan 2024 – Apr 2027
GPA 3.9 / 4.0Machine Learning, Predictive Analytics, Data Mining (R), High-Performance Computing (PySpark), Database Design (SQL & NoSQL), Data Visualization, Big Data Analysis, Project Management.
Publication
IEEE · Dec 2022
Co-authored research on large-scale vehicle recognition; curated a 36,000+ image dataset across 29 classes; trained ResNet50, YOLOv5, and OSNet in PyTorch to 99.29% classification accuracy.
Read on ResearchGate →I'm open to internships, collaborations, and conversations about data engineering, ML, and applied AI. Reach out — I'd love to hear from you.