Data Science & Analytics · Seneca Polytechnic

I build data pipelines and AI systems that turn raw data into decisions.

Full-stack data scientist working across ETL orchestration, machine learning, and applied AI — from cleaning 34M+ raw records to shipping RAG-powered chatbots in production.

View projects Get in touch
Python · Airflow · SQL · PySpark · FastAPI · pgvector · LLM
Sam Alavi
About

Data science, ML engineering, and a bias toward shipping.

I'm an Honours Data Science & Analytics student at Seneca Polytechnic (GPA 3.9/4.0), currently a Data Analyst Intern at Manulife building AI agents and PySpark workflows on Databricks. My work spans the full stack: orchestrating ETL pipelines in Airflow, training and evaluating ML models at scale, and building the APIs and interfaces that put them in front of real users.

I'm especially drawn to applied AI — retrieval-augmented generation, agentic workflows, and NLP — and to the engineering discipline that makes those systems reliable in production. Past work includes a co-authored IEEE paper on large-scale vehicle recognition and a top-10 national sales performance while balancing full-time study.

Featured projects

Three systems, one thread: data in, decisions out.

01

Stock Watch — financial-news RAG assistant

An Airflow-orchestrated backend extracts financial news, generates embeddings into a vector database, and serves a chatbot UI that answers natural-language questions grounded in retrieved context — deployed behind a FastAPI service. My most complete full-stack build: orchestration, retrieval, and a conversational frontend working together.

Airflow RAG Vector DB LLM FastAPI Python

02

Track TTC — real-time transit data platform

An end-to-end ETL platform ingesting live TTC GTFS transit data, enriching it with static schedule data, and loading curated records into PostgreSQL for downstream analytics — orchestrated with Apache Airflow and containerized with Docker for reproducible deploys.

Airflow PostgreSQL MinIO Docker ETL

03

NYC Taxi Fare Analysis — 34M+ trip records

Cleaned, validated, and transformed 34M+ NYC Yellow Taxi trip records with Python, Polars, and SQL, then engineered features and trained XGBoost and TensorFlow models to predict fares (RMSE ≈ 1.06, MAPE ≈ 4.2%). Results shipped as Tableau dashboards for non-technical stakeholders.

Python Polars SQL XGBoost TensorFlow Tableau
Skills

Breadth across the stack, depth in AI.

Programming

Python SQL PySpark R

Machine Learning & AI

PyTorch TensorFlow scikit-learn NLP AI Agents

Data Engineering

Databricks Apache Airflow ETL / ELT Docker

Cloud & Tools

GCP AWS FastAPI Power BI
Experience

Where I've put this to work.

May 2026 — Present

Data Analyst Intern · Manulife

Power BI dashboards for KPI tracking; an AI agent in Copilot Studio that turns natural-language questions into data queries; PySpark/Databricks workflows over large-scale conversational datasets.

Mar 2026 — May 2026

AI Research Assistant · Seneca Polytechnic

Built an AI pipeline that extracts structured data from RFQ emails and FastAPI endpoints to orchestrate extraction, pricing, and workflow automation; agent-based workflows with Pydantic AI.

May 2025 — Dec 2025

Software & Data Engineer · Sowcial Experiences

Deployed FastAPI services on Google Cloud Run; designed SQL schemas and relational models in Cloud SQL; delivered backend features in an Agile team.

Jul 2024 — Present

Manager on Duty / Sales · Rogers Communications

Ranked top 10 nationally for sales performance for three consecutive quarters; used Power BI and Excel to track trends and support decisions.

Dec 2024 — May 2025

Data Analyst (Volunteer) · Seneca Hackathon Organizer

Managed registration data in SQL databases and built weekly Power BI reports supporting planning and logistics under tight timelines.

Education

Honours Bachelor, Data Science & Analytics

Seneca Polytechnic · Jan 2024 – Apr 2027

GPA 3.9 / 4.0

Machine Learning, Predictive Analytics, Data Mining (R), High-Performance Computing (PySpark), Database Design (SQL & NoSQL), Data Visualization, Big Data Analysis, Project Management.

Publication

SIVD: Dataset of Iranian Vehicles for Real-Time Multi-Camera Tracking

IEEE · Dec 2022

Co-authored research on large-scale vehicle recognition; curated a 36,000+ image dataset across 29 classes; trained ResNet50, YOLOv5, and OSNet in PyTorch to 99.29% classification accuracy.

Read on ResearchGate →
Contact

Let's connect.

I'm open to internships, collaborations, and conversations about data engineering, ML, and applied AI. Reach out — I'd love to hear from you.

Email me LinkedIn GitHub