← ALL PROJECTS
Ontario University Metrics mascot

Ontario University Metrics

A public web app where students can explore real university admissions data across 600+ Ontario programs.

FRONTEND
Next.jsTypeScript
DATABASE
PostgreSQL
DATA PIPELINE
TypeScript ETLWeb scrapingToken-set matching

The Problem

High school students applying to Ontario universities have almost no reliable way to know what grades actually get someone into a program.

Crowdsourced Reddit spreadsheets are the best available resource, but entries are often missing a program code or name, making it hard to search and get accurate results, with no historical view or fast way to compare across programs or years.
University pages publish their own estimated minimum cutoffs, but these don’t always reflect what students are actually getting in with.

Students are left manually cross-referencing scattered spreadsheets and official pages just to guess their odds.

Approach

The underlying data was sourced from two places: crowdsourced admissions spreadsheets, which track real self-reported admission results and feed the core data and visualizations, and official university pages, which publish targeted averages representing each program’s estimated minimum cutoff. These served as a benchmark, allowing the platform to show how far official estimates were from actual admitted grades. Both sources came in inconsistent formats with no shared structure across schools or years. Before any analytics could be built, this raw data had to be reconciled into a single consistent dataset and cross-referenced, so the platform could show reported admission averages alongside official published ranges.

The pipeline runs in three stages: scrape, clean, and normalize.

Key Challenge

The hardest problem was matching the data correctly. The same program could appear under different names depending on the source:

Flipped word order (e.g. “Computer Engineering” vs. “Engineering, Computer”)
Abbreviations or added specializations
Inconsistent formatting
Missing codes, missing names, or misspellings

An exact string match against a canonical program list only worked about 75% of the time, dropping or mismatching a quarter of the data.

To fix this, program and university names were normalized and broken into word tokens, then matched using token-set similarity.

This process raised match accuracy from 75% to 95%.

Outcome

10,000+
users total
1,500+
users in the first 24 hours
600+
programs tracked
21
universities

The platform launched on Reddit and quickly gained traction among students researching admissions. Feedback highlighted a limitation: since the data is self-reported, students with higher grades are more likely to post their results, which can skew average admitted grades higher than reality. This is a known bias in crowdsourced admissions data, and future versions could address it by weighting or flagging self-reported entries differently from verified sources.

Personal Takeaways

First live deploy

Learned to ship to the real public and keep it reliable, not just demo-ready.

Users shape products

Users surfaced issues I’d missed, like self-reporting bias, and kept me iterating.

Solved a problem I had

Built the admissions tool I’d wished for in high school, which hooked me on fixing problems.

NEXT PROJECT
ConvoSpark AI
→
TORONTO, ONGITHUB · LINKEDIN · EMAIL
← BACK TO LINE