All selected work
DATA SYSTEMS · SEMANTIC CLUSTERING

Semantic deduplication.

A scraping, cleaning, and semantic deduplication workflow for large research datasets, developed in collaboration with Mridul Joshi, Researcher at Stanford.

CONTEXT

Research collaboration

WHEN

Aug — Dec 2025

MY CONTRIBUTION

Independently designed a MongoDB-backed data pipeline using HDBSCAN clustering and LLMs to identify semantic duplicates.

AT A GLANCE

7M+ duplicates removed in the deduplication project.

MongoDBHDBSCANLLMs
SEMANTIC DEDUPLICATION
DATA SYSTEMS · SEMANTIC CLUSTERING
Scrape & cleanHDBSCAN clustersDeduplicated data
↳7M+ duplicates removed in the deduplication project.

Illustrative workflow · explore the engineering above

01

The problem

Exact matching could not detect differently worded records that repeated the same meaning. These semantic duplicates reduced the quality and reliability of data used for downstream research and ML.

02

What I built

Engineered a scraping and cleaning workflow backed by MongoDB. Used HDBSCAN clustering to group semantically similar records and language models to verify duplicate meaning before removing redundant records.

03

The engineering thinking

Similarity alone does not establish duplication. Clustering groups candidate records; LLM verification helps distinguish repeated meaning from records that are merely related.

04

The outcome

Processed 33M+ news records and removed 7M+ duplicates, improving data quality for downstream ML training and research. This project is separate from the smartphone-use field experiment.

EXPLORE ANOTHER PROJECTCTS-71
Let’s talk about the details