DS 410 Course Project

GitHub

Summary

Topic modeling project using PySpark that is able to filter through and process multiple terebytes of Common Crawl data. Works with Slurm and can be parallelized across multiple nodes and CPUs.

Contribution

Led a major refactor of the code to increase maintainablity and enabled bugfixes through responsible coding that involved actually looking at the codebase and convincing my team to do so as well. Actually learned how to use PySpark and Slurm by reading the docs. Enabled team to responsibly use coding agents to accelerate development while still keeping maintainability.