DS 410 Course Project
Summary
Topic modeling project using PySpark that is able to filter through and process multiple terebytes of Common Crawl data. Works with Slurm and can be parallelized across multiple nodes and CPUs.
Contribution
Led a major refactor of the code to increase maintainablity and enabled bugfixes through responsible coding that involved actually looking at the codebase and convincing my team to do so as well. Actually learned how to use PySpark and Slurm by reading the docs. Enabled team to responsibly use coding agents to accelerate development while still keeping maintainability.