Skip to content

ETL Data Pipeline Improvements - #149

Open
iansargent wants to merge 10 commits into
mainfrom
ETL-pipeline-improvements
Open

iansargent wants to merge 10 commits into
mainfrom
ETL-pipeline-improvements

Conversation

@iansargent

Copy link
Copy Markdown
Collaborator

This PR fixes some ETL pipeline bugs and inefficiencies.

  1. Added concurrency to data_collection/ processes to run in parallel (decreases overall time significantly)
  2. Removes legacy .parquet file reading/writing logic
  3. Ensures API key is there before hitting the Census API
  4. Scraping census variables from the .json format (instead of using BeautifulSoup for html parsing) --> Faster!
  5. Improves Census ACS-5 table processing into tidy format (replaces the slow looping .iterrows() function with pd.DataFrame() vectorized calculations)
  6. Changed data_collection/flood.py script to pull from the FEMA official API instead of VCGI (they no longer support this table)

@iansargent
iansargent requested a review from duaneatat October 5, 2026 18:13

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant