VICTORY!!! from here it’s a matter of refining the scraper to make sure there are no duplicates / useless parts / specs are good. the final pieces of the puzzle were removing JSDOM caching to reduce memory usage (since i already check for duplicates reasonably well before the request is made), adding some more filtering for irrelevant products, manually garbage collecting with global.gc() (!!), and allowing for errors in product parsing. i could probably make it more robust by forcing all the categories to be evaluated before products.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.