Added scanning of robots.txt and sitemap.xml
CHECKPOINTTT
I finished building the first stage of the web app by coding the second and third ways way discover legal document links – scanning robots.txt and sitemap.xml using Requests, Python and BeautifulSoup, and filtering them against a list of keywords such as “privacy”, “terms”, and “cookies” and removing duplicates. Now I have three ways of finding links, and can move onto actually reading what’s inside them!
Yay :)
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.