Finally done with the first ml classifier.
What it does: basically, it takes an url, extract characteristics from it and and predict wether it belongs to 4 of the classes.
Dataset used: here
Pipeline: raw csv -> keep url + type -> normalize url -> remove duplicate url (if there are any) -> stratified sampling.
Next, I’m going to add character-level TF-IDF so that the model doesnt happens to predict google.com as 99.05% phishing 💀💀
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.