Devlog 7
This project is about to be completedddd!!! im so excited
and btw, after running some tests on the model in ver 4, it usually predict every websites as “phishing”. Even though the overall acurracy is about 0.95. I found this very concerning because it shouldn’t have happened. After I dig deeper, I found that “has_https” caused a lot of bias in the model (its scale was about 6 or smt) but bascially it contributes a lot toward the final prediction. Then i’ve removed it and then the phishing biased percentages reduced by a bit. I’d figured that training a new model would be better rather than going around fixing some parameters in the predictor script. In the v5, I’ve tried to add a threshold testing. Basically instead the model immediately saying phishing prob = .7 => phishing, I experimented with <.6 -> not high-risk, .6-.9 -> suspicious and >.9 -> malicious. This pretty much didn’t do nothing much cuz I traded of some accuracy for nothing and google.com was still being predicted as .99 phishing. After that, I retrained the model while removing the URL scheme from the TF-IDF input. So instead of the tf-idf seeing: https://google.com, it sees approx google.com. The idea here was to prevent the TF-IDF from learning: https://, ://, s://, ttps as phishing indicator. Running all those tests for about 2 hours and the results aren’t significant. I’ve decided to add a trusted domain list and then moving toward designing the extension and the local API. 💀💀💀💀
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.