stardance has been extended another month! the new deadline is october 31 :)

You are browsing as a guest. Sign up (or log in) to start making projects!

itzanhzz

@itzanhzz

Joined August 16th, 2026

  • 12Devlogs
  • 2Projects
  • 2Ships
  • 0Votes
Ship Changes requested

Scam buzzer is a Chrome extension that checks the pages you visit against a small ML model running on your own machine, then it warns you when a URL looks like a scam or phishing attempt. The whole thing here runs locally!
To check this out, read the README and run it yourself!

  • 11 devlogs
  • 17h
Try project → See source code →
Open comments for this post

3h 0m 58s logged

Devlog 7

This project is about to be completedddd!!! im so excited
and btw, after running some tests on the model in ver 4, it usually predict every websites as “phishing”. Even though the overall acurracy is about 0.95. I found this very concerning because it shouldn’t have happened. After I dig deeper, I found that “has_https” caused a lot of bias in the model (its scale was about 6 or smt) but bascially it contributes a lot toward the final prediction. Then i’ve removed it and then the phishing biased percentages reduced by a bit. I’d figured that training a new model would be better rather than going around fixing some parameters in the predictor script. In the v5, I’ve tried to add a threshold testing. Basically instead the model immediately saying phishing prob = .7 => phishing, I experimented with <.6 -> not high-risk, .6-.9 -> suspicious and >.9 -> malicious. This pretty much didn’t do nothing much cuz I traded of some accuracy for nothing and google.com was still being predicted as .99 phishing. After that, I retrained the model while removing the URL scheme from the TF-IDF input. So instead of the tf-idf seeing: https://google.com, it sees approx google.com. The idea here was to prevent the TF-IDF from learning: https://, ://, s://, ttps as phishing indicator. Running all those tests for about 2 hours and the results aren’t significant. I’ve decided to add a trusted domain list and then moving toward designing the extension and the local API. 💀💀💀💀

0
0
10
Open comments for this post

3h 7m 4s logged

Devlog 6

Meaning running testing and refactoring some of the training python scripts, I’ve used FastAPI to create the api for the ML model. There was a schemas that the api expect and also predict (on /predict POST endpoint)

0
0
17
Open comments for this post

29m 37s logged

Devlog 5

Quick one here, TF-IDF and the lexical feature classifier achieve the highest accuracy here. I’ll be using that for the chrome extension.

0
0
5
Open comments for this post

3h 39m 16s logged

Devlog 4

I’ve grouped some of the domain features into 3 main groups:

  • Identity: How long the domain is, how many subdomain labels there are, etc… it’d help us to distinguish from google.com and secure-login-account-verification.example.com
  • Composition: digits, hyphens or special characters in the domain, as well as ratio and Shannon entropy. The entropy calculation was written from scratch rather than pulling from scipy => it kept the dependencies light. High entropy like x7Qhd6agz1 look very different from low entropy one like google
  • Suspicion indicator: binary flags for punycode, numeric TLDs, long subdomains, and high-entropy domains.
    I’ve also updated the hybrid model training script to compare TF-IDF alone, TF-IDF plus some lexical features, TF-IDF with new 3 grouped domain features, and TF-IDF with both.
0
0
9
Open comments for this post

46m 21s logged

Quick devlog

After combining the TF-IDF and handcrafted features, the model seems to performs worse… I might need to check some scaling problem or something? The teacher says to me that my hand-crafted features are highly correlated or like they are having redundant signals…

0
0
5
Open comments for this post

1h 33m 54s logged

Devlog 3

I’m finally done with researching and implementing character TF-IDF for this classifier. I’m trying to combine 2 of the classifier method into a hybrid to see if the model still predict that google.com is phishing or not. (hopefully not)

0
0
11
Open comments for this post

1h 23m 19s logged

Devlog 2:

I’ve just doen with working on the TF-IDF model. Tho, the training time is still long (like whole 5 minutes or something for the entire pipeline). But the accuracy is great (96.95%)!!

0
0
9
Open comments for this post

30m 12s logged

Finally done with the first ml classifier.

What it does: basically, it takes an url, extract characteristics from it and and predict wether it belongs to 4 of the classes.
Dataset used: here
Pipeline: raw csv -> keep url + type -> normalize url -> remove duplicate url (if there are any) -> stratified sampling.
Next, I’m going to add character-level TF-IDF so that the model doesnt happens to predict google.com as 99.05% phishing 💀💀

0
0
8

Followers

Loading…