Devlog 4
I’ve grouped some of the domain features into 3 main groups:
- Identity: How long the domain is, how many subdomain labels there are, etc… it’d help us to distinguish from google.com and secure-login-account-verification.example.com
- Composition: digits, hyphens or special characters in the domain, as well as ratio and Shannon entropy. The entropy calculation was written from scratch rather than pulling from scipy => it kept the dependencies light. High entropy like x7Qhd6agz1 look very different from low entropy one like google
- Suspicion indicator: binary flags for punycode, numeric TLDs, long subdomains, and high-entropy domains.
I’ve also updated the hybrid model training script to compare TF-IDF alone, TF-IDF plus some lexical features, TF-IDF with new 3 grouped domain features, and TF-IDF with both.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.