finally, benchmarking has begun.
The benchmark was re-evaluated meticulously based on Qwen3.5-4B as a ~50% guideline for benchmarking.
Then, I set up my now dead macbook pro (dead display, dead battery) to a portable monitor to use as basically a cloud server and started benchmarking. This is the fourth night of trying (after benchmarking failed over the night due to the computer not plugging in and losing all charge (day 1), automation logic not working (day 2), LFM weights getting stuck (day 3) so hopefully it works today with the 17 consecutive models to be benchmarked.
using 17 models, including:
MiniCPM5-1B (0.57GB)
LFM2.5-2.6B GGUF (1.59GB)
Granite3B (2.09G8)
G9V3 (1.77GB)
Nanbeige (2.50GB)
Granite8B (4.98GB)
Ling(4.58GB)
Falcon (4.28GB)
LFM8B (4.51GB)
Gemma E2B(4.04GB)
Gemma E4B (4.79GB)
Gemma12B (5.30GB)
Qwen3.5-9B (4.35GB)
Qwen3-4B-2507 (2.1GB)
Qwen3.5-2B (1.60GB)
Qwen3.5-4B (2.83GB)
DeepSeek-R1-Qwen3-8B (4.29GB)
Note, all models were selected to fit within 6-8GB vram total so it could be evaluated on my local hardware.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.