Slopmark
- 1 Devlogs
- 22 Total hours
building a benchmark which basically identifies how much of what ur agent ouputs is slop with the most absurd ways of benching ai
building a benchmark which basically identifies how much of what ur agent ouputs is slop with the most absurd ways of benching ai
ok so quick context before the actual product dump
ive been shipping this under hack club macondo this whole time. i had multiple journals (devlogs) posted there everything was fine until i tried to ship it which then hit me with the ‘u r ‘old’ unc’ so i had to shift programs and here i am
that means im not gonna re-post every old macondo update one by one. this single post is the catch-up. everything we built so far in one place. after this normal session-by-session posts continue on stardance
live: https://slopmark.vercel.app
repo: https://github.com/nothariharan/benchmark
slopmark = the honest slop detector
public ai benches are cooked. contamination. saturation. goodhart. vibes scoring. llm-as-judge lying to itself
so we built an open eval platform where every model eats the same harness and a rule-based verifier decides pass/fail. never another llm judging another llm
spine is dumb on purpose:
novel task → fixed harness → behavioral verifier → score + receipt
if a model sounds smart but misses the contract it fails. that is the whole product
next.js on vercel · eval in lib/harness + lib/eval · providers: openrouter / aiml / fireworks / byok · ~364 seed tasks · rule verifiers only
deploy trick: vercel has no sqlite so curated results are committed as results.json and the site just reads them. challenges = curated forever sprints. sessions = ad-hoc byok runs that dont pollute the scoreboards
ok architecture done. sessions are the real work
revisitable sprints. fixed tasks. fixed models. fixed harness. run the grid. commit the receipt. anyone can reopen later and see who failed what
https://slopmark.vercel.app/challenge/niche-sprint-v1
6 aiml models x 10 mixed niche tasks = 60 runs. zero_context. procedural + extract + html + 8-word constraint + json + regex
gemini 2.5 flash 90% · deepseek v3 70% · gpt-4o mini + llama 3.1 8b 60% · ministral + qwen 50%
every model ate dirt on the 8-word constraint. that one task justified the whole sprint
https://slopmark.vercel.app/challenge/drawing-contest-v1
text models only. raw svg. no image models. html_contract verifier + live gallery
10 draws (house cat snowman bicycle rocket etc). 5 fireworks models
gpt-oss 120b + glm 5.1 = 100%. rest at 90%. common miss: rocket flame color. looking fine is not passing
https://slopmark.vercel.app/challenge/game-night-v1
same fireworks lineup. party-game instruction contracts: no-e pizza lipogram, pirate wifi, questions-only db explain, taboo cat, 12-word black holes, printer haiku with jam, ALL CAPS chef, haunted house review, internet in one word
gpt-oss 90% · deepseek 60% · kimi 30% · glm 5.2 20% · glm 5.1 10%
drawing said everyone can draw kinda. game night said almost nobody can shut up and follow a party rule under zero context
bench + byok · build your own sprint (/challenges/new) · sessions wall · realshot duels · goal games · hall of shame · leaderboard
started as arena, pivoted to verifier-first slopmark, stripped auth so vercel just works
shipped: verifier engine, 3 public sessions with committed receipts, byok + ad-hoc sprints, realshot/goal/shame/leaderboard, vercel json deploy
this post is the macondo → stardance handoff. product is slopmark. program is now stardance. building continues from here