devlog 3:
it’s been even longer but I’ve finally done enough to call this a “devlog”.
basically everything has been fully rewritten in python. A few (quite detrimental) bug fixes are:
- nonexistent error recovery system (i.e. pausing benchmark + invalid responses)
- scoring rubric weights ignored in prompt sent to judge
- save files mislabelling
- A/B scoring only having one slider for better model
- several UI patches (floating navbar -> broken margins)
- added leaderboard from the original index.html currently visible on the github under “Leaderboard”, displaying top scores.
To work on:
- largely inconsistent and seemingly inflated scoring
- check if benchmark can actually be completed, finalise benchmarking criteria + process
- benchmark the rest of the models
i hope this can be done within 2 more devlogs.
Image 1:
much more simplified run system (compared to before).
Image 2:
the og graph, now in v2!
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.