You are browsing as a guest. Sign up (or log in) to start making projects!

3h 2m 22s logged

devlog 3:
it’s been even longer but I’ve finally done enough to call this a “devlog”.

basically everything has been fully rewritten in python. A few (quite detrimental) bug fixes are:

  • nonexistent error recovery system (i.e. pausing benchmark + invalid responses)
  • scoring rubric weights ignored in prompt sent to judge
  • save files mislabelling
  • A/B scoring only having one slider for better model
  • several UI patches (floating navbar -> broken margins)
  • added leaderboard from the original index.html currently visible on the github under “Leaderboard”, displaying top scores.

To work on:

  • largely inconsistent and seemingly inflated scoring
  • check if benchmark can actually be completed, finalise benchmarking criteria + process
  • benchmark the rest of the models
    i hope this can be done within 2 more devlogs.

Image 1:
much more simplified run system (compared to before).
Image 2:
the og graph, now in v2!

0
2

Comments 0

No comments yet. Be the first!