deleting the score
biggest change this session was taking things out. round mode had points, a running total and a combo multiplier. all three gone!!!
the argument lowk fell out as soon as i said it aloud. the number at the top of the results was the run time, and the quickest way thru a round is to answer everything wrong immediately. so the headline figure was paying you for the one behaviour the app exists to catch. points and combos were the same information twice over, worked out from how fast and how often you were right, then shown next to how fast and how often you were right
whats there instead weights every answer by how little help it had. warm up counts once, round 1 three times, round 2 five, round 3 six, and saying a concept in your own words in round 4 is worth 150. speed isnt in the formula anywhere and two tests hold that down: the same answers given at wildly different speeds come out identical, and fast-and-wrong lands below slow-and-right
it reads +615 of 2,120. a bare number says nothing when a short session and a long one have different ceilings
round 3
the chip tray used to have wrong words mixed in, so the round was a hunt for what to avoid. now it holds the sentence’s own pieces and nothing else
grading stopped being fussy abt list order. seven tests, four of them cases that still have to fail: “sugar traps light” for “light traps sugar”, a scrambled non-list, a missing chip, a changed verb
round 4
cant be skipped any more. there was a “stop here and see the results” button sitting right in front of the only moment in the session with nothing on screen to lean on, which is where bailing is most tempting
the interface stopped hiding
the header used to empty itself during a round, on the theory that chrome at the edges competes with retrieval. thats true for one person at a desk. round mode gets passed round a room with several people reading one screen, so it stays up now
locking the page to the viewport turned up two bugs. one of them clipped the verdict off the bottom on a short window, and the verdict is the line that tells you the answer
the run clock was a nice one. it measured wall clock from the start of the run, so it counted loading and the between-round screens, and it disagreed with the time the results reported. header said 3:20 against a reported 1:08. built off the same splits the results use now
the cold open prompt
trimmed it after noticing it asked the model to write explanation lines nobody reads. a/b against the real proxy over 10 topics took median completion tokens from 397 to 266, abt 33%, on the one call every student sits and waits for
setTimeout
i overrode it in the harness to freeze auto-advance long enough to screenshot a verdict, and it swallowed the harness’s own sleeps. script hung and i captured the wrong state. keep a reference to the real function before you replace it
status
the assembled round is verified, which closes the open item from the last log. full ladder played end to end, in dev and against a production build
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.