You are browsing as a guest. Sign up (or log in) to start making projects!

5h 46m logged

Debate Mode, and the source that did not exist

The biggest change this session was a second mode. Round Mode’s ladder tops out at saying a thing in your own words. Debate asks something else: hold a position for four speeches against an opponent trying to take it off you. Its own route, not a sixth round.
The rating is arithmetic and never the model’s. The judge returns a winner, a margin and ten dimension scores, and the elo change is computed in integer maths the model never sees. A model asked to update an elo returns a number that looks like arithmetic, and the failure is invisible because every value it gives is plausible.
Two tabs that never mix. Open debate is judged on whether the argument holds up. Tournament prep is judged against a tournament bar, where a dropped argument is conceded and 60 means mediocre, not good for a beginner. Format is required, never defaulted: a Public Forum ballot given to someone practising Lincoln-Douglas is worse than no ballot, because they will act on it.
A real user ran a topic-only session on Tritoflex, a spray-on rubber roofing compound, and was repeatedly told it was torch-applied. He caught it because he installs the stuff. A student studying something new would not have. Asked cold about an obscure term, the model returns a protein, an alpha helix, signal transduction, implicated in cancer. Not one word true, and none of it looks invented. That is the failure worth naming: not that it does not know, but that it does not say so.
The worst of it was not a missing disclaimer. In a topic-only session the Round 4 marking screen printed “Contradicts the source” under a heading reading “The notes”, when the student had pasted nothing and the server blanks every citation on the way out. Someone who wrote the true sentence was told they contradicted evidence that was never there. The app exists to close the gap between feeling like you know something and knowing it, and that screen widened it. It now says “the model disagrees”.
The header badge went from “AI-generated” to “AI, unchecked”. Authorship is the fact a student already has, since they typed the topic in. Nothing checked this is the one that costs them. “AI” stayed because the label outlives the moment: screenshots get shared.
Runs used to die with the tab: honest about the plumbing, wrong about the pedagogy, since retrieval practice works through spacing. The session already computed the thing worth keeping, a map of which ideas you can recognise but cannot say, showed it once, then threw it away. One standing per concept persists now, and the front door leads with what is still unsaid.
The method note, and it is the best yet: my first two boundary assertions both passed against the exact bugs they were written for. One compared trimmed pieces, which cannot distinguish a mid-word cut from a clean one, because trimming erases the evidence. The other tested a sample that had collapsed to one chunk, so there was no join in it to be wrong. Both were testing nothing, confidently, and only revealed themselves when I reverted the fix to watch them fail and they stayed green. A test you have not seen fail is not a test.
Closing an item from the last log: the green “strong” band is no longer unverified. It has been seen in a browser by someone who knew the answers, not by a driver answering option 1 every time.
The README’s privacy section had gone from cautious to false, because “nothing is stored beyond the current session” stopped being true the moment runs started leaving a record. Second time a claim here has drifted as the product changed, so anything that stores something now triggers a README check in the same commit. Still no accounts: standings are per device, and clearing site data erases them.
Built with Claude Code. AI-assisted commits carry Co-Authored-By trailers.

0
4

Comments 0

No comments yet. Be the first!