This part of the project was mostly about making the final test harder to fool. I preregistered the analysis, froze the generator and sealer, and sealed the held-out split before looking at any confirmatory outcome. The important result here is a constraint rather than a score: if I can still change the test after seeing results, the result is not trustworthy. The evaluation itself was not run at this point, so this was preparation for evidence, not evidence that SL-LLM-R works. I would rather keep the experiment locked than make the final result easier to obtain.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.