005: the window armed, and the run that went dark
Planned three things: catch the scheduler firing against a real clock, get a photo of a populated changes feed, and if there was room, add a retention limit to the scans table. The retention part landed clean. The clock-watching part did not, since the whole attempt crashed unexpectedly partway through and came back much later with only a partial trail left behind.
Retention is simple and real: keep the last hundred completed scans per target, prune right after a scan finishes, never let a failed cleanup corrupt a scan that actually succeeded. Deleting an old scan correctly drags its stored history with it, since nothing on screen ever looks back further than the limit already allows anyway.
The near miss is worth writing down. A log line looked exactly like proof the real scheduler behavior had finally been caught live. It was not, a different unrelated row fired for an unrelated reason, and it nearly got written up as the real finding before a second look caught it. Worse, the machine’s own clock had been wrong for the entire stretch, quietly off by a meaningful margin, so no timing claim from that window can be trusted. Every timing tool now stamps both the wall clock and a separate monotonic clock side by side, so drift shows up honestly instead of lying silently. Someone else picked up the real observation properly, on a trustworthy clock, the same day.
One real bug got caught and fixed regardless: a visual check of the live pages turned up literal broken source text rendering in production on every page, leftover from unrelated commits, plus a stray constant pasted into the middle of an unrelated import. Both fixed and disclosed rather than skipped past.
The actual lesson: the product held up fine through all of it. The environment around it, crashes, clock errors, lost caches, was what actually failed this time.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.