You are browsing as a guest. Sign up (or log in) to start making projects!

2h 48m 47s logged

Implemented a huge list of fixes

  • Lazy page-ID tensor resolution after quantization

  • Storage-alias-aware VRAM accounting

  • Correct affine INT4 partial-group quantization

  • KV sequence continuity checks

  • Free-page geometry invalidation

  • Canonical LM-head bias and final-softcap handling

  • Certified shortlist logic with exact fallback

  • Gemma-style attention scaling/softcap fixes

  • Removal of hot-path CUDA scalar tensor allocations

  • Device normalization for “cuda”, “cuda:N”, and torch.device

  • Profile separation has been reconnected and different profiles select different kernels.

  • Paged KV storage exists through SlabbedKVCache.

  • A prior paged-attention slowdown was caused by GPU synchronization from .item() calls in SlabbedKVCache.append(), not by paged addressing itself.

  • Logical page ownership is now tracked CPU-side to avoid those synchronizations.

  • Paged attention has measured approximately:

    • Materialized: 55.1 tok/s
    • Paged: 57.5 tok/s
      on a relevant smaller models
0
52

Comments 0

No comments yet. Be the first!