Implemented a huge list of fixes
-
Lazy page-ID tensor resolution after quantization
-
Storage-alias-aware VRAM accounting
-
Correct affine INT4 partial-group quantization
-
KV sequence continuity checks
-
Free-page geometry invalidation
-
Canonical LM-head bias and final-softcap handling
-
Certified shortlist logic with exact fallback
-
Gemma-style attention scaling/softcap fixes
-
Removal of hot-path CUDA scalar tensor allocations
-
Device normalization for “cuda”, “cuda:N”, and torch.device
-
Profile separation has been reconnected and different profiles select different kernels.
-
Paged KV storage exists through SlabbedKVCache.
-
A prior paged-attention slowdown was caused by GPU synchronization from .item() calls in SlabbedKVCache.append(), not by paged addressing itself.
-
Logical page ownership is now tracked CPU-side to avoid those synchronizations.
-
Paged attention has measured approximately:
- Materialized: 55.1 tok/s
- Paged: 57.5 tok/s
on a relevant smaller models
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.