You are browsing as a guest. Sign up (or log in) to start making projects!

5h 11m 5s logged

ThinTensor Improvements

  • Universal semantic-IR runtime for mixed attention, dense and MoE architectures.
  • Per-tensor precision, kernel and residency planning instead of layer-wide quantization.
  • Production FP8, INT4, MXFP4, Q2 and Q1 quantizers with calibration-aware scoring.
  • Hard ≥0.995 cosine certification with exact routing, greedy-token and KV checks.
  • Qwen3.5 hybrid linear-attention, shared-expert and packed rank-3 MoE support.
  • Durable grouped-INT4 expert sidecars with bounded-RAM, mmap-based construction.
  • Fixed four-matrix Triton projection, cuda/cuda:0, layer-zero and converter-role bugs.
  • Improved fused kernels, KV caching, expert routing and decode hot paths.
  • Strict cache identity across model, tokenizer, GPU, CUDA, kernels and calibration.
  • Remaining goal: prove ≥30 tok/s on Qwen3.5-35B-A3B at certified ≥0.995 cosine.
  • Getting way better speed on 27B dense models at a cosine that is also way better
0
9

Comments 0

No comments yet. Be the first!