#context-window

Wiki 1

  • KV Cache Sizing KV cache memory per token = 2 × KV-heads × head-dim × layers × bytes/element; halving via q4_0 lets you double context for free