Device
Memory
Context in use
KV cache precision
Memory limit
Spec. decode
Sort models
Every setting per model
Weight quantPerplexity vs Q8 Weights GBKV GBTotal GB Decode tok/sPrefill tok/s AgentsPer agent