Cohere Β· 35B Β· Dense
Optimized for retrieval-augmented generation Check if your GPU or Mac can run Command R 35B locally β 19.6 GB min, 32.6 GB recommended.
2024-03128K context
| Quant | Bits | VRAM | Quality | Status |
|---|---|---|---|---|
| Q2_K | 2 | 11.7 GB | low | β |
| Q3_K_M | 3 | 16.2 GB | moderate | β |
| Q4_K_M | 4 | 18.4 GB | good | β |
| Q5_K_M | 5 | 22.9 GB | good | β |
| Q6_K | 6 | 27.4 GB | excellent | β |
| Q8_0 | 8 | 36.4 GB | excellent | β |
| F16 | 16 | 72.2 GB | lossless | β |
About this model
Command R is a generative model optimized for long context tasks such as retrieval-augmented generation (RAG) and using external APIs and tools. As a model built for companies to implement at scale, Command R boasts:
- Strong accuracy on RAG and Tool Use
- Low latency, and high throughput
- Longer 128k context
- Strong capabilities across 10 key languages
There are currently two versions of Command R:
- Original release tagged v0.1
- August 2024 update tagged 08-2024
References
Can I run Command R 35B locally?
- Can I run Command R 35B locally?
- Command R 35B needs about 19.6 GB of memory at a minimum and 32.6 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
- How much VRAM does Command R 35B need?
- At Q4_K_M, Command R 35B uses about 18.4 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.