Microsoft Β· 3.8B Β· Dense
Lightweight reasoning model Check if your GPU or Mac can run Phi-4 Mini Reasoning locally β 2.1 GB min, 3.5 GB recommended.
| Quant | Bits | VRAM | Quality | Status |
|---|---|---|---|---|
| Q2_K | 2 | 1.7 GB | low | β |
| Q3_K_M | 3 | 2.2 GB | moderate | β |
| Q4_K_M | 4 | 2.4 GB | good | β |
| Q5_K_M | 5 | 2.9 GB | good | β |
| Q6_K | 6 | 3.4 GB | excellent | β |
| Q8_0 | 8 | 4.4 GB | excellent | β |
| F16 | 16 | 8.3 GB | lossless | β |
About this model
Phi 4 mini reasoning is designed for multi-step, logic-intensive mathematical problem-solving tasks under memory/compute constrained environments and latency bound scenarios. Some of the use cases include formal proof generation, symbolic computation, advanced word problems, and a wide range of mathematical reasoning scenarios. These models excel at maintaining context across steps, applying structured logic, and delivering accurate, reliable solutions in domains that require deep analytical thinking.
The graph compares the performance of various models on popular math benchmarks for long sentence generation. Phi-4-mini-reasoning outperforms its base model on long sentence generation across each evaluation, as well as larger models like OpenThinker-7B, Llama-3.2-3B-instruct, DeepSeek-R1-Distill-Qwen-7B, DeepSeek-R1-Distill-Llama-8B, and Bespoke-Stratos-7B. Phi-4-mini-reasoning is comparable to OpenAI o1-mini across math benchmarks, surpassing the modelβs performance during Math-500 and GPQA Diamond evaluations. As seen above, Phi-4-mini-reasoning with 3.8B parameters outperforms models of over twice its size.β―
References
Can I run Phi-4 Mini Reasoning locally?
- Can I run Phi-4 Mini Reasoning locally?
- Phi-4 Mini Reasoning needs about 2.1 GB of memory at a minimum and 3.5 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
- How much VRAM does Phi-4 Mini Reasoning need?
- At Q4_K_M, Phi-4 Mini Reasoning uses about 2.4 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.