iPhone model evaluation
Everyday English quality
100 tasks, repeated three times at the app’s default response settings. Incomplete evaluations receive no final score.
| Model | Catalog | Quality | Reviewed |
|---|---|---|---|
| SmolLM2 135M Instruct | Current catalog | 8.0/100 Below threshold | 300/300 |
| Qwen 2.5 0.5B Instruct | Current catalog | 45.42/100 Below threshold | 300/300 |
| Llama 3.2 1B Instruct | Current catalog | 41.67/100 Below threshold | 300/300 |
| DeepSeek R1 Distill Qwen 1.5B | Current catalog | 11.0/100 Below threshold | 300/300 |
| LFM 2.5 1.2B Instruct | Current catalog | 64.33/100 Below threshold | 300/300 |
| Gemma 3 1B IT | Current catalog | 58.5/100 Below threshold | 300/300 |
| MiniCPM5 1B | Current catalog | 55.25/100 Below threshold | 300/300 |
| Granite 4.0 H 1B | Current catalog | 75.67/100 Below threshold | 300/300 |
| MiniCPM5 2B | Current catalog | 86.67/100 Below threshold | 300/300 |
| Ministral 3 3B Instruct | Current catalog | 68.83/100 Below threshold | 300/300 |
| Granite 4.2 3B | Current catalog | 85.33/100 Threshold met | 300/300 |
| Phi-4 Mini Instruct | Current catalog | 59.67/100 Below threshold | 300/300 |
| SmolLM2 1.7B Instruct | Current catalog | 56.42/100 Below threshold | 300/300 |
| Qwen2.5 Coder 1.5B Instruct | Current catalog | 70.17/100 Below threshold | 300/300 |
| Qwen 3.5 0.8B | Evaluation candidate | 61.17/100 Below threshold | 300/300 |
| Qwen 3.5 2B | Evaluation candidate | 81.17/100 Below threshold | 300/300 |
How scoring works
- Instructions: 25%
- Grounding: 20%
- Conversation: 20%
- Reasoning: 15%
- Writing: 10%
- Honesty: 10%
Each task has four criteria defined before testing. Exact answers and JSON are checked automatically. Writing and other subjective tasks are reviewed by the Codex assistant against a fixed rubric with model labels and timing hidden. The reviewer has run-order context; this is not independent or double-blind human validation. The score combines raw rubric points using the category weights; all three repetitions count.
Tests include corrections, conversation memory, changed instructions after Stop, summaries grounded in supplied text, and questions with missing information. Coding and speech recognition are separate from this score.
Eligibility requires at least 80 overall, at least 70 in every category, and at least 85 in instructions and conversation. A default candidate needs at least 85 overall. Passing quality thresholds does not establish release readiness.
Performance and reliability
Load time, first-token latency, generation speed, memory and thermal behavior are reported separately. Download recovery, offline operation and cancellation must pass independently. Results describe the tested artifact and app configuration; they do not establish performance on other iPhones.