← All models

iPhone model evaluation

Everyday English quality

100 tasks, repeated three times at the app’s default response settings. Incomplete evaluations receive no final score.

Evaluation target: iPhone 17 Pro Max
ModelCatalogQualityReviewed
SmolLM2 135M InstructCurrent catalog8.0/100
Below threshold
300/300
Qwen 2.5 0.5B InstructCurrent catalog45.42/100
Below threshold
300/300
Llama 3.2 1B InstructCurrent catalog41.67/100
Below threshold
300/300
DeepSeek R1 Distill Qwen 1.5BCurrent catalog11.0/100
Below threshold
300/300
LFM 2.5 1.2B InstructCurrent catalog64.33/100
Below threshold
300/300
Gemma 3 1B ITCurrent catalog58.5/100
Below threshold
300/300
MiniCPM5 1BCurrent catalog55.25/100
Below threshold
300/300
Granite 4.0 H 1BCurrent catalog75.67/100
Below threshold
300/300
MiniCPM5 2BCurrent catalog86.67/100
Below threshold
300/300
Ministral 3 3B InstructCurrent catalog68.83/100
Below threshold
300/300
Granite 4.2 3BCurrent catalog85.33/100
Threshold met
300/300
Phi-4 Mini InstructCurrent catalog59.67/100
Below threshold
300/300
SmolLM2 1.7B InstructCurrent catalog56.42/100
Below threshold
300/300
Qwen2.5 Coder 1.5B InstructCurrent catalog70.17/100
Below threshold
300/300
Qwen 3.5 0.8BEvaluation candidate61.17/100
Below threshold
300/300
Qwen 3.5 2BEvaluation candidate81.17/100
Below threshold
300/300

How scoring works

  • Instructions: 25%
  • Grounding: 20%
  • Conversation: 20%
  • Reasoning: 15%
  • Writing: 10%
  • Honesty: 10%

Each task has four criteria defined before testing. Exact answers and JSON are checked automatically. Writing and other subjective tasks are reviewed by the Codex assistant against a fixed rubric with model labels and timing hidden. The reviewer has run-order context; this is not independent or double-blind human validation. The score combines raw rubric points using the category weights; all three repetitions count.

Tests include corrections, conversation memory, changed instructions after Stop, summaries grounded in supplied text, and questions with missing information. Coding and speech recognition are separate from this score.

Eligibility requires at least 80 overall, at least 70 in every category, and at least 85 in instructions and conversation. A default candidate needs at least 85 overall. Passing quality thresholds does not establish release readiness.

Performance and reliability

Load time, first-token latency, generation speed, memory and thermal behavior are reported separately. Download recovery, offline operation and cancellation must pass independently. Results describe the tested artifact and app configuration; they do not establish performance on other iPhones.

Read the score summary data