Llama
As of 2026-09-05 14:00:03 UTC, the Llama family's 7-day experience score is 56.7/100, based on 12 qualifying community feedback items. Confidence: Medium.
Data window: 2026-08-29 14:00:03 to 2026-09-05 14:00:03 UTC · Scoring method: experience_score_v2.
30-day experience score
By experience dimension
| General text | not enough data | n=2 | |
|---|---|---|---|
| Coding | not enough data | n=1 | |
| Reasoning | not enough data | n=2 | |
| Image / vision | not enough data | n=0 | |
| Video generation | not enough data | n=0 | |
| Roleplay / creative | not enough data | n=0 | |
| Safety & refusals | not enough data | n=0 | |
| Speed & latency | 63.9 | n=6 | |
| Local deploy | not enough data | n=1 |
Top positive drivers
- speed quality Speed & latency ×3
- local deploy feasibility Local deploy ×1
- performance issue Reasoning ×1
- reliability preference Speed & latency ×1
Top negative drivers
- competitor flight Coding ×1
- creativity down General text ×1
- release rollout General text ×1
What users are saying
Selected from the latest 7 days, prioritizing recency, evidence quality, confidence, and engagement.
“有人拿2018年的MacBook Pro跑Llama 3 8B,比某些2024年的新款轻薄本还顺,就因为它有16GB内存。8GB内存的机器,不管你NPU多强,跑个3B的小模型也就4到6个token每秒,得等。”
via Zhihu · Speed & latency · Chinese · 2026-09-05
“The biggest pain point I have with intel gpus would be the prefill speed dropping to a tenth of the 0 context level, going from 1k to 150 as context grows”
via Reddit · Speed & latency · 2026-09-04
“Their chip runs Llama 3.1 8B at 17k tok/s, and even if llama 3.1 is dated, I can think of many problems I could use it for, especially at those speeds.”
via Hacker News · Speed & latency · 2026-09-04
“thats how i reached 27 tok/s. on llama cpp i was getting 17, unusable”
via Reddit · Speed & latency · 2026-09-04
“llama.cpp on a 16GB M2 Pro, 27B Q1_0, 15.1 tok/s. Cloud can go down. The laptop does not.”
via Reddit · Speed & latency · 2026-09-03