Llama
截至 2026-09-05 14:00:03 UTC,Llama 模型家族近 7 日体感分为 56.7/100,基于 12 条合格社区反馈,置信度:中。
统计窗口:2026-08-29 14:00:03 至 2026-09-05 14:00:03 UTC · 评分方法:experience_score_v2。
30 天体感分曲线
分维度体感
| 通用文本 | 样本不足 | n=2 | |
|---|---|---|---|
| 编程 | 样本不足 | n=1 | |
| 推理 | 样本不足 | n=2 | |
| 图像/视觉 | 样本不足 | n=0 | |
| 视频生成 | 样本不足 | n=0 | |
| 角色扮演/创作 | 样本不足 | n=0 | |
| 安全与拒答 | 样本不足 | n=0 | |
| 速度与延迟 | 63.9 | n=6 | |
| 本地部署 | 样本不足 | n=1 |
主要好评因素
- 速度表现 速度与延迟 ×3
- local deploy feasibility 本地部署 ×1
- 性能问题 推理 ×1
- reliability preference 速度与延迟 ×1
主要差评因素
- 转投竞品 编程 ×1
- creativity down 通用文本 ×1
- 发布与上线 通用文本 ×1
用户原声
精选最近 7 日评论,优先考虑时间、证据质量、置信度与互动信号。
“有人拿2018年的MacBook Pro跑Llama 3 8B,比某些2024年的新款轻薄本还顺,就因为它有16GB内存。8GB内存的机器,不管你NPU多强,跑个3B的小模型也就4到6个token每秒,得等。”
来自 知乎 · 速度与延迟 · 2026-09-05
“The biggest pain point I have with intel gpus would be the prefill speed dropping to a tenth of the 0 context level, going from 1k to 150 as context grows”
来自 Reddit · 速度与延迟 · 英文 · 2026-09-04
“Their chip runs Llama 3.1 8B at 17k tok/s, and even if llama 3.1 is dated, I can think of many problems I could use it for, especially at those speeds.”
来自 Hacker News · 速度与延迟 · 英文 · 2026-09-04
“thats how i reached 27 tok/s. on llama cpp i was getting 17, unusable”
来自 Reddit · 速度与延迟 · 英文 · 2026-09-04
“llama.cpp on a 16GB M2 Pro, 27B Q1_0, 15.1 tok/s. Cloud can go down. The laptop does not.”
来自 Reddit · 速度与延迟 · 英文 · 2026-09-03