Why I find it surprising is I can run Gemma 4 12b and Gemma 4 26b both QAT versions with some offloading at the 20-25 tok/sec range. Ling 3 Tiny should surely be much faster than them.
LFM2.5-2.6B - It turned out significantly better for me. I tested it a bit, though, specifically in terms of extracting and finding the desired text note in obsidian and ling 3.0 tiny shows more tool calls issues
I have had good luck using ling-3.0-tiny for the OpenZim MCP server for example, that model will work its way through the damn encyclopedia and pull out the answers (and I can verify where they came from). Sure Qwen3.8 does it just as well,
I find Ling 3.0 tiny particularly interesting as it looks really nice for a tiny model with 7.9B total parameters, with only 1.3B parameters activated per token
社区评论
精选近 7 天的不同意见,摘录条数不代表真实好评比例。
8 条精选摘录
展开原文
6gb here. Put Ling tiny in a harness and go wild
展开原文
Why I find it surprising is I can run Gemma 4 12b and Gemma 4 26b both QAT versions with some offloading at the 20-25 tok/sec range. Ling 3 Tiny should surely be much faster than them.
展开原文
Ling-3.0-Tiny gives me only 30 t/s while Ling-mini-2.0 gives me 50-60 t/s on CPU-only inference itself. Same with GPU-CUDA. 70-80 t/s vs 150+ t/s
展开原文
LFM2.5-2.6B - It turned out significantly better for me. I tested it a bit, though, specifically in terms of extracting and finding the desired text note in obsidian and ling 3.0 tiny shows more tool calls issues
展开原文
Ling-tiny is really cool
展开原文
I have had good luck using ling-3.0-tiny for the OpenZim MCP server for example, that model will work its way through the damn encyclopedia and pull out the answers (and I can verify where they came from). Sure Qwen3.8 does it just as well,
展开原文
I find Ling 3.0 tiny particularly interesting as it looks really nice for a tiny model with 7.9B total parameters, with only 1.3B parameters activated per token
展开原文
smarter than many 30b parameter models
近 7 天暂时没有符合此筛选条件的精选评论。