Model lifecycle

Kimi K3

Released 2026-07-16 · Back to Kimi

Exact-version signals observed 2026-08-07–2026-09-03 · n=481

Data window: 2026-08-07 00:00:00 to 2026-09-04 00:00:00 UTC · Scoring method: experience_score_v2.

Trend history still collecting The current window is ready; 2 of 4 independent segments are available.

Lifecycle updated daily · latest complete UTC day 2026-09-03

Fixed 14-day baseline
60.6
n=156
Latest 28 days
62.7
n=481
Change vs baseline
+2.1
90-day slope
pts / 30 days

14-day rolling experience

Each daily point summarizes the previous 14 complete UTC days. The dashed line is the fixed release baseline; the verdict still uses independent non-overlapping 14-day segments.

14-day rolling experience from 2026-08-13 to 2026-09-03; latest score 64.3, fixed baseline 60.6.40506070Fixed baseline 60.62026-07-30–2026-08-13 · 60.6 · n=1562026-07-31–2026-08-14 · 60.5 · n=1832026-08-01–2026-08-15 · 61.8 · n=1922026-08-02–2026-08-16 · 61.1 · n=2072026-08-03–2026-08-17 · 60.5 · n=2282026-08-04–2026-08-18 · 61.3 · n=2422026-08-05–2026-08-19 · 61.0 · n=2682026-08-06–2026-08-20 · 61.2 · n=2802026-08-07–2026-08-21 · 62.6 · n=2812026-08-08–2026-08-22 · 62.2 · n=2792026-08-09–2026-08-23 · 60.8 · n=2702026-08-10–2026-08-24 · 61.9 · n=2662026-08-11–2026-08-25 · 62.5 · n=2622026-08-12–2026-08-26 · 64.2 · n=2452026-08-13–2026-08-27 · 64.0 · n=2162026-08-14–2026-08-28 · 64.8 · n=1932026-08-15–2026-08-29 · 64.5 · n=1992026-08-16–2026-08-30 · 65.5 · n=1902026-08-17–2026-08-31 · 64.2 · n=1952026-08-18–2026-09-01 · 61.7 · n=2072026-08-19–2026-09-02 · 64.0 · n=1982026-08-20–2026-09-03 · 64.3 · n=20008-1308-1708-2108-2508-2909-0209-03

2 of 4 independent 14-day segments ready for trend judgment.

What explains the change

Experience change and discussion-mix change are separated. These are observational contributions, not proof of cause.

Largest negative experience contributions: Speed & latency.

Category Baseline → current Category change Weight share Experience contribution Discussion-mix contribution
Speed & latency Baseline → current34.9 → 34.1n=30 → 95 Category change−0.8 Weight share21.4% → 20.7% Experience contribution−0.16 Discussion-mix contribution+0.18
Coding Baseline → current63.6 → 68.5n=20 → 60 Category change+5.0 Weight share13.4% → 13.1% Experience contribution+0.65 Discussion-mix contribution−0.02
Reasoning Baseline → current80.3 → 86.4n=36 → 97 Category change+6.1 Weight share24.1% → 20.8% Experience contribution+1.36 Discussion-mix contribution−0.77
General text Baseline → current57.6 → 65.2n=53 → 189 Category change+7.7 Weight share31.4% → 37.7% Experience contribution+2.64 Discussion-mix contribution+0.07
Image / vision Baseline → currentnot enough datan=0 → 3 Category change Weight share Experience contribution Discussion-mix contribution
Local deploy Baseline → currentnot enough datan=7 → 18 Category change Weight share Experience contribution Discussion-mix contribution
Roleplay / creative Baseline → currentnot enough datan=4 → 9 Category change Weight share Experience contribution Discussion-mix contribution
Safety & refusals Baseline → currentnot enough datan=6 → 10 Category change Weight share Experience contribution Discussion-mix contribution
Video generation Baseline → currentnot enough datan=0 → 0 Category change Weight share Experience contribution Discussion-mix contribution

Unresolved contribution from categories below the sample threshold: −1.86 pts.

Lifecycle scores use the same experience-signal weights as the main index, without sample shrinkage after n=30. They measure public user perception, not model capability or backend causes.