
TL;DR
Kimi K3, a 2.8 trillion parameter model from Moonshot AI, ranks second on the AA-Briefcase benchmark, achieving an Elo score of 1543, just behind Fable 5. Despite its high performance, it is more expensive to run than Opus 4.8 and takes nearly an hour per task.
✦ Why It Matters
Evaluate Kimi K3 for your next project requiring advanced analytical capabilities in knowledge work.
Key Takeaways
Full Summary
Kimi K3 is a new AI model developed by Moonshot AI, featuring 2.8 trillion parameters and scoring 57 on the Artificial Analysis Intelligence Index. It achieved an Elo score of 1543 on the AA-Briefcase benchmark, which evaluates models on realistic knowledge work tasks like spreadsheets and presentations.
This score represents a significant improvement of 727 points over its predecessor, Kimi K2.6. However, Kimi K3 is among the most expensive models to operate, costing $10.57 per task and averaging 56.4 minutes to complete each task.
Its performance is strong in analytical quality but weaker in presentation quality compared to competitors. The high cost and time per task are driven by the model's token pricing and output requirements, making it a consideration for organizations looking to implement advanced AI solutions.
Related