Hacker News

dhorthy
Benchmarking Fable, Sol, and Kimi K3 on SlopCodeBench github.com

Bolwin4 hours ago

Two things

1. If we're using native harnesses, I'd have preferred you use kimi code, not opencode 2. The variation in the two kimi providers just shows how you can't trust n = 1 trials

hn-front (c) 2024 voximity
source