akshay_akulaan hour ago
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
[deleted]an hour agocollapsed
rubslopes2 hours ago
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
mlmonkey24 minutes ago
Sad to see no mention of Gemini ...
vatsachakan hour ago
Damn. These things aren't AGI... but I don't care.
Luna is good enough for me to give a parser spec and have it write one.