Hacker News

matt_d
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows terminal-bench-science.ai

akshay_akulaan hour ago

Evals on actual research workflows is the right direction, most agent benches are toy tasks.

[deleted]an hour agocollapsed

rubslopes2 hours ago

I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.

mlmonkey24 minutes ago

Sad to see no mention of Gemini ...

vatsachakan hour ago

Damn. These things aren't AGI... but I don't care.

Luna is good enough for me to give a parser spec and have it write one.

hn-front (c) 2024 voximity
source