Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> The tests are not visible to the model

> The model is allowed to write and run code in an agentic loop and iteratively self-correct during evaluation

What does this mean? How does the model iteratively self-correct without seeing the tests? Can it see the test results?



It isn’t allowed to see the final evaluation test (used in calculating its pass/fail), but it can run code and see the output of its own code in order to understand what doesn’t work. If it ends up creating tests as part of that based on the original problem statement then presumably that’s allowed.


Is this speculation or do you work at Anthropic? It would be cool to see the prompts used for this.


What he is describing has become the 'standard' way to run that kind of benchmark, so he is almost certainly correct. SWE Bench [1] is the best open source benchmark.

[1] https://www.swebench.com/




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: