Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

(senior-swe-bench.snorkel.ai)

35 points | by matt_d 2 hours ago ago

15 comments