← All threads

Repository-Level Agent Benchmarks

Benchmarks that evaluate AI agents on building entire software repositories from requirements documents.

1 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.