BitsExplorer
Preview of Barakhsin Github Io

Barakhsin Github Io

Agent oversight assumes that when a system fails, the record it left is enough to tell you so. I measure whether that holds. Established LLM judges report a failure on clean agent traces between 34% and 78% of the time.

Description read from the page itself, because the listing it was found in gave none.

Open it Listing