Saturday, August 8, 2026

Video on AI benchmark testing maintenability

 


New benchmark that is far from being saturated.  Called slop code bench.  It tests error accumulation and maintainability as agents continue working on same code base.

This is the sort of benchmark that hopefully sees great progress.   If models cannot be left alone for long without the code devolving their usefulness cannot expand to take on more roles.

No comments:

Post a Comment