The week’s clearest result is that reliable agents require runtime constraints, process-aware evaluation, and targeted correction—not stronger instructions alone. ‌ ‌ ‌ ‌ ‌