Simon Willison · 29 сентября 2026
Quoting Anthropic Frontier Red Team
We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview …
Почему это важно
Краткий обзор собран из открытых источников. Полные детали — по ссылке на первоисточник.
Читать первоисточник — Simon Willison
Мы не копируем чужой контент. Эта страница — честная подборка с прямой ссылкой на оригинал.