
Models & ResearchThere's An AI For That · 58m ago
Benchmark study shows models struggle to apply newly discovered rules
Researchers evaluated language models in synthetic code and logic environments featuring rules intentionally omitted from training data. Although the models improved when allowed to experiment independently, their accuracy in executing newly learned rules remained highly inconsistent across testing iterations.
Ming Zhang
Read the original