1 story in this blend

Researchers evaluated language models in synthetic code and logic environments featuring rules intentionally omitted from training data. Although the models improved when allowed to experiment independently, their accuracy in executing newly learned rules remained highly inconsistent across testing iterations.