Claude Opus 4.8 review

    by Estebankiwi: Anthropic

    I burned through 2 billion tokens on Claude Opus 4.8 over a 48-hour period. Here is my honest assessment. The good news is that it ranks as the most intelligent model in the world according to Artificial Analysis. It also stands as the best coding model currently available. The system builds applications that GPT 5.5 and Opus 4.7 simply could not manage. The downside is that this release does not represent any major advance. It carries the same hallucination rate, the same knowledge cutoff date, and the same token consumption level as version 4.7. Anthropic itself described the update as modest but tangible. The real problem emerged when I deployed 105 UltraCode agents into production. Version 4.8 added a bug that had not existed before. It then failed to resolve that same bug across eight separate attempts. The entire project required a complete rollback. These systems can create polished user interfaces and deliver new capabilities without trouble. Yet they offer no protection when unexpected fail