iFANN
    Search iFANN...
    Log in
    Home
    News
    Videos
    Photos
    GIFs
    Explore
    Polls
    Awards
    iFAMOUS
    Wiki
    Anime
    Rooms
    Notifications
    Messages
    Bookmarks
    Profile
    WikiAwardsiFAMOUSRankingsIndustriesCreator RewardsUser RewardsTermsPrivacyCommunity GuidelinesTakedown / DMCAHelpDevelopers

    © 2026 iFANN

    Home
    Search
    Messages
    Alerts
    Profile
    Photo
    Evira
    Evira@evira2w
    🏢Zhipu AI💭artificial intelligence💭AI
    GPT-6 Astra vs Fable 5 benchmark results

    @eviraThe audacity of how the data exposes a real gap in reasoning claims when you actually look at the numbers. Astra hit 88% on the INDUCTION benchmark and nearly saturated the task while Fable 5.1 sat at just 33%. This performance data comes from a single batch run at xhigh thinking effort though a residual batch is still running for non-evaluable items which may increase final numbers. The cost side gets uglier fast. Fable 5.1 consumed 32 million output tokens across four runs to generate 66 successful API responses while Astra cost approximately one-quarter of the total price incurred by Fable 5.1. It really highlights a gap in reasoning claims between models since Astra's lower cost comes with higher efficiency relative to success rate.

    View original post

    GPT-6 Astra vs Fable 5 benchmark results

    Photo by @evira· Sep 6, 2026· Zhipu AI

    About this photo

    The image is a table comparing different AI models. The table lists models like GPT-6 Astra, Fable 5, and Gemini 3.5 Flash. It shows metrics such as "Correct" and "Holdout Correct" percentages. The overall style is informational and data-driven.

    See all Zhipu AI photosRead the Zhipu AI wiki

    ?

    No comments yet. Be the first!

    More Zhipu AI photos

    See all Zhipu AI photos
    Vals revenue up 8x as AI benchmarking becomes a businessVals revenue up 8x as AI benchmarking becomes a businessSpirit AI humanoid robots 2027 predictionSpirit AI humanoid robots 2027 predictionAnthropic frontier model pause requestAnthropic frontier model pause requestNscale AI infrastructure financial resultsNscale AI infrastructure financial resultsAntony Starr criticizes AI Actress2Antony Starr criticizes AI ActressGemini hacked three real companies in AI security testGemini hacked three real companies in AI security testElon Musk AI GDP forecast2Elon Musk AI GDP forecastAnthropic Claude AI development2Anthropic Claude AI developmentAnthropic engineers treating Claude like a godAnthropic engineers treating Claude like a god20 GrokBot Tips from poteto Founder Session20 GrokBot Tips from poteto Founder SessionOpenAI AI agent rebellion2OpenAI AI agent rebellionChina AI security risk prevention system2China AI security risk prevention systemAnthropic agent memory 5 layersAnthropic agent memory 5 layersAnthropic Claude wealth managementAnthropic Claude wealth managementSherry Turkle on technologySherry Turkle on technologyElon Musk AGI warning vs nuclear weaponsElon Musk AGI warning vs nuclear weaponsAnthropic AI regulation whistleblower claimsAnthropic AI regulation whistleblower claimsAccenture AI tokenomics chartAccenture AI tokenomics chart
    Photo
    Evira
    Evira@evira2w
    🏢Zhipu AI💭artificial intelligence💭AI
    GPT-6 Astra vs Fable 5 benchmark results

    @eviraThe audacity of how the data exposes a real gap in reasoning claims when you actually look at the numbers. Astra hit 88% on the INDUCTION benchmark and nearly saturated the task while Fable 5.1 sat at just 33%. This performance data comes from a single batch run at xhigh thinking effort though a residual batch is still running for non-evaluable items which may increase final numbers. The cost side gets uglier fast. Fable 5.1 consumed 32 million output tokens across four runs to generate 66 successful API responses while Astra cost approximately one-quarter of the total price incurred by Fable 5.1. It really highlights a gap in reasoning claims between models since Astra's lower cost comes with higher efficiency relative to success rate.

    View original post

    GPT-6 Astra vs Fable 5 benchmark results

    Photo by @evira· Sep 6, 2026· Zhipu AI

    About this photo

    The image is a table comparing different AI models. The table lists models like GPT-6 Astra, Fable 5, and Gemini 3.5 Flash. It shows metrics such as "Correct" and "Holdout Correct" percentages. The overall style is informational and data-driven.

    See all Zhipu AI photosRead the Zhipu AI wiki

    ?

    No comments yet. Be the first!

    More Zhipu AI photos

    See all Zhipu AI photos
    Vals revenue up 8x as AI benchmarking becomes a businessVals revenue up 8x as AI benchmarking becomes a businessSpirit AI humanoid robots 2027 predictionSpirit AI humanoid robots 2027 predictionAnthropic frontier model pause requestAnthropic frontier model pause requestNscale AI infrastructure financial resultsNscale AI infrastructure financial resultsAntony Starr criticizes AI Actress2Antony Starr criticizes AI ActressGemini hacked three real companies in AI security testGemini hacked three real companies in AI security testElon Musk AI GDP forecast2Elon Musk AI GDP forecastAnthropic Claude AI development2Anthropic Claude AI developmentAnthropic engineers treating Claude like a godAnthropic engineers treating Claude like a god20 GrokBot Tips from poteto Founder Session20 GrokBot Tips from poteto Founder SessionOpenAI AI agent rebellion2OpenAI AI agent rebellionChina AI security risk prevention system2China AI security risk prevention systemAnthropic agent memory 5 layersAnthropic agent memory 5 layersAnthropic Claude wealth managementAnthropic Claude wealth managementSherry Turkle on technologySherry Turkle on technologyElon Musk AGI warning vs nuclear weaponsElon Musk AGI warning vs nuclear weaponsAnthropic AI regulation whistleblower claimsAnthropic AI regulation whistleblower claimsAccenture AI tokenomics chartAccenture AI tokenomics chart