BENCHBAIT
WEEK OF 2026-08-17

REAL AI NEWS. UNNECESSARY HEADLINES.

THE WIRE

GROK + X + WEB · 10 CURRENT · EVERY JOKE HAS RECEIPTS

LEAD STORY · blog.google · REFRESHED 2026-08-22gemini 3.7 flash shipped three weeks after 3.6 at half price. the changelog is now a weekly circular.READ THE RECEIPT ↗

Ten new headlines per research pass. Grok searches current X posts first, then the web; uncited jokes do not make the front page.

OPEN THE ARCHIVE · 15 EARLIER DISPATCHES
ornith.ai · ARCHIVED 2026-08-22ornith-1.5 writes its own training tests. teachers everywhere have requested the source code.SOURCE ↗api-docs.deepseek.com · ARCHIVED 2026-08-22deepseek v4 flash grew eyes. screenshot benchmarks have begun dressing more carefully.SOURCE ↗claude.com · ARCHIVED 2026-08-22claude mythos 5 joined the security desk. the vulnerability scanner now has a security clearance.SOURCE ↗developer.nvidia.com · ARCHIVED 2026-08-22nvidia avo gave opus 5 a perfect arc-agi-3 score. the benchmark has asked for a recount.SOURCE ↗openai.com · ARCHIVED 2026-08-22openai paused frontier training for two weeks. even the models are being told to touch grass.SOURCE ↗claude.com · ARCHIVED 2026-08-22claude computer use is generally available. the office mouse is updating its resume.SOURCE ↗x.ai · ARCHIVED 2026-08-14grok 4.6 joined github copilot. the model picker now needs a model picker.SOURCE ↗x.ai · ARCHIVED 2026-08-12Grok 4.6 matched Sol on a benchmark, then doubled the free trial. science settled.SOURCE ↗x.ai · ARCHIVED 2026-08-11grok bot got its own computer. the browser tabs have requested collective bargaining.SOURCE ↗openai.com · ARCHIVED 2026-08-06GPT-5.6 Sol got a focus slider. alignment now has a volume knob.SOURCE ↗blog.google · ARCHIVED 2026-08-04google shipped three gemini agent models in july. the recap required its own agent.SOURCE ↗ai.meta.com · ARCHIVED 2026-08-01meta reached muse spark 1.2. llama 4 has been moved to the heritage collection.SOURCE ↗ai.meta.com · ARCHIVED 2026-07-09muse spark 1.1 got a million-token context window. meta finally found somewhere to put the terms of service.SOURCE ↗anthropic.com · ARCHIVED 2026-06-30Claude Fable 5 returned from export controls with the most frontier benchmark: paperwork.SOURCE ↗deepseek.com · ARCHIVED 2026-05-22deepseek v4 activated 49 billion parameters. the other 1.55 trillion are on a break.SOURCE ↗

THE AI LEADERBOARD NOBODY WILL AGREE WITH

PICK YOUR MODEL.
START A WAR.

Benchmarks set the starting grid. You get 1B tokens every day. Spend them all on one model and help it take the throne.

YOUR DAILY 1B IS READY

TOKEN THRONE · LIVE CAMPAIGN

THE THRONE

1 vote = 1B imaginary tokens

  1. Claude Fable 5Anthropic

    shut down by export controls. came back with paperwork.

    MYTHOS-CLASS SWE (ALLEGEDLY)
    DEFEND THE CROWN
    44B0B FROM FANS
  2. 02
    Claude Opus 5Anthropic

    the affordable Fable. still not affordable.

    POLITENESS-WEIGHTED IQ
    4B TO PASS #1
    41B0B FROM FANS
  3. 03
    GPT-5.6 SolOpenAI

    frontier intelligence that scales with your ambition (20% off).

    PRICE PER AMBITION
    4B TO PASS #2
    38B0B FROM FANS
  4. 04
    Grok 4.6SpaceXAI

    merged with a rocket company to feel something.

    PARAMS (TRILLION, TEASED)
    6B TO PASS #3
    33B0B FROM FANS
  5. 05
    Gemini 3.1 ProGoogle DeepMind

    still reading your email, now 3.1x faster.

    PRODUCT-INTEGRATION INDEX
    3B TO PASS #4
    31B0B FROM FANS
  6. 06
    Kimi K3 MaxMoonshot AI

    open weights, ships at 3am, beats the closed guys.

    LIVEBENCH (NOBODY ELSE'S)
    6B TO PASS #5
    26B0B FROM FANS
  7. 07
    Muse Spark 1.2Meta

    renamed Llama so we would forget Llama.

    REBRAND VELOCITY
    4B TO PASS #6
    23B0B FROM FANS
  8. 08
    DeepSeek V4DeepSeek

    #1 in usage, priced like a rounding error.

    OPENROUTER ADOPTION
    2B TO PASS #7
    22B0B FROM FANS
  9. 09
    GLM-5.3Z.ai

    shipped twice this month. leaderboards cannot rotate fast enough.

    RELEASE-CADENCE INDEX
    4B TO PASS #8
    19B0B FROM FANS
  10. 10
    Qwen3.8-MaxAlibaba

    3 billion downloads. still runs half of GitHub CI.

    DOWNLOADS (SELF-REPORTED)
    3B TO PASS #9
    17B0B FROM FANS
  11. 11
    Gemini 3.7 FlashGoogle

    newer, faster, already in everything you use.

    TOKENS PER FREE-TIER APOLOGY
    6B TO PASS #10
    12B0B FROM FANS
  12. 12
    Claude Opus 4.7Anthropic

    one of eight Anthropic rows in the top 20.

    PORTFOLIO CROWDING
    3B TO PASS #11
    10B0B FROM FANS
  13. 13
    Llama 4Meta

    presumably fine. nobody has checked since Muse Spark.

    OPEN WEIGHTS (VINTAGE)
    6B TO PASS #12
    5B1B FROM FANS

METHODOLOGY

RECEIPTS

01

BENCHMARKS START IT

Public benchmark standings seed each weekly bankroll. The seed is directional, not scientific.

02

FANS FINISH IT

Every browser gets one daily allocation. Give the full billion to one model. No splitting.

03

MONDAY RESETS IT

A fresh fight begins every Monday at midnight Pacific. Yesterday's dynasty becomes history.