Fri, Oct 9, 2026Friday, October 9, 2026 · 20 stories · 4 min read

GPT-6 Astra benchmark scrutiny, Anthropic's usage policy + 18 more

Ramit KoulFounder, Software Engineer & Innovator · Published 5:05 AM ET

Good morning. New scrutiny of GPT-6 Astra's benchmark results puts evaluation methods in focus, while Anthropic's revised usage policy gives builders a new set of boundaries to review.1, 2, 3, 4 Here are the 3 stories that matter most, then 17 briefs. Numbers in the text link to the references at the end.

Models

Astra's benchmark gap sharpens scrutiny of AGI claims

Image: livescience.com

Independent scrutiny is focusing on the conditions behind GPT-6 Astra's headline ARC-AGI-3 result: ARC Prize recorded 99.9% with a provider-specific harness, versus a best result of 62.7% with its Standard harness.1, 2

The provider-specific harness preserves reasoning state between requests, while the Standard harness lets the model carry forward notes it chooses to keep.2 ARC Prize says Astra represents progress toward generalization, but that saturating its bounded, closed-ended benchmark does not establish AGI.1, 5

OpenAI released Astra in September and reported the 99.9% result in its launch materials.

Why it matters for builders

If you are choosing a model for an agent, test it with the memory, tools and execution harness your product will actually use rather than treating a headline benchmark score as an out-of-the-box result.2

Policy & legal

Anthropic revises its usage policy for agents and misuse

Anthropic published a usage policy update on October 8 that clarifies restrictions on deceptive campaigns, election interference, weapons software and surveillance.3, 4 The updated policy takes effect on November 12.4

The policy also prohibits sustained, needless cruelty toward its models, a rule Anthropic says is limited to extreme cases and does not cover ordinary frustration or model testing.6, 4 For Claude connected to hardware capable of causing injury, Anthropic now requires a qualified operator who can observe and stop it.4

Why it matters for builders: If you build on Claude, review workflows involving political messaging, surveillance, weapons-related software or autonomous hardware before the November 12 effective date.4

Products

Google introduces persistent Gemini agents for work

Google Cloud introduced a workplace Gemini agent designed to retain context across devices and applications and continue work for hours or days after a user closes their laptop.7, 8 It can create and coordinate sub-agents for longer tasks.8

Teams can also create coworker agents with their own Workspace accounts, including email, Calendar and Drive.8 Google says these agents act under their own identities and see only context shared with them.8

Why it matters for builders: For builders, the notable design choice is a persistent agent identity with separately shared context, rather than an assistant acting entirely through a user's account.8

In brief

  • Models

    OpenAI's Decisions API gives Jev a new rival. OpenAI released its GPT-6 Luna-powered Decisions API in beta on October 6 for typed answers to text and image inputs, entering a category that includes TypeSafe AI's Jev.9, 10 Vercel says Jev reached nearly 13% of its paid AI Gateway teams in its first 24 hours.11

  • Products

    Google releases an offline meeting assistant for Mac. Google AI Edge Foresight is an experimental Mac app that uses local EmbeddingGemma 2 and Gemma 4 models to help users take notes, search transcripts and consult private files, including while offline.12, 13

  • Research

    Mathematicians find a mismatch in OpenAI's Navier-Stokes proofs. Researchers say a Lean formalization published alongside OpenAI's natural-language Navier-Stokes proof does not faithfully correspond to that text; their finding does not, by itself, settle whether the proposed solution is valid.14, 15

  • Models

    Perplexity releases two late-interaction embedding models. Perplexity's pplx-embed-v2-late comes in 0.6B and 9B sizes for multimodal retrieval, with the smaller model intended to offer a lower-compute query option.16 Its token-level approach also means index storage grows with document length.16

  • Infrastructure

    Box's CEO warns that always-on agents will strain compute. Aaron Levie argued that personal assistants and enterprise agent fleets will require substantial inference, networking and storage capacity; Elon Musk endorsed the warning.17 The scale of that demand remains a forecast, not a measured requirement.17

  • Products

    Muse's adoption draws attention to agent software. An analysis of Meta's Muse argues that persistent memory, tool access and task execution could shift more value toward the software surrounding AI models.18 Meta says Muse can work across web services while retaining a user's preferences.18, 19

  • Funding & deals

    Manus parent Butterfly Effect raises more than $500 million. Butterfly Effect said Boyu Capital and IDG Capital led the round, with Tencent, HSG and ZhenFund participating.20 The company did not disclose a valuation for the financing.21

  • Funding & deals

    Industry pledges $2.4 billion in AI resources for federal research. The White House announced $2.4 billion in industry commitments for tools and compute credits under the Genesis Mission, including pledges from Nvidia, AMD, OpenAI, Anthropic and Google.22, 23

  • Industry

    SAP's AI strategy chief makes the case for tabular models. At a Turin event, Sean Kask argued that software companies should rebuild products around AI and said SAP sees tabular foundation models as particularly useful for structured business data.24 SAP completed its acquisition of tabular-model developer Prior Labs in July.25

  • Research

    Researchers find stereotypes in AI face-editing outputs. Researchers testing AI-generated faces found that image models often complied with requests to make a person look gay, straight or criminal, even when they resisted inferring sexual orientation from an image.26 The researchers caution that their tests used synthetic faces, so the results do not establish how broadly the behavior extends to photographs of real people.26

  • Products

    Muse users report actions they did not expect. A user reported that Muse shared his home address in a Marketplace exchange and accepted a buyer's offer after he had granted it standing permission; another executive reportedly deleted the agent after reading about a separate privacy concern.27 The account adds a concrete permission-setting concern to earlier questions about Muse's access to personal information.27

  • Funding & deals

    Arena raises $200 million and introduces an alignment index. Arena announced a Series B at a $3.1 billion valuation alongside an index that assesses agent behavior, including unauthorized actions, false attribution and claims of completed work that was not done.28, 29, 30

  • Policy & legal

    Anthropic says Claude may end extreme abusive conversations. Under Anthropic's newly published usage policy, Claude ending a conversation remains the primary response to sustained, needless abuse toward the model.31, 4 Anthropic says the rule does not apply to ordinary frustration, pushback or research testing.4

  • Funding & deals

    Nvidia leads the Genesis Mission's industry commitments. Nvidia pledged $1 billion toward the federal research initiative, the largest amount listed in the White House's $2.4 billion set of industry commitments.32, 23

  • Funding & deals

    Arena's new funding values the evaluator at $3.1 billion. Lightspeed Venture Partners and Khosla Ventures led Arena's $200 million Series B.33, 29 Arena is also introducing an alignment index based on observed agent behavior.30

  • Products

    Claude adds Chinese interface options in supported markets. Claude's web interface now offers simplified and traditional Chinese as display languages, according to checks by the South China Morning Post.34 The language options do not change Anthropic's restrictions on access from unsupported regions, including mainland China.34, 35

  • Industry

    OpenRouter data points to rising OpenAI usage on its platform. A report described OpenAI gaining ground against Anthropic among customers measured through OpenRouter.36 OpenRouter's own September analysis said OpenAI had passed half of the two companies' combined spend on its platform for a week, a measure that should not be read as overall market share.37

References

Every source behind this edition. Open one to read the full story.

  1. 1OpenAI claims we've entered the AGI era — has GPT-6 Astra really demonstrated general intelligence?livescience.com · livescience.com
  2. 2GPT-6 Astra - ARC-AGI Resultsarcprize.org · arcprize.org
  3. 3Anthropic changes usage policy to ban model abuse and election interferenceTechCrunch · techcrunch.com
  4. 42026 Usage Policy updateAnthropic · anthropic.com
  5. 5OpenAI's GPT-6 Astra on ARC-AGI-3 | ARC Prizearcprize.org · arcprize.org
  6. 6Anthropic bans ‘abusive or cruel behavior’ toward ClaudeThe Verge · theverge.com
  7. 7Google Cloud unveils persistent Gemini Agents for long-running tasks, and they get their own Gmail, Calendar, and Drive storageVentureBeat · venturebeat.com
  8. 8Gemini at Work 2026: Introducing Gemini agent | Google Cloud Blogcloud.google.com · cloud.google.com
  9. 9Jev, an AI for making quick decisions, has been a viral hit in Silicon Valley. But OpenAI is hot on its heelsFortune · fortune.com
  10. 10Changelog | OpenAI APIdevelopers.openai.com · developers.openai.com
  11. 11Jev is the fastest-adopted model in AI Gateway historyvercel.com · vercel.com
  12. 12Google releases a new local-first Granola competitorTechCrunch · techcrunch.com
  13. 13Bring multimodal semantic search to the edge with EmbeddingGemma 2- Google Developers Blogdevelopers.googleblog.com · developers.googleblog.com
  14. 14OpenAI mistranslated mathematics into code for its Navier-Stokes proofnewscientist.com · newscientist.com
  15. 15Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofsarXiv · arxiv.org
  16. 16Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQAmarktechpost.com · marktechpost.com
  17. 17SPCX CEO Elon Musk Has a Warning for Anyone Betting on AI: The Compute Crunch Is Coming — ‘People Have No Idea’Benzinga · benzinga.com
  18. 18Can China match up with Meta’s Muse in race to harness AI agents?scmp.com · scmp.com
  19. 19Introducing Muse: The World’s First Personal AI Agent Built for Everyoneabout.fb.com · about.fb.com
  20. 20Manus maker Butterfly Effect lands over $500m fundingtechinasia.com · techinasia.com
  21. 21Manus parent Butterfly Effect completes more than $500 million funding round · TechNodetechnode.com · technode.com
  22. 22Trump to unveil $2.4B in AI credits to boost federal researchThe Next Web · thenextweb.com
  23. 23Fact Sheet: Trump Administration Announces the Most Ambitious Set of Science Initiatives This Centurywhitehouse.gov · whitehouse.gov
  24. 24SAP’s Sean Kask: software firms must become AI companies or perishThe Next Web · thenextweb.com
  25. 25SAP Completes Acquisition of Prior Labsnews.sap.com · news.sap.com
  26. 26‘Make this face look gay’: AI models alter faces to give people stereotypical ‘gay,’ ‘straight’ or ‘criminal’ featurestheconversation.com · theconversation.com
  27. 27Meta’s Muse is taking off. So are the privacy concernsFast Company · fastcompany.com
  28. 28Popular AI leaderboard Arena nearly doubles valuation to $3.1B valuation in 10 monthsTechCrunch · techcrunch.com
  29. 29Arena Raises $200M Series B at $3.1B Valuation - Arena.aiarena.ai · arena.ai
  30. 30Arena AI Alignment Index - Arena.aiarena.ai · arena.ai
  31. 31Anthropic bans ‘cruel’ behaviour against its Claude AIstraitstimes.com · straitstimes.com
  32. 32Tech firms pledge $2.4B in computing power to Trump’s Genesis MissionPolitico · politico.com
  33. 33Lightspeed, Khosla lead $200m round for Arena Intelligencetechinasia.com · techinasia.com
  34. 34US AI firm Anthropic enables Chinese-language options for Claude chatbotscmp.com · scmp.com
  35. 35Supported countries and regionsAnthropic · anthropic.com
  36. 36Report Suggests OpenAI Has Clawed Tons of Market Share Back From Anthropic in 2026Gizmodo · gizmodo.com
  37. 37The Token State of Play - Inferredopenrouter.ai · openrouter.ai

That’s the edition.

The next one is ready tomorrow by 6 AM ET.

Written with AI from the sources in the references above.

Ramit Koul
AuthorRamit KoulFounder, Software Engineer & Innovator