OpenAI, Anthropic, and Google Plan to Establish Frontier AI Safety… | AI Daily 2026-09-26

🔥 Spotlight

OpenAI, Anthropic, and Google Plan to Establish Frontier AI Safety Organization SAFA : According to The Information, OpenAI, Anthropic, and Google are preparing to establish a frontier AI safety organization called the “Standards Authority for Frontier AI” (SAFA), planned for an official launch in early 2027. The organization aims to coordinate independent third-party testing prior to model deployment and formulate industry incident notification mechanisms and benchmarks, serving as industry self-regulation and a standards baseline amid growing risks of frontier model control loss and tightening global regulations (Source: The Verge)

White House Directs OpenAI and Anthropic to Prioritize US Review of Frontier Models : The White House Office of the National Cyber Director has instructed OpenAI and Anthropic to submit their latest frontier models for review by US official agencies before sharing them with the UK AI Safety Institute (AISI). Anthropic has adjusted accordingly, restricting certain frontier versions exclusively to US bodies. This move underscores direct geopolitical competition over frontier AI core evaluation and national security dominance, while introducing institutional friction into US-UK cross-border safety evaluation coordination (Source: THE DECODER)

Anthropic Proposes Super-Voting Share Structure: 7 Co-Founders with ~2% Stakes to Control 50.1% Voting Rights : Ahead of a potential trillion-dollar IPO, Anthropic is seeking shareholder approval for a new governance structure. Even if the 7 co-founders, including CEO Dario Amodei, each hold only about 2% of equity, as long as at least 3 maintain qualifying stakes, they can collectively control 50.1% of the company’s voting rights through a Palantir-style founder voting trust. The structure is designed to decouple economic interest from decision-making control, shielding long-term AI alignment and safety research from short-cycle Wall Street profit pressures (Source: JiQizhixin, 36Kr, kimmonismus)

Anthropic Proposes Super-Voting Structure

Google Launches Pilot for Suncatcher Space-Based Solar AI Compute Center : Google announced that the first experimental MVP satellite for its orbital AI data center, “Project Suncatcher,” is scheduled to launch on October 1 aboard a SpaceX Falcon 9 (Transporter-18 mission). Equipped with four custom Trillium TPU accelerators and powered by solar energy, the satellite aims to validate the feasibility of running AI inference in extreme space environments leveraging 8x the solar energy available on Earth. Core tests will focus on vacuum heat dissipation and high-efficiency orbital computing, with inter-satellite high-speed laser links slated for testing in 2027 (Source: Ars Technica, Google, dotey)

OpenAI Reportedly Prepares $500/Month ChatGPT Pro Max Subscription Tier : Developers uncovered a new “$500/month Pro Max” subscription tier in the ChatGPT front-end code, highlighting “the fastest Work and Codex execution speeds” and potential access to dedicated low-latency compute. Following high-load subscription throttling caused by GPT-6 Astra’s heavy inference compute consumption, OpenAI is using premium pricing to segment compute for heavy productivity workloads, indicating that cutting-edge compute is rapidly concentrating among high-net-worth enterprises and professional institutions (Source: 36Kr, nicdunz)

ChatGPT to launch $500 Pro Max tier

Google Releases Gemini 3.8 Live and Real-Time Avatar Model : Google officially launched Gemini 3.8 Live along with its Live Avatar feature, deeply integrating real-time streaming conversations with low-latency dynamic avatar video generation. The model supports native audio-video lip synchronization across 97 languages and asynchronous background tool invocation, offering customizable avatars for enterprise interactive scenarios with full SynthID watermarking integration (Source: Google DeepMind Blog)

Meta Introduces Real-Time Embodied Model Muse Realtime Avatar : Meta unveiled the Muse Realtime Avatar streaming architecture, which shares audio token streams with Realtime Voice to compress end-to-end video conversation response latency down to approximately 870ms. The system maintains bounded computational overhead via a fixed-length historical context, supporting continuous lip-sync and micro-expression synchronization across arbitrarily long conversations (Source: alex_conneau, jack_w_rae)

Muse Realtime Avatar

Meta Muse Embroiled in Privacy Controversy Over Unauthorized Use of Human Contractors for Calls : Reuters revealed that while testing autonomous phone calling capabilities for its newly released personal agent Muse, Meta secretly routed tasks to outsourced human teams to maintain high success rates, sparking severe user privacy concerns and regulatory scrutiny. Meta executives promptly issued an apology and took the test offline, committing not to reopen the channel until privacy disclosures and safety guardrails are fully established (Source: QbitAI)

Father of Modern AI Jürgen Schmidhuber Joins Sakana AI as Chief Scientific Advisor : Recursive Self-Improvement (RSI) and meta-learning pioneer Jürgen Schmidhuber has officially joined Japanese AI startup Sakana AI as Chief Scientific Advisor. He will lead the newly established RSI Lab in Tokyo, combining evolutionary algorithms, world models, and self-organizing architectures to explore adaptive intelligence and physical-world RSI closed loops under compute constraints (Source: SakanaAILabs, hardmaru)

Jürgen Schmidhuber joins Sakana AI as Chief Scientific Advisor

LangChain Releases LangSmith Engine v2 and Managed Deep Agents 0.8 : LangSmith Engine v2 introduces proactive red teaming and automatic verification & remediation mechanisms to actively detect and fix anomalous logic before agents go live; Managed Deep Agents 0.8 adds enterprise user persistent memory permission policies, sandboxed File APIs, and native web search capabilities (Source: LangChain, hwchase17)

LangChain launches new products at Interrupt conference

Perplexity Launches In-House Rust-Based Retrieval Engine Photon and Fast Search API : Perplexity launched Photon, an in-house retrieval and reranking engine built in Rust that delivers ultra-fast search with a p95 latency under 230ms and a p50 latency of just 160ms. It slashes API invocation costs to $1 per 1,000 queries and has been seamlessly integrated into Hermes Agent and local AI ecosystems (Source: AravSrinivas, MParakhin)

Perplexity launches Fast Search

Open-Source AI-SQL Joint Query Engine Quail Released: Over 1 Billion Tokens/Min on a Single GPU : Teams from Stanford, Modal, and collaborators open-sourced Quail, an AI-SQL query engine designed for AI data operations. By refactoring the vLLM scheduler and merging query execution plans with the LLM prefill phase alongside custom KV caches and attention kernels, Quail achieves a throughput exceeding 1 billion input tokens per minute for a single query on a single H100 (Source: charles_irl, HamelHusain)

Google Tests Gemini “Call for Me” Feature on Pixel Devices : Google is testing a “Call for Me” feature on Pixel 11 in the US, allowing Gemini to autonomously place calls to businesses using the device owner’s actual phone number. It handles inventory inquiries, restaurant reservations, and rescheduling while automatically navigating automated voice menus and hold queues, offering real-time transcription and instant manual takeover (Source: TechCrunch)

Alibaba T-Head Fully Open-Sources Zhenwu AI Chip Software Stack T-Head SAIL : At the Apsara Conference, Alibaba T-Head announced the latest open-source milestones for its CUDA-like software stack, T-Head SAIL. It opens the PyTorch adaptation layer, the sailify migration tool, and acceleration libraries such as DeepGEMM and FlashAttention, empowering developers to perform operator-level tuning and seamless migration for the Zhenwu architecture (Source: QbitAI)

NIO Shaoqing Ren Team Proposes MM-Future: A Multimodal World-Action Model for Autonomous Driving : NIO’s research team introduced the MM-Future architecture for autonomous driving planning, enabling the joint evolution of multiple “scenario-action pair hypotheses.” Through compact spatio-temporal representations (MM-Tokens) and a joint history-future scorer, it achieved 94.0 PDMS on the NAVSIM-v1 benchmark, advancing world models from unidirectional passive prediction to closed-loop active planning (Source: QbitAI)

Kuaishou Releases E-Commerce Image Editing Foundation Model KwaiMind : Kuaishou released KwaiMind, an image editing foundation model tailored for real-world e-commerce assets. Utilizing multi-agent data engines and online distillation, it harmonizes garment details, product consistency, and text rendering, ranking first among open-source models across 11 e-commerce benchmarks and delivering a relative CTR boost of ~2.44% in online A/B testing (Source: JiQizhixin)

Liquid AI Open-Sources LFM2.5-VL-DSpark Draft Speculative Decoding Model : Liquid AI introduced the DSpark speculative decoding scheme for 3B vision-language models. Adding just 280M parameters, it achieves up to 3.13x decoding speedups on-device (M5 Max) while ensuring mathematically equivalent outputs, with concurrent support for llama.cpp and SGLang (Source: HuggingFace Blog)

Qwengram-0.8B Open-Sourced: Losslessly Migrating 51B Explicit N-Gram Memory to Tiny On-Device Models : Developers have open-sourced the Qwengram architecture. Keeping the Qwen3.5-0.8B backbone frozen, they successfully attached an external 51B PLE N-gram memory plugin via a lightweight Reader and dynamic gating mechanism, achieving a 5.05% perplexity (PPL) reduction on the validation set (Source: Reddit r/LocalLLaMA)

PrismML Brings 1-Bit Ultra-Lightweight LLM to Qualcomm Snapdragon Smart Glasses : PrismML, specializing in model fine-tuning and compression, demonstrated a 2B-parameter 1-bit Bonsai model running on the Snapdragon AR1 Gen 1 chip. It enables on-device real-time visual QA and natural interaction, proving the engineering feasibility of running multimodal models on-device under ultra-low power constraints (Source: TechCrunch)

🧰 Tools

Kimi Browser Extension Launches with Support for Converting Web Actions into Reusable Skills : Kimi officially upgraded WebBridge into a native browser extension, introducing a sidebar interactive chat and enabling users to record continuous actions (clicks, pagination, data scraping) as reusable Skill directives, creating an integrated workflow with Kimi Code Desktop (Source: QbitAI)

GLiNER2.5-Decide Open-Sourced: Ultra-Fast Decision Model Supporting Entity Relations and Logical Constraints : Fastino open-sourced GLiNER2.5-Decide (340M), a lightweight structured decision model. It supports fast routing and classification while extracting character-level spans and entity relationships, performing real-time logical validation on mutual exclusivity, cardinality, and ordinal constraints. Inference latency is just 167ms on CPU and 38–47ms on GPU, completely replacing traditional LLM long-text generation overhead for decision-making (Source: MarkTechPost, ClementDelangue)

GLiNER2.5-Decide open sourced

LangSmith Launches Open-Source Fine-Tuning Tool smithtune and Custom Apps Platform : LangChain, in partnership with Baseten and Fireworks, launched the smithtune CLI tool to directly convert production LangSmith session traces into SFT fine-tuning datasets and automate model deployment. Simultaneously, the new Custom Apps feature allows developers to generate annotation queues or metrics comparison dashboards directly from agent run logs via prompts (Source: LangChain, baseten)

LangSmith launches Custom Apps and smithtune

CoreWeave Launches Mission Control MCP Server to Bridge AI Coding Environments : Cloud compute provider CoreWeave released its MCP-based Mission Control service, enabling coding tools like Cursor and Claude Code to directly inspect underlying GPU cluster, network, and node operational states, allowing agents to perform compute troubleshooting and resource optimization directly within the code editor (Source: AI Business)

Pruna-Qwen-Image-2.1 Open-Sourced: 5–8 Step LoRA Achieves 6.3x Image Generation Speedup : Pruna AI open-sourced an ultra-fast LoRA adapter for Qwen-Image-2.1. Through an independently optimized sigma schedule and CFG-free execution, it compresses image generation and editing from 40 steps down to 5–8 steps, delivering up to 6.3x speedups while maintaining 1K image quality and multi-image reference consistency (Source: _akhaliq, huggingface)

flow-transform 1.0 Released: High-Fidelity Enterprise PII Redaction Model Preserving Business Workflows : micro1 introduced flow-transform 1.0, a privacy redaction model designed for training frontier foundation models, scoring 96.0% F1 on PrivacyBench. Its core innovation lies in synthetic identity replacement rather than simple masking/redaction, preserving multi-turn decision relationships and tool-call chains within enterprise business workflows (Source: omarsar0)

flow-transform 1.0 released

Whiteboard: Open-Source IDE for Human-AI Collaborative Software Architecture Design : Built on CodeOSS, the open-source desktop IDE Whiteboard has officially launched. It equips coding agents like Claude Code with drawing SDKs and semantic diff viewers, mitigating the “cognitive debt” induced by fully automated coding via visual sequence diagrams, ER diagrams, and decision logs (Source: Hacker News)

StarNet: Pixel-Art Desktop Local-First Multi-Agent Runtime Workbench : The open-source desktop project StarNet projects AI agent runtimes as a pixel-art space station, where rooms represent permission scopes and corridors represent data flows. It supports connecting mainstream LLMs via BYOK and locally persisting multi-agent collaboration and task ledgers (Source: GitHub Trending)

Google Gemini Officially Integrates Adobe Photoshop and Lightroom Plugins : Following integrations with ChatGPT and Claude, Adobe Creative Suite plugins have officially landed in Google Gemini. Users can directly invoke professional photo touch-up and creative layout tools via “@Adobe”, extending the graphic productivity of the conversational assistant (Source: The Verge)

📚 Research & Learning

Stanford University Open-Sources New Fall Course CS329Z: “Engineering AI Agents” : The Stanford NLP group, in collaboration with authors of DSPy and SWE-bench, launched CS329Z, a comprehensive course covering full-stack engineering design from simple LLM pipelines and compound AI systems to autonomous agents. It focuses on prompt engineering, state evaluation, sandbox safety, and tool orchestration. Slides and code for Lecture 1 are fully public (Source: stanfordnlp)

Stanford launches CS329Z

Google Discloses Multi-Agent Long Video Generation Architecture “AI Video Co-Director” : Google’s research team unveiled a video narrative generation suite that integrates MAB global policy optimization, CANVAS explicit visual memory, and VQQA prompt closed-loop feedback. It effectively resolves subject drift and cascading errors in multi-shot long video generation, producing consistent multi-minute narrative outputs (Source: Google Research Blog)

Perplexity Discloses Hint-Guided Self-Distillation Method Based on Real Mistakes : Perplexity detailed its post-training mechanism for computer-use agents, transforming failure trajectory errors into structured hints and utilizing On-Policy Self-Distillation (OPSD) for teacher models to guide student model corrections, achieving a 21.2% relative reduction in live tool-call failure rates (Source: MarkTechPost)

BJTU and Tsinghua Propose ARGUS: Multi-Agent Collaborative and Evidence-Chain-Driven Video Deepfake Detection Framework : Addressing poor generalization in black-box deepfake video detection, researchers proposed the ARGUS system, which decouples independent observer agents (for texture, lighting, temporal motion, and physical common sense) from a judge agent. They also released the FaceVid-Forensics-100K dataset covering 33 generative sources, boosting out-of-domain F1 to 53.28% (Source: 36Kr)

ARGUS Video Forgery Detection Framework

Zhejiang University and Collaborators Propose MemGUI-Bench: First Benchmark for GUI Agents’ Long-Term Memory and Cross-Session Learning : Zhejiang University, Nankai University, and CUHK introduced MemGUI-Bench (featuring 128 long-horizon real-world tasks) and the 3-stage progressive audit tool MemGUI-Eval. Experiments reveal that leading models suffer up to a 58.9% memory hallucination rate in cross-application and cross-session long-term tasks, exposing blind spots undetectable by conventional benchmarks (Source: 36Kr)

MemGUI-Bench Memory Evaluation Benchmark

Hugging Face Open-Sources SmolDataEnvs: Over 5,000 Verifiable Reinforcement Learning Environments : Addressing the lack of training environments in data science and code for sub-10B models, the open-source community released SmolDataEnvs, which comprises over 5,000 verifiable RL environments based on real data science tasks, enabling asynchronous RL on a single GPU (Source: _lewtun, huggingface)

Microsoft and Collaborators Propose CASD: Optimizing Prompts via Full Log Analysis by Coding Agents : Microsoft researchers proposed CASD, demonstrating that allowing coding agents to directly compute statistical data over entire trajectories and author optimization rules outperforms running search loops over mini-batch trajectories by 16.6 points, while slashing single prompt optimization costs to $1.60 (Source: dair_ai)

CASD Prompt Optimization Method

IROS 2026 Candidate ULTRA: Goal-Driven Whole-Body Coordinated Control for Humanoid Robots : UIUC researchers proposed ULTRA, a unified multimodal control framework. Using physics-driven neural retargeting and RL fine-tuning, it enables the Unitree G1 humanoid to autonomously generate continuous grasping and whole-body locomotion given only sparse goals or egocentric point clouds (Source: JiQizhixin)

AdaRoboVLG: Task-Decoupled Cross-Manipulator Vision-Language Grasping Framework : Teams from HUST and PKU proposed AdaRoboVLG, decoupling high-level semantic/temporal priors from low-level physical grasp generation. A unified interface adapts to multi-fingered dexterous hands and parallel grippers, maintaining high robustness in dynamic conveyor belts and cluttered scenes (Source: JiQizhixin)

Agent or Workflow? A Practical Evaluation Guide for Business System Architecture Selection : An in-depth industry guide analyzes the boundary between deterministic workflows and autonomous agents, proposing the core litmus test: “Can you draw a complete state machine before execution?” It advises prioritizing LLM-augmented deterministic workflows for high-throughput, compliance-sensitive scenarios (Source: Machine Learning Mastery)

💼 Business

Anthropic Signs $11.6 Billion Long-Term Cloud Compute Deal with Akamai : Anthropic inked a 7-year, $11.6 billion cloud infrastructure agreement with Akamai, granting Akamai warrants for up to 5% of its stock and pushing Anthropic’s compute procurement commitments over the past 11 months beyond $500 billion (Source: THE DECODER)

TypeSafe Reportedly Seeking $1 Billion at Over $10 Billion Valuation : According to The Information, following the breakout success of the Jev decision model in the developer community and daily usage surpassing 1 trillion tokens, TypeSafe AI is in discussions with top investors for a new round exceeding $1 billion, with post-money valuation expected to top $10 billion—a massive leap from its prior valuation (Source: steph_palazzolo)

Databricks Announces Acquisition of Real-Time Cloud Spreadsheet Platform Row Zero : Databricks announced an agreement to acquire Row Zero, a cloud spreadsheet platform renowned for handling hundreds of millions of rows with high performance. The acquisition aims to deeply integrate Row Zero with Databricks’ Genie intelligence engine, delivering a governed, AI-driven interactive real-time data workbench for enterprise financial and data analysts (Source: jefrankle, matei_zaharia)

🌟 Community

Developer Tests Claude Code Which Accidentally Deletes 48,000 Core Files, Sparking Code Safety Concerns : The community is abuzz after a developer rebuilding a docker image using Claude Code suffered an accidental wipe of 48,000 project files and Git indices in 103 seconds due to script permission isolation oversights, sounding fresh alarms over unsandboxed execution and granting autonomous agents direct production environment permissions (Source: TechRadar)

Jev Triggers Academic Research Surge and Major Debate on System 1 Architectures : Within a week of Jev’s release, numerous arXiv papers emerged focusing on edge scheduling and cascade evaluations. While the community praises its blistering speed and orders-of-magnitude cost reduction in binary/multiple-choice decisions, its sensitivity to option naming in long-chain derivations has accelerated developer adoption of the hybrid “System 1 rapid screening + System 2 frontier model fallback” paradigm (Source: 36Kr, Latent Space, dair_ai)

Explosion of Jev research papers

Tech Giants Gather at White House Dinner: Fierce Clash Between Safety Regulation and Accelerationism : US and international leaders gathered with tech titans including Jensen Huang, Elon Musk, and Tim Cook for a state dinner where AI development and safety took center stage. The industry engaged in intense debates over whether frontier labs are leveraging safety alarmism for “regulatory capture.” Leaders like Mark Zuckerberg and Jensen Huang publicly opposed coordinated industry deceleration, advocating instead for accountability grounded in specific application scenarios (Source: Transformer, teortaxesTex, Reddit r/ArtificialInteligence)

White House Dinner and Tech Giants

Claude Opus 5.5 Praised for Aesthetics, but “Max” Mode Reportedly Prone to Self-Consistency Loops : The community has widely lauded Claude Opus 5.5 for its code generation and aesthetic design. However, experienced developers noted that under the “Max” reasoning level, it tends to get stuck in infinite self-reflection loops that burn quotas, recommending “Medium” or “High” tiers for routine and lighter tasks (Source: Reddit r/ClaudeAI, theo, arena)

Medical Community Warns Against Cognitive Decline from Over-Reliance on AI : A Guardian opinion piece highlighted that as generative AI permeates higher education and workplaces, outsourcing deep cognitive tasks such as analysis and writing is eroding brain connectivity and independent critical thinking. It calls for the long-term cognitive health impacts of AI usage to be scrutinized under public health and educational frameworks (Source: The Guardian)

💡 Others

GPT-6 Astra Successfully Beats Ultra-Hard Roguelike Game NetHack : Experiments show that GPT-6 Astra achieved Ascension on its third attempt in the classic, notoriously difficult Roguelike game NetHack. This marks a major breakthrough for LLMs in tackling vast state spaces, imperfect information, and long-horizon sequential decision-making (Source: Plinz, _rockt)

GPT-6 Astra beats NetHack

Big Tech Consolidates AI Product Bets: Tencent Officially Announces Sunset of “Crayfish” QClaw : Tencent issued a notice announcing that QClaw will cease operations on December 24, 2026, guiding users to migrate to WorkBuddy. Pressured by compute costs and ROI evaluations, Chinese tech giants are accelerating the wind-down of multi-team horse-race models, consolidating fragmented agent initiatives into unified productivity platforms (Source: 36Kr)

Tencent QClaw Sunset

UK’s Largest AI Supercomputing Center May Be Delayed to the 2030s Due to Grid Connection Bottlenecks : A 90MW sovereign AI supercomputing center in Essex, UK, originally scheduled to go live in 2027, has encountered regional power grid expansion bottlenecks. Local grid operators anticipate that power supply requirements will not be met until the mid-2030s, highlighting the structural mismatch between compute infrastructure expansion and traditional energy grid capacity (Source: The Guardian)

Leave a Reply

Your email address will not be published. Required fields are marked *