GPT-5.6 Enters Self-Optimization Mode, Reducing Inference Service… | AI Daily 2026-07-31

🔥 Focus

GPT-5.6 Enters Self-Optimization Mode, Reducing Inference Service Costs by 20%: GPT-5.6 Sol has been successfully deployed in OpenAI’s production environment, autonomously rewriting and optimizing GPU kernels to reduce end-to-end service costs by 20%, and improving token generation efficiency by over 15% through optimizing speculative decoding models. This progress demonstrates the practical implementation of Recursive Self-Improvement (RSI) at the system infrastructure level, with AI beginning to actively participate in transforming its own runtime environment. (Source: OpenAI News)

Meta Releases Q2 Earnings: Capital Expenditure Soars, Zuckerberg Predicts Era of Personal Agents: Meta released its Q2 2026 financial results, reporting revenue of $60.8 billion. However, due to increased AI infrastructure investment and litigation expenses, free cash flow plummeted by 91% to $784 million, causing the stock price to drop sharply after hours. CEO Mark Zuckerberg predicted that billions of people will have personal AI agents running 24/7 within the next five years, and stated that Meta will continue to increase AI capital expenditure to drive the integration of agents and hardware ecosystems. (Source: TechCrunch)

Silicon Valley Erupts in Major Debate Over “Open Source vs. Closed Source” Routes: With over 70 giants including NVIDIA, Microsoft, and OpenAI co-signing the “Open Weights and American AI Leadership” open letter, Silicon Valley’s open-source camp has gained significant momentum. Meanwhile, Anthropic, as the only closed-source giant that did not sign, saw its CEO publish an open letter arguing that they do not advocate for banning open weights, but emphasizing concerns over national security and malicious distillation. This move triggered a backlash in Silicon Valley, with some startups beginning to pivot to open-source models out of fear of being “reverse-consumed” by the closed-source ecosystem. (Source: ZDNet)

OpenAI Launches “Academic Researcher Program,” Opening Flagship Models to 100,000 Scientists for Free: OpenAI announced that it will offer free access to its flagship models (including the GPT-5.6 series) to 100,000 university researchers worldwide, providing enterprise-grade privacy protection, an expanded version of Deep Research, and 75 specialized life science skills. This move aims to deeply bind scientific research pipelines to the ChatGPT ecosystem, locking in the academic community’s long-term AI usage habits through free access. (Source: OpenAI News)

Lilian Weng Returns to OpenAI at Lightning Speed to Lead Self-Evolution Research: Co-founder Lilian Weng, who recently announced her departure from Thinking Machines Lab due to health reasons, has confirmed she is rejoining her former employer OpenAI. She will lead a high-level “Self-Evolution” (RSI) research team, focusing on using AI models to develop and improve next-generation models. This indicates that OpenAI is accelerating the transition of AI self-improvement from the experimental stage to systematic engineering. (Source: QbitAI)

Google Releases Lyria 3.5 Music Generation Model, Introducing Selective Section Painting: Google DeepMind has launched Lyria 3.5 in Flow Music, enhancing the naturalness and expressiveness of melodies, lyrics, and vocals. The new version introduces the “Selective Section Painting” feature, allowing users to fine-tune specific parts of an audio track or extend a short melody into a full song without recreating the entire track. (Source: Google DeepMind Blog)

Moke Robotics’ LJM World Model Ranks Second Globally on WorldArena: Chinese startup Moke Robotics has introduced its LJM (Latent Joint-conditional Model) world model, securing second place globally in the embodied world model evaluation WorldArena. The model adopts a “dual-brain collaboration” architecture, separating physical interaction reasoning from video rendering, allowing the robot to first understand the physical semantics of interaction in the latent space before generating predicted frames. (Source: Synced)

Google DeepMind Releases Gemini Robotics 2, Unlocking Whole-Body Control and Collaborative Capabilities: Google has released its next-generation embodied AI brain, Gemini Robotics 2, closing the loop from simulation training to real-world deployment. The model unlocks intelligent whole-body control, advanced dexterous manipulation, and multi-robot collaboration, enabling robots of various morphologies to achieve long-term autonomous operation through simulation training without relying on real-world data. (Source: )

Tencent WorkBuddy Upgrades to “Human-AI Co-Writing” Collaborative Editing: Tencent’s productivity agent tool WorkBuddy has released version V5.3.5, launching the “Human-AI Co-Writing” feature in collaboration with Tencent Docs. Users can now perform real-time collaborative editing and modifications with AI directly within Word, Excel, PPT, and other documents, achieving multi-terminal synchronization across “human-to-human, human-to-machine, and machine-to-machine” for the first time, pushing office software into the AI-native collaboration era. (Source: QbitAI)

Tencent Video Beta-Tests “WorkSolo” Lightweight AI Creation Platform: Tencent Video’s Intelligent Creation Platform Department is beta-testing WorkSolo, a creation platform integrating “AI short dramas, interactive video games, and free canvas.” Unlike WorkRally, which focuses on the industrialized production of premium animated dramas, WorkSolo targets the rapid production of lightweight short dramas and interactive content, utilizing a credit-based computing power model with free trials and member subscriptions. (Source: 36Kr)

Google Earth Introduces Nano Banana for AI Geographic Reimaginings: Google has integrated the Nano Banana 2 image generation model into Google Earth, allowing users to generate AI images based on satellite and 3D imagery. Users can modify real-world buildings or generate geographic information charts via text prompts, though testing shows it still has limitations in handling precise geographic logic and text generation. (Source: The Verge)

OpenAI Hardware Roadmap Exposed: Smart Speaker First, Phone Second: Reports suggest that one year after acquiring Jony Ive’s company, OpenAI has established a clear hardware roadmap. The first product is expected to be a screenless smart speaker launched in early 2027, equipped with a camera and mechanical structures. Additionally, OpenAI is developing an “AI Agent Phone” powered by a custom MediaTek chip, planned for mass production in the first half of 2027. (Source: 36Kr)

Elon Musk Releases Grok Voice Think Fast 2.0, Topping Agent Test Leaderboard: SpaceXAI has launched its next-generation voice model, Grok Voice Think Fast 2.0, ranking first in the Tau Voice agent test. The model features an average first-audio response time of just 0.70 seconds, supports streaming computation for “reasoning while speaking,” reduces token consumption by 60% compared to its predecessor, and is priced at $0.08 per minute. (Source: 36Kr)

🧰 Tools

Replit Launches AI Design Tool Replit Design, Featuring Post-Prompting Interaction: Replit has released the AI design tool Replit Design, highlighting “post-prompting” interaction. The tool does not require users to input complex prompts or design languages; instead, it uses Ambient Intelligence to actively recommend the next design step at each stage, supporting one-click adoption and significantly lowering the application design barrier for non-developers. (Source: amasad)

Cisco Releases Free AI Model Provenance Kit to Prevent Open-Source Model Compliance Vulnerabilities: Cisco has released a free “AI Model Provenance Kit,” cataloging fingerprint characteristics of nearly 900 open-source models. By comparing architectural metadata with weight-level signals, the tool bypasses user-filled labels to accurately identify the true lineage and derivative relationships of open-source models, helping enterprises prevent compliance and security vulnerabilities. (Source: VentureBeat)

Token Saver: An Open-Source MCP Extension to Drastically Reduce Token Costs for PDF Analysis: Developers have open-sourced Token Saver, an MCP extension plugin for Claude Desktop. Running a lightweight hybrid RAG system locally (combining BM25 and semantic search), the tool avoids uploading the entire PDF to the cloud and only pushes the most relevant paragraphs to the model, reducing token costs for long document analysis by 90% to 99%. (Source: MarkTechPost)

Hint: AI Home Assistant App Co-Founded by Martha Stewart Launches: Hint, an AI home assistant app co-founded by former Casper engineering head Kyle Rush and lifestyle guru Martha Stewart, has launched. After users input their address and upload home contracts or invoices, the app combines public geographic, soil, and weather data to automatically generate home maintenance plans, and answers home management questions like insurance claims and energy optimization via an AI chatroom. (Source: TechCrunch)

📚 Learning

DeepLearning.AI Partners with Qodo to Launch Free “AI Code Review” Course: Andrew Ng’s DeepLearning.AI has launched a new short course teaching how to build and optimize AI code review agents. The course covers performing code reviews before PRs, providing full repository context, building context engines using vector retrieval, and collaborating with security and standards expert agents to reduce false positives. (Source: )

CtrlBench-Rec Paper: An External Controllability Evaluation Framework for Black-Box Recommendation Systems: A team from the Institute of Information Engineering, Chinese Academy of Sciences, and other institutions published a paper proposing the CtrlBench-Rec evaluation framework. Without accessing the model’s internal parameters, the framework utilizes collaborative multi-agents as probes to measure the response capabilities of black-box recommendation systems in dimensions such as content discovery, interest profiling, and long-tail bias mitigation through simulated external interactions like searches and clicks. (Source: HuggingFace Daily Papers)

Tencent Open-Sources AngelSpec Speculative Decoding Training Framework, Achieving 2x Decoding Acceleration: Tencent has open-sourced the AngelSpec speculative decoding training framework, supporting Multi-Token Prediction (MTP) and block-parallel DFlash architectures. Featuring a shared-parameter, multi-depth design, the framework achieves a nearly 2x end-to-end decoding speedup on Qwen3 and Hy3 models through target model rollout and prefix-adaptive attention mechanisms. (Source: MarkTechPost)

YC Paper Club Discusses Multi-GPU Kernels and GPU-Accelerated Game Engines: YC Paper Club shared several cutting-edge systems papers, including the Parallel Kittens framework for simplifying multi-GPU AI kernels, the “Intelligence per Watt” metric for measuring local vs. cloud AI energy efficiency, and Madrona, a high-throughput game engine running entirely on GPUs to accelerate reinforcement learning training. (Source: )

💼 Business

Tokens Infinity AI Completes New Funding Round of Hundreds of Millions of Yuan: Tokens Infinity AI, an AI Agent infrastructure company for enterprise software engineering, announced the completion of a new funding round led by Linxin Investment, with participation from existing shareholder Huakong Fund, bringing its total funding within a year to hundreds of millions of yuan. Founded by Yang Ping, the former head of ByteDance’s MarsCode, the company focuses on fully autonomous, highly secure, and self-evolving enterprise-grade agent foundations. (Source: QbitAI)

Dili Raises $21.7 Million to Solve Infrastructure Compliance Challenges with AI: AI compliance startup Dili announced the completion of a $15 million Series A funding round, bringing its total funding to $21.7 million. The round was led by Khosla Ventures, with participation from Y Combinator and others. Dili uses AI to parse contracts, payrolls, and ERP data for large infrastructure projects like buildings and data centers, automatically verifying compliance with federal acts such as the Davis-Bacon Act through deterministic rules. (Source: TechCrunch)

Atlassian Introduces “AI Wallet” Mechanism to Limit Employee AI Spending: Software giant Atlassian has introduced an “AI Wallet” mechanism for its R&D teams, setting a monthly AI usage limit of $500 to $2,000 per employee to cope with the IT cost pressures brought by skyrocketing token usage. This move contrasts with Silicon Valley’s previous “tokenmaxxing” trend, which encouraged employees to use AI without limits. (Source: The Guardian)

🌟 Community

GCC Announces Ban on AI-Generated Code Contributions of Legal Significance: The GNU Compiler Collection (GCC) steering committee has announced a new policy. Due to considerations of copyright ownership, license compliance, and legal liability, it will refuse to accept any code contributions of “legal significance” generated by large model agents starting this year. This has sparked widespread discussion in the open-source community regarding the attribution of responsibility for AI-generated code. (Source: Hacker News)

Meta Smart Glasses Mocked as “Pervert Glasses,” Raising Privacy Concerns: As videos of pranks and secret recordings in public spaces using Meta Ray-Ban smart glasses spread on social media, the hardware has been mockingly dubbed “pervert glasses” in the community. The head of Instagram stated that accounts and videos violating user privacy or suspected of harassment will be banned and cleared. (Source: Reddit r/ArtificialInteligence)

💡 Others

36Kr Partners with PureblueAI to Release the Second “2026 Consumer Brand AI Recommendation Power List”: 36Kr, in partnership with PureblueAI, has released a new edition of the list, evaluating the recommendation performance of five major categories—including mobile phones, automobiles, and home appliances—across mainstream AI platforms. Data shows that Xiaomi retained its top spot in new energy sedan recommendations, while iQOO took the lead in high-performance smartphones. AI recommendation mechanisms are gradually becoming a new battlefield for brands in specific consumer scenarios. (Source: 36Kr)

Leave a Reply

Your email address will not be published. Required fields are marked *