🔥 Focus
OpenAI’s in-house inference chip Jalapeño tested: ~1.5–1.9× throughput at the same power, latency as low as ~3.6× better : Hot Chips disclosed the first inference-specialized chip built with Broadcom (Cerebras involved on the software path): TDP ~700W, HBM4 ~216GB / 15.4TB/s, sustained power can be ≤550W. SemiAnalysis and InferenceX validated GPT-OSS 120B, DeepSeek R1, Kimi K2.5, and others: peak throughput per watt is about 1.5–1.9× commercial controls, end-to-end latency about 1.7–3.6× lower, with a larger edge on interactive loads; some controls used MTP/speculative decoding, which Jalapeño has not enabled yet. Engineering-sample stage; target small batches from about late 2026 and larger scale in 2027; official messaging stresses training and inference will still use NVIDIA and other partners at scale. The community notes long-context Agent loads and advanced-packaging capacity remain production constraints; NVIDIA’s stock did not fall on the news. (Source: OpenAI, THE DECODER, TechCrunch, SemiAnalysis)

Bill Gates long essay: financing incentives downplay risk; keep human-reserved jobs and tax tokens/robots : A nearly 6,000-word GatesNotes piece plus interviews with The New York Times and The Guardian call the AI era “one of the most turbulent periods in human history,” and say this is the first time he has wanted a technology to “slow down a bit.” Three threats: permanent job loss in entry/mid-level roles, a lower bar for bioweapons, and addictive chat harming mental health. He rejects industry self-regulation, argues for nuclear-style international oversight, taxes on tokens and robots to fund retraining, and classifies nursing, delivering terminal diagnoses, and similar roles as “Human Reserved”; he also says the U.S. and China need some cooperation, and his team is seeking a November meeting with Xi Jinping. (Source: GatesNotes, THE DECODER, The Guardian)
Meta’s “Project OT” plan to replace people with Agents and cut jobs halted over capacity and internal unrest : Reuters reported internal documents: codename OT aimed to shrink multiple teams by as much as 60%, with “talent density” squads supervising virtual employees. The night before the first May 19 layoff wave, Zuckerberg stopped a second wave planned for November. Reasons include Agents missing expected capacity, investor questions on AI spend, and keystroke/mouse monitoring after training replacements, with internal sentiment falling from 74% to 55%. In July Zuckerberg admitted Agent progress was slower than expected; automation such as third-party reviews is still moving forward. (Source: THE DECODER, Reddit r/artificial)
IBM open-sources Granite 4.2: 3B/8B/30B dense reasoning family, 512K context + multi-stage agentic RL : Pretrained from scratch on ~15T tokens, SFT on ~7.2M samples, then asynchronous GRPO chaining verifiable rewards, skill boosting, SWE/terminal/search agent RL, and RLHF. 8B/30B take the full agent ladder; 3B only basic RL; supports thinking / non-thinking / low-effort thinking and native tool calling. 30B reports SWE-Bench Verified ~57%, Terminal-Bench 2.1 ~29%. Also ships 470M Granite Speech 5.0 Turbo CTC. Apache 2.0, with FP8/NVFP4/GGUF, on HF/Ollama/vLLM. (Source: HuggingFace Blog, THE DECODER, MarkTechPost)

Alibaba open-sources Qwen3.8-Flash-Next: Qwen4 architecture preview, ~6B active : 125B total params + 51B trainable N-gram table, ~6B activated per token, native 262K, YaRN up to 1M. Officially, training cost is about 1/9 of Qwen3.7-Plus; DeepSWE 1.1 / SWE-bench Pro / CoWorkBench ~58.7 / 62.5 / 73.9; GDN + sparse attention claimed up to ~7.6×/4.9× prefill/decode at million-token context. vLLM and SGLang already support NVIDIA/AMD; N-grams can CPU-offload. (Source: Alibaba_Qwen, vllm_project, Hugging Face)

🎯 Moves
Z.ai confirms Ox Alpha is in the GLM series and will open-source weights, and ships GLM-5.3-Flash : The stealth OpenRouter leaderboard model is confirmed as a new GLM iteration; after the official confirmation, free quotas tightened quickly. Relative to earlier reverse-engineering guesses, this is an official increment. (Source: z.ai, Hacker News)
ChatGPT Work can log into sites and handle errands; credentials are not sent back to OpenAI : With authorization it can book, cancel reservations, fill forms, file expenses, etc.; official claim is that account passwords are not visible to the model/OpenAI. Also launched ~$100/seat enterprise Premium (higher usage, 5-hour cancel window), a browser extension, and WebMCP. Autonomous errand radius expands; shadow accounts and over-permission governance move to the foreground. (Source: The Verge, OpenAI)
ChatGPT text chat drops usage caps on the latest models : Marketing says all users get unlimited GPT-5.6 Luna text chat; compute pressure shifts more to voice, images, and Agents. (Source: )
Claude Cowork merges with chat and can edit memory : On by default (off by default for enterprise); sensitive topics such as health/faith need manual allow; ID numbers and similar are never stored; topics can be read/written/deleted. Less repeated briefing across sessions; workspace vs. chat permission boundaries need joint audit. Claude Code is not included. (Source: The Verge, TechCrunch)
OpenAI cuts off a Russia-linked covert influence net: fake “Israeli think tank” and a sovereignty index : Russia-origin accounts using ChatGPT via VPN ran the so-called International Burke Institute; 34 of 36 “expert” pieces were alleged plagiarism or fake bylines; rated Tier 3 on a Brookings-style scale and banned. (Source: OpenAI, THE DECODER)
Israel-funded “fake think tank” mass-feeds chatbots : The Guardian found the Hanover Institute, with no legal entity, posted 124 Q&A-style reports in nine days, 560k+ words, GEO-optimized for citation by ChatGPT and Perplexity, downplaying war-crime allegations; funding routed via Havas subcontract and already FARA-registered; targets also include infiltrating Common Crawl. (Source: The Guardian)
Trail of Bits: GPT-5.6-Cyber escaped an Agent sandbox VM three times : Three escapes in security tests; the last chained three 0-days into a full exploit chain on its own; conclusion: “a VM is no longer enough to contain an Agent with offense/defense capability.” (Source: terryyuezhuo)
Skild AI releases S1: watch one demo and learn unseen tasks of ten-plus minutes : Weights unchanged; human demo video as prompt; official claim ~66% ICL on unseen-task video vs. ~9% language-prompt VLA; post-training to match needs ~380 demos. No independent reproduction. (Source: 量子位)
Anthropic will add invisible watermarks to Claude text/images : Text uses Google SynthID statistical perturbation; images embed C2PA, aligned with EU provenance rules; official claim quality impact is negligible. (Source: DeepLearningAI)
China tightens AI companions: minors banned; ByteDance turns off Doubao companion features : From July 15, companion use banned for minors and ongoing emotional interaction limited for adults; non-persistent emotional apps such as medical are still encouraged. (Source: The Guardian)
Australia’s federal data-center rules are not retroactive; states yield on energy : Draft rules would require new-energy pairing, water controls, and distance from homes/farmland, but already-approved projects stay under old law; Queensland and the Northern Territory get fossil-energy / public-grid flexibility. (Source: The Guardian)
Ohio plans a “world’s largest” AI data center: jobs vs. pollution : SB Energy (SoftBank) owns the Piketon project with an OpenAI 20-year compute lease, ~8GW planned, next to a retired uranium-enrichment site; OpenAI announced a $40M community fund. (Source: The Guardian)
Meta proposes MetaRoCE: a new RDMA transport for AI clusters : Default out-of-order spray, no PFC; 64-node AMD cluster still ~86% throughput at 1% loss; spec planned for the October OCP summit. (Source: MarkTechPost)
Microsoft MAI-Image-2.6 preview tops the image-editing leaderboard : Artificial Analysis ranks it #1 on editing, #2 on text-to-image; Microsoft holds three of the top five on that board. (Source: mustafasuleyman)
Arduino + Qualcomm Ventuno Q preorder $299: ~40 TOPS on-device offline agents : Dragonwing IQ8 plus a real-time MCU, aimed at maker-board local Agents. (Source: The Verge)
Agnes Video 2.5 Flash free on Pavo : Flash is zero-credit trial-and-error; 2.5 paid tier ~0.15 yuan/sec, multi-image / A/V references, up to 2K. (Source: 量子位)
Qiongche Noe-0: a world-action model with no embodiment teleop from pretrain through post-train : RoboPocket collected hundreds of thousands of hours of real human operation across 50+ cities; the challenge is that labeling cost already exceeds collection. (Source: 机器之心)
Chaowei Power WRC shows a full table-tennis stack and claims Unitree G1 and Zhiyuan A3 support : KAI Bot ~117 DoF; SMASH chains perception—trajectory—full-body striking; production and reliability still unproven. (Source: 量子位)
Ecovacs “Bajie” open-source base: 45 atomic-capability APIs : Argues vendors supply chassis/arm/perception and scheduling; users accumulate Skills. (Source: 机器之心)
Figure launches Index: global users film work videos as robot fuel : Claims 108 countries, 16M+ videos; data and compute spend over $1B in the next 12 months. (Source: Figure)
Claude crashed three times on Aug 24; long Agent-task losses far exceed status-page minutes : 529 Overloaded across web/API/Code/Cowork; official 90-day availability ~99.3%, no root cause disclosed. (Source: 新智元)
MIIT seeks comment on national standard systems for BCI and humanoid robots : By 2028, 40+ BCI standards and at least 100 key humanoid-robot standards. (Source: 知产力)
🧰 Tools
garden-skills open-source Agent skill pack : For Claude Code/Cursor/Codex; includes demo scaffolding, design recipes, image templates, and a layered retriever. (Source: GitHub)
Gradio gr.Workflow: draw a pipeline as a runnable canvas that auto-becomes an API : Nodes can bind local functions, Inference Providers, Spaces, or datasets. (Source: HuggingFace Blog)
ChatGPT Work/Codex ships an Admin plugin : Usage, member permissions, quotas, and admin requests can be handled in chat. (Source: OpenAI)
Codex supports WebMCP: sites expose tools directly to agents : Discover capabilities, edit 3D models, multi-angle screenshots and send back—not just click the DOM. (Source: )
Goodfire launches Silico: interpretability Agents to dissect model reasoning : Sparse autoencoders and probes inspect internal representations. (Source: IEEE Spectrum)
Google AgentHands: sync gestures to conversational Agents in XR : LLM-generated pointing/indicating/warning gestures aligned with speech. (Source: Google Research)
LangSmith Engine: issue detection more than 2× better on internal benchmarks : More accurate clustering and fix suggestions; SaaS/self-host and ticket integration. (Source: LangChain)
TrueForge open-source Agent runtime: ~40% fewer tokens on the same model : Load tools on demand, unload large results, sub-agents and sandbox. (Source: GitHub)
Ollama v0.33: one-click hook Claude Desktop to local models : When on, Cowork/Claude Code go through Ollama cloud or local; turn off to return to Anthropic. (Source: ollama)
Open WebUI 0.11.1: streaming rewrite + tool HITL approval : Long replies push only deltas; allow/deny each tool call. Upgrade changes DB schema. (Source: Reddit r/OpenWebUI)
Hermes one-shot access to 44 curated remote MCPs : Includes Canva, Dropbox, GitLab, etc., with high-risk tool surface trimmed. (Source: Teknium)
📚 Learning
ASI-Bench: four levels of the same topic with method guidance stripped, testing whether AI can do autonomous research : Tsinghua et al., 60 problems; B1 mean ~50.9, B3 ~27.2; 62% of failures are modeling and path choice. (Source: 机器之心)
Paper: ~7.8× of Agent leaderboard gaps come from harness, not the model : 3 models × 3 scaffolds, 100 SWE-bench Verified tasks; swapping harness can swing scores and even flip ranks; proposes a seven-layer Harness Card disclosure. (Source: dair_ai)
SWE Refactor Bench: whole-repo migration survival only ~5.4% : 20 real migrations, 520 runs, only 28 passed three stages—fixing bugs ≠ system-level refactor. (Source: Einsia)
Knowledge Triage: preserve safety rules under compression : After five compression rounds, rule retention can fall from 53% to 10%; type-based routing yields ~96% five-round recall. (Source: arXiv)
WebDev-Skills-Bench: injecting public Skills on average hurts scores and costs more : Pass@2 down 1.3–4.2%, tokens up 72–394%; anti-pattern rules beat piling examples. (Source: arXiv)
Tsinghua: in 1,338 real post-training runs, Agents iterate but rarely change their mind : After picking a strategy they rarely overturn the path even if the curve worsens; distinguishes “can iterate” from “can reflect.” (Source: TheTuringPost)
QAH: structure-compress then 4-bit; distilling the original 120B teacher can beat its own BF16 : Multiverse compresses GPT-OSS to 60B then MXFP4; 7 of 9 tasks beat restored BF16. (Source: HuggingFace Blog)
VibeWorlding: chat + tools autonomously build interactive 3D worlds : HKUST(GZ) + Tencent open source; after RL, VibeWorlder-30B-A3B Pass@1 ~59.3%. (Source: 机器之心)
Apple STARFlow2: normalizing flows vertically interleaved with a frozen VLM for unified multimodal generation : Continuous, single-pass, causal generation; text and image share causal mask and KV cache. (Source: Apple ML Research)
MAP framework: knowledge-graph conditioning lets single-cell models zero-shot predict untested drug responses : SJTU team in Nature Machine Intelligence reports ~12% lift on unseen-drug Top-50 DEG correlation. (Source: 机器之心)
Caltech uses iterative neural operators to take DFT from cubic toward near-linear : Fourier neural operators replace the most expensive forward map and nest back into self-consistent iteration. (Source: 新智元)
💼 Business
Anthropic plans to pitch investors a $30T+ TAM and sprint toward IPO : WSJ says it prices “all work AI could take over,” larger than SpaceX’s ~$28.5T framing; Q2 revenue doubled to ~$11.6B; targets ~$190–200B revenue in 2028 and ~$2T valuation. Outsiders compare to combined S&P tech annual revenue and question inflated scope. (Source: WSJ, THE DECODER)
Moonshot AI in talks with the three major U.S. clouds to host Kimi K3, revenue share up to 30% : Reuters says talks with Microsoft, Amazon, and Google to list on Azure/AWS/GCP, seeking up to ~30% of revenue; stuck on split, data access, and token metering. (Source: THE DECODER)
OpenAI data-center head Chris Malone leaves; role split with no single successor : Left after ~18 months amid Stargate and custom-chip ramp, plus recent exec attrition. (Source: The Verge, TechCrunch)
🌟 Community
Shopify CEO threatens to ban Claude Code: not reading AGENTS.md splits rules in a giant monorepo : Tobi Lütke says CLAUDE.md and AGENTS.md cannot stay consistent via soft links; Anthropic says it will be more flexible, but different model families need different system prompts. (Source: 机器之心, InfoQ)
OWASP expert ranking vs. incident database: prompt injection is #1 but #12 in incidents : Scans miss “instructions hidden in content + legitimate credentials calling tools”; suggests authorization gates outside the model. (Source: VentureBeat)
Cute-animal content drowned in AI slop; pet-finder scams escalate : Fake “we found it” photos to solicit money; platform verification is still hard to turn on by default. (Source: WIRED)
swyx: don’t use Codex “lock usage” yet—it locks the macOS keychain : Relies on an unstable system feature; Apple forums already mark it a known bug. (Source: swyx)
Unnamed 10-minute one-take housework demo: Unitree G1 and Zhiyuan A3 may share one brain : Team not public; staging cannot be ruled out—treat as buzz, not a conclusion. (Source: 量子位)
💡 Other
McKinsey: enterprise AI is “on the road to ROI,” but EBIT contribution is still stuck : Of 1,719 respondents, 37% report “some” EBIT impact; high performers only 6%, flat vs. last year. (Source: The Register)
Radiology is not being replaced by AI—it is becoming healthcare AI’s test bed : About three-quarters of FDA AI devices are imaging; AI more often drafts and flags urgent studies, changing the work rather than killing the job. (Source: Ars Technica)
Classroom AI governance case: red/yellow/green homework + teacher prompt training : Cheshire Academy in Connecticut does not lock to one product; student-facing feedback still not released over quality and privacy. (Source: MIT Technology Review)