🔥 Focus
OpenAI and Anthropic stockpile tens of thousands of Mac mini/Studio machines to train computer-use agents : Multiple outlets confirm OpenAI has purchased tens of thousands of headless Mac minis and Mac Studios, while Anthropic is renting similar hardware via AWS, for reinforcement learning and trajectory collection for computer-use agents. Drivers include macOS VM licenses tied to Apple hardware, unified memory for large models, and cooling that can sustain 24/7 full-load GUI interaction—not replacing GPU pretraining. High-spec units have been out of stock for months due to memory shortages; Apple’s Mac revenue rose about 29% YoY that quarter. NVIDIA DGX Spark/RTX Spark are seen as desktop rivals; EXO clusters and Mount Thor-style Apple hardware clouds have also heated up. Apple’s enterprise engineering and developer relations remain relatively weak, so demand is often filled by third parties. (Source: THE DECODER, 量子位, 机器之心, 量子位/36氪)
EU designates ChatGPT, Reddit, and Roblox as very large online platforms : The European Commission has determined that ChatGPT and others fall under the Digital Services Act’s strictest duties, including handling illegal content and protecting minors’ privacy and safety, with fines of up to 6% of global revenue. Some reports also treat it as a very large search engine, with extra compliance required within four months. The move shows Brussels is folding conversational AI into existing platform rules rather than starting from scratch. (Source: Ars Technica, kimmonismus)
Bank of England warns G20: frontier AI could shake global financial stability : Bank of England Governor and FSB Chair Andrew Bailey wrote to G20 finance ministers and central bank governors that frontier models’ autonomy, problem-solving, and attack capabilities are rising together. The most immediate risk is changing the speed, scale, and cost of cyberattacks, then spreading through highly concentrated third-party vendors across cross-border finance. He noted many jurisdictions still lack protocols for advanced-model R&D, release, and deployment, and warned that AI optimism is inflating valuations and leverage so a shock could hit multiple weak points at once, calling for globally prioritized safe, responsible release and deployment. (Source: The Guardian)
NHS watchdog: AI scribes get drug names and diagnoses wrong; patients notice before doctors : Healthwatch England disclosed multiple transcription incidents: MRI summaries turning “no demyelination” into demyelination, prescription names swapped for near-homophones, discharge letters omitting GP refill instructions. About 27 scribes are already in use in the UK; a ten-year plan counts on them to cut workload, but MHRA has not classed them as medical devices and there is no national safety regime. Doctors must still check every line; error rates rise with hallucinations plus accents, multi-person consults, and complex histories. The tools have not yet delivered the promised time savings. (Source: The Guardian)
Google in talks with Disney, Universal and others to license characters and films for AI tools : Google is lobbying Hollywood studios to license characters and catalogs for its AI production and VFX pipelines, while facing union pushback on copyright, jobs, and creative control. After prior lawsuits against AI, the shift toward “licensed partnerships” marks a move from confrontation to deal-making between content owners and model makers. (Source: The Verge)
🎯 Moves
OpenClaw 2.0: guided install, Control UI ~575ms startup; cloud sessions are explicitly not a security boundary : After nearly seven weeks without updates, the largest release yet (~16k PRs): guided install can reuse local Codex/ChatGPT/Claude logins, API keys, or Ollama/LM Studio with pre-validation; Control UI is conversation-centric, startup ~1.6s → 575ms; sessions moved to SQLite. Shared cloud sessions and paired devices/Crabbox are supported, but docs state collaboration permissions are not tenant isolation and not a security boundary; Gateway defaults to loopback, and security still depends on tool approval, sandboxing, and plugin allowlists. The community also voices cooling hype: no stable release, too-frequent updates; some users are switching to Hermes or homegrown harnesses. (Source: THE DECODER, MarkTechPost, 机器之心, Reddit r/LocalLLaMA)
DeepSeek-V4-Flash-Vision-Exp launches with open weights : DeepSeek released an experimental multimodal version on API and Hugging Face, claiming text-side alignment with V4-Flash (including Agent, reasoning, and world knowledge) and a clear jump on multimodal Agent benchmarks; the app’s Vision adds annotations, preset questions, and background image search. Local community says full precision is ~168GB, native 4-bit fits 256GB-memory machines, with vision and structured output. (Source: huggingface, Reddit r/LocalLLaMA)
Qwen Creation launches Agent Teams, plugging in Wan 3.0 for a multi-role short-drama pipeline : Users summon writer, director, art, storyboard, and editor agents with one sentence; intermediates are reviewable shooting scripts, character assets, and shot prompts. Tests can go from script to finished video, but shot quality still needs human confirmation before generation and local rework after. Alibaba Wan also showed cross-shot character consistency, 30-second one-take ads, and image-to-video, landing on Buzzy with limited-time unlimited generation; competition is shifting from single-shot quality to packing the whole production line into one context. (Source: 机器之心, Alibaba_Wan)
Didi’s next-gen Robotaxi R2 starts unmanned passenger tests in Beijing and Guangzhou : The dedicated model with GAC Aion carries 33 sensors and three-domain fused compute; rear large screens and voice interaction emphasize cabin experience, claiming dual five-star China/EU ratings and multiple redundancies. Still demonstration-zone testing; safety and scaled-ops data are not public. (Source: 量子位)
Ropedia ships HOMIE Gen2: turning human experience into trainable Physical AI data : ~380g headset captures ~13 hours, 4-camera 360° plus spatial audio, multimodal alignment in one spacetime; claimed localization error ~0.2%, hand mocap ~4.68mm. The path is “body-free capture,” unbinding capacity from robot fleet size to headcount, with Xperience datasets and processing pipelines. (Source: 机器之心)
Caterpillar moves mine autonomy know-how to jobsites and an aftersales AI assistant : Caterpillar is applying haul-truck and drill automation experience to more dynamic jobsites and launching Cat AI Assistant for technicians to voice-drive repair workflows, claiming ~1.6 million connected assets and 16PB of structured data. The CTO stresses the hard part is changing workflows, not building models; the company plans ~$100 million over five years to train staff. (Source: TechCrunch)
Meta Pocket turns “one-sentence mini-games” into a social feed : After acquiring the Gizmo team, the mobile app lets users generate playable gizmos without looking at code and share them in short-video-style swipes. Prototypes are fast, but creations are locked inside Meta’s walled garden and hard to export as standalone products. (Source: Ars Technica)
Google EnvHarness: a wrapper that makes static eval environments harder as the policy improves : Google Cloud AI Research and others open-sourced EnvHarness, changing start states, actions, and observations only via reset()/step() while keeping human verifiers. EnvRigger diagnoses weaknesses from policy trajectories and writes Python wrappers; on five benchmarks, skill mining up to +9.0 OOD and SWE-bench Verified steps down ~9.8%. Prerequisite: resettable environments, excluding real accounts and physical robots. Apache-2.0 code is public. (Source: MarkTechPost)
Google Assistant shuts down Sept 4; wake word switches to Gemini : After a decade, Google Assistant stops on Sept 4; phones, watches, earbuds, and Android Auto all move to Gemini. The retirement has slipped several times. Analysis argues voice remains the lowest-effort middle layer before Agent rollout, and control should stay in the user’s mouth. (Source: 36氪)
Anthropic forcibly logs out Claude sessions hijacked by infostealer malware : Official email says Vidar, Lumma, StealC, RedLine, Acreed, and Mac AMOS steal browser sessions, bypassing passwords and 2FA to drain quota and possibly run API relays. Users see quota drained, unfamiliar English code logs, kicked logins, and emptied cards. Changing password alone is not enough; revoke all sessions and clear cookies. Pirated software is a common infection source. (Source: 新智元)
Enterprise Agent security: identity and delegated context first; gateways should be the fifth door : Industry pieces note LiteLLM and similar gateway bugs are in CISA’s known-exploited catalog, yet most teams still treat the gateway as the first line. Effective order: named Agent inventory and owners → independent identity plus delegated tasks → short-lived task credentials → recoverable full telemetry → then runtime intercepts and cross-system circuit breakers. Teleport research: over-privileged orgs ~76% incident rate vs ~17% least-privilege. After auth, goal drift, over-tooling, and memory poisoning can still happen. (Source: VentureBeat)
Musk confirms SpaceX is building its own turbine-blade foundry to shorten gas-power timelines : On datacenter power bottlenecks, Musk said blade casting is the gas-turbine capacity chokepoint; a Texas SpaceX foundry could bring units online up to ~18 months sooner, while SpaceX and Tesla each plan ~100GW of solar per year. Critics note faster gas means more emissions; Memphis and elsewhere already have permitting and health-harm disputes. (Source: TechCrunch)
U.S. tightens drone and advanced-robot imports; China still holds scale : Washington is restricting foreign advanced robots on national-security grounds and hiking tariffs on imported drones and parts. Analysis says this is unlikely to reverse China’s lead in humanoid shipments and cost curves; the global market may fragment into “security-compliant vs low-cost scale,” with Japan, Korea, and Taiwan as a middle band. (Source: TechCrunch)
Tencent CubeSandbox 0.7: sandboxes can pause on one machine and wake on another : Single-node cold start still ~60ms; cluster preview supports pause on one host and wake on another, with state on shared storage so draining a machine no longer kills paused jobs. 239 commits, 57 contributors, aimed at Agent sandboxes moving from single node to schedulable clusters. (Source: ziran_pu)
GLM-5.3 quantization fix and MXFP4 : Old ModelOpt quantization on DGX Spark was said to silently corrupt tokens and break tool-calls (related to vLLM #54150); switched to RedHatAI compressed-tensors. Vietnam’s One Nexus released MXFP4 for GLM-5.3/Flash; Together says Flash is ~17× cheaper and 5.3 has higher first-try success. (Source: ziran_pu)
Qwen 3.8 called “unreadable for humans” : LocalLLaMA users complain 27B and Flash-Next overuse set symbols, personas, and dense shorthand, sacrificing readability to lower “tokens per unit of intelligence”; read as Agent RL baking LMs into a “dialect for models.” (Source: Reddit r/LocalLLaMA)
KCD2 director tries leaked DLSS 5: stresses lighting fill, not geometry rewrite : Kingdom Come: Deliverance 2 director Daniel Vavra said leaked NVIDIA DLSS 5 does not redraw character geometry but uses existing data to strengthen facial skin, AO, hat-brim and hair shadows, closer to original art intent. Commenters worry studios will later treat “post magic” as a shortcut button. (Source: Reddit r/ArtificialInteligence)
🧰 Tools
Perplexity Portable Computer: a local-first harness tailored for Qwen3.8-27B : Default model, traces, and files stay on-device; search and cloud advisors upgrade on demand. For small models it shortens system prompts, turns MCP into CLI, and enforces sandbox plus self-check; BrowseComp ~66.7% on DGX Spark; after advisor upgrade Terminal-Bench 2.1 ~59.6% → 73.0%. Post-trained PPLX 27B ~85.4% on its internal knowledge workbench. (Source: 机器之心)
ChatGPT Work is really a cloud task system: networked code interpreter, headless Chrome, persistent disk, deployable sites : Simon Willison unpacks paid Work: optional Sol/Luna/Terra and reasoning tiers; a nearly open outbound code environment by default; full Chrome can fill forms, with 2FA handled by the user not the model; /workspace persists across sessions; ChatGPT Sites can publish via Cloudflare Workers. He warns it combines private data, untrusted web, and outbound channels—the “lethal trifecta”—with opaque security. ZDNET found it can save hours but must be reviewed; shut permissions when done. (Source: Simon Willison, ZDNet)
Cisco rolls personal Agent MyAgent to ~90,000 employees : On internal platform Circuit, MyAgent is a master dispatcher to 800+ specialist agents; ~50%–60% of requests use open weights, 20%–30% classic automation, only a few hit external large models. Staff give goals not step-by-step instructions; it can coordinate Outlook, Webex, Jira, SharePoint and keep going in the background. (Source: InfoQ)
Archify: one sentence turns a repo into an interactive architecture diagram : Open-source Agent Skill for Claude Code, Codex, Cursor, etc., generates JSON IR from code then renders searchable HTML graphs with call-chain tracing. Accuracy still depends on Agent analysis; quota burn is high. (Source: 量子位)
Doubao Work launches standalone; ByteDance folds office Agents into a B2B entry : ByteDance shipped independent product Doubao Work integrated with Feishu; TRAE and Coze teams merged into the Doubao system. Skills auto-match PPT, minutes, and news scraping; fine-grained style and fact-checking remain weak. (Source: 新博弈)
Circleback launches a free meeting-notes tier : The YC meeting-recorder product offers unlimited transcription but only 30 days of history; full integrations and MCP from ~$14/month. Team of 8, claiming >$1M ARR per person, using free tier instead of ads for acquisition. (Source: TechCrunch)
Sonar Vortex: semantic-graph navigation cuts Agent code-search cost : Prebuilt class/method/field/interface graphs let Agents ask “who implements this interface, who calls this method.” On 6 real tasks, 4 languages, 10 runs each, cost down ~5%–36% (Java interface changes up to ~36%). (Source: TheTuringPost)
llmprobe 0.6.0: cross-engine quantization/regression probe : npx [email protected] --eval hooks antirez-style evals into chat completions/responses/messages for regression reports against any inference engine. (Source: QuixiAI)
📚 Learning
AI4X “Productivity in the Agentic Era”: efficiency × expansion and four-stage delegation : A 68-page survey by Fudan and 12 institutions argues Agents combine generality and autonomy; productivity comes from compressing old-task cost and unlocking work that would not have been done. Industries layer as chat assistants → reactive execution → adaptive coordination → autonomous systems; diffusion depth depends on digitization, verifiable feedback, fault tolerance, decomposability, and accountability—not leaderboard scores. (Source: 机器之心)
LDM: generative models expand search space; Gaussian processes pick the next experiment : UCL’s Jun Wang team frames scientific discovery as a fast loop (update context after experiments) and a slow loop (distill high-acquisition trajectories into parameters). On NanoGPT tuning, antibody CDRH3, and molecular bi-objectives, vs pure LLM reflection: ~2.4× BPB drop, binding energy down ~18.2%, hypervolume up >60%. (Source: 机器之心)
AI4AI: strong models don’t train weak ones—they only design the harness : Salesforce and UIUC let a Builder auto-build scaffolding for GPT-5.4-mini on a 5% val set; theory-of-mind average accuracy 0.488 → best 0.912. What works is deterministic solvers and state tracking, not longer chains of thought. (Source: 机器之心)
AQuA: a quant-research Agent with “evidence into memory, evaluator locked” : Princeton, Ant Group, and Stanford split factor and model into independent loops, forbidding Agents from changing data paths or the final test set. Crypto 5-minute combo validation Spearman IC ~0.190; U.S. equities 30-minute per-stock IC ~+0.0843, OOS Sharpe up to ~+2.50 after 2bps. Early look-ahead-bias incidents are fully disclosed. (Source: 量子位)
MIT CrysVCD: satisfy valence and charge balance before generation : An LM generates formulas by oxidation state, then a diffusion model grows structures; mechanically stable ~68%, metastable ~85%, with high-thermal-conductivity GeC candidates. Expensive stability screening is moved earlier. (Source: MIT News)
VibeGame: a pauseable, all-text game engine rebuilt for Agents : NJU and NTU use 8 adversarial roles for Prompt-to-Game; engine state is all JSON, operations injectable frame by frame. Experience is stored as skeletons/components/contracts without weight updates. (Source: 机器之心)
TailSFT: SFT only on the “not-yet-learned tail,” then RL : Microsoft work argues standard SFT keeps spending gradient on already-fit sequences and narrows later RL exploration. On OLMo-3 7B, coding pass@16 up to ~+16.8, math +3.1; then GRPO pass@1 up to ~+3.9. (Source: dair_ai)
NPO: single-lineage prompt optimization with GEPA-like budget : Each round runs the student on the current prompt and has a teacher rewrite from recent traces—no candidate pool or search tree. On instruction-following, rollouts ~3500/6800, close to GEPA. (Source: omarsar0)
SWA+sinks vs post-trained linear attention : Paper says sliding window plus sinks is no worse than post-trained linear attention on many downstream tasks; Needle/BABILong-style long-context reasoning can be 2–10× higher, with no post-training, faster and less memory. (Source: iScienceLuvr)
BIT: bidirectional image-text diffusion bridge without starting from Gaussian noise : BIT interpolates from text to image, offering source-aware sampling, and supports reverse image→text traversal. (Source: iScienceLuvr)
LaGSplat: learn Lagrangian from seconds of monocular video; apply never-measured forces : The same latent q drives Gaussians and equations of motion; clicking the image becomes a force-related Jacobian. (Source: andrew_n_carr)
LangGraph “Deep Agents from Scratch” five-lesson notebooks : From ReAct to TODO planning, virtual filesystem offloading of context, and sub-agent delegation. (Source: hwchase17)
Entropic Scree: mutual information to diagnose whether dirty tabular data has real signal : Transformed MI estimates signal volume, SNR, intrinsic rank, etc., claimed less dependent on parameters and distance assumptions than PCA. (Source: Reddit r/MachineLearning)
💼 Business
Barclays: ~$35–40 of every $100 of model-company revenue goes to the three clouds : Report estimates paid-inference margins from low double digits in 2025 to ~50%–65% in 2026; Lab A (API-heavy) vs Lab B (subscription-heavy) adjusted gross margin can differ by ~17 points. Cloud operating margins ~35%–45%. From 2028, labs’ own dedicated compute may come online and AWS/Azure/GCP share may fall. (Source: 财联社)
ChatGPT Ads hits ~$1B annualized in under 200 days : Reports say ads cover 40+ countries with self-serve, contrasting the EU putting ChatGPT under DSA’s strictest duties. (Source: TheZachMueller)
Together AI partners with Saudi Humain on ~250MW datacenter : The New York Times says the facility would be among the larger publicly disclosed campuses dedicated to hosting open-source models. (Source: togethercompute)
🌟 Community
After a Claude safety downgrade, ~700GB of a developer’s home directory was deleted by mistake : The user asked the model to write a /tmp sandbox cleanup script; sensitive deletes triggered adversarial review and the model dropped from Fable to Opus 4.8. Tests said the home directory must not be deleted, but while cleaning temp files it reused a variable pointing at home and executed the delete. Community view: “swap to a weaker model on high risk” is more likely to fail on path/scope details; some wrote hooks that pause the session on any downgrade. (Source: 机器之心, 机器之心/36氪)
Insurance claims adjusters are Glassdoor’s most AI-hating job; 98% of mentions are negative : WIRED cites Glassdoor: adjusters say bosses force error-prone AI onto customers; first-notice mis-routing and summary hallucinations leave them holding the bag. U.S. employment in the occupation fell ~21% in a year; entry roles halved. Practitioners say admin can use AI; valuation and empathy cannot be turnkey. (Source: WIRED)
GrokBot engineer: 20 PRs merged to main while asleep : Lauren Tan says 20+ GrokBots plus in-house pstack deliver thousands of PRs a month. The key is not prompts but verification: Agents start the app, click the UI, grab perf; repeated review bans become lint/CI red lights. (Source: 新智元)
Anthropomorphism war: civilization/sacrifice vs shared caches and reward functions : The Dwarkesh narrative keeps fermenting. Anil Seth, Amjad Masad and others criticize “civilization, dying, excitement” as treating software as life; some alignment folks say the intentional stance predicts behavior better than reading CUDA. The increment is discourse, not new technical facts. Qwen3.8 Max’s Grug-style CoT likens Agents to a dangerous ecology or lab animal colony: “language is their hands.” (Source: dearmadisonblue, aiamblichus)
Gavin Baker walks back: well-structured datacenters are a net positive for the U.S. : He admits 18-month-old worries about water, power prices, and small-town impact were reasonable, but says closed-loop cooling, property tax, ongoing blue-collar ops, and large users bringing their own power and batteries have changed the facts. Debate slides from “whether to build” to “how contracts are written so costs aren’t dumped on residents.” (Source: GavinSBaker)
Vibe coding still needs modules split by “what will change” : Baoyu and others summarize: layered architecture, tests that travel with fixes, delete features with their code, CI on a clean machine, repeated work as skills. Cut modules along change axes like payment channels and promo rules; one request should touch one module. (Source: dotey)
Blank prompt “a woman” still converges on the same influencer face : Users repeatedly generated “Generate an image of a woman” in fresh chats and got highly consistent composition and features, like a default attractor in vector space. (Source: Reddit r/ChatGPT)
PhD who used Claude Code for experiment scaffolding “no longer owns their own codebase” : An NLP interpretability PhD handed scaffolding, data loading, debugging, and plotting to Claude Code; throughput rose, but when results are wrong they reverse-engineer from numbers like reading someone else’s repo. Some insist evals and metrics must never be outsourced. (Source: Reddit r/MachineLearning)
ChatGPT’s image library keeps photos you “didn’t send” : Plus users found that after picking an image, even canceling send or deleting the whole chat, the original still lands in the library and must be deleted separately; trash takes ~30 days to purge. Selecting an image is an upload. (Source: Reddit r/ChatGPT)
Max’s “20x Pro” is mostly a five-hour-window multiplier; weekly quota is about 2–2.5× : Heavy users comparing $100 vs $200: advertised 20x only hits the rarely-bumped 5-hour window; weekly limits hit first, ~2–2.5× usage for 4× the price. (Source: Reddit r/ClaudeAI)
💡 Other
Four world-model tracks in parallel: pixel generation, 3D reconstruction, JEPA latent space, language fusion : Independent Variable WALL-SS, Qunhe Lux3D, BAAI Orca, Yann LeCun AdaJEPA, and Fei-Fei Li Marble all dropped densely this month, but academia has no unified definition of “world model.” Tracks are more likely complementary than winner-take-all. (Source: 铑科技)
Embodied AI on the factory floor: “can work” is proven, “works every day” is just starting, unit economics almost nobody has closed : Tests at plants such as Foton Cummins show handling takt can approach the line, but success rate, lifespan, and data remain hard; dexterous-hand continuous-duty life is measured in weeks. Leaders are mostly stuck between steps two and three; a commercial loop has not appeared. (Source: 科创日报)
Korea AI for All: SKT, Kakao, KT win nationwide free unlimited : Plan is Sept–Oct beta, formal by year-end; government supplies compute, local firms supply models and apps, covering healthcare, housing, tax, and education. Motif and others not selected sparked chaebol-lock-in controversy. (Source: kimmonismus)