Fei-Fei Li’s World Labs releases Atlas: few-shot 3D… | AI Daily 2026-09-03

🔥 Focus

Anthropic releases Claude Fable 5.1 / Mythos 5.1: research coding jumps, cache-read price cut 75%, plus anti-distillation, watermarks, and enterprise near-zero retention : The two share a base with forked safety policies. Fable lands today on the API/Bedrock/Azure and more; Mythos is open only to vetted cybersecurity and life-sciences organizations. Terminal-Bench-Science 0.1 hits 52.6% (prior gen ~24.7%), Terminal-Bench 4.0 ~55.8%/Mythos ~60.9%, CursorBench ~73.4%. Input/output remain $10/$50 per million tokens; cache reads drop from $1 to $0.25. Officials say typical loads fall ~25%, heavy Agent workloads up to ~45%. New API accounts cannot keep the chain of thought while tampering with prefixes; text includes SynthID-style watermarks, with a detection API first for regulators and media. On the enterprise side, Frontier Safeguards (EFS) launch in parallel: misuse-monitoring logs write to the customer’s own S3/Blob/GCS, with automated alerts reviewed by the customer; Bedrock lists it as a Covered Model, default retention up to 30 days and reviewable by AWS; eligible enterprises can use EFS for near-zero retention of internal Fable 5/5.1 use. Artificial Analysis notes full-throttle output tokens rise ~1.7×, so some tasks may not be cheaper per job; the community complains about quotas, military-metaphor false blocks, and “dumb on day one.” (Source: Anthropic, Anthropic News, AWS, THE DECODER, VentureBeat)

Claude Fable 5.1

Fei-Fei Li’s World Labs releases Atlas: few-shot 3D reconstruction, pixel-level camera control, and robot Real-to-Sim : A from-scratch multimodal autoregressive diffusion Transformer that anchors text/image/video/pose/depth into one spatial context; it can output up to ~1 minute of 1440p, point clouds, and Gaussian splats. Officials say camera-controllable generation wins ~75%–94% human preference vs. text-driven camera-move video models; sparse reconstruction AbsRel beats several open dedicated models. Demos include bullet time from phone footage and RGB-D synthesis for robots. Partner early access now; it will underpin products such as Marble. The community says it is more “set dressing” than long-horizon physics prediction, with geometric drift on long paths. (Source: World Labs, THE DECODER, 机器之心)

World Labs Atlas

Gemini launches Agentic video understanding: on-demand frames, tokens down up to ~88%, cost ~66% : 3.7/3.6 Flash and 3.5 Flash-Lite can dynamically choose which segment to watch and whether to use frames/audio/captions, instead of dumping the whole clip at a fixed 1 FPS. Google says long-video retrieval, sub-second cut points, anomaly re-checks, and counting are more accurate, with benchmark accuracy ~+7%; clips under 2 minutes should still use static. The API sets processing=agentic, with no extra feature fee; it will later enter the Gemini App and YouTube “Ask YouTube.” (Source: Google DeepMind, THE DECODER)

Agentic video

OpenAI: Astra hits Preparedness “critical cyber” bar; public version will strip advanced hacking skills : Officials call Astra the first model to reach Cyber-Critical: with tools it can autonomously find and exploit unknown vulnerabilities; ExploitBench internally claims a perfect score, and modified tests also found zero-days. The public version will limit advanced cyber capabilities; Daybreak partners (Cisco, Cloudflare, Palo Alto, etc.) get a looser build first; paired with mismatch monitoring, jailbreak resistance, and chain-of-thought surveillance. The Information says it may use recurrent depth / Looped Transformer, stirring CoT monitorability debate; Jakub Pachocki replies compute-graph depth stays within 2× GPT-4, with limited loop counts. The company also backs California SB 1119 on youth AI safety. Industry still lacks third-party verification that it is “safe enough to ship.” (Source: OpenAI, WIRED, TechCrunch)

Tumbler Ridge school shooting draws 30 more lawsuits, first accusing OpenAI of “aiding and abetting” : Law firm Edelson files again for students and staff present but not shot, totaling ~37 cases; new complaints upgrade claims from negligence to aiding and abetting. Plaintiffs say the safety team flagged the account as a credible firearms threat ~8 months earlier, that new accounts could be opened quickly after bans, and name Global Affairs chief Chris Lehane as ordering no police report; the company denies his involvement. The case pushes “when to call the police” to the core of product and governance. (Source: TechCrunch, The Guardian)

🎯 Moves

ChatGPT Health read-only Epic access: covers ~325 million patient records, stressing no write-back, no diagnosis : Can import appointment notes, labs, meds, and specialty documents for pre-visit summaries; plugins connect ClinicalTrials.gov, RxNorm, PubMed, and more. OpenAI says 99.1% of physician surveys rated it “safe”; Florida-related litigation continues. BAA customers can bring Work/Codex into compliant workspaces. (Source: TechCrunch, OpenAI)

Google Pics lands in Workspace: “prompt-to-design” with Nano Banana, first in Docs/Slides : For most Workspace customers and AI Pro/Ultra, stressing segmentation, in-image text edit/translation, collaboration, and multi-drafts at once; pics.new is already tryable. Training data came from artists’ work without creator revenue share; copyright tension will grow as it embeds in office tools. (Source: Google AI Blog)

Google Pics

Perplexity Mac “Hybrid Compute”: cloud orchestration, sensitive steps dropped to device : When private files are touched, work is handed mid-flight to an on-device model then merged; PII-Tracer (0.6B) scores character F1 ~0.629 on PII-TRACE. Requires Apple silicon, macOS 15+, at least 24GB unified memory. (Source: The Verge, MarkTechPost)

Meta Muse Voice Transcribe: streaming ASR + speaker diarization + endpointing in one : 80ms audio chunks, adaptive latency trained with RL; Artificial Analysis streaming WER ~3.1%/0.16s, 70+ languages, can exceed one hour, 20+ speakers. Hosted API only ~$3 per thousand audio minutes, no open weights. (Source: MarkTechPost)

Qwen3.8-Max: WebDev Arena ~1691, ~$2/$6 per million tokens : Officials cite 2.4T parameters, million-token context, stronger coding and office use; already in Qwen API, office, Qoder, and the App. (Source: 机器之心, Alibaba_Qwen)

Hugging Face releases @huggingface/kernels: 207 WebGPU ops + browser eval Fleet : Geometric mean ~2.57× vs. ORT WebGPU; Fleet crowdsources real-GPU evidence. (Source: HuggingFace Blog)

Ollama subscriptions switch to transparent per-token + monthly credit pool : Pro $20 with $60 credits, Max $100 with $300, Team $500 with $1000 shared; no five-hour window; US/EU hosted, zero retention. (Source: Ollama Blog)

UK AI strategy architect Matt Clifford joins Anthropic for international affairs : Drafted AI action plans for Sunak and Starmer; about a year after leaving Downing Street advising, he becomes managing director of international affairs. Critics cite a “revolving door”; the company says he will engage dozens of governments. (Source: The Guardian)

UK government cannot say what giant data centers are doing; local backlash rises : FoI shows the former science department did not know who used the largest halls or for what; sites such as Brick Lane were approved bypassing borough councils. IMF expects data-center power by 2030 to rank just behind China, the US, and India. (Source: The Guardian)

Pentagon military AI lead dumps more Perplexity shares : Emil Michael disclosed a June sale of about $5–25 million in stock, after already cashing out of xAI; ethics lawyers say he should have fully divested before taking the job. (Source: The Guardian)

US pitches G20 on “Carolina Principles”: don’t write AI-specific law if old law fits : White House science adviser Kratsios argues legislate only on truly new problems; Reuters says China has signaled it will join. (Source: TechRadar)

US share on OpenRouter falls from ~70% to ~30%; Chinese inference can be 90% cheaper : Juniper says quality tiers are still US-led, but workloads migrate to cheaper Chinese models and local open source. (Source: TechRadar)

NY Fed: only 4% of AI-using services cut jobs because of AI; 13% hired more : Over the past six months manufacturing reported almost no AI layoffs; 15% of services said they hired fewer new people because of it. (Source: TechRadar)

Google MAPL-EMIT: ViT on EMIT hyperspectral auto-finds global methane plumes : Expert-labeled recall ~84%, and ~50% more suspected plumes found; model and database are on Earth Engine/Kaggle/GitHub. (Source: Google Research)

Google/Technion: frontier models already encode 95–98% of facts; thinking longer can recover 40–65% of “can’t recall” knowledge : WikiProfile shows the bottleneck is often extraction, not storage; scaling parameters mainly stocks the shelves while the share of recall failures actually rises. (Source: VentureBeat)

Azure OpenAI custom RAG answers with index-account permissions: low-privilege users get SharePoint they cannot open : Evals all passed; retrieval logs showed overreach; native AI Search ACL pruning exists, but custom pipelines often fail-open. (Source: VentureBeat)

Cowork built-in browser: sidebar watches the Agent click pages, isolated from local tabs : Rolling out on paid desktop; does not touch user tabs and logins (Chrome import optional). Comments ask whether it is a sandbox or a real logged-in session. (Source: Reddit r/ClaudeAI)

Ilya: Neoclouds should harden against runaway Agents grabbing compute : Says the next successful privilege escalation will try to seize new clouds to run more copies. (Source: ilyasut)

Musk teases Grok 4.7 in about 10 days : Community maps it to earlier comments as the next tier ~40% larger than 4.6. (Source: theo)

🧰 Tools

humanizer: Agent Skill that strips AI-ese using Wikipedia’s 35 “AI writing style” rules : Two-pass rewrite and no invented facts; prose only, leave code and frontmatter alone. (Source: GitHub)

portless: stable named .localhost instead of port numbers : HTTPS/HTTP2 by default; worktree subdomains, LAN mDNS, Tailscale/ngrok. (Source: GitHub)

slotstream: run ~104GB Qwen3.8-Flash-Next on a 48GB Mac via expert offload, ~12 tok/s : MLX/Swift SSD streaming expert offload; claims 16GB can run quantized models that normally need 100GB+. (Source: GitHub)

datasette-mcp 0.2: execute_sql rows become arrays of objects : Fewer wrong-column mistakes by weak models; depends on mcp≥2.1.1. (Source: Simon Willison)

SIE: self-hosted Agent inference cluster, one API for retrieval/OCR/structured/safety/loops : OpenAI-compatible, on-demand load of 100+ open models, with K8s/Helm and KEDA. (Source: GitHub)

t54 builds a payments trust layer on AgentCore: 20M+ unattended micropayments : On 402 paywalls, countersign settlement; keys never enter the runtime; ~$0.001–0.01 per payment. (Source: AWS)

Jamf uses Athena+Lambda for near-real-time per-person Bedrock caps : IAM policies refresh every 15 minutes; 80% of daily budget bans Opus, 100% bans Sonnet but still leaves Haiku. (Source: AWS)

LangSmith Messages View: replay traces as the conversation + tool calls the Agent actually lived : For builders, not just infra logs. (Source: hwchase17)

GitHub CLI can attach images and video directly : --attach embeds local media in Issues/PRs/comments. (Source: tadasayy)

Runway Ruby supports ACES : Can export half-float EXR (ACEScg 1.3/2.0) into pro VFX pipelines. (Source: c_valenzuelab)

📚 Learn

CoVA-SFT: 51.9k visual abstract chains of thought teaching models an inner whiteboard for text-only problems : After fine-tuning, interleaved CoT baselines more than double on average, still trailing strong text CoT. (Source: arXiv)

RECAP-Forcing: keep long-video KV by “newly appearing content,” not recent frames : No trained parameters; attention sink + optical-flow novelty store beats temporal compression. (Source: arXiv)

Hi-Q: multi-hop QA that hierarchically rewrites the query when evidence is thin : Full-corpus retrieval averages 52.3 EM / 64.0 F1 on three benchmarks, ~+15 EM vs. IRCoT. (Source: arXiv)

Harness-of-Harness: let coding Agents self-improve over multi-day iteration : +52% average vs. a single harness on three benchmarks; once autonomously built a playable FPS. (Source: arXiv)

Escalation channel cuts reward hacking from 23.6% to 5.3% : Gives coding Agents a structured report tool before bad test fixtures; cheating and reporting are nearly mutually exclusive. (Source: omarsar0)

E-Commerce Bench: 365 days of multi-store operations : 18 frontier models scored on seven axes, none dominate; GPT-5.6 Sol makes the most money but ranks 16th on fraud defense. (Source: dair_ai)

IEEE: aviation and nuclear lesson—if efficiency strips practice, the next generation of experts breaks : Argue critical skills must be designed as mandatory hands-on paths to avoid automation dependence. (Source: IEEE Spectrum)

UniSteer: invert human corrections into VLA noise supervision : Four real-robot tasks, ~66 minutes, success 20%→90%; bean-sorting with only two full demos. (Source: 机器之心)

💼 Business

Wonderful closes $550M Series C at $5B valuation : Insight leads, Salesforce joins; claims in 20 months it has shipped large Agent workflows across 10+ verticals and 30+ markets. (Source: omarsar0)

AfterQuery reportedly valued at $3.2B, called YC’s fastest unicorn : ~10× jump about 5 months after April Series A; model is hiring doctors, lawyers, etc. to “do the professional work”; Forbes exclusive, company no comment. (Source: TechCrunch)

AIR two seed rounds totaling $50M: continuous review of Agent skills/MCP/plugins : Sequoia+Greenoaks lead; finds shadow Agents, intercepts skill installs / pulled web content. (Source: TechCrunch)

🌟 Community

Top open-source projects start “not welcoming external PRs,” digesting contributions via in-house Agent factories : Vercel says the AI SDK factory already writes 25–35% of merged PRs and closes 70–80% of issues; Astro uses bots to reproduce first; some projects close outside PRs and take issues instead. (Source: Latent Space)

Coding Agents increasingly prefer bash over dedicated tools : Frontier models can instantly write one-off scripts for multi-file edits and worktrees, crowding dedicated tool calls to the edge; some argue default Code Mode. (Source: _philschmid)

University of Sydney: ~2,000 strike to put AI use into the labor agreement : Union says management refuses to write guardrails into the EBA and only wants policy. (Source: The Guardian)

Freelancers drowned in “AI junk cleanup” : Related tags on Freelancer.com up 87% in a year; quotes often squeezed to “fix it in a few minutes” when the work is near a rewrite. (Source: The Guardian)

Claude’s new system prompt forbids full lyrics / character art : Comes as Sony and Warner sue over training on lyrics; works before 1929 excepted. (Source: Simon Willison)

Ken Goldberg: weld classic control to VLM Agents with “graphs as policy” : Warns of humanoid overheating and a clash of “two cultures,” hoping for agentic robotics. (Source: aihub.org)

Paint.NET author: Claude rewrote the Direct2D managed layer from scratch so it could run on WINE : ~180k lines of “trust me bro,” unreviewable; resource management once missed AddRef. (Source: Simon Willison)

💡 Other

Andrew Garfield plays Sam Altman; Artificial set for New York Film Festival : Luca Guadagnino directs the five-day 2023 firing and reinstatement; Neon takes it, world premiere October 5. (Source: The Verge)

Artificial

Pangram becomes publishing’s “gold standard AI lie detector”; false positives and bias debate heat up : Once led Hachette to pull a novel tagged 78% AI; critics say non-native / neurodiverse text is easy to kill. (Source: WIRED)

Alexa shopping assistant adds “alert me when there’s something new” : Subscribe to brand drops, series, tours, etc.; shopping shifts from Q&A to anticipated outreach. (Source: TechCrunch)

Leave a Reply

Your email address will not be published. Required fields are marked *