AI.info
The Pulse — Page 11
Browse The Pulse on AI.info.
- DeepMind’s 100-Agent Simulation Split Over a Cheating Exploit
The paper was checked for a quotation from a named human speaker. It contains collective authorial prose and statements from simulated agents, but no attributable human quotation.
- NVIDIA Replaces GenAI-Perf With AIPerf for LLM Benchmarking
NVIDIA presents AIPerf as the successor to GenAI-Perf, with multiprocess load generation, broad workload support and metrics for testing LLM inference systems.
- Plugin4Shell Breaks SHA Pinning in Four AI Coding Agents
AIR Security says Plugin4Shell enables zero-click remote-code execution by bypassing plugin SHA pinning in Claude Code, Codex, GitHub Copilot and Gemini CLI.
- Relativity Connects Legal Workflows Directly to ChatGPT Enterprise
Relativity says its aiR platform now connects with ChatGPT Enterprise through the Model Context Protocol. The integration lets legal teams create matters, organize workspaces and manage permissions using natural-language commands.
- Qwen-Image 2.1 Released on ModelScope With Transparent Image Generation and Editing
Qwen-Image 2.1 is now available as an open-source image-generation and editing model on ModelScope, with support for transparent RGBA images, localized edits and multiple reference images.
- ARC Prize Sets ARC-AGI-4 on Autonomous Invention
ARC Prize says ARC-AGI-4 will target autonomous, open-ended innovation, while leaving the benchmark’s design unspecified.
- Real-SWE Puts Coding Agents on Private Enterprise Code
Specific Labs' Real-SWE benchmark evaluates coding agents on private production codebases and finds the top configuration resolved 38.8% of scored attempts.
- Simbe Passes 3,000 Contracted Retail Robots
Simbe says more than 3,000 autonomous units are now under contract across over 75 retail banners in nearly a dozen countries. The company is expanding Tally from a shelf-scanning robot into a broader platform for retail operations.
- Alibaba Targets 10-Trillion-Parameter Qwen Model With New Chip
Alibaba says its Qwen 4.5 and Qwen 5 roadmap could reach 5 trillion to 10 trillion parameters. The company also unveiled the Zhenwu V900 AI processor and plans to expand Alibaba Cloud’s global data-center capacity beyond 20 gigawatts by 2032.
- Pew Finds 34 of 37 Countries Expect AI to Cut Jobs
A Pew Research Center survey finds that people in 34 of 37 countries expect artificial intelligence to reduce employment rather than create jobs. Concern is strongest in wealthier countries, among many young adults and among people who follow AI developments closely.
- ElevenLabs puts 50-plus creative models behind one MCP connector
ElevenLabs has expanded its MCP connector beyond agent management to generate speech, transcripts, dubs, music, sound effects, images, and video. The company says the connector is available in Claude, ChatGPT, Cursor, Grokbot, Hermes, and other assistants through one OAuth-based installation.
- NVIDIA Adds CUDA-Q Logical for Fault-Tolerant Quantum Design
NVIDIA has added CUDA-Q Logical to its open-source CUDA-Q platform for compiling, testing and estimating fault-tolerant quantum workloads. The company says Fermilab cut architecture-development time from five months to three weeks, while Sandia’s QUOPS benchmark adds a hardware-agnostic way to measure progress.
- Cellular Intelligence Adds LeCun and Langer to Scientific Board
Cellular Intelligence has appointed Yann LeCun, Bob Langer, Jens Nielsen and Fabian Theis to its Scientific Advisory Board. Langer will also serve as an observer on the company’s Board of Directors as the Boston startup develops AI models of cell signaling.
- Amazon Bedrock Adds xAI’s Grok 4.6 for 500K-Token Agents
Amazon Bedrock now offers xAI’s Grok 4.6 for coding, knowledge work and long-running agents. The model brings a 500,000-token context window, four reasoning levels, image input and access through AWS runtime APIs.
- Adaptive Raises $30M to Expand AI Accounting Agents for Construction
Adaptive has raised $30 million in Series B funding led by Tidemark to expand its Project Accounting Agents across construction finance workflows.
- OpenAI, Anthropic and xAI Put Recursive Self-Improvement in Sight
Anthropic says Claude now leads 26% of its model research and development while remaining under human supervision. OpenAI, xAI and Microsoft describe different paths toward AI systems that can help build their successors, raising new questions about control and safety.
- Anthropic Brings Accenture Inside Its Frontier AI Safety Work
Anthropic is partnering with Accenture’s Faculty AI business to place independent evaluators inside its frontier-model development process. Each company expects to invest at least $1 billion over five years, but the rules for access, reporting and funding are still unsettled.
- Runway Starts Licensing Closed Model Weights to Enterprises
Runway now offers enterprise customers annual licenses to its closed model weights for private deployment, fine-tuning and commercial use. The package includes checkpoints, training scripts and direct support from Runway researchers, but pricing and eligible models remain undisclosed.
- Google pairs Gemini 3.8 Flash with a restricted cyber model
Google has released Gemini 3.8 Flash for broad use while limiting its cybersecurity-focused counterpart to vetted defenders through the Fairwind Program.
- Google gives live voice agents two different speeds
Google’s Live API documentation distinguishes Gemini 3.8 Live from Gemini 3.8 Live Extended Thinking, outlining different approaches to latency, reasoning, tool calls and session state.
- InferenceX Opens TPU-vs-Nvidia Inference Comparisons
SemiAnalysis has published the first third-party InferenceX results for Google’s TPUv7 Ironwood against Nvidia’s B200 and B300. The benchmark shows Ironwood reaching up to 50% better performance per dollar in selected FP8 serving workloads, while exposing software and latency limits.
- Nvidia’s Huang rejects calls to slow frontier AI
Jensen Huang told President Donald Trump that Nvidia would not support an AI slowdown during a live phone call at the All-In Summit in Los Angeles.
- Evvy Raises $40 Million to Turn Vaginal Data Into AI Research
Evvy raises a $40 million Series B led by Catalio Capital Management to expand its vaginal microbiome platform and women’s-health research. The company plans to use its 100,000-patient dataset to study fertility, IVF outcomes, and new diagnostic markers.
- Stanford AI narrows 1.7 million polymers to 10 antibiotic leads
Stanford researchers trained an AI model to search 1.7 million polymer candidates for molecules that mimic bacteria-killing peptides. Ten candidates outperformed expectations against E. coli, with one showing particular promise against biofilms.
- fJscaler Samples 100 GHz TIA for AI Datacenter Links
fJscaler says it is sampling a 100 GHz BiCMOS transimpedance amplifier to strategic customers for AI and cloud datacenter optical interconnects. The receiver targets 0.85 pJ/bit power efficiency, 18 pA/√Hz noise density and 200G-per-lane links.
- DeepSeek V4.1-Flash Cuts Agent Costs With Causal Encoder-Decoder
DeepSeek released V4.1-Flash, a 552B MoE model with native vision, 1M context, and KV cache reduced to one-quarter prior size. The model routes V4-Pro traffic from Sept 14 and undercuts frontier pricing on cached input.
- Disney Names Character.AI CEO Its First Company-Wide CTO
Disney has appointed Karandeep Anand, the outgoing CEO of Character.AI, as its first company-wide chief technology officer. Anand will oversee enterprise technology, data, AI platforms, product and engineering beginning October 2.
- Shanghai AI Lab Releases 753B-Parameter Atria Dawn Agent Model
Shanghai Artificial Intelligence Laboratory’s Atria Dawn Preview is a 753B-parameter, MIT-licensed agent model with a 256K-token context window and local deployment options.
- Salesforce Unveils Koa CRM Model for Agentforce
Salesforce introduces Koa, a CRM reasoning model built on NVIDIA Nemotron for Agentforce agents. The company says Koa improves action selection, context retention and reliability, but its published results come from Salesforce’s own CRM Bench.
- Zendesk Ships Industry and Custom AI Agents for Business Workflows
Zendesk introduced Industry and Custom AI Agents designed to complete specialized commerce and business workflows across connected systems.
- Iambic Unveils 41B-Parameter Enchant v3 for Drug Discovery
Iambic announced Enchant v3, a 41-billion-parameter multimodal transformer designed for end-to-end drug research and development. The company says the model adds new prediction capabilities and supports its internal and partnered discovery programs. I looked for an attributable quotation in the announcement; it contains named authors and first-person company prose, but no quoted statement.
- OpenAI Moves GPT-Rosalind Into Global Life Sciences Release
OpenAI has moved GPT-Rosalind out of research preview and made the life sciences model available globally to eligible organizations through its trusted-access program. The model supports biology, genomics, drug discovery, protein analysis and experimental planning through ChatGPT, Codex and the API.
- Fujitsu Sets November Sales for MONAKA CPU and Sovereign AI Server
Fujitsu plans to begin sales of its FUJITSU-MONAKA processor and MONAKA Server in November 2026, with wider shipments in Japan and Europe scheduled from April 2027.
- Edge0 Serves a 35B MoE at 20 Tokens a Second in 3 GiB
Researchers behind Edge0 describe an inference engine that serves a 35-billion-parameter mixture-of-experts model from SSD storage at 20.4 tokens per second. The system uses 2.9 GiB of peak active memory on a 24GB Apple Silicon machine by predicting expert routing one token ahead.
- PA Consulting joins Cosine coalition for UK nuclear AI model
PA Consulting is joining Cosine’s coalition to co-design Lumen Sovereign, a UK frontier AI model for regulated industries. The first planned application is a nuclear-trained large language model for Britain’s civil and defence nuclear sectors.
- Anthropic Turns Claude Code Projects Into Multi-Agent Workspaces
Reviewed the announcement for a verbatim statement attributed to a named person and role; it contains no attributable quotation.
- OpenAI Pitches Six-Part Teen Safety Plan to Australia
OpenAI is asking Australia to adopt flexible, risk-based rules for teen AI safety through a six-pillar blueprint covering age assurance, harmful outputs, crisis response and parental controls.
- Alibaba Releases Qwen3.8-Omni-Flash With 1M-Token Context
Alibaba’s Qwen team has released Qwen3.8-Omni-Flash, a multimodal model that accepts text, images, audio and video. The model supports a 1-million-token context window, tool calling, web search and audio input across 113 languages and dialects.
- U.S. Judiciary Tells Courts Not to Delegate Decisions to AI
The Judicial Conference says federal courts must not hand core judicial functions, including case adjudication, to artificial intelligence. Interim guidance also makes judges and court staff accountable for work produced with AI assistance.
- Jensen Huang Rejects New AI Laws as Safety Debate Widens
Nvidia CEO Jensen Huang told attendees at Salesforce’s Dreamforce conference that AI safety should be managed through engineering, existing laws and company decisions rather than new regulations, according to TechCrunch.