‹ 2026-09-16 06:34Z · 10 citations ›

AIINT BRIEF — 2026-09-16

BLUF

Google has released Gemini 3.8 Live and 3.8 Live Extended Thinking, introducing speech-to-speech models with interruption handling that mirror the OpenAI GPT-Live architecture [1]. Meta has launched a WhatsApp Business MCP server, enabling AI coding agents to automate business setup and messaging workflows [2]. On the infrastructure side, llama.cpp v0.4.1 and recent commits have added support for Maple 20B-A1B and Tencent Hy 4, while fixing stateful decoding and GPU MoE inference for OpenVINO [3] [4].

Developments

Google Gemini 3.8 Live: Speech-to-Speech with Interruption Handling

Meta WhatsApp Business MCP Server

llama.cpp v0.4.1 and OpenVINO Optimisations

Anthropic SDK v1.6.0: Managed Agent Permissions and Compaction

JustFit: 200K-Token LLM Serving on 24 GiB Laptops

Trending

Assessment confidence

Corpus covers releases from Google, Meta, Anthropic, and llama.cpp, plus recent arXiv papers on inference optimisation and benchmarking; does not cover security incidents or non-AI ecosystem news.

Sources

  1. Gemini Live audioSimon Willison · 2026-09-15 · corpus #35106https://simonwillison.net/2026/Sep/15/gemini-live/
  2. Meta now lets AI agents handle the boring parts of WhatsApp Business setupTechCrunch AI · 2026-09-15 · corpus #35074https://techcrunch.com/2026/09/15/meta-now-lets-ai-agents-handle-the-boring-parts-of-whatsapp-business-setup/
  3. ggml-org/llama.cpp v0.4.1llama.cpp releases · 2026-09-14 · corpus #34893https://github.com/ggml-org/llama.cpp/releases/tag/v0.4.1
  4. ggml-org/llama.cpp b10981llama.cpp releases · 2026-09-15 · corpus #35050https://github.com/ggml-org/llama.cpp/releases/tag/b10981
  5. anthropics/anthropic-sdk-python v1.6.0Anthropic python SDK releases · 2026-09-15 · corpus #35065https://github.com/anthropics/anthropic-sdk-python/releases/tag/v1.6.0
  6. JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State ManagementarXiv cs.AI · 2026-09-15 · corpus #35180https://arxiv.org/abs/2609.17475v1
  7. Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM CompressionarXiv cs.LG · 2026-09-14 · corpus #34977https://arxiv.org/abs/2609.15838v1
  8. Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder FamiliesarXiv cs.CL · 2026-09-14 · corpus #35156https://arxiv.org/abs/2609.16391v1
  9. ECHO: A Matched-Contrast Benchmark for Context-Sensitive Turn-Taking in Full-Duplex DialoguearXiv cs.CL · 2026-09-15 · corpus #35223https://arxiv.org/abs/2609.17360v1
  10. RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken DialoguearXiv cs.CL · 2026-09-15 · corpus #35143https://arxiv.org/abs/2609.16614v1

1 of 29 feeds silent · these sources have not been collected recently, so briefs may be missing their coverage: