‹ 2026-08-14 06:32Z · 17 citations ›

AIINT BRIEF — 2026-08-14

BLUF

The dominant theme this week is the maturation of agentic infrastructure and the economic reality of model serving. DeepSeek V4 Pro has launched via API, offering a new high-end reasoning option, while Ollama and llama.cpp are aggressively optimising local inference with NVFP4 support, speculative decoding defaults, and new quantisation types. Concurrently, rigorous new benchmarks (VAKRA, QuoteBench, SciFigBench) are exposing the fragility of agent tool-use and the hidden costs of memory and embedding pipelines, signalling a shift from raw capability to reliability and efficiency.

Developments

DeepSeek V4 Pro Launches via API

Agentic Reliability and Tool-Use Benchmarks

The Cost of Embeddings and Agentic Memory

llama.cpp and Ollama Optimise for Speculative Decoding and Quantisation

Claude Code Enhances Multi-Session and Subagent Management

Language-Conditional Dequantization for Multilingual Models

Trending

Assessment confidence

Corpus coverage is strong for local inference tooling (llama.cpp, Ollama) and recent benchmark papers, but limited for major lab model releases beyond DeepSeek V4 Pro. No significant security-specific AI news was found in the corpus.

Sources

  1. DeepSeek V4 Pro 0813 (on OpenRouter)Simon Willison · 2026-08-12 · corpus #1444https://simonwillison.net/2026/Aug/12/deepseek-v4-pro-0813/
  2. VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use PoliciesarXiv cs.AI · 2026-08-12 · corpus #1454https://arxiv.org/abs/2608.12282v1
  3. QuoteBench: How Matched Scores Can Hide Command-Path FailuresarXiv cs.AI · 2026-08-13 · corpus #1630https://arxiv.org/abs/2608.13547v1
  4. How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific FiguresarXiv cs.AI · 2026-08-13 · corpus #1673https://arxiv.org/abs/2608.13267v1
  5. The Embedder's Dilemma: LLMs Are Better, but at What Cost?arXiv cs.CL · 2026-08-13 · corpus #1708https://arxiv.org/abs/2608.12875v1
  6. Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory SystemsarXiv cs.CL · 2026-08-12 · corpus #1514https://arxiv.org/abs/2608.11879v1
  7. ggml-org/llama.cpp b10417llama.cpp releases · 2026-08-13 · corpus #1612https://github.com/ggml-org/llama.cpp/releases/tag/b10417
  8. ggml-org/llama.cpp b10415llama.cpp releases · 2026-08-13 · corpus #1614https://github.com/ggml-org/llama.cpp/releases/tag/b10415
  9. ggml-org/llama.cpp b10414llama.cpp releases · 2026-08-13 · corpus #1615https://github.com/ggml-org/llama.cpp/releases/tag/b10414
  10. ggml-org/llama.cpp b10413llama.cpp releases · 2026-08-13 · corpus #1616https://github.com/ggml-org/llama.cpp/releases/tag/b10413
  11. anthropics/claude-code v2.1.232Claude Code releases · 2026-08-13 · corpus #1619https://github.com/anthropics/claude-code/releases/tag/v2.1.232
  12. ggml-org/llama.cpp b10423llama.cpp releases · 2026-08-13 · corpus #1622https://github.com/ggml-org/llama.cpp/releases/tag/b10423
  13. ggml-org/llama.cpp b10419llama.cpp releases · 2026-08-13 · corpus #1623https://github.com/ggml-org/llama.cpp/releases/tag/b10419
  14. ollama/ollama v0.32.11Ollama releases · 2026-08-14 · corpus #1624https://github.com/ollama/ollama/releases/tag/v0.32.11
  15. anthropics/claude-code v2.1.231Claude Code releases · 2026-08-13 · corpus #1583https://github.com/anthropics/claude-code/releases/tag/v2.1.231
  16. anthropics/claude-code v2.1.229Claude Code releases · 2026-08-12 · corpus #1442https://github.com/anthropics/claude-code/releases/tag/v2.1.229
  17. Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English LanguagesarXiv cs.CL · 2026-08-12 · corpus #1525https://arxiv.org/abs/2608.11786v1

1 of 29 feeds silent · these sources have not been collected recently, so briefs may be missing their coverage: