AIINT BRIEF — 2026-09-25
BLUF
The open tooling stack is tightening security and expanding hardware support: the Model Context Protocol (MCP) TypeScript SDK 2.1.0 introduces request-time OAuth scope challenges and DPoP sender-constrained tokens, whilellama.cpp v0.5.0 and recent patches add AMD RDNA3/4 Vulkan int8 matmul, Ling 3.0 VL support, and multi-address HTTP binding. On the evaluation front, new benchmarks (EnigmaForge, Era by Eon, PASTABench) are shifting focus from static retrieval to proactive agent safety and hidden-knowledge reasoning, revealing significant gaps in current model intuition and multi-step risk management.
Developments
MCP SDK 2.1.0: Request-time OAuth and DPoP
- The
modelcontextprotocol/typescript-sdkreleased v2.1.0, adding request-time OAuth scope challenges viascopeChallengecallbacks and DPoP (RFC 9449) sender-constrained access token support. - This changes how developers secure tool and resource access, moving from static permissions to dynamic, per-request scope verification and cryptographic proof-of-possession, reducing token theft risks in agent workflows.
- Sources: [1], [2], [3], [4]
llama.cpp v0.5.0 and Hardware Optimisations
llama.cppv0.5.0 focuses on backend performance, adding HRM-Text (DFM Mimir 1B) support, Metal MoE/SSM_CONV fusion, and multi-address HTTP binding for the server.- Recent patches (b11160, b11175) add int8 coopmat1 matmul for AMD RDNA3/4 GPUs and q5_k quant type support for Hexagon, expanding efficient inference options for specific consumer and edge hardware.
- Sources: [5], [6], [7], [8]
Anthropic Claude Code Gateway Updates
anthropic/claude-codev2.1.281 adds support for newer Claude Desktop keys indesktoppolicy blocks,assume_rolefor Bedrock upstreams via STS, and configurable guardrails on Bedrock upstreams.- This provides finer-grained control over agent permissions and security policies when routing through AWS Bedrock, allowing developers to enforce guardrails and manage cross-account IAM roles more robustly.
- Sources: [9]
New Benchmarks: Hidden Knowledge and Proactive Safety
- EnigmaForge and Era by Eon are new benchmarks that test models on "hidden knowledge" reasoning, where answers are derived from implicit data rather than explicit questions, exposing a 22x spread in intuition among frontier models.
- PASTABench introduces Decoupled Proactive Safety Monitoring for agents, evaluating not just step-level actions but trajectory-level risk accumulation and timely intervention opportunities.
- These benchmarks signal a shift in evaluation from static QA to dynamic, multi-step reasoning and safety, highlighting weaknesses in current models' ability to handle implicit context and proactive risk management.
- Sources: [7], [10], [11]
KV-Cache Compression and Quantization Research
- New papers propose risk-controlled KV-cache eviction based on deployment risk targets rather than average quality, and tensor decomposition of KV caches showing low-rank structure in token/feature modes.
- Another study on post-training quantization (PTQ) formulates configuration selection as a pre-deployment cost problem, predicting degradation before deployment.
- These developments offer concrete methods for managing memory and accuracy trade-offs in long-context inference, crucial for scaling agent deployments.
- Sources: [12], [13], [14]
Trending
- Agent Safety Evaluation: Benchmarks like PASTABench and FDE-Bench are gaining traction, focusing on proactive risk monitoring and deployment environment configuration rather than just task completion. [11], [15]
- Hardware-Accelerated Inference:
llama.cppis rapidly expanding support for specific hardware (AMD RDNA, Hexagon) and quantization formats (q5_k, int8 coopmat), driving down inference costs for edge devices. [6], [8], [7] - Security in MCP: The addition of DPoP and request-time OAuth challenges in the MCP SDK reflects a growing emphasis on securing agent-to-tool communication against token theft and scope escalation. [1], [3]
Assessment confidence
Corpus coverage is limited to the provided items; no major model releases or ecosystem ownership changes were reported in the last 24-48 hours. Security stories are excluded unless directly impacting AI evolution (e.g., MCP security features).Sources
- modelcontextprotocol/typescript-sdk @modelcontextprotocol/node@2.1.0https://github.com/modelcontextprotocol/typescript-sdk/releases/tag/%40modelcontextprotocol/node%402.1.0
- modelcontextprotocol/typescript-sdk @modelcontextprotocol/server@2.1.0https://github.com/modelcontextprotocol/typescript-sdk/releases/tag/%40modelcontextprotocol/server%402.1.0
- modelcontextprotocol/typescript-sdk @modelcontextprotocol/core@2.1.0https://github.com/modelcontextprotocol/typescript-sdk/releases/tag/%40modelcontextprotocol/core%402.1.0
- modelcontextprotocol/typescript-sdk @modelcontextprotocol/client@2.1.0https://github.com/modelcontextprotocol/typescript-sdk/releases/tag/%40modelcontextprotocol/client%402.1.0
- ggml-org/llama.cpp v0.5.0https://github.com/ggml-org/llama.cpp/releases/tag/v0.5.0
- ggml-org/llama.cpp b11156https://github.com/ggml-org/llama.cpp/releases/tag/b11156
- EnigmaForge: The Question Is Hidden in the Storyhttps://arxiv.org/abs/2609.30144v1
- ggml-org/llama.cpp b11175https://github.com/ggml-org/llama.cpp/releases/tag/b11175
- anthropics/claude-code v2.1.281https://github.com/anthropics/claude-code/releases/tag/v2.1.281
- Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledgehttps://arxiv.org/abs/2609.30055v1
- PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safetyhttps://arxiv.org/abs/2609.28197v1
- Risk-Controlled KV-Cache Eviction: From Memory Budgets to Risk Targetshttps://arxiv.org/abs/2609.27981v1
- Tensor Decomposition of Transformer Key-Value Caches: Spectral Structure and Format Comparisonhttps://arxiv.org/abs/2609.28029v1
- Predicting Quantization Price for Selecting PTQ Configurations Before Deploymenthttps://arxiv.org/abs/2609.28270v1
- FDE-Bench: Evaluating LLM Agents for Deployment Environment Configurationhttps://arxiv.org/abs/2609.27571v1
