Skip to content

AI Daily — 2026-09-16

Sep 16, 20261 min
TL;DRSalesforce bundles Agentforce 360's seven named agents with an AI Control Plane governance layer; the same day, a study finds agent compromises stay invisible to standard safety dashboards; Cathay Financial Holdings builds identity/permission and audit-trail controls before declaring 'Agent First'; South Korea's KISA survey finds 82% of firms have unidentified shadow AI agents; Mistral anchors an $11B funding week; Gemini 3.8 Live voice mode ships at less than half OpenAI's price

🌏 中文版

The One-Line Take

The gate for enterprise agent adoption is shifting from "which model" to "can we govern it" — Salesforce, Cathay Financial Holdings, and South Korea's government all proved the same point today: without identity permissions and audit trails, a capable agent still shouldn't be scaled up.

Deep Dive: Governance, Not Capability, Is the Real Bottleneck for Agent Deployment

I think today's events point to the same transaction-cost problem: the cost of trusting an agent enough to hand it real authority is now higher than the cost of training or picking a stronger model.

Salesforce's Agentforce 360 is the most direct evidence: alongside seven named functional agents, it ships an AI Control Plane and a six-pillar "Trusted Enterprise AI Harness" governance framework — the governance layer is bundled with the capability layer, not an optional add-on. (source)

But an academic study published the same day shows this trust mechanism itself isn't reliable yet: attacks can succeed at an agent's planning, memory, or tool-call layer while the standard enterprise safety monitoring sees nothing wrong, because the final output still looks clean. (source) That means while the market is selling governance frameworks, whether those frameworks can actually catch an agent going wrong is still an open question.

Taiwan's example confirms the same logic: Cathay Financial Holdings declared today it's moving from "Cloud First" to "Agent First," unveiling three digital coworkers still in testing — but alongside that it built five governance mechanisms: identity/permissions, system integration, security monitoring, and audit trails. Governance came first, before scale. (source) South Korea has gone further, making "Security for AI" a national initiative: KISA's AI Security Guide v2.0 explicitly targets misuse of agent execution permissions, and a CSA survey found 82% of companies have AI agents inside their organizations that they hadn't even identified. (source)

What this means for practitioners: when evaluating an agent platform going forward, the first question shouldn't be "how capable is the model" but "can we see and trace what happens when it fails." For Taiwanese enterprises specifically, that means auditing for shadow agents your own IT department doesn't know about before expanding deployment — not rushing to scale.

Today's Developments

Vendor Moves

Salesforce: Launched Agentforce 360, adding seven named functional agents (Casey, Paige, Carter, and others), and trained a CRM reasoning model, Koa, on NVIDIA Nemotron 3 Super, claiming an error rate three times lower than mainstream models on its own benchmark. (source · Koa source)

Google DeepMind: Shipped Gemini 3.8 Live, whose voice mode can listen and speak simultaneously while calling tools in parallel, topping the Speech-to-Speech leaderboard at less than half OpenAI's GPT-Live-1 pricing. (source)

Apple: The rebuilt Siri now runs on Google's Gemini models under the hood. Early testers praised multi-step command handling and on-screen context understanding, though hallucinations remain, and the EU market isn't getting it yet. (source)

Models & Infrastructure

Azure SQL Database is rising in coding-agent database choices: A third-party study had real coding agent CLIs — Claude Code, Codex, Cursor — pick their own databases across 356 runs; Azure SQL Database ranked second, behind only Neon. (source)

Technical Progress

Today's Arxiv Digest features three papers that all puncture the same assumption — that a good-looking score means an agent system is trustworthy: 57.5% of conversations LLM judges rated "satisfied" actually failed the task; swapping harnesses without swapping the model shows no stable advantage yet costs more; and a one-shot debugging judge stops searching too early and misses the real root cause.

IBM Research: Published a reproducibility framework on Hugging Face examining whether agents can consistently reproduce the same successful outcome after completing a task — directly useful for teams pushing agents into production. (source)

Security Incidents

PraisonAI hit by two high-severity CVEs: The open-source multi-agent framework was found to have an auth-bypass flaw (CVSS 8.2, where MCP's security policy doesn't consistently check credentials) and a sandbox-escape flaw (CVSS 7.6). (source)

Monitoring blind spot: Separately, a study finds agent compromises stay invisible to routine safety monitoring — see the Deep Dive above for detail. (source)

Regulation & Governance

Industry and the White House split over the Amodei-led slowdown call: A former Anthropic employee warned AI could risk human extinction within a decade, prompting Amodei, Altman, and Hassabis to jointly call for slowing down — but Trump and Vance publicly pushed back against regulation, and Cohere's CEO called the move "a cartel in different packaging." (source)

Regional Roundup

China

Regulators now require AI payment agents to undergo "Know Your Agent" review modeled on KYC, with fund clearing remaining the responsibility of licensed institutions. (source)

Officials also pushed back on the "malicious competition" framing in response to the US industry's recent slowdown calls; analysts note a US-China agreement on AI governance is "near impossible." (source)

Taiwan

Cathay Financial Holdings' technology conference declared a shift from "Cloud First" to "Agent First," unveiling three AI digital coworkers still in testing (project management, tech governance review, legal contract review), backed by identity permissions, security monitoring, and audit-trail mechanisms. (source)

Japan/Korea

South Korea's KISA is drafting AI Security Guide v2.0, targeting new attack surfaces like misuse of agent execution permissions, sensor interference, and verification of autonomous decisions; a CSA survey found 82% of companies have unidentified AI agents inside their organization, and 65% experienced an agent-related security incident in the past year. (source)

Southeast Asia

Singapore's financial sector proposed the non-mandatory SAFR framework as a governance reference for institutions deploying AI agents. (source)

India

Indian firms are consulting lawyers to revisit contract terms as agentic AI shifts from making suggestions to making autonomous decisions, seeking clarity on liability when autonomous systems make mistakes. (source)

Europe

With the EU AI Act's enforcement provisions now in effect, the AI Board convened in Brussels to discuss global governance frameworks for frontier models, making Europe the first major jurisdiction to actually levy large-scale penalties for AI misuse. (source)

Middle East

Saudi Arabia's SDAIA hosted a global AI ethics forum in Riyadh and announced SAMAI 2, targeting professional upskilling in energy, healthcare, industry, and education. (source)

Africa

A report finds African AI startups' bottleneck is the lack of first-cheque seed funding in the $100K–$200K range, meaning many genuinely AI-native companies disappear before they're even counted in funding statistics. (source)

Latin America

In GTIPA's global AI policy report, Argentina's chapter highlights three pillars: a lighter-touch regulatory stance, compute infrastructure investment, and an increasingly active developer ecosystem. (source)

Oceania

Australia unveiled an AI Action Plan that avoids a single dedicated AI law, instead standing up a new AI Safety Institute (backed by A$29.9M) to test frontier models, while requiring companies to build accountability mechanisms before scaling agent deployment. (source)

Business Cases / Funding

A Mistral-anchored $11B funding week: 18 rounds totaling $11B between Sep 7–13, with Mistral's $3.5B Series D+ the single largest. (source)

Exein: The Italian physical-AI security startup raised $270M at a $1.7B valuation, becoming Italy's newest unicorn. (source)

Euclyd: The Dutch inference-chip startup raised over €200M in a Series A co-led by Samsung Electronics, with its first systems not shipping to customers until 2028. (source)

AlphaPai: The Shanghai institutional investment-research workstation closed its third funding round in a year, a $50M Series B bringing cumulative funding to $92M — see today's funding brief.

Jack & Jill: The London-based dual-sided agent recruiting startup raised a $40M Series A led by Air Street Capital — see today's funding brief.

Tools & Ecosystem

alibaba/open-code-review: Alibaba open-sourced a code review CLI combining a "deterministic engine + LLM agent" hybrid architecture, beating pure Claude Code review on Precision/F1 at the same model while using only 1/9 the tokens — today's #1 on GitHub Trending. See today's GitHub Digest.

edgar-mcp: An MCP server connecting to SEC EDGAR that lets agents precisely read a single section of a 10-K instead of the entire 300-page filing. See today's tool recommendation.

Amazon Bedrock AgentCore: Added a managed OAuth consent portal, demonstrating GitHub and Slack authorization-code flows and letting teams inspect end-user authorization activity via CloudTrail. (source)

Key Numbers

ItemNumberSource
Gemini 3.8 Live pricingLess than half OpenAI GPT-Live-1aichatdaily
Weekly AI funding total$11B (18 rounds)StartupHub.ai
Share of Korean firms with unidentified AI agents82%Seoul Economic Daily
"Satisfied" ratings that were actually failed tasks57.5%GAUGE (arXiv)
Exein valuation$1.7BTechCrunch

Today's Digest Roundup

Tomorrow's Watch

  • After Salesforce's Trusted Enterprise AI Harness governance framework ships, will other CRM/SaaS vendors follow with comparable governance-layer products?
  • Once South Korea's AI Security Guide v2.0 is finalized, could it become a template other Asia-Pacific regulators (like Singapore's SAFR) reference?
  • After Gemini 3.8 Live's steep price cut, will OpenAI adjust its voice-mode pricing in response?

Today's Takeaway

I used to think agent security risk mainly came from "being attacked by outside hackers." Today I realized the bigger risk is that companies don't even know how many agents are running inside their own walls — South Korea's 82% shadow-agent figure, and Cathay Financial Holdings deliberately building audit trails before scaling deployment, point to the same conclusion: the next round of IT governance priorities will be "inventory," not "defense."

References