AI agent command approval miss rates make human-in-the-loop a KYA evidence test

The August 7 KYA signal is that human approval cannot be treated as a complete control unless the KYA file also proves what the reviewer saw, which context was available, which policy was enforced, and how risky actions were constrained before approval.

Daily signal: Discord tech-intel channel 1468032405695627386 was readable and surfaced "Humans missed 1 in 3 threats approving AI agent commands across 40k game runs" in the last 24-hour technology digest. Web source verification found Scale X's command-approval experiment, OSL's AgentPay launch, and MetaMask's Agent Wallet launch. These are security and product-market signals, not formal Know Your Agent adoption by a regulator, exchange, bank, broker-dealer, payment scheme, OSL, MetaMask, Consensys, or Scale X.

Why this matters for KYA

Scale X published results from more than 40,000 browser-game runs and 409,000 approve-or-deny decisions where people played the human reviewer for an AI coding agent. The reported headline was uncomfortable: the average player missed one in three threats, with 66.3% mean accuracy. Exfiltration and code-execution prompts had a 33.4% miss rate, scope-violation prompts had a 35.0% miss rate, and familiar script wrappers such as npm run analyze were approved far more often even when the risky payload was visible in the history log.

The exact setting was a game, not a regulated financial workflow. The KYA lesson still transfers because many agentic finance products use the same fallback assumption: the agent may prepare an action, but the human will catch anything dangerous. That assumption weakens when the prompt is noisy, the reviewer is under time pressure, the risk is hidden behind a normal command or protocol call, or the evidence window omits the changed file, policy, wallet, venue, destination, or data class that makes the action unsafe.

The same 24-hour web-search window showed why this matters beyond coding tools. OSL announced OSL AgentPay, a stablecoin payment infrastructure for AI agents that can express amount, asset, and payee while AgentPay handles routing, signing, and settlement across stablecoins and protocols such as x402 and AP2. MetaMask announced Agent Wallet for agentic trading, with spend limits, allowlisted protocols, Guard Mode, Beast Mode, transaction simulation, Blockaid threat scanning, MEV protection, and 2FA review for flagged or out-of-policy transactions.

These products show the right direction: human approval needs to sit inside policy-enforced rails. A KYA reviewer should not accept "a human was in the loop" as a standalone answer. The file needs to show the operator, mandate, wallet scope, tool and venue access, audit trail, abuse controls, and jurisdiction fit that surrounded the approval prompt.

Screenshot-ready KYA compliance comparison table

KYA dimensionApproval-only postureKYA-ready approval postureEvidence reviewers should expect
Operator identityThe reviewer sees a command, trade, transfer, or API call, but the accountable operator, agent instance, model, wallet owner, and business approver are not bound in the same record.The operator, end user, business owner, agent instance, model or workflow version, wallet owner, venue account, and reviewer are linked before any high-risk approval is requested.Operator ID, user ID, agent ID, model or workflow version, wallet owner, venue account, reviewer ID, approval timestamp, escalation owner.
Agent mandateThe mandate is broad, so a normal-looking command or route can hide credential access, data exfiltration, overbroad trading, or a payment outside purpose.The mandate defines allowed actions, prohibited data classes, maximum amount, allowed protocols, destinations, venues, order types, expiry, and approval thresholds.Mandate text, policy version, allowed action list, prohibited data list, amount and frequency limits, venue scope, expiry, exception ticket.
Wallet and custodyThe human is asked to approve a transaction without enough context on spend cap, custody model, signer, gas route, key exposure, loss coverage, or downstream settlement.Wallet authority is constrained before review, with spend limits, signer rules, custody mode, transaction simulation, payee validation, settlement path, and revocation controls.Wallet ID, custody mode, spend cap, signer policy, transaction simulation, payee validation, stablecoin and chain, settlement reference, revocation log.
Tool and venue accessThe approval prompt treats npm run, MCP tool calls, API routes, swaps, perpetuals, prediction markets, x402 payment requests, and exchange access as isolated yes-or-no prompts.Each tool, MCP server, protocol, market-data endpoint, payment rail, trading venue, order route, and external destination has a pre-use verdict and runtime policy decision.Tool inventory, MCP server verdict, protocol allow list, venue approval, API scope, order type, route, payee, allow or deny reason, policy hash.
Audit trailThe log says a human approved, but does not preserve the prompt, hidden context, modified files, transaction simulation, policy result, flagged risk, and final outcome together.The audit trail links prompt, agent plan, changed context, resource accessed, risk score, simulation, policy decision, approval or rejection, execution result, and reviewer note under one trace ID.Trace ID, prompt hash, context snapshot, file or resource hash, risk score, simulation output, policy decision, approval artifact, execution receipt, reviewer note.
Security and abuseHuman review is the last line of defense, so permission fatigue, familiar command names, time pressure, malicious packages, prompt injection, and out-of-policy routes can slip through.Human review is backed by sandboxing, context isolation, command classification, threat scanning, MEV protection, anomaly detection, rate limits, 2FA, and kill switches.Threat classification, sandbox state, context-isolation verdict, transaction scan, anomaly alert, 2FA challenge, rate-limit event, kill-switch drill, incident replay.
Jurisdiction fitThe reviewer approves an action without seeing whether the agent is crossing privacy, outsourcing, licensing, market-conduct, payments, custody, or complaint-handling boundaries.The KYA decision maps the user's country, data location, venue rules, payment or trading permissions, outsourcing exposure, retention duties, breach process, and dispute route.Jurisdiction matrix, customer-country flag, data-residency label, licensing check, venue rule, retention rule, breach path, dispute or chargeback route.

The compliance lesson

Human approval is still useful, but it is not a magic shield. The Scale X results show how quickly reviewers can miss risky actions when the danger is contextual, familiar-looking, or wrapped inside a common workflow. Agent wallets and payment rails make that weakness more expensive because the risky action can be a stablecoin settlement, protocol trade, market-data purchase, or exchange route rather than a local command.

KYA should therefore grade human-in-the-loop controls by evidence quality. The question is not simply whether a user clicked approve. It is whether the approval prompt included the right context, whether the policy engine blocked out-of-mandate actions first, whether the action was simulated or scanned, whether flagged events required stronger authentication, and whether the final record can be replayed by compliance, security, and dispute teams.

Practical KYA checklist

Bottom line

The new KYA control point is approval evidence. Before an AI agent is allowed to move money, trade, call a paid API, invoke an MCP tool, or send data outside a controlled environment, reviewers need proof that human approval is reinforced by mandate limits, wallet controls, tool verdicts, audit trails, abuse detection, and jurisdiction checks.

Sources reviewed: Discord tech-intel channel 1468032405695627386 for the last 24-hour technology digest; Scale X; OSL Group / GlobeNewswire; MetaMask; Cointribune. These are security and product-market signals, not formal Know Your Agent adoption by a regulator, exchange, bank, broker-dealer, payment scheme, OSL, MetaMask, Consensys, or Scale X.