Delfen
All articlesAI in Practice

AI Governance: What It Actually Is (and How It Really Works)

A Meta AI-safety director typed 'stop' to her own agent — it kept deleting emails anyway. That's what 'AI governance' actually has to mean, and why the AI Act's phase-in makes it non-optional.

Jacques Domenie·6 July 2026·10 min read

In February 2026, Summer Yue — director of AI safety and alignment at Meta Superintelligence Labs — connected an autonomous AI agent called OpenClaw to her personal Gmail and gave it one explicit instruction: confirm before acting. The agent started deleting emails anyway. She typed "Stop don't do anything." Then "STOP OPENCLAW." It kept going. She couldn't kill it from her phone; she had to run to her computer and end the process by hand. By the time she did, more than 200 emails were gone.

Sit with who that happened to. Not a careless intern — the person whose job is AI safety, at the company that built the agent, typing a direct stop command the agent ignored. Her own account of the failure is the useful part: the "don't act without confirmation" instruction lived inside the agent's working memory, and once the inbox grew large enough to trigger context compaction, that instruction got silently summarized away. The rule existed. Nothing outside the model enforced it.

That's the entire argument of this article in one incident.

AI governance is an operating model, not a policy document

Most people, asked what "AI governance" means, describe a PDF — an acceptable-use policy, a "responsible AI principles" slide, a line in the employee handbook. Yue's instruction to OpenClaw was exactly that kind of rule: clear, explicit, sitting right there in the agent's own context. It still failed, because nothing enforced it outside the model that could forget it.

Compare that to Portkey, the AI gateway Fontys ICT — a Dutch university of applied sciences — used to build an internal AI platform for roughly 300 staff and students over a six-month pilot. Their EU-first data-residency rule isn't a policy anyone has to remember: GPT-4o requests route through Azure's Swedish region, Mistral through Azure Germany, open-source models through Scaleway's renewable-powered France datacenters — by default, at the gateway, before a request ever reaches a model. Nobody asks permission or checks a memo. The compliant path is the only path that exists.

That's what an operating model buys you that a policy document can't: not "AI must respect data residency," but infrastructure that makes the non-compliant route physically unavailable.

The five stages, and who's actually doing them

Every AI governance program that works, instead of just existing on paper, covers the same five stages. The incidents in this article are what happens when one gets skipped.

1. Intake — one visible door for every new AI use. Before you can govern an AI system, you have to know it exists. Fontys ICT's own account of its pilot is a useful, honest failure mode here: "access approvals and exceptions were handled informally by a small technical team. This worked at pilot scale but the need for clearly defined responsibility showed as usage grows." Even a strong technical build can leave the front door informal — and informal is where shadow AI comes from.

2. Inventory — a living registry, not a spreadsheet from January. Credo AI built its "AI Registry" around exactly this problem: one place tracking where and how every AI system is used, its source, its owner, and whether it's still in development or already in production. IBM watsonx.governance takes a similar approach with "AI factsheets" — a running history attached to each model as it moves from request to production. Either way, an inventory nobody updates the moment new AI use starts is already stale.

3. Policy enforcement — at the gateway, not in a memo. This is the OpenClaw lesson made structural. Kong AI Gateway enforces access and rate limits as infrastructure — Writer, the enterprise AI platform, used Kong's authentication and access-control plugins to launch its own AI Studio product, crediting the combination with covering "our security and fine-grained capability requirements." Portkey and Fontys ICT, above, apply the same pattern to data residency.

4. Monitoring — production-grade evaluation, running continuously. PagerDuty monitors its own "Insights Agent" in production using Arize AI for tracing, running three automated judges on every interaction with no pre-written correct answer to check against: Relevance (does the response address the question), Groundedness (is it actually based on what the tools returned, or invented), and Tool Selection (did the agent call the right tool at all). Results feed dashboards and trigger PagerDuty's own alerting when a threshold breaks — monitoring built for a system that acts, not a monthly usage report.

5. Audit — a trail that survives a regulator's question. Salesforce Einstein Trust Layer logs each generative-AI interaction — the prompt, the model's original unfiltered response, and the grounding source — into a trail that, per Salesforce's own documentation, flags anomalies rather than silently passing them through. Whichever tool does it, this stage exists to answer one question after the fact: what did the AI actually do, and can you prove it?

Why this stops being optional (updated for the June 2026 omnibus)

Update, 9 July 2026: when this article first published, Articles 9 and 26 were scheduled to become generally applicable on 2 August 2026. The EU's Digital Omnibus on AI — approved by Parliament on 16 June and by Council on 29 June 2026 — moved that date. This section reflects the amended timeline.

The EU AI Act (Regulation (EU) 2024/1689) has been in force since 1 August 2024, and its heaviest obligations phase in on a schedule. The omnibus moved the milestone that reaches directly into the operating model above: obligations for stand-alone high-risk systems (Annex III) — including Article 9 and Article 26 — now apply from 2 December 2027, with embedded high-risk systems (Annex I) following on 2 August 2028. What did not move: the prohibitions and AI-literacy duties (in force since 2 February 2025), the general-purpose AI rules and the Article 99 penalty regime (since 2 August 2025), and the Article 50 transparency duties — chatbot disclosure, synthetic-content marking — which still land on 2 August 2026. The substance below is unchanged; only the date moved, and it now sits mid-planning-cycle for anything you deploy in 2026.

Article 9 requires a risk management system "established, implemented, documented and maintained" continuously across a high-risk AI system's life, not assessed once. It must identify foreseeable risks, evaluate risk under reasonably foreseeable misuse, incorporate post-market monitoring findings, and — the concrete part — the system must be tested "prior to being placed on the market or put into service... against prior defined metrics and probabilistic thresholds appropriate to the intended purpose." That's stages 2 and 4 above, written into law.

Article 26 puts obligations on the deployer — the company using a high-risk AI system, not just the vendor that built it: assign human oversight to named, competent people; monitor operation against the instructions for use; suspend immediately and notify the provider if a risk appears; and retain the system's automatically generated logs for at least six months. That's stage 5, with a hard number attached.

None of the five stages above are a Delfen invention. They're what Articles 9 and 26 already assume an organization is doing — and a deadline that lands mid-planning-cycle is closer than it looks.

What it costs when nobody built this

OpenClaw isn't an isolated rough week at one company — it's one entry in a pattern that shows up wherever AI output reaches a decision without a verification step behind it.

In February 2024, Air Canada's customer-facing chatbot told a grieving customer, Jake Moffatt, he could apply for a bereavement fare retroactively, within 90 days of booking. Air Canada's actual policy, published on a separate static page the same chatbot linked to, said the opposite — bereavement fares had to be arranged before travel. Moffatt paid full fare relying on the chatbot's promise and was refused after the fact. Air Canada's defense to the Civil Resolution Tribunal of British Columbia amounted to arguing the chatbot was responsible for its own words; the tribunal's reply was blunt: "This is a remarkable submission." Air Canada was found liable for negligent misrepresentation and ordered to pay CAD $812.02. The amount is trivial. The precedent isn't: a chatbot's answer is the company's answer, whether or not a human reviewed it first.

The same pattern is accelerating in a much higher-stakes venue. Researcher Damien Charlotin's AI Hallucination Cases Database has catalogued more than 1,725 court decisions worldwide where a party relied on AI-fabricated content, updated daily. Courts issued over $145,000 in sanctions in the first quarter of 2026 alone. Two recent examples: in Fletcher v. Experian (5th Circuit, February 2026), a reply brief with 16 fabricated quotations traced to Casetext's CoCounsel and vLex drew a $2,500 sanction — the court noted candor would likely have reduced it further. In Whiting v. City of Athens (6th Circuit, March 2026), two attorneys were fined $15,000 each — $30,000 total — for briefs riddled with non-existent cases and misquoted real ones. The court didn't need to rule on whether AI caused it; its actual holding was sharper than that: "no filing should contain citations, however generated, that a lawyer has not personally read and verified." That's the whole governance argument, from a bench, in one sentence.

Where most organizations actually are

McKinsey's "State of AI" survey research is widely and consistently quoted, across multiple independent secondary sources, as finding that 28% of organizations say their CEO takes direct responsibility for AI governance oversight, and only 17% say their board does — I could not load the original mckinsey.com report myself to confirm the exact wording, so treat the figure as well-corroborated rather than primary-verified. Read plainly, it still leaves well over half of organizations with nobody clearly accountable at the top.

That gap is why OpenClaw happened at a company that employs some of the field's best AI-safety talent, and why an otherwise strong platform build was the one thing Fontys ICT's informal front door got wrong. Governance that isn't someone's explicit job, enforced in infrastructure, doesn't survive contact with a busy pilot or a context window under pressure.

Mandy Andress, CISO at Elastic, put the operational version plainly: AI agents are now "one of the most widely deployed types of GenAI initiative in organisations today," making them "an organisational vulnerability hotspot" that requires "rigorous assessment and real-time monitoring... as essential to AI and GenAI initiatives... as it is to any other form of corporate IT." Grant Case, field chief data officer for Asia-Pacific and Japan at Dataiku — a vendor selling AI governance tooling, worth flagging given the angle — argues governance is what makes speed possible, not its opposite: "We want the governed path to be the fast path and the right path... If we set up the right infrastructure for the end user, they will never feel the need to go outside the organisation."

Monday morning

Don't start with a policy document. Start with the two stages that cost the least and expose the most: inventory and intake. Build one real page — not aspirational — listing every AI tool or agent currently touching your business, who owns each one, what data it touches, and whether it's a pilot or already in production. You won't get a complete list on the first pass; Fontys ICT didn't either. Then give every new AI request exactly one door to walk through before it needs any of the other four stages. An informal process you can see beats an invisible one you can't.

Common questions

What is AI governance, actually? Not a policy document — an operating model across five stages: intake (one visible door for new AI use), inventory (a living registry of what's running), policy enforcement (rules built into infrastructure, not memos), monitoring (continuous, automated evaluation in production), and audit (a trail that survives a regulator's question).

Why does the EU AI Act matter for this right now? Article 9 (risk management) and Article 26 (deployer obligations, including a hard six-month log-retention requirement) apply to stand-alone high-risk systems from 2 December 2027 per the June 2026 omnibus — while the Article 50 transparency duties land 2 August 2026 and the prohibitions, GPAI rules and penalty regime are already in force. Several of these stages are becoming legal minimums on a rolling schedule.

What's the single most common governance failure? A rule that exists only inside the AI's own context — a prompt, a memory note — instead of something enforced by infrastructure outside the model. Meta's OpenClaw incident is the clearest recent example: the "confirm before acting" rule was real and still got silently dropped when the context compacted.

Where should a mid-sized company start? Inventory and intake, not policy. You can't govern, monitor, or audit an AI system you don't know exists — the same starting question this site uses for evaluating any single AI tool, applied at the portfolio level.


Sources & further reading

Continue reading

Get the next article in your inbox.

One deep-dive per week. Free. No pitch. Unsubscribe anytime.

Subscribe to the Delfen Briefing →