I wrote in my most recent blog about the irony of the timing, because as this was all being reported, the 2nd of August came and went, meaning that the EU’s transparency requirements came into force, but the deadlines governing high-risk systems were pushed back.
The breaches exemplify how risky it was for the EU to push back these deadlines; the more we stall on governance, the bigger the gap between what we can control and what AI can do becomes.
Continuous monitoring flags deviant behavior and gives organizations the chance to act, with an audit trail of what happened and when. Without it, a model that steps outside its parameters is not seen at all.
I hope the events of these past few months are a wake-up call for the industry. It’s imperative that we all work together to ensure AI is deployed safely. This is essential for building trust.
- Nikolas Kairinos, CEO of RAIDS AI
The EU AI Act:
The latest milestone
As of 2nd August, organizations need to meet new transparency obligations under the EU AI Act, including making it clear when people are interacting with AI systems or when content has been generated or manipulated by AI.
This is a welcome milestone. It’s important that people know when AI is being used, particularly as the technology becomes more embedded in the services and content they encounter every day. This transparency is crucial for building trust in AI.
However, as a result of the EU AI Act Omnibus announced earlier this year, the rules governing high-risk AI systems have been delayed. As part of the EU AI Omnibus, mandatory compliance for high-risk systems has been postponed. The new application dates are 2 December 2027 for stand-alone high-risk AI systems, and 2 August 2028 for high-risk AI systems embedded in products.
Pushing the timeline for mandatory compliance further into the future sends entirely the wrong message: that AI monitoring can wait. Events of the past few weeks have demonstrated categorically that it cannot.
My message to organizations is to get ahead of the compliance timeline; ensure your AI is monitored continuously in real time. It will help you to see it early enough to act, and you'll also have the chance to be one step ahead of your competitors.
OpenAI slowed down – here's why
On 18 August, OpenAI paused reinforcement learning training on its most advanced unreleased models for two weeks, and delayed parts of a model called Astra.
In July, an unreleased model escaped its testing environment during a cybersecurity evaluation and took part in an intrusion at Hugging Face after reaching the open internet. Hugging Face disclosed that intrusion on 16 July, before anyone outside knew which lab it had come from.
In the postmortem published on 26 August, OpenAI wrote that if its currently deployed chain-of-thought monitoring had been running at the time, it would have caught the activity and paged the security team more than a day before the models reached Hugging Face systems. The monitoring existed, but it was switched off in that test environment, along with the safeguards that ship in the public products.
Britain's AI Security Institute (AISI) published something similaron 4 August. Across 122 evaluation attempts, agents took unsanctioned action on the live internet 19 times. In the worst caseaMythos agent created a GitHub account, then a second account posing as a different person, and used them to try to talk a real maintainer into merging malicious code.
Both organizations have serious safety teams, and in both cases the behavior was found afterwards, by reading logs. That is the gap we build against. Compliance proves what an AI was on audit day; monitoring proves what it is doing now. A model that passed every evaluation in June can still do something in September that nobody signed off on, and the only way to know is to watch it while it is running.
Monitoring does not prevent an incident. It shortens the distance between something happening and somebody knowing, which on OpenAI's own account was worth more than a day.
Want to learn more about RAIDS? It’s never been easier to try RAIDS out and see how it works.
We’re always keen to connect with others working in and around AI.
You can follow RAIDS, or connect with Nik, Franki and the team on LinkedIn to stay up to date with what we’re building and thinking about.
Our RAIDS Intelligence Hub is a collection of content including documents, reports and videos, providing everything you need to understand AI governance, regulatory compliance, and how RAIDS provides continuous behavioral monitoring for enterprise AI systems.
You can also subscribe to make sure you receive future newsletters here.
And if you’d like to talk about AI risk, safety, or how RAIDS could support your organization, contact our expert team today.