ilert seamlessly connects with your tools using our pre-built integrations or via email. ilert integrates with monitoring, ticketing, chat, and collaboration tools.
See how industry leaders achieve 99.9% uptime with ilert
Organizations worldwide trust ilert to streamline incident management, enhance reliability, and minimize downtime. Read what our customers have to say about their experience with our platform.
On July 27, 2026, the Digital Omnibus on AI, Regulation (EU) 2026/1744, entered into force and rewrote the EU AI Act's timeline. The high-risk obligations that were due to apply on August 2, 2026, including Article 73 serious-incident reporting, now apply from December 2, 2027 for standalone high-risk systems (Annex III) and August 2, 2028 for AI embedded in regulated products (Annex I).
Not everything moved. Since August 2, 2026, Article 50 transparency obligations apply and the Commission can fine providers of general-purpose AI models, and the GPAI obligations in force since August 2025, including serious-incident reporting for systemic-risk models, stay in force.
Deferred, not cancelled. The extra 16 months are the window to build serious-incident reporting into your operations, not a reason to shelve it. Here is what changed, what didn't, and what your incident process needs to be able to do.
What the Digital Omnibus changed
The European Commission proposed the Digital Omnibus on AI on November 19, 2025, after it became clear the compliance ecosystem would not be ready for the original August 2026 deadline: harmonised standards from CEN-CENELEC were running late, conformity-assessment bodies weren't designated, and many member states had not stood up their market surveillance authorities. Parliament and Council reached political agreement on May 7, 2026, formally adopted the text in June, and published it in the Official Journal on July 24, 2026. It entered into force three days later, six days before the original deadline.
One detail matters for planning: the new dates are fixed, not conditional. The Commission's original proposal tied the delay to the readiness of standards and guidance, with 2027/2028 as backstops. The co-legislators replaced that with hard dates. Dr. Nils Rauer of Pinsent Masons called the agreement "a pragmatic step towards greater legal certainty", there is no scenario where the obligations land earlier, and no realistic one where they slip again.
The EU AI Act timeline, as it stands now
Feb 2, 2025 (in effect): Prohibited AI practices (Article 5); AI literacy duty (Article 4)
Aug 2, 2025 (in effect): Obligations for general-purpose AI (GPAI) model providers, incl. systemic-risk duties (Chapter V)
Aug 2, 2026 (in effect): Article 50 transparency obligations; Commission fining powers over GPAI providers (Article 101)
Dec 2, 2027: High-risk obligations for standalone Annex III systems, incl. Article 12 logging and Article 73 serious-incident reporting
Aug 2, 2028: High-risk obligations for AI embedded in regulated products (Annex I)
What applies in 2026 regardless of the delay
If you build or deploy AI systems in the EU, two things are live now:
Transparency (Article 50), since August 2, 2026. Users must be told when they are interacting with an AI system, and AI-generated or manipulated content must be labeled and machine-detectable. Not an incident duty, but the first AI Act obligation most engineering teams will actually be audited on.
GPAI enforcement, since August 2, 2026. The transparency, documentation, and copyright obligations for GPAI providers have applied since August 2025, what is new is the Commission's power to fine non-compliance, at up to 3% of global annual turnover or €15 million, whichever is higher. If you provide a GPAI model with systemic risk, serious-incident tracking and reporting is already expected under Article 55(1)(c), with the Code of Practice setting deadlines that mirror Article 73, plus a 5-day window for serious cybersecurity breaches.
Article 73: what serious-incident reporting actually requires
Article 73 is the AI Act's core incident-response obligation, and it is worth getting the numbers right, it is frequently misreported as a GDPR-style 72-hour rule. The actual windows for providers of high-risk AI systems:
Immediately after establishing a causal link between the AI system and the incident, or the reasonable likelihood of one, and no later than 15 days after becoming aware of a serious incident.
No later than 2 days for a widespread infringement or a serious and irreversible disruption of critical infrastructure.
No later than 10 days in the event of a person's death.
A "serious incident" means an incident or malfunction that directly or indirectly leads to death or serious harm to health, a serious and irreversible disruption of critical infrastructure, an infringement of fundamental-rights obligations, or serious harm to property or the environment.
Three operational details engineers should note:
An incomplete initial report is explicitly allowed, followed by a complete one. The regulation expects you to notify while the investigation is still running, which means your reporting workflow can't depend on the postmortem being finished.
You must investigate, assess risk, and take corrective action, and you may not alter the AI system in a way that could affect the evaluation of the incident's causes before informing the authorities. Your evidence trail has to survive the response.
Deployers are in the loop too. Under Article 26(5), a deployer that identifies a serious incident must immediately inform the provider (then the importer/distributor and market surveillance authorities). Provider-deployer communication paths are part of compliance, not a nice-to-have.
The Commission has published draft guidance and a standard reporting template for Article 73; final versions are still pending. And the Act itself already de-duplicates: where a high-risk system falls under a sector regime with equivalent reporting duties, critical infrastructure under NIS2, medical devices under the MDR, Article 73(9)–(10) limit AI Act reporting to fundamental-rights incidents, with everything else routed through the sector rules. Translation: one well-built incident process serves multiple regulators.
Why build now, not in November 2027
The case for using the deferral rather than enjoying it:
The dates are now certain. Fixed deadlines removed the ambiguity that justified waiting.
The duties are process-heavy, not paperwork-heavy. Tamper-evident logging, cross-functional notification, evidence preservation, and reporting-while-investigating are operational muscles. Teams that add them late in 2027 will be scrambling to catch up; teams that run them now get a year of practice.
The same process already pays off elsewhere. NIS2's 24-hour early warning and 72-hour notification, GDPR's 72-hour breach window, and DORA's incident reporting all run on their own clocks, none of them was delayed. A single incident workflow with automatic timelines, role-based escalation, and audit-ready postmortems satisfies several regimes with one evidence set.
Article 73 readiness checklist
What your incident process has to be able to do before December 2027 and what pays off now under NIS2, DORA, and GDPR:
Classify at intake. Triage asks whether an incident could meet the Article 3(49) definition of "serious", so the AI Act clock is recognised on day one, not in the postmortem.
Timestamp awareness. The 2/10/15-day windows run from the moment you become aware. Capture it automatically, not from memory.
Preserve evidence before you fix. No changes to the AI system that could affect causal analysis until authorities are informed (Article 73(6)). Logs, model versions, and inputs must survive the response.
Report while investigating. An initial, incomplete report is allowed. The workflow can't wait for the root cause.
Wire the deployer → provider path. Deployers must inform providers immediately (Article 26(5)). Agree the channel and template in advance.
Notify legal and comms in parallel with engineering. The person who files the report is rarely the person fixing the system.
Map the regime per system. Know which clock applies: AI Act, NIS2, DORA, GDPR or whether a sector rule absorbs the AI Act duty under Article 73(9)–(10).
Export the timeline. Every notification is a document. Produce it from the record, not from recollection.
Article 73 readiness checklist
How ilert maps to the AI Act's incident duties
ilert doesn't make you AI Act compliant, no tool does. What it does is make the operational duties routine instead of heroic:
Audit-ready timelines (Article 12, Article 73(6)). Every alert, escalation, acknowledgment, and action is captured automatically in the incident timeline and can be exported, evidence that survives the incident without anyone writing it down mid-outage.
Reporting-grade speed (the 2/10/15-day windows).Multi-channel alerting, on-call schedules, and escalation policies mean awareness of an incident is measured in minutes. The hard part of Article 73 isn't the notification deadline, it's having the causal analysis and documentation ready. Automated timelines and AI-assisted postmortems produce that documentation as a by-product of response.
Cross-functional loops (Article 26(5), Article 73(1)). Playbooks notify engineering, legal, and comms in parallel, and status pages give stakeholders verified updates without pulling responders into meetings.
Monitoring-to-incident (Article 55(1)(c), Article 72). Threshold breaches from your monitoring stack, output drift, inference-error spikes, latency, open an incident automatically through ilert's alert sources, so post-market monitoring and incident tracking run on the same record.
EU-grade data handling. ilert is built and hosted in Germany, ISO 27001 certified, and GDPR-compliant, with AI inference kept in-region, so the tooling you use to respond to AI incidents doesn't create a new data-transfer problem.
Governance over autonomy. The AI Act's through-line is human oversight of AI. ilert AI SRE is built on the same principle: the agent investigates, correlates, and drafts, a responder reviews and decides what happens next. AI runs the investigation. You make the call.
See the whole flow, alerting, timeline, postmortem, in ilert's weekly live demo.
FAQ
Did the EU AI Act get delayed?
Partly. The Digital Omnibus on AI (Regulation (EU) 2026/1744, in force July 27, 2026) moved high-risk obligations to December 2, 2027 (Annex III) and August 2, 2028 (Annex I). GPAI obligations and Article 50 transparency kept their original dates.
What applies from August 2, 2026?
Article 50 transparency obligations, and the Commission's power to fine GPAI providers. GPAI obligations themselves, including serious-incident reporting for systemic-risk models, were already in force.
When does Article 73 serious-incident reporting apply?
It attaches to the high-risk regime: December 2, 2027 for standalone Annex III systems, August 2, 2028 for Annex I embedded systems. GPAI models with systemic risk already have a parallel reporting duty under Article 55(1)(c).
How fast must serious incidents be reported under the AI Act?
Immediately once a causal link (or its reasonable likelihood) is established, and no later than 15 days after awareness, 2 days for widespread infringements or critical-infrastructure disruption, 10 days for a death. Not 72 hours; that's GDPR.
Does the delay affect NIS2, DORA, or GDPR?
No. Those regimes and their reporting clocks are unchanged. (The wider Digital Omnibus package proposes amendments to GDPR and NIS2, but as of publication those are still under negotiation and are not law.) The AI Act itself routes most incidents involving sector-regulated systems through the sector rules (Article 73(9)–(10)), making one unified incident process the sensible hedge.
This article is practical incident-process guidance, not legal advice.
Bleemeo monitoring now connects natively to ilert, linking threshold detection to on-call management and alerting. DevOps, SRE, and IT operations teams get a direct path from a breached threshold to the phone of the engineer who can fix it, and back to a clean slate once the problem is gone.
What is Bleemeo?
Bleemeo is a cloud monitoring platform that covers servers, Docker containers, Kubernetes clusters, AWS and Azure resources, VMware environments, and SNMP network devices, all from one SaaS console. It also handles uptime checks, application performance, and logs, so teams can watch their infrastructure and the services running on it in one place.
Data collection runs through Glouton, Bleemeo's open-source agent. Alerting rules are threshold-based with separate warning and critical levels, and teams that want finer control can express conditions in PromQL.
Why connect Bleemeo to ilert?
Bleemeo is good at detecting that something is wrong. What happens next, who gets woken up, and what follows if they don't answer, is where ilert takes over.
Paging is where ilert starts, not where it ends. The same alert can become a declared incident with its own channel and incident commander, a status page update for customers, and a postmortem once it's closed.
Two details matter day to day. First, resolution is automatic: when the issue clears in Bleemeo, ilert closes the corresponding alert, so what's open reflects the real state of your systems instead of a backlog someone has to tidy up. Second, grouping happens per monitored metric, or per agent for agent-level events. Repeated notifications about the same problem update the existing entry rather than piling up new ones, so a struggling host wakes you once, not fifty times.
Two European platforms, one alerting chain
Bleemeo is a French company, founded in Toulouse in 2015, with metrics and logs stored in the EU. ilert is a German company in Cologne, hosted in EU data centres and ISO 27001 certified. For teams working under GDPR, NIS2, or DORA scrutiny, that means a monitoring-to-paging chain where both counterparties contract under European law, one fewer subprocessor conversation in the security review.
Set up the integration in three steps
1. In ilert:
Go to Alert sources and create a new one
Search for Bleemeo and select it
Give it a name, assign teams, choose an escalation policy, and set your grouping preference
Finish setup and copy the integration key
2. In Bleemeo:
Open Administration → Integrations and click Add Integration
Select ilert, paste the key, and add a descriptive name such as "Production On-Call"
3. Route your alerts:
Go to Notifications, create a new rule or edit an existing one
In the Targets step, select your ilert integration
Because ilert ships a purpose-built source for Bleemeo, there is no field mapping to configure. Alerts flow inbound, from Bleemeo into ilert, and setup takes a few minutes end to end.
When you get paged at 3am, it takes about 30 seconds for the notification to reach you and maybe two minutes until you're in front of a laptop, awake enough to read. What you see then is usually a raw alert. A metric name, a threshold, a link to a dashboard. Then the ritual starts: open the dashboard, check what deployed in the last few hours, grep the logs for the first error, ask in Slack whether anyone touched the database.
Most of that time is search. Once you've found the needle in the haystack, the fix is usually fast, and most of the time it's a rollback. It's the search that takes 20, 45, sometimes 60 minutes of your night.
The goal for ilert AI SRE was simple to state: by the time you open the laptop, the search should be done. We opened the closed beta about a year ago, back then under the name ilert Responder, with a few teams. Since then we've rebuilt the agent three times, and each rebuild taught us something the previous version didn't know: where an agent actually helps, what a good harness looks like, and what "production-ready" means when the agent is on call next to you. Today it's generally available on every paid plan and during trials.
What it does
Give AI SRE an alert from any of our 150+ integrations and it investigates the way you would, just faster. It looks at logs, metrics and traces. It looks at what changed, because change is the number one cause of incidents: deployments, config changes, merged pull requests, and, if you've connected your repository, the actual diff. If 50 alerts fire at once, it first triages them into clusters and investigates each cluster on its own.
A few minutes later, you're not looking at a blank alert. You're looking at a root cause analysis.
Every investigation has the same shape. A hypothesis for the root cause, with a confidence label: high, medium or low. Five or six key findings, and every finding links to its evidence: the deployment it's referring to, the log line that produced the out-of-memory error, the metric that broke first. A short list of things it ruled out. And, where it makes sense, a proposed action: roll back to the last healthy version, double the memory on that pod, restart the service.
Then it waits for you. You make the call. Nothing changes the state of your system without an engineer clicking approve.
At GA, you start the investigation. There are three ways in: from an alert, from an incident, or by describing what you're seeing in the agent's chat. Plain language is enough; "customers are reporting that metrics on their status pages are stale" is a complete brief, and it's how the incident in the next section was handed to the agent. Investigations that start on their own the moment an alert fires, so the analysis is already waiting when you pick up the phone, are the next step, not this one.
A real one
A while ago we had an external penetration test. It found a blind SSRF in the metrics feature of our status pages: customers can point a status page at their Datadog or Prometheus to display API response times, and an attacker could try to make that fetcher call internal URLs instead. We fixed it the same week with a Kubernetes network policy. The network policy was too broad. The metrics service could no longer talk to its own database, and metrics on customer status pages went stale.
I like this incident as a test case for two reasons. First, no runbook on earth covers "after a pentest fix, status page metrics go stale." Novel incidents don't have runbooks. Second, the symptom is ambiguous: tell an agent "customers say metrics stopped working" and there are a dozen internal things called metrics it could chase, including our entire Prometheus setup. The agent had to find candidate pods, find the logs showing a service that couldn't reach its database, then walk the recent changes until it landed on the network policy. That's the kind of search that eats an hour of a human's night, and it's exactly the kind of search the agent is good at.
Why there is no autonomous mode at GA
We think about autonomy in levels. Observe only: the agent has read-only access and produces an analysis. Propose: the agent suggests an action and a human approves it. Pre-approved actions: a class of low-risk actions the agent may run on its own when its confidence is high. Fully autonomous: you only get paged when the agent is stuck.
GA ships the first two. I want to be direct about why, because the usefulness of an agent grows with its autonomy. A self-driving car that needs you watching the road is helpful; one that doesn't is a different product. I believe we'll get there, and I think production agents will reach the maturity coding agents reached over the last year. But I don't have a way to do full autonomy safely today, so I wouldn't run it today, and I don't think you should either. We run demos where the agent does everything end to end, from investigation to status page update to fix, and we always add: don't do this at home.
Observe-and-propose is not a consolation prize. Read-only access means that even if the agent goes badly wrong, the blast radius is an incorrect document. And most of the time you spend on an incident is the root cause analysis. Cutting that from 45 minutes to a few is the bulk of the value, before the agent ever touches anything.
What it reads, and what it deliberately doesn't
The best documentation of how a system actually works is the code and the live telemetry. Everything else goes stale. So AI SRE doesn't start from your Confluence, your Notion or your runbooks. Pointing an agent at a wiki with hundreds of runbooks, the last of which was updated three years ago, causes more damage than it produces results.
What it does read is everything you already send to ilert, plus what you connect: logs from Elastic or CloudWatch, dashboards and metrics from Grafana, Prometheus or InfluxDB, your Kubernetes cluster, your GitHub repositories and CI/CD pipeline. That breadth is the point. Every observability vendor is building an AI SRE right now, and each one sees its own slice. We sit on top of all of them, we stay vendor-neutral, and if your observability tool produces its own RCA, we take it as one input among many. Nobody wants another siloed tool that only does RCA for one data source.
Beyond what you connect, the agent runs a discovery phase when you set it up, and again whenever it's been idle for a while. It builds a service topology from live tracing data, so it understands your dependencies without anyone maintaining a service catalog by hand. That topology is also what allows our agent to do cost-effective alert triage before running an investigation with powerful (but more expensive) reasoning models. If you don't have tracing instrumented, and most companies don't, you can drop an eBPF collector into your cluster and get most of the way there without touching code.
Over time the agent keeps a lightweight long-term memory of the tribal knowledge it picks up, the "this service always hits its limits at midnight because of the batch jobs" kind of thing. It's plain text. There's no vector database to keep in sync.
How we test it
There is no recipe for root cause analysis, which makes it hard to test. We do three things.
First, a public benchmark. OpenRCA, from Microsoft Research, is the hardest public test of root cause analysis I know of: 335 real failure cases from three production systems, 68 GB of telemetry, scored all-or-nothing per task. What counts as a full point depends on the task: some require the right component, reason, and time; others are scored on a smaller subset, sometimes just the component. At paper release, the best published baseline, an agent on Claude 3.5 Sonnet, solved about 11% of the cases. By the time we ran our own benchmark, later leaderboard results had moved on, Opus 4.6 sat at roughly 36%, so treat that 11% as the paper's number, not today's frontier.
On a 50-task sample across all three systems, our first prototype solved 26% under that strict rule, and matched the ground truth in 30% of cases by the broader measure, at roughly nine minutes per investigation. I'd treat the number as a floor, and the sample size as the caveat it is, but two things we learned matter more than the number. Handing the agent a plain shell on the raw telemetry nearly doubled the strict pass rate compared to handing it our eight production-shaped Grafana and Prometheus tools, 26% against 14%, with fewer tool calls. Tool design is not a detail. And OpenRCA is telemetry only: no deployment history, no config changes, no alerts. In production, changes are the agent's strongest signal, so this benchmark measures it with one hand tied behind its back. In the same run, Claude Code carrying our investigator prompt scored 38%, ahead of our own harness. We publish that because it's the test we hold ourselves to internally: if a general coding agent with your prompt beats your harness, your harness has work to do.
Second, chaos. We inject real faults into a real environment and hand the agent only what an on-call engineer would see: the symptoms. It's never told what we broke. Across a dozen scenarios covering bad deploys, config regressions, resource exhaustion, dependency failures:
Correct root cause identified in all runs
Median time from alert to first finding: 194 seconds
Proposed remediation matched what our engineers would have done in 90% of runs
Wrong hypothesis presented with high confidence: Zero. This is the number we watch most closely.
Third, replay. Every production investigation is recorded: every tool call, every response, the final RCA. When we change a model or a prompt, we replay those recordings and compare the outcome, with an LLM as judge and with embedding similarity. This catches regressions when a new model comes out. It doesn't catch everything: a recording only covers the tools the original run happened to use.
The agent is sometimes wrong. Not hallucinating-wrong; in a highly contextualized environment we see very little of that. Plain wrong: a confident line from the symptoms to the wrong cause. The confidence label and the evidence links exist so you can see that in under a minute, and either redirect it with a follow-up question or do the analysis yourself.
What it doesn't do
It won't act on its own. Every action requires approval.
It only sees what you connect. Most of the value of an agent is in the context it gets. If it can't see that something deployed, it can't blame the deploy.
It won't fix not having observability. Some people hope the agent lets them skip that step. It doesn't, at least not yet.
It's not great at backing out of a wrong path on its own. When the search space is large and the starting point is vague, it can get stuck. That's when you ask for a follow-up or take over.
What this changes for the team
The conversations we're having with engineering leaders about this are different from the ones we used to have about MTTR. A pattern I keep seeing: a small, strong SRE team runs first-line on-call across many product teams, with the agent doing the triage and the first investigation. Fair rotations normally force every service team to staff around five people just to cover on-call. Centralizing first-line on an AI-augmented team removes that constraint per team. Engineers spend more of their time shipping and less of it in 25-person incident bridges. Same headcount, more services, lower cost per incident.
That's the version of "AI in operations" I find worth building: not fewer engineers, but engineers who were hired to build things getting to build things. Nobody was hired to be on call full-time.
Your data
All AI workloads run on dedicated infrastructure in our EU regions, Frankfurt and Stockholm, and models are called through regional endpoints. We don't train on your data and have opted out of training with every model provider we use. Read-only API keys are all the agent needs to investigate. Personal and user-level data isn't shared with external models. If your security team wants to bring your own model API keys behind your own guardrails, we support that. Details: https://docs.ilert.com/trust-center
Availability
AI SRE is on for every paid plan and for trials starting today, running on the AI credits included in your plan: 250 per month on Pro, 1,000 on Scale, 5,000 on Enterprise. You can see your baseline and decide what happens when you run out.
Existing customers: nothing to migrate. Enable it under Account settings → AI features, create an agent and connect it with your tools, and launch your first investigation from the next incident.
Where this goes
The end state we're building toward is that you're not paged at 3am at all. You wake up to a report: the agent stopped the bleeding, verified for an hour that the symptoms were gone, and there's a pull request for the real fix. That's not GA. GA is the search taking a few minutes instead of an hour, started with one click, with the fix one click behind it. The trust for the rest gets earned one investigation at a time, and this is where it starts.
If you have a scenario you think it will get wrong, I want to hear about it.