Go-live is treated as the end of the AI story. The dashboard is on. A few people in the original team know how to refresh it. Then the output is wrong on a Tuesday, or the feed is stale, or operations cannot tell whether to trust the recommendation. Someone has to be paged. Someone has to decide whether to roll back. Someone has to tell the business what to do until the system is honest again.
If those three names are missing, the organisation does not have production AI. It has a deployed artefact with no operator. That gap is not the same as failing to hand a working pilot into production in the first place. Handoff is how work enters production. Incident ownership is who runs it after it is there.
Name the operator, not the platform
“The data science team will support it” is not a page. “Operations will notice” is not a page. A production operator is a role that can be woken, that can take the system off the live path, and that knows who in the business must be told. The operator may sit in technology. They may sit in the process that uses the output. What they cannot be is an unnamed group, or the person who built the notebook and has since moved on.
| Question after go-live | When it is not named |
|---|---|
| Who is paged when the output is wrong or missing? | The first person who notices improvises, or nobody notices until a customer does |
| Who may roll back or take the model off the live path? | The system stays live because no one will take the blame for switching it off |
| Who talks to the business while it is dark? | Process owners keep using a number they no longer trust, or they silently revert to the old method |
| Who records what failed and what changed? | The next incident starts from rumour because the last one was never written down |
This is not general governance
Governance still matters. It is the reason unclear ownership of data, model change, and rollback stalls delivery. Incident ownership is more specific. It is the named human on the other end of a page after the system is live. You can have a policy that “rollback is owned by the model owner” and still have nobody who answers at 21:00. The policy is not the pager.
What the operator must be allowed to do
An operator without rollback rights is a messenger. An operator who cannot speak for the business is a technician. The role only works when the person who is paged can take the live path offline, restore the last known-good behaviour, and give the process owner a sentence they can use with their team. If legal or architecture must join every incident before anything is switched off, the page will wait until morning and the business will have already worked around the system.
- Write the page path before go-live: who is contacted, in what order, and who is skipped if they do not answer.
- Give the operator authority to roll back or disable the live path without waiting for a steering slot.
- Name the business counterpart who will hear the incident in language the process can use, not in model metrics.
- Run one rehearsal on the real path — a planned disable and restore — before the first unplanned failure.
- When the original builder leaves, the operator role transfers as a named person, not as a shared inbox.
If you cannot name who is paged, who rolls back, and who talks to the business, you did not finish go-live. You left a system running without an operator.
Finish go-live with an operator, not with a demo
Production incident ownership belongs in the same strategy conversation as the rest of the operating model. The commercial follow-on is naming the production operator on the strategy-vision page — who is paged after go-live — rather than another pilot that nobody will be responsible for when it fails in use.
Limitations and corrections
Company operating argument about named production-incident ownership after go-live. It does not quantify incident volume and is not a legal or on-call staffing standard.
To request a correction, contact Luminaracorp.