
The Company’s standing procedure for every incident at the KOM Oman AI Factory — physical, infrastructure and cyber. Classification, response targets, command structure, notification, root-cause discipline and the cyber playbooks. A controlled internal document, released to counterparties where the engagement protocol requires it.

This plan governs the handling of every incident affecting the KOM Oman AI Factory, from a failed fan tray to a confirmed intrusion. It deliberately uses one classification scheme and one command structure for physical, infrastructure and cyber events. A facility that runs two parallel incident processes discovers, at the worst moment, that nobody is certain which one applies — a cooling failure caused by a compromised building-management controller is both, and the response cannot pause to decide.
The plan is owned by the CISO jointly with the COO: the COO owns availability of the facility, the CISO owns the integrity of the systems that run it. Where the two conflict during an incident, section 05 states which authority prevails and when.
| Status | An internal policy of Prima Artificial Intelligence LLC. It states how the Company runs its own incident response; it is not a marketing document and is not written to a third party’s template. |
|---|---|
| Authority | Issued under the Corporate Governance Charter. Owned jointly by the CISO and the COO; material amendment requires the approval of both. |
| Controlled copies | Executive Leadership Team; Operations Manager and Shift Leads; NOC; Information Security; General Counsel. |
| Release outside Prima | Released to customers, offtakers, lenders, insurers and certification bodies where an engagement or contractual protocol requires evidence of the Company’s incident-response capability. Release is logged; the document is not published. |
| Review cycle | Annually, and after any P1, any security incident, or any material change to the estate. |

Site and building security events; power, cooling and mechanical faults; fabric and platform faults; compromise of Prima-operated IT and OT systems; loss or exposure of Prima-held data; any event triggering a customer notification obligation.
Events confined to a customer's own workload, operating system, containers or data — the customer's incident process governs those. Prima responds on request as smart hands and provides telemetry, but does not command the customer's response.
The boundary is the leaf-switch customer port and the rack door. Prima owns everything up to it; the customer owns what runs behind it. Where an incident crosses that boundary — a Prima fault that corrupts a customer job, or a customer workload that causes a thermal excursion — section 05 provides for a joint incident, run by a Prima incident commander with a named customer counterpart.
| Incident | Any unplanned event that degrades, interrupts or threatens the service, the facility or the security of either — whether or not a customer notices it. |
|---|---|
| Security incident | An incident in which the confidentiality, integrity or availability of a system or of data has been, or is credibly suspected to have been, compromised. |
| Breach | A security incident in which unauthorised access to, or disclosure of, data is confirmed rather than suspected. Every breach is a security incident; not every security incident is a breach. |
| Incident commander | The single person accountable for the response at any moment. The role moves with escalation; it is never held by two people at once. |
| Time to respond | From the earlier of NOC detection or customer report, to acknowledgement by the accountable owner — not to resolution. |
| Time to restore | From the same start point to the service being returned to its committed state, verified by the NOC. |
Incidents are classified on impact, never on cause. This matters more than it sounds: classifying on cause invites argument at the moment argument is most expensive, and it systematically under-rates novel failures because they do not match a known category. A total loss of compute is a P1 whether the cause is a transformer, a firmware bug or an intruder.
| Level | Impact definition | Respond | Update | Target restore |
|---|---|---|---|---|
| P1 | CriticalCommitted capacity unusable | 15 min | 60 min | 4 h |
| P2 | HighDegraded, or redundancy lost — service on a single path | 1 h | 4 h | 24 h |
| P3 | MediumNon-service-affecting fault — single node, port or component | 4 h | Daily | 5 bd |
| P4 | LowRequests, smart hands, information | NBD | On close | As agreed |

| Level | Example event | Domain |
|---|---|---|
| P1 | Loss of utility feed with generator failure to start on load | Power |
| P1 | Coolant supply outside the committed envelope for more than 15 minutes | Cooling |
| P1 | Confirmed unauthorised entry into a data hall | Physical security |
| P1 | Ransomware executing on any system with a path to OT or to customer data | Cyber |
| P1 | Confirmed exfiltration of customer data or model artefacts | Cyber · breach |
| P2 | Loss of one of two UPS strings, load carried on the survivor | Power |
| P2 | Chiller failure with N+1 holding the envelope | Cooling |
| P2 | Credential compromise on an administrative account, no confirmed lateral movement | Cyber |
| P2 | Camera or reader offline on a controlled layer boundary | Physical security |
| P3 | Single compute node down, capacity within committed tolerance | Platform |
| P3 | Vulnerability disclosed on an in-scope system, no active exploitation | Cyber |
| P4 | Scheduled rack work, cable moves, documentation requests | Operations |

Every incident enters through the Network Operations Centre, whatever its origin. There is no second queue: an incident reported to an engineer in a corridor is not an incident until the NOC holds it, and staff are trained to route rather than absorb. This single-intake rule is what makes the response clock and the availability record trustworthy.
| Source | What it watches | Interval |
|---|---|---|
| Infrastructure monitoring | Power chain telemetry, CDU and BMS sensors, fabric port state, rack reachability probes | 30 seconds |
| Platform monitoring | Node health, GPU state, scheduler queue health, storage and filesystem telemetry | 30 seconds |
| Security monitoring | Log and event correlation across IT and OT, endpoint telemetry, authentication anomaly | Continuous |
| Access and surveillance | Door-forced and door-held alarms, failed credential patterns, camera and reader health | Continuous |
| Customer report | Named contacts via NOC hotline, ticket portal or email | 24×7 |
| Staff and vendor report | Any person on site, through the NOC — including the security provider | 24×7 |
| External notification | Vendor advisories, CERT bulletins, threat intelligence, law enforcement | On receipt |
Triage is a NOC function and is not delegated. The engineer on shift may open at any severity, page any escalation level and summon the security path without seeking permission — a triage function that must ask before escalating does not escalate in time.
At triage the NOC asks one additional question of every incident: could this be caused by someone rather than something? Where the answer is yes or unknown, the incident takes the security path in parallel with the operational one — because the operational instinct is to restore quickly, and restoring quickly frequently destroys the evidence an investigation needs.
| Triggers the fork | Unexplained configuration change; authentication anomaly; simultaneous unrelated faults; any OT anomaly; anomalous outbound traffic; boundary alarm without a matching entitlement. |
|---|---|
| Immediate effect | CISO on-call paged in parallel with the operational owner. Evidence preservation begins before remediation — section 11. |
| Who decides | The NOC engineer opens the fork. Only the CISO or delegate closes it. |
| Restoration | Restoration proceeds unless the CISO holds it. Where availability and evidence conflict, section 05 governs. |

Every incident runs the same six phases. Small incidents pass through some of them in seconds; a P1 may hold in containment for hours. The value of a fixed lifecycle is that at any moment every participant can state which phase the incident is in, and therefore what is and is not being attempted.
| Phase | Objective | Exit criterion |
|---|---|---|
| 01 DetectNOC | Establish that something is wrong, with a defensible start time | Ticket open, severity set |
| 02 TriageNOC | Determine scope, name the owner, start the escalation clocks | Owner acknowledged |
| 03 ContainOwner · CISO | Stop the spread — isolate the fault domain; preserve evidence | No further degradation |
| 04 RestoreOwner | Return the service to its committed state, by repair or by failover | NOC verifies against baseline |
| 05 CloseOwner · NOC | Confirm stability, quantify availability impact, notify the customer of closure | Stable for the hold period |
| 06 LearnRCA owner | Establish cause, issue the written analysis, commit dated corrective actions | Actions accepted and tracked |
Containment precedes restoration in every case, strictly so on the security path: restoring a compromised system without containing the compromise reinfects it, destroys the evidence and starts the incident again with less telemetry than before. On the operational path containment is often instantaneous — the design is concurrently maintainable, so isolating a failed UPS string is the containment step.
| Severity | Hold period before an incident may be closed | Verified by |
|---|---|---|
| P1 | 4 hours of stable operation, telemetry matching the pre-incident baseline | NOC + owner |
| P2 | 2 hours stable, with redundancy demonstrably restored | NOC |
| P3 | Verified on completion | NOC |
| P4 | Verified on completion | Requester |
A security incident additionally requires the CISO's written concurrence to close. Where root cause is unestablished, the incident closes operationally but the security case stays open — tracked separately, so a restored service is never mistaken for a resolved compromise.
Re-opens the original ticket rather than creating a new one. The availability impact accumulates against the original RCA, and the corrective action is treated as failed.
Escalates automatically to the COO regardless of severity, with a written explanation of why the previous corrective actions did not hold.
Any fault recurring in three consecutive months enters the quarterly service review as a standing item until two clear months pass.

Escalation is time-based and automatic. It does not wait for a request, and it is not a judgement about competence — it is a standing rule that puts progressively more authority next to a problem that is taking progressively longer. A P1 unresolved at two hours reaches the COO whether or not anyone asked for help.
On the security path the ladder runs in parallel: CISO on-call at detection, CISO at 15 minutes, CEO and General Counsel at 2 hours where a breach is confirmed or credibly suspected. A P1 unresolved at four hours produces a formal incident notice to the affected customers within the same hour.
| Role | Accountability during the incident |
|---|---|
| Incident commanderShift Lead, then Ops Manager, then COO | Owns the response. Decides sequence, allocates people, and is the only role that may change severity or declare restoration. Does not perform hands-on work. |
| Technical leadNamed per domain | Directs diagnosis and repair in one domain — electrical, mechanical, platform, network, security. Reports to the commander, not to the customer. |
| ScribeNOC Engineer | Maintains the timeline in the ticket in real time: every observation, decision and action with its timestamp. The scribe is never also the technical lead. |
| Communications leadOps Manager or delegate | Sole channel to customers and to internal stakeholders. Prevents the commander being consumed by status requests. |
| Security leadCISO or delegate | Owns the security path: evidence preservation, containment of compromise, forensic decisions, notification assessment. |
| Vendor liaisonFacilities or Engineering Manager | Single point of contact to OEM and managed-service vendors; holds the contractual response entitlements. |

During a security incident the fastest route to restoration and the correct route to containment sometimes diverge. That conflict is resolved in advance rather than negotiated under pressure.

Customers are notified on a schedule, without asking. A customer who has to chase for status during an outage is having two problems, and the second one is ours. Notification is the communications lead's sole responsibility so that it continues while the technical work does.
| Severity | First notification | Then | On restoration |
|---|---|---|---|
| P1 | Within 60 min of classification | Every 60 min | Within 30 min |
| P2 | Within 4 h of classification | Every 4 h | Within 4 h |
| P3 | In the monthly report, or on request | — | Monthly report |
| P4 | On completion | — | On completion |
A security incident affecting a customer's environment or data is notified within 24 hours of confirmation, irrespective of its operational severity, and thereafter as material facts change rather than on a fixed cycle — section 12 sets out the regulatory obligations that run alongside.
| Reference | Ticket identifier, severity, and when the incident started — not when we noticed. |
|---|---|
| Impact | What the customer is experiencing, stated in the customer's terms — racks, GPUs, halls affected — not in ours. |
| Status | Current lifecycle phase per section 04, and what is being attempted now. |
| Expectation | Next update time and best estimate of restoration with an honest confidence level. Where no estimate is possible, that is said plainly. |
| Action required | Anything the customer needs to do, or explicit confirmation that no customer action is required. |
| Contact | Named person and channel for the customer's own escalation. |
24×7 hotline and ticket portal. Can raise severity, request status, or dispute a classification.
Named contact per account. Reached directly where NOC response is unsatisfactory.
Named executive contact. Available for any P1, and on request for a disputed incident.

A restored service is not a resolved incident. Root-cause analysis is a deliverable with a deadline, not an aspiration, and its output is a set of dated commitments that are tracked to completion in the same register as the incident itself.
| Trigger | Scope of analysis | Issued within |
|---|---|---|
| Every P1 | Full written analysis with timeline, cause, contributing factors, availability impact and corrective actions | 5 business days |
| P2 on requestor on recurrence in-quarter | Full written analysis, same structure | 10 business days |
| Every security incident | Full analysis including how detection performed and whether containment held | 10 business days |
| Third recurrenceAny severity | Analysis of why previous corrective actions failed, presented to the COO | 10 business days |
| Three environmental excursions in a month | Root-cause review of the cooling or air path concerned | 10 business days |
Analysis follows a structured causal method — successive why questions from the observed failure to the conditions that permitted it, with contributing factors recorded separately from the proximate cause. Two rules give the method its value:

| Form | Each action has a named owner, a dated commitment and a definition of done. Actions without all three are not accepted into the register. |
|---|---|
| Classification | Immediate (in place before the RCA issues) · Short-term (within 30 days) · Structural (dated, may extend across a build phase). |
| Tracking | Open actions are reviewed monthly by the Operations Manager and reported in the monthly service report until closed. |
| Verification | Closure requires evidence, not assertion — a test result, a procedure revision, a configuration record. |
| Overdue | An action past its committed date escalates to the COO automatically and appears in the quarterly service review. |
Customers affected by a P1 receive the sanitised RCA: full timeline, cause, availability impact and corrective actions, with other customers' identities, third-party commercial terms and security-sensitive detail removed. Sanitisation removes information that would harm another party — it does not remove information that is unflattering to Prima. Where a security-sensitive control is at issue, the RCA states that a control was involved and that it has been remediated, without describing it in a way that would assist an attacker.

A GPU facility is an unusual target: it concentrates very high-value data and operational technology whose compromise has immediate physical consequence. The threat model below drives the playbooks in sections 09 and 10, and is reviewed twice yearly and after any material change to the estate.
| Asset | Why it is attractive | Consequence |
|---|---|---|
| Customer model artefactsWeights, checkpoints, datasets | Directly monetisable and irreplaceable — years of a customer's investment in one file tree | Catastrophic |
| OT — BMS, EPMS, ACSBuilding, power, access | Compromise produces physical effect: thermal event, power interruption, door release | Severe |
| Control planeProvisioning, scheduler, IPMI/BMC | Access to every tenant at once; below the operating systems customers can defend | Severe |
| Identity systemsIAM, directory, PAM | The route to everything else; the target of choice in a patient intrusion | Severe |
| Surveillance and access recordsCCTV, entitlement logs | Reconnaissance for a physical attack; and the evidence trail an attacker wants erased | High |
| Corporate ITEmail, documents, finance | Commercial intelligence, contract terms, and the usual ransomware target | High |
Motivated by the compute itself and by what runs on it. Patient, well-resourced, targets identity and supply chain rather than the perimeter. Assumed to be interested in a sovereign AI facility in the Gulf as a matter of course.
Ransomware and extortion, increasingly with data theft first. Targets corporate IT for access and OT for leverage — a threat to interrupt cooling is a powerful negotiating position.
Holds legitimate credentials and physical access. Addressed through separation of duties, entitlement recertification and dual authorisation on destructive operations.
OEM remote access, managed-service accounts, firmware in the delivery path. Legitimate access, weaker control environment — the route most often under-defended.

Each playbook states the first action — where incidents are won or lost, and the one decision that must not fall to whoever happens to be nearest.
| First action | Isolate at the network, do not power off. Powering off destroys volatile memory that holds keys, process state and the attacker's tooling. Isolation stops spread and preserves the machine. |
|---|---|
| Contain | Segment the affected zone; disable compromised accounts; block command-and-control at egress; verify OT segmentation is intact — assume the OT boundary is the objective. |
| Assess | Establish whether data was exfiltrated before encryption. Modern extortion steals first; treating the event as availability-only understates it and delays notification. |
| Eradicate | Rebuild from known-good images. Restoration of a compromised system to service is not permitted without CISO sign-off, whatever the availability pressure. |
| Recover | Restore from immutable, offline-verified backups. Backup integrity is tested before restoration, not assumed — see the exercise programme, section 13. |
| Position on payment | Prima does not pay ransoms. The decision is reserved to the Board and does not sit with the incident commander, so that no one under operational pressure can be induced to make it. |
| First action | Suspend the account and invalidate its active sessions. A password reset without session invalidation leaves the attacker in place — the most common error in this playbook. |
|---|---|
| Contain | Revoke tokens, API keys and certificates; force re-authentication across privileged systems; check for new accounts and altered entitlements. |
| Investigate | Reconstruct everything the identity touched from first anomaly to suspension. Where the identity held OT or control-plane access, escalate to P1 without waiting for confirmation of misuse. |
| Physical cross-check | Compare the digital timeline against access-control records. A credential in use while its holder was not on site, or vice versa, is a strong indicator — and the reason both systems are monitored together. |
| Recover | Re-issue credentials in person with identity re-verification. Recertify the identity's full entitlement set rather than restoring it wholesale. |
| First action | Preserve, then block. Capture the flow evidence before severing the channel — a blocked channel with no record leaves the scope of loss unknowable, which is worse than the loss. |
|---|---|
| Contain | Block the destination at egress; isolate the source; suspend the identities and services involved; freeze the affected storage from further modification. |
| Scope | Determine what left, when, in what volume and whose, from flow records, storage logs and endpoint telemetry. Scope is stated as established fact and bounded uncertainty — never a reassuring guess. |
| Notify | Affected customers within 24 hours of confirmation, with scope as then understood and a commitment to update. Regulatory assessment runs in parallel — section 12. |
| Customer artefacts | Where customer model artefacts are implicated, the customer is treated as a participant in the investigation, not a recipient of its conclusions. Prima shares telemetry in real time under the incident. |

These three playbooks share a property that changes the response: compromise produces physical effect, not merely data loss — cooling stopped, breakers opened, doors released. The response therefore prioritises reverting to manual control over investigating in place.
| First action | Assert local manual control of the plant before touching the network. Confirm cooling and power can be run independently of the compromised system, and that a human is watching them. Only then investigate. |
|---|---|
| Contain | Sever the OT zone from external and corporate connectivity; disable remote vendor access wholesale, not selectively; retain internal monitoring so the plant does not go dark to the NOC. |
| Verify physical state | Confirm plant condition by direct observation and independent instrumentation. A compromised BMS may report normal while conditions are not — the displayed value is the least trustworthy thing in the incident. |
| Assess intent | Distinguish reconnaissance from setpoint manipulation. Any evidence of changed setpoints, altered alarm thresholds or suppressed alarms escalates immediately to the CEO and to affected customers. |
| Access control | Where ACS is implicated, revert affected boundaries to manual verification with officer presence. Entitlements are not trusted until the system is proven clean. |
| Recover | Firmware and configuration restored from verified known-good baselines held offline. Controllers are not patched in place during an active incident. |
| First action | Determine whether the event is volumetric or authenticated. A flood and a compromised administrative interface look similar in a dashboard and require opposite responses. |
|---|---|
| Contain — volumetric | Engage upstream scrubbing; shed non-essential services to protect the control plane; preserve out-of-band management access as the last channel to be surrendered. |
| Contain — authenticated | Treat as identity compromise per 9.2 with immediate P1 severity, because control-plane access reaches every tenant simultaneously. |
| Protect the substrate | IPMI and BMC interfaces are never exposed beyond the management network. Any evidence of reachability from outside it is itself a P1 finding, regardless of whether access occurred. |
| Customer effect | Loss of the management plane is customer-visible even when compute continues. It is notified as a P1 rather than reported later as a footnote. |
| First action | Disable the vendor's access entirely — every account, not the one implicated. Selective revocation assumes knowledge of the intrusion's scope that does not exist in the first hour. |
|---|---|
| Contain | Review every session that vendor conducted in the exposure window from the recordings; identify configuration changes; verify against the change register. |
| Firmware and images | Quarantine suspected units and verify firmware against vendor-published hashes. No further units from the same batch are deployed pending clearance. |
| Notification received | A vendor-disclosed compromise is opened as a Prima incident at P2 minimum from the moment of disclosure, whether or not the vendor believes Prima was affected. |
| Restore access | Re-enabled only after the vendor evidences remediation, and then under enhanced monitoring for a defined period. Prima's own control does not depend on the vendor's assurance. |

Evidence decisions are made in the first minutes and cannot be revisited, so the rule is simple enough to apply under pressure: preserve before you repair, and where the two genuinely conflict, escalate rather than choose alone.
| Seq | Evidence class | Why the order matters |
|---|---|---|
| 01 | Volatile memoryRAM, process state, network connections | Lost on power-off; holds encryption keys, injected code and live sessions — the reason no compromised host is powered down |
| 02 | Network stateActive flows, ARP and routing tables | Lost within seconds to minutes of isolation |
| 03 | Logs on the hostAuthentication, application, system | May be actively deleted by the attacker; central copies verified against the local ones |
| 04 | Disk imagesFull forensic copies | Durable but must be captured before rebuild or restoration |
| 05 | Central telemetrySIEM, flow records, BMS series | Retained independently; least likely to be lost, most likely to be needed for scope |
| 06 | Physical recordsCCTV, access-control logs | Retained per the Physical Security Specification; correlate the digital timeline to a person |
| Custodian | The CISO or a named delegate holds custody of all evidence from the moment of collection. Custody is not shared. |
|---|---|
| Record | Every item is logged with what it is, when and by whom it was collected, its cryptographic hash, and every subsequent access. |
| Integrity | Images and log exports are hashed on collection and re-verified before any analysis. Analysis is performed on copies; originals are not mounted. |
| Storage | Held encrypted, in a location separate from the affected environment, with access restricted to the custodian and named investigators. |
| Retention | Seven years for evidence relating to a security incident; longer where litigation or a regulatory process is reasonably anticipated. |

| Recipient | Trigger | Timing |
|---|---|---|
| Affected customers | Security incident affecting the customer's environment or data | 24 h of confirmation |
| Affected customers | Operational P1 or P2 | Per section 06 |
| Data-protection authorityOman Personal Data Protection Law | Personal-data breach meeting the statutory threshold | As prescribed |
| National CERTOman | Significant incident affecting critical information infrastructure | As prescribed |
| Insurer | Any incident potentially within cyber or property cover | Per policy terms |
| Lenders | Incident meeting a reporting threshold in the finance documents | Per facility terms |
| Board of Directors | Any confirmed breach; any P1 with material commercial exposure | Within 24 h |
The Company holds no breach history. That statement requires context to be worth anything, and the context is this: the KOM Oman AI Factory is pre-operational, with ready-for-service in December 2026. A facility that has not yet carried a customer workload has had no opportunity to accumulate an incident record. A nil return is a statement of fact about elapsed time, not evidence of operational resilience, and it is recorded here as such.

A response plan that has never been rehearsed is a document, not a capability. The exercise programme is the mechanism by which this plan is known to work, and its findings are tracked in the same corrective-action register as real incidents.
| Exercise | Scope | Frequency |
|---|---|---|
| Shift drill | Triage and classification against scripted scenarios; escalation paging verified | Monthly, per shift |
| On-load generator run | Transfer under real load; thermal ride-through observed and recorded | Monthly |
| Tabletop — cyber | Ransomware, OT compromise and exfiltration scenarios with the full command team | Quarterly |
| Backup restoration test | Restoration from immutable backup proven end to end, not merely verified as present | Quarterly |
| Black-building test | Full loss of utility and controlled recovery of the facility | Annually |
| Penetration testIT and OT scope | Independent third party; remediation tracked to closure with retest | Annually |
| Notification rehearsal | Customer, regulator and insurer notification paths exercised against the clock | Annually |
Exercise findings are raised as corrective actions with owners and dates, tracked identically to those arising from real incidents. An exercise that produces no findings is treated as insufficiently demanding and its scenario is revised.

Time to respond against target, by severity, with every miss explained individually.
Time to restore against target, and availability impact in minutes.
Share of incidents found by monitoring rather than reported by a customer.
Corrective actions closed on time; count and age of those overdue.
Of these, detection quality is watched most closely. Response and restoration times measure how well the team performs once it knows; the share of incidents a customer had to tell us about measures whether we knew at all.
Owner of this policy. Questions on classification, the notification commitments, the exercise programme, or a request to release this document outside Prima. Version 1.0 will fix the statutory notification periods on written advice from Omani counsel.