A Tier 3 data centre is certified to 99.982% uptime. Expressed in time, that is 1.6 hours of allowable downtime across an entire year. Every maintenance decision the facilities team makes either protects that number or chips away at it.
Preventative maintenance in a tier-rated environment is not about keeping the equipment clean and the log sheets current. It is about ensuring that every system in the facility is operating within its rated parameters, that every redundant component is genuinely capable of carrying the full load if the primary fails, and that the maintenance programme is designed around what tier compliance actually demands on the ground, not what the budget allows on a spreadsheet.
This guide covers what preventative maintenance for server rooms and data centres requires in South Africa in 2026, drawing from 22 years of managing tier-rated environments and from the specific lessons those engagements have produced. It addresses the most common gaps Boron finds when auditing facilities on contract takeover, the compliance standards that apply, and what to look for when choosing a maintenance provider who will actually protect your tier rating rather than simply visit your site on a schedule.
What are the key responsibilities of a data centre facilities manager in South Africa?
The facilities manager of a data centre or server room is responsible for six things: the power infrastructure (UPS, generators, electrical distribution), the cooling infrastructure (chillers, CRAC or CRAH units, cooling towers), the fire detection and suppression systems, physical access control and security, the maintenance programme across all of the above, and the documentation that proves compliance with the facility’s tier certification and regulatory obligations.
In South Africa, that responsibility set now includes an additional layer that did not exist with the same weight five years ago: managing the facility’s operating environment against a grid that is structurally unreliable. A facilities manager who is running a sound preventative maintenance programme on the power and cooling infrastructure but has not reviewed the maintenance intervals to account for actual generator runtime during load shedding is managing against a plan that has drifted out of alignment with operational reality.
The human dimension matters too. When a long-tenured facilities professional retires or changes roles, the operational knowledge they hold, knowing which piece of equipment has a recurring fault, which sensor has been running slightly out of calibration for years, which transfer switch sometimes hesitates, leaves with them. A well-maintained documentation system and a maintenance log that records not just what was done but what was observed captures that knowledge in a form that does not depend on any individual being present.
Tier 3 means concurrently maintainable. Every critical component has a redundant counterpart that is already running and capable of carrying the full load. You can take any single component offline for maintenance without affecting the facility’s uptime.
What this means in practice is that if you have two UPS systems each rated at 200 kilowatts, and the facility load is 200 kilowatts, each UPS must be running at 50% of its capacity during normal operation. If one UPS needs to come offline, the other picks up the full 200 kilowatts without any manual switching, without any perceptible change to the tenants, and without the facility dropping below its tier standard during the maintenance window.
The same principle applies to chillers, generators, and cooling towers. You need enough redundancy that any single component can be taken out of service without the remaining components running above their safe operating capacity. A facility with two chillers, each running at 90% capacity to carry the full load, does not have genuine Tier 3 cooling redundancy because if one chiller fails, the remaining chiller cannot carry the load. The design meets the letter of dual-component redundancy but not its spirit.
“When it says Tier 3, in layman’s terms, the system is concurrently maintainable. You can do maintenance on a system while it is running without affecting it. So you always have a redundant path. For everything.” (Nathan Chirwa, Director, Boron Facilities Management).
What happens when a facility slips out of tier compliance without realising it?
The moment any redundant component is offline for any reason, the tier rating is temporarily compromised. The facility is operating without its safety net, and any further failure during that window has no fallback.
This is the scenario that a sound maintenance programme is designed to prevent. If a chiller is offline for maintenance, the maintenance window should be scheduled at a time when the remaining chiller capacity is sufficient to carry the full load with thermal reserve. If a UPS is offline for battery replacement, the transfer to the redundant UPS must be tested, not assumed. If a generator is offline for service, the service window should be during a period of low outage probability, and the remaining generator must be tested under full load before the maintenance window opens.
Facilities that schedule maintenance visits on calendar dates without reference to the facility’s current operational state are at risk of creating exactly this scenario. A monthly visit on the first Tuesday regardless of what else is running, is an administrative schedule, not an engineering one.
“In my 22 years of managing data centres, we have not once come to a point where we risked the tiering oversight. That’s why we have an in-house engineer who does the tiering. He stresses to us: this site is certified for Tier 3, and we need to maintain it as a Tier 3.” (Nathan Chirwa, Director, Boron Facilities Management)
The scalability design mistake that creates the most expensive preventative maintenance problems
The most common and most expensive problem Boron finds when auditing existing facilities is not ageing equipment or deferred maintenance. It is equipment that was sized correctly for the original load and is now approaching or exceeding its rated capacity because the load grew without a corresponding infrastructure upgrade.
The scenario plays out the same way across many different facilities. A data centre is designed for 200 kilowatts. The UPS is sized at 200 kilowatts. The cooling plant is sized for 200 kilowatts. Two years later, the facility has taken on additional tenants, and the load has crept to 240 kilowatts. The UPS is now running at 120% of its design capacity, which means any load spike pushes it into bypass mode. The cooling plant is struggling to maintain the supply temperature that the equipment requires. The redundancy that was built in at 200 kilowatts no longer functions as designed at 240 kilowatts.
The engineering fix is to design the core infrastructure for the load you expect to reach within five years, and to install only the components you need for the current load. A transformer can be rated for 500 kilowatts while the connected load is 200 kilowatts. When the load grows, the transformer does not need to be replaced. Cabling can be sized for future load. The modular UPS architecture discussed earlier allows capacity to be added in increments without infrastructure changes. The capital cost of this approach at design time is a fraction of the cost of infrastructure replacement at full load in a live facility.
Designing redundancy into a facility from the start adds cost to the original project. Retrofitting redundancy into a live facility adds cost, risk, and downtime to the operating business. For any facility where continuous uptime is commercially or contractually required, the upfront design investment almost always represents the lower total cost.
When you retrofit a redundant power path into a live system, the coupling point between the new path and the existing infrastructure requires the existing system to go offline briefly. In a data centre with tenants on SLA contracts carrying financial penalties for downtime, that brief offline period is a commercial event. It needs to be communicated, scheduled, and managed as a planned outage, which means coordination with every tenant, potential service credits, and the reputational cost of an outage that was caused by a maintenance project rather than by equipment failure. Getting the redundancy right at the design stage eliminates that entire category of risk.
What a preventative maintenance programme actually covers across power, cooling, fire, and access
A full preventative maintenance programme for a data centre or server room covers four infrastructure layers: power, cooling, fire detection and suppression, and physical access control. Each has its own maintenance schedule, its own failure modes, and its own compliance requirements.
For power: quarterly UPS battery inspections, quarterly generator load bank tests adjusted for actual runtime hours during load shedding, monthly thermographic scanning of electrical distribution boards, six-monthly inspection of automatic transfer switches and manual transfer panels, and annual testing of earthing systems and lightning protection.
For cooling: monthly inspection of CRAC and CRAH units, including filter condition, refrigerant pressure, and supply temperature calibration; quarterly inspection of chillers and cooling towers, including belt and bearing condition, water treatment checks, and thermal imaging of compressors; and annual load testing of the cooling plant under conditions representative of peak summer load in the facility’s location.
For fire: quarterly inspection and functional testing of all detection and suppression system components, annual discharge simulation for gas suppression systems, and a full compliance review against the applicable SANS and NFPA standards at each contract anniversary.
For access control and CCTV: monthly functional testing of all readers, locks, and alarm triggers; six-monthly full system audit including camera coverage review, access log analysis, and credential database audit; and annual inspection of physical barriers and turnstiles for mechanical wear.
How does load shedding impact the preventative maintenance strategy in South Africa?
Load shedding has moved the basis for maintenance scheduling from calendar time to runtime hours. Equipment that used to accumulate 500 hours of runtime in twelve months now accumulates it in six to eight weeks during Stage 6. Maintenance programmes that have not been adjusted for this are systematically overdue.
The practical adjustment is straightforward. Instead of scheduling a generator service every six months, schedule it at 500 hours of runtime, whichever comes first. Track runtime hours through the generator’s hour meter or through the BMS if the facility has one, and trigger the service when the threshold is reached rather than when the calendar says it is time. The same principle applies to cooling plant bearing inspections, UPS battery capacity checks, and any other maintenance activity that is driven by operating cycles rather than by time alone.
The relevant compliance frameworks are: the Uptime Institute tier certification standards, the Occupational Health and Safety Act (Act 85 of 1993) as it applies to electrical and mechanical work on occupied facilities, SANS 10142 for electrical installations, the SANS 10400 series for building and fire safety requirements, and ISO 9001:2015 as the quality management standard against which a maintenance provider’s processes should be audited.
ISO 14001:2015 and ISO 45001:2018 are relevant to the environmental and occupational health management systems of any provider doing maintenance work on a facility.
A provider certified against all three of these standards through an accredited certification body has had their quality, environmental, and safety management systems independently assessed. That assessment is different from the provider having written a policy document. An audit certificate means the processes have been observed in operation and found to conform to the standard.
Boron Facilities Management holds ISO 9001:2015, ISO 14001:2015 and ISO 45001:2018 certification as an Integrated Management System through QAS International, issued in June 2026 and valid to June 2027. Certificate number SAP1021IMS. CIDB Grade 5EB (Electrical) and 5ME (Mechanical). Level 1 B-BBEE. Uptime Institute Tier Designer accredited.

How to conduct a preventative maintenance audit for a South African server room or data centre
A preventative maintenance audit has two phases: a document review and a physical inspection. The document review covers the maintenance log, the service interval schedule, the equipment register, the last generator load bank test result, and the last UPS battery inspection report. The physical inspection covers every critical component against the findings in the document review.
The most revealing part of a takeover audit is the comparison between the maintenance log and the equipment runtime data. If the log shows a generator was serviced six months ago but the hour meter shows 900 hours of runtime since that service, the service interval has been missed in practice even if it appeared on the schedule on paper. If the log shows quarterly UPS battery inspections but the internal resistance measurements were not recorded, the inspection was visual only and cannot confirm battery bank performance under load.
A professional FM audit firm produces a written report that covers each infrastructure layer, rates the condition of each system against its design specification, identifies deferred maintenance and its risk classification, and provides a prioritised capex list for remediation. That report is useful to the facilities manager, to the financial director who needs to budget for infrastructure maintenance, and to the insurance underwriter who needs to assess risk exposure.
Benefits of hiring a professional FM audit firm in South Africa
The primary benefit is independence. An audit produced by the same provider responsible for the maintenance programme has a structural conflict of interest. An audit produced by a different, qualified provider gives the client a view of the facility’s condition that is not influenced by the commercial relationship with the incumbent maintenance contractor.
The secondary benefit is qualification. Not every FM audit firm has CIDB-graded engineers on the team or ECSA-registered professionals making the technical judgements. An audit conclusion that “the UPS is in good condition” from a team without formal electrical engineering credentials is not the same as the same conclusion from a team with a professional engineer who carries personal liability for the assessment. For facilities under SLA contracts with financial penalty clauses, the qualification level of the audit team directly affects how much weight the audit findings carry in any subsequent commercial dispute.
A B-BBEE credential should be part of the evaluation, but it is one credential among five, not a substitute for any of the others. A Level 1 B-BBEE provider that cannot answer the engineering credentials question is not a sound choice for a tier-rated data centre, regardless of how well the empowerment scorecard looks. Both matter. Neither substitute for the other.
If your server room or data centre has not had an independent maintenance audit in the past twelve months, Boron offers a complimentary infrastructure assessment. Our team reviews your power, cooling, fire, and access control systems and provides a written findings report. No cost. No obligation. Turnaround within five business days of the site visit.
