Artigo

Data Centers: executive and technical guide to design, planning, and construction

EnQ Digital·04 de setembro de 2026

Designing a data center means managing consequences. The final product is not the building, but the continuous, secure delivery of computing capacity, connectivity, and storage. This guide covers the full chain of decisions — from the business case and risk analysis to thermal and electrical engineering, telecommunications, protection, construction, commissioning, and operation.

The sequence is always the same: the business case defines criticality; criticality defines requirements; requirements define the design; the design is proven through testing; and operation preserves the performance achieved. Skipping a step transfers uncertainty to the construction phase, where changes cost more and jeopardize schedule.

About this material: the architectures presented are reference examples and do not replace an executive design, technical report, professional liability, or consultation with the current editions of applicable standards. Redundancy is not automatically synonymous with availability: the outcome depends on physical independence, selectivity, automation, procedures, integrated testing, and operational competence.

Executive summary

Each part of the guide answers a specific decision. Use the table as a reading map: managers should start with the first four rows; engineering should move through the technical disciplines; operations should focus on commissioning, documentation, and indicators; procurement should use the matrices and checklists to structure the RFP, technical equalization, and acceptance.

PartContentMain decision
FundamentalsObjective, typologies, and life cycleBuild, contract colocation, edge, or hybrid?
Requirements and risksBIA, criticality, threat, and toleranceWhich interruption is acceptable and at what cost?
AvailabilityClasses, topologies, maintenance, and failuresWhich architecture meets the service?
PlanningCapacity, area, density, and expansionHow much, when, and where to build?
ThermalAirflow, DX, chilled water, and liquid coolingHow to remove heat with stability and efficiency?
ElectricalGrid, GMG, UPS, batteries, ATS, PDU, and protectionHow to deliver safe power up to the IT load?
TelecomRoutes, MMR, cabling, and topologiesHow to eliminate hidden dependencies?
Protection and controlFire, access, CCTV, BMS, and DCIMHow to detect, contain, and respond?
Construction and deliveryMaterials, stages, QA/QC, and commissioningHow to prove the design works?
Operation and futureSLA, maintenance, PUE, AI, and high densityHow to sustain reliability over time?

What is a data center

Technical aisle between two rows of racks, with perforated raised floor
Technical aisle with two rows of racks. Photo: Robert Harker/Wikimedia Commons, CC BY-SA 3.0.

A data center is the coordinated set of spaces, utilities, controls, and processes that sustains technology and telecommunications equipment. Power, cooling, and networks are service chains; each link needs capacity, protection, monitoring, and a maintenance path.

LayerExamplesTypical failure
BusinessServices, customers, legislation, revenuePoorly defined criticality
IT and telecomCompute, storage, network, platformsDemand exceeding capacity
FacilitiesPower, cooling, fire, securitySingle point of failure
OperationsPeople, processes, parts, and suppliersHuman error or delayed response

Principle: critical infrastructure must be designed around the service that needs to survive, not just the equipment that will be installed.

The full life cycle runs from strategy and requirements to the preliminary study, conceptual design, basic and executive design, contracting, construction, testing, operation, optimization, and finally expansion or decommissioning.

Types and deployment models

ModelAdvantageAttention point
EnterpriseControl and integration with the businessCAPEX, scale, and specialized team
ColocationSpeed, connectivity, and shared infrastructureContract, demarcation of responsibilities, and cross-connects
HyperscaleScale, standardization, and automationHigh demand for power, water, and land
EdgeLow latency and proximity to the userDistributed sites, maintenance, and physical security
Modular/prefabricatedSchedule and repeatabilityCivil interfaces, logistics, and expansion
HybridFlexibility across sites and cloudsEnd-to-end operational architecture and observability

The build-versus-contract comparison must incorporate total cost, schedule, regulatory risk, available power, connectivity, expansion capacity, useful life, cost of capital, staffing, obsolescence, and contractual flexibility. A seemingly cheap in-house room may require disproportionate electrical and thermal retrofitting; colocation can reduce lead time, but requires an SLA, audit, and exit plan.

  • Map loads that cannot leave the site and those that can be distributed.
  • Model at least three growth scenarios and one stress scenario.
  • Include the cost of downtime, not just CAPEX and monthly fees.
  • Define who is responsible for each segment: utility, data center, rack, equipment, and application.

Requirements in balance

A viable project balances three groups: business defines what needs to continue; engineering translates that into capacity, topology, and performance; finance sets the investment and operating envelope. This balance must be documented in a Basis of Design (BoD), with assumptions, exclusions, and acceptance criteria.

GroupEssential questions
BusinessWhich services? Who is affected? What window? What contractual or legal obligation?
TechnicalInitial and future load? Density? Autonomy? Routes? Maintainability?
FinancialCAPEX, OPEX, contingency, cost of capital, energy, maintenance, and renewal?
OperationalStaff, 24x7 coverage, spare parts, suppliers, procedures, and training?

The Owner's Project Requirements (OPR) document consolidates this decision and must contain:

  • Project objectives and scope.
  • Load profile and growth horizon.
  • Availability, security, and efficiency criteria.
  • Environmental conditions and operating limits.
  • Metering, alarms, integration, and data retention.
  • Testing, documentation, training, and warranty.

BIA and business risk

The Business Impact Analysis identifies critical processes, dependencies, maximum tolerable downtime, and progressive impacts. The RTO guides the recovery time; the RPO guides the maximum acceptable data loss. They do not replace the facilities analysis, but they prevent infrastructure from being sized based on perception.

ImpactExample evidenceTreatment
FinancialLost revenue, penalties, reworkQuantify per hour and per event
OperationalBacklog buildup, productivity lossModel degradation and recovery
Legal/regulatoryDeadlines, data retention, securityValidate with legal and compliance
ReputationChurn, press, trustUse scenarios and indicators
Life/safetyMedical services, emergency, controlMinimum tolerance and formal response

Probability and consequence must be supported by history, asset condition, exposure, and existing controls. Residual risk is what remains after mitigation — accepting it is a decision by the risk owner, not an omission in the design.

Example: if a power failure can interrupt R$ 300 thousand per hour and the likely recovery time is 4 hours, the direct reference impact is R$ 1.2 million — before penalties, image, and operational recovery.

Internal and external risks

DomainThreatsVerification
LocationFlood, landslide, external fire, aircraft crashMaps, history, exposure radius, and inspection
UtilitiesWeak power grid, fuel, water, telecomFeasibility letters and physical routes
BuildingStructural load, water above critical room, vibrationSurvey and testing
FireCombustible load, compartmentalization, detectionLegal design and cause/spread analysis
PeopleImproper access, error, lack of trainingSegregation, procedure, and training
Supply chainLead time, exclusive part, single supplierCritical list and strategic stock
Cyber-physicalExposed BMS/DCIM, weak credentialsSegmentation, hardening, backup, and logging
  • Risk register with owner, deadline, and evidence.
  • FMEA for failure modes, effects, and controls.
  • HAZOP or What-if for process deviations.
  • Fault Tree for combinations leading to the top event.
  • Bow-tie to visualize prevention, event, and mitigation.

Availability without oversimplification

Industrial diesel generator installed in a data center technical room
Diesel generator at a data center. Photo: Mikael Häggström/Wikimedia Commons, CC0 (public domain).

Availability is the proportion of time the service is able to perform its function. For a repairable component, a common approximation is A = MTBF / (MTBF + MTTR). In a data center, however, common dependencies, human failures, maintenance, controls, and operational response make systemic behavior far more complex than the formula suggests.

N exact capacity, 1 path SOURCE MODULE CRITICAL LOAD no alternative path N+1 reserve in the same path SOURCE N +1 CRITICAL LOAD tolerates failure of 1 module 2N two independent systems SOURCE A SOURCE B MODULE A MODULE B CRITICAL LOAD dual-cord, 2 physical paths 2(N+1) 2 systems, each with reserve SOURCE A SOURCE B MODULE A (N+1) MODULE B (N+1) CRITICAL LOAD tolerates failure + maintenance Red = single point of failure · Dashed green = reserve module · Navy/purple = independent A/B paths
Redundancy topologies: N (single path), N+1 (reserve in the same path), 2N (two independent systems), and 2(N+1) (two systems, each with reserve). Conceptual representation, not executive.
ConceptPractical meaning
NCapacity needed to support the design load
N+1One additional module beyond what is needed
2NTwo complete and independent systems
2(N+1)Two complete systems, each with an additional reserve
Concurrent maintenanceRemove a component or path in a planned manner without stopping the load
Fault toleranceA failure does not cause immediate interruption of the critical load
  • Two devices fed by the same panel do not form independent paths.
  • Logical redundancy without physical separation can fail from a single event.
  • A single-cord load requires an STS or a specific architecture; dual-cord loads must have consistent A/B distribution.
  • Rated battery autonomy varies with load, temperature, age, and discharge regime.

Classes and reference standards

The ISO/IEC 22237 series classifies infrastructure by availability, security, and efficiency across the life cycle. The Uptime Institute uses Tier criteria oriented to topology and operational performance. Levels and classes from different reference standards should not be treated as automatically equivalent: the contractual requirement must indicate the standard, edition, scope, and method of proof.

IntentTypical architectureValidation question
Basic capacitySingle path and sized componentsDoes a planned maintenance require shutdown?
Redundant componentsN+1 in parts of the chainIs there still a single path?
MaintainabilityMultiple paths and isolationCan each component be maintained with an active load?
Fault toleranceSegregation and automatic responseDoes an isolated failure interrupt the load?

Caution: do not advertise certification, class, or Tier based solely on a diagram. Certification depends on the program, scope, documentation, construction, and/or operation as evaluated by the relevant authority.

Dual-path electrical architecture

The A/B principle creates two power paths up to the critical load. To be effective, it must consider source, generation, transformation, switching, UPS, distribution, cables, PDU, and equipment connection. Shared components must be identified and justified — not discovered during a failure.

GRID A utility GRID B independent source GMG A generator + fuel GMG B generator + fuel UPS A battery + bypass UPS B battery + bypass PDU A metering and protection PDU B metering and protection CRITICAL LOAD dual-cord A/B rack
Dual path A/B: two independent paths from the utility to the rack. Conceptual representation, not executive.
  • N capacity during maintenance and under the defined failure condition.
  • Selectivity and protection coordination across all operating modes.
  • Testable interlocks and transfer logic.
  • Physical separation against fire, water, and intervention.
  • Metering points from the utility to the rack.
  • Safe bypass plan and return to normal condition.

From IT inventory to design power

Sizing begins with an inventory of equipment, measured or declared power, coincidence factor, usage profile, growth, and reserves. Nameplate power is not average consumption — but the design also cannot ignore peaks, transients, simultaneous boot, battery recharge, and degradation.

StageReference calculation
Initial IT loadSum per equipment × utilization and coincidence factor
GrowthAnnualized scenarios or by installation waves
Rack capacityElectrical, thermal, physical, and floor limits — the lowest applies
IT thermal loadApproximately equal to the electrical power dissipated in steady state
InfrastructureIT + cooling + losses + lighting + auxiliaries
ReserveExplicit margin, without duplicating conservative factors

Example: 100 racks × 8 kW = 800 kW of IT load. With a design PUE of 1.45, the average total reference power would be 1.16 MW. The electrical system must still account for failure modes, maintenance, startups, and rated capacity.

Density and rack strategy

Density can be expressed in kW/rack, kW/m² of room, or per computing unit. The average hides hotspots: an 8 kW/rack room can contain 30 kW islands. The design must separate zones, define per-rack limits, and reserve evolution paths for liquid cooling.

Indicative rangeTypical applicationThermal strategy
Up to 5 kW/rackConventional IT or low occupancyOrganized aisles and bypass control
5 to 15 kW/rackVirtualization and storageContainment and units with modulation
15 to 30 kW/rackDense compute and moderate acceleratorsIn-row/rear-door or highly controlled air
Above 30 kW/rackHPC and high-density AIEvaluate direct-to-chip, CDU, and technology water

The ranges are indicative. The decision depends on the required flow rate, equipment ΔT, available pressure, layout, thermal tolerance, and manufacturer specification.

  • Width and depth compatible with equipment and cabling.
  • Aisles for circulation, maintenance, and component removal.
  • Cable management without blocking exhaust.
  • Blanking panels and sealing of openings.
  • Grounding, bonding, and labeling.

Site selection

CriterionDue diligence questions
PowerFirm capacity? Voltage? External redundancy? Connection lead time? Quality?
TelecomHow many carriers? Truly diverse routes? Opposite entrances?
EnvironmentalClimate, water, noise, emissions, permits, flooding, and neighborhood
CivilSoil, load, vibration, height, access, and equipment logistics
OperationsLabor, fuel, parts, security, and response time
ExpansionArea, utilities, phases, interferences, and continuity during construction

In an existing building (brownfield), perform an as-built survey, scanning when applicable, load testing, panel and cable inspection, structural analysis, and verification of water above or beside critical areas. The invisible constraint often defines the design: shaft, exhaust route, ceiling height, floor load, fire protection, or rigging access.

Thermal fundamentals

CRAC cooling unit installed next to rows of racks
CRAC unit installed in a data center environment. Photo: Robert Harker/Wikimedia Commons, CC BY-SA 3.0.

Nearly all electrical energy consumed by IT converts into heat. The thermal system's task is to capture that heat at the source, transport it, and reject it to the environment, keeping the equipment's inlet within the defined range. Average room temperature does not reveal recirculation, bypass, or hotspots.

PhenomenonEffectControl
RecirculationHot air returns to the rack inletContainment, sealing, and proper pressure
BypassCold air does not pass through the ITFlow adjustment, blanking panels, and layout
MixingReduces ΔT and efficiencyPhysical segregation of flows
Low flowInlet temperature risesCapacity, fans, and clear paths
Excessive flowWasted energy and undue pressureDemand-based modulation
Aisle between rows of enclosed racks with perforated raised floor for cold air supply
Cold aisle between rows of enclosed cabinets, with perforated raised floor for air supply. Photo: Robert Harker/Wikimedia Commons, CC BY-SA 3.0.

Environmental parameters

Environmental classes must be defined according to the equipment installed. As a widely used reference, ASHRAE TC 9.9 indicates a recommended range of 18 °C to 27 °C for classes A1 through A4; allowable limits vary by class and should not be confused with a permanent operating target. Humidity must be managed by dew point, condensation risk, electrostatics, and corrosion.

ParameterWhy measure itBest practice
Inlet temperatureThe air actually received by the ITSensors per rack and height, not just on the wall
Humidity and dew pointCondensation, ESD, and corrosionTrend and consistent alarms
Differential pressureConfirms flow directionMeasure between relevant zones
Particulates and corrosionElectronic reliabilityFiltration, sealing, and coupons when necessary
Water leakageEarly detectionSensors per zone and periodic testing

Tape equipment, batteries, and power components may have different limits. The manufacturer's specification and warranty conditions prevail.

Cooling technologies

SolutionWhen it makes sensePoints of attention
Direct expansion (DX)Small, modular, or distributed sitesRefrigerant, distance, redundancy, and partial efficiency
Chilled waterMedium and large scale, centralizationPumps, chiller, tower/dry cooler, water, and controls
In-rowHotspots and proximity to the loadHydraulic or refrigerant distribution in the room
Rear-doorRetrofit and high densityWeight, hoses, condensation, and maintenance
Direct-to-chipHigh-density CPU and GPUCDU, water quality, pressure, detection, and interfaces
ImmersionSpecific loads and very high densityFluid, maintenance, compatibility, and operation

Compare total cost, local climate, water availability, density, modularity, part-load efficiency, concurrent maintenance, area, noise, refrigerants, leak risk, and team capability. The best COP of an isolated unit does not guarantee the best system performance.

Thermal sizing and controls

Capacity must be verified for outdoor design conditions, sensible load, losses, redundancy, and degraded modes. Adding up rated capacities without confirming curves, water and air temperature, altitude, and flow rates creates false confidence.

VerificationAcceptance question
Thermal balanceHave all heat sources and scenarios been included?
PsychrometricsIs there control against condensation and excessive humidification?
HydraulicsAre pressure drop, valve authority, and balancing proven?
ControlsHave the sequence of operation and setpoints been reviewed?
Failure and maintenanceDoes the temperature remain acceptable under the defined event?
Dynamic testingHave load steps and transfers been simulated?

Metric: installed capacity is not enough. Track available, used, and committed capacity per zone, always considering the limiting component.

Critical electrical chain

The typical chain includes the utility, medium voltage, transformation, panels, generation, transfer, UPS, distribution, and rack. Each stage must be analyzed under normal condition, maintenance, failure, emergency, and return to normal.

ComponentFunctionCritical risk
Entry and MVReceive, switch, and protectArc flash, coordination, and external outage
TransformerAdjust voltage and isolateCapacity, temperature, harmonics, and fire
GMGSupply prolonged outageStartup, fuel, emissions, and heat rejection
ATS/STSTransfer sources and pathsLogic, synchronism, and switching failure
UPSCondition and sustain transitionBypass, battery, overload, and common-mode failure
PDU/RPP/buswayDistribute and meterSelectivity, expansion, and connection error

Generator sets and fuel

The generator must meet steady-state load and transient steps, including motors, UPS, recharge, cooling, and auxiliaries. The analysis must consider startup sequence, voltage and frequency dip, short-circuit capacity, altitude, temperature, fuel, and emissions.

TopicCriterion
AutonomyDefined by risk, refueling contract, and crisis scenario
TanksUsable capacity, containment, detection, transfer, and fuel quality
StartupRedundant batteries and chargers, pre-heating, and testing
ParallelingControl, synchronism, load sharing, and protection
MaintenanceAccess, parts, load bank, and planned downtime
Exhaust and noiseBack pressure, temperature, dispersion, acoustics, and permits

Meaningful test: a weekly no-load startup does not prove the chain. The program must include tests under load, source failure, transfer, stability, operational autonomy, and return, with calibrated instruments and clear criteria.

UPS, batteries, and energy storage

TechnologyAdvantageAttention point
Double conversionLoad isolation and broad applicationLosses, bypass, and partial efficiency
ModularScalability and per-module maintenanceCommon control and bus failures
VRLAMaturity and initial costTemperature, life, weight, and monitoring
Lithium-ionService life, cycles, and smaller area/weightBMS, thermal propagation, and fire strategy
FlywheelHigh power and many cyclesShort autonomy and integration with generation

To size autonomy, consider the actual discharge curve, constant power, minimum voltage, temperature, aging, manufacturing tolerance, and growth margin. Required autonomy should not be chosen by habit: it covers the interval between the loss of source and generation stabilization, or a controlled shutdown strategy.

Safety: battery rooms and systems require specific evaluation of ventilation, detection, containment, access, PPE, and emergency response according to the technology and applicable regulations.

Protection, selectivity, and power quality

The electrical study must include load flow, short circuit, coordination and selectivity, voltage drop, grounding, surge protection, harmonics, and, where applicable, arc-flash incident energy. The studies need to reflect all operating modes, including on generator.

ProblemSymptomDesign treatment
HarmonicsHeating and distortionMetering, filters, transformers, and suitable conductors
Low power factorOccupied capacity and penaltiesCoordinated correction and resonance analysis
TransientsReset or damageCascaded surge protection and bonding
Lack of selectivityTrips a larger area than the faultCurves, settings, and protection testing
ImbalanceHeating and loss of capacityPhase distribution and monitoring

Plan meters at the utility, generation, UPS, cooling, distribution, and rack. Without granularity, PUE cannot be calculated with confidence, nor can loss, imbalance, or idle capacity be located.

End-to-end connectivity

Having two carrier contracts does not guarantee diversity. The fibers may share a pole, duct, box, building entrance, or room. The analysis must map physical routes, points of presence, meet-me rooms, cross-connects, active equipment, and power feed.

CARRIER 1 north physical entrance CARRIER 2 south physical entrance MEET-ME ROOM A segregated room MEET-ME ROOM B segregated room CORE A routing and edge CORE B routing and edge LEAF / ToR A-B uplinks to both cores
Real physical diversity: independent entrances, meet-me rooms, and cores down to the leaf. Conceptual representation, not executive.
  • Separate and documented physical entrances.
  • Segregated meet-me rooms when criticality requires it.
  • Internal A/B routes down to rooms and racks.
  • Sized, grounded, and protected trays and ducts.
  • Permanent labeling and as-built documentation.
  • Optical and copper testing with acceptance reports.

Cabling and network architecture

LayerOptionsCriteria
Campus and carrierSingle-mode fiberDistance, route, splicing, availability, and SLA
Internal backboneOS2, OM4, or OM5 depending on applicationSpeed, distance, and evolution
HorizontalCopper or fiberLength, density, PoE, EMI, and management
Core/spineHigh capacity and redundancyDomain failure, convergence, and automation
Leaf/ToR/EoRServer connectionCabling, ports, oversubscription, and maintenance
OOBOut-of-band networkIndependence, security, and access during failures

Size fibers, ports, trays, and patch panels with explicit growth. The data rate depends on the standard, transceiver, distance, optical loss, and connectors. The design must also separate production, management, storage, security, and building automation networks.

Fire protection

The strategy combines prevention, early detection, compartmentalization, material control, suppression, and response. No single technology eliminates the risk. The design must be approved by the relevant authorities and coordinated with architecture, electrical, cooling, and operations.

LayerExamplesObjective
PreventionHousekeeping, thermography, maintenance, materialsReduce probability
DetectionSpot, aspirating, thermal, and alarmsIdentify early and locate
PassiveCompartmentalization, sealing, and doorsLimit spread
SuppressionSprinkler/pre-action, water mist, clean agentControl or extinguish
ResponseDefined EPO, brigade, fire department, continuityProtect people and service

Sealing: when agent effectiveness depends on retention within the enclosure, the design must provide integrity, sealing, pressure relief, and a leak-tightness test per applicable criteria.

Physical security and asset protection

ZoneTypical control
PerimeterBarriers, lighting, detection, and surveillance
ReceptionIdentification, authorization, logging, and escort
Technical areasAccess profiles, two-factor authentication, and interlocking
IT roomLeast privilege, logs, and customer segregation
Racks and cagesLock, seal, sensor, and inventory
Media and freightChain of custody, inspection, and secure disposal

For CCTV and access control, define coverage, resolution, retention, time synchronism, redundancy, privacy, and investigation. Access events must correlate person, credential, door, camera, and work order. Anti-passback and mantraps need contingency and life-safety procedures.

Controllers, cameras, BMS, and DCIM are digital assets: they require segmented networks, individual accounts, MFA where supported, hardening, patch management, tested backups, and logs sent to a protected repository.

BMS, EPMS, and DCIM

SystemFocusExamples
BMSBuilding and environmental automationCooling, leaks, pressure, doors
EPMSPower and power qualityMeters, breakers, UPS, GMG, trends
DCIMData center capacity and assetsRack, port, power, space, workflow
ITSM/NOCIncidents, changes, and serviceTickets, escalation, SLA, and knowledge
  • Each alarm needs priority, delay, hysteresis, and procedure.
  • Consequential alarms should be grouped to avoid alarm storms.
  • Time synchronism is mandatory for causal analysis.
  • The operator needs to know what to do, not just what happened.
  • Periodic tests must prove sensor, communication, logic, and notification.

Useful data: historical trending turns monitoring into management — anticipating efficiency loss, battery degradation, thermal saturation, and load imbalance.

Construction models

ModelCharacteristicRisk to manage
Design-bid-buildDesign completed before construction contractingInterfaces and late changes
Design-buildIntegrated responsibilityRequirements governance and verification independence
EPC/turnkeyDelivery concentrated in a single contractorScope, performance, and cost transparency
Construction managementPackages coordinated by the managerIntegration and responsibility across contracts
Modular/prefabricatedOff-site manufacturingFAT, transport, interfaces, and assembly

The design must anticipate construction sequencing, lifting, equipment entry, storage, temporary areas, energization, cleaning, asset protection, and future replacements. In an active environment, method of procedure, window, rollback, and isolation become part of the design.

Materials, execution, and QA/QC

Materials must meet mechanical, electrical, thermal, fire, corrosion, and maintenance performance. "Equivalent" substitutions can only be accepted after formal comparison of specification, interfaces, certifications, pressure drop, efficiency, curves, and warranty.

DisciplineInspection points
Civil and architectureLevels, load, waterproofing, sealing, finish, and cleaning
ElectricalTorque, labeling, insulation resistance, grounding, protection, and parameterization
MechanicalCleaning, flushing, leak-tightness, insulation, balancing, and drainage
TelecomBend radius, segregation, certification, labeling, and as-built
AutomationPoints, scales, alarms, fail-safe, integration, and history
FireZones, detectors, interfaces, discharge/flow, and legal documentation

Use Inspection and Test Plans with hold points, witness points, criteria, evidence, and responsible parties. Do not leave pending items unclassified: the punch list must record criticality, owner, deadline, acceptance blocking, and closure evidence.

Tiered commissioning

Commissioning is a performance assurance process that begins with the requirements, not a test done on the last day. It verifies documentation, installation, startup, controls, capacity, and interaction between systems.

LevelExample activity
L0 — requirements and designReview of OPR, BoD, sequences, and testability
L1 — factoryFAT, certificates, inspection, and configuration
L2 — installationPhysical inspection, torque, cables, piping, and labeling
L3 — startupIndividual startup and functional testing
L4 — systemCapacity, alarms, redundancy, and operating modes
L5 — integratedChained failures, utility loss, black building, and recovery

An integrated system test (IST) must use an approved script, pre-conditions, calibrated instruments, risk analysis, abort criteria, defined roles, communication, and synchronized data collection. The goal is to demonstrate behavior and recovery, not trigger an uncontrolled failure.

Documentation and handover

DeliverableMinimum content
As-builtDrawings, diagrams, routes, lists, and final revisions
O&MManuals, routines, limits, parts, and contacts
StudiesShort circuit, selectivity, thermal, hydraulic, and risk
ConfigurationsSetpoints, firmware, backups, and communication matrix
TestsFAT, SAT, functional tests, IST, and open items
TrainingContent, attendance, evaluation, and simulations
AssetsTag, model, serial, warranty, criticality, and spare parts

Golden rule: do not accept a system that merely "turns on." Accept a system whose capacity, protection, control, alarms, failures, and recovery are documented and demonstrated.

BIM models, the asset base, and DCIM data only add value if there is update governance. Define the official source, responsible party, frequency, integration, and audit.

Critical operation

Design reliability deteriorates when changes, maintenance, and knowledge are not controlled. Mature operation combines governance, procedures, training, communication, and continuous improvement.

ProcessEssential control
ChangeRisk assessment, approval, MOP, rollback, and post-validation
MaintenancePlan by criticality, window, spare parts, and evidence
IncidentCommand, communication, escalation, logging, and RCA
CapacityCurrent, committed, reserved, and forecast per domain
SupplierSLA, access, competence, parts, and performance
TrainingQualification, simulation, and internal recertification

The SOP describes normal operation; the MOP, a planned intervention; the EOP, the response to an abnormal event. All must have prerequisites, roles, numbered steps, stopping points, validation, and version control.

Indicators and efficiency

PUE is the ratio between the total energy of the data center and the energy of the IT. It is useful when boundaries, period, metering, and conditions are consistent. It does not measure availability, carbon, water, IT productivity, or a single server's isolated efficiency.

IndicatorUseCaution
PUEInfrastructure energy efficiencyCompare with equivalent boundaries and periods
WUEWater use associated with operationDefine methodology and water context
CUEEmissions associated with energyDepends on the emission factor
AvailabilityTime able to provide serviceDefine exclusions and scope
MTTRRecovery speedDo not hide recurrence
Stranded capacityIdle resource due to another limitAnalyze power, thermal, space, and network
  • Establish a baseline and metering quality.
  • Address bypass, recirculation, and auxiliary loads before raising setpoints.
  • Use sequencing and variable speed.
  • Test changes in waves, with limits and rollback.
  • Convert savings into avoided cost and released capacity.

AI, HPC, and liquid cooling

Immersion cooling tank with servers and colored cabling
Real immersion cooling system for servers. Photo: Rolf Brink/Wikimedia Commons, CC BY-SA 4.0.

AI loads concentrate power, network, and heat into a few racks. The challenge stops being just total capacity and comes to include transients, power quality, physical distribution, flow rate, return temperature, CDU redundancy, and maintenance of liquid circuits close to the electronics.

TopicDesign questions
PowerWhat peak, ramp, and diversity per cluster?
NetworkWhat topology, latency, cabling, and space for optics?
LiquidWhat water class, temperature, pressure, quality, and responsibility?
CDURedundancy, heat exchange, pumps, control, and bypass?
RiskDetection, containment, dripless connectors, and response?
OperationHow to drain, purge, maintain, and expand without stopping?

Even if the first phase is air cooled, reserve routes, technical area, structural load, hydraulic points, electrical capacity, and instrumentation for high-density pods. Improvised retrofits tend to create dependencies and downtime.

Deployment roadmap

PhaseVerifiable outputs
0. StrategyBIA, build/buy options, site shortlist, and business case
1. Preliminary studyLoad, area, risks, concept, schedule, and budget
2. Conceptual designArchitectures, diagrams, criteria, and estimate
3. Basic and executive designCalculations, details, specifications, BoQ, and test plans
4. ContractingRFP, equalization, responsibilities, and milestones
5. ConstructionQA/QC, safety, changes, and evidence
6. CommissioningTests by level, IST, training, and punch list
7. OperationHandover, baseline, SLA, maintenance, and improvement

Each phase must end with approval criteria. The gate does not exist to add bureaucracy, but to prevent critical uncertainties from advancing and turning into expensive changes. Unmet requirements must be corrected, formally accepted as risk, or removed from scope with the impact recorded.

Checklist for RFP and contracting

BlockItems
ScopeBoundaries, interfaces, inclusions, exclusions, and responsibilities
PerformanceCapacity, autonomy, efficiency, availability, and environment
DesignStandards, deliverables, review, BIM/CAD, and change control
SupplyAccepted brands, equivalence, FAT, logistics, warranty, and parts
ConstructionPlan, safety, QA/QC, ITP, schedule, and active environment
TestingSAT, functional tests, load bank, IST, and acceptance criteria
OperationTraining, O&M, as-built, support, and SLA
  • Compare technical exceptions, not just total price.
  • Normalize efficiency, usable capacity, redundancy, and autonomy.
  • Evaluate lead time, obsolescence, and local support.
  • Require an item-by-item compliance matrix.
  • Tie payment to milestones and performance evidence.

Operational readiness checklist

QuestionExpected evidence
Does the team know the sequence of operation?Recorded training and simulation
Do procedures reflect the as-built?Reviewed and version-controlled SOP, MOP, and EOP
Do alarms have a defined response?Alarm matrix and playbooks
Are critical parts available?Inventory, preservation, and replenishment
Have configuration backups been tested?Proven restoration
Do contracts cover the 24x7 scenario?SLA, contacts, and escalation
Is there a baseline before occupancy?Power, thermal, network, and environment recorded

Result: the data center is ready when people, processes, documentation, tools, and suppliers can operate and recover the infrastructure — not merely when construction ends.

Quick formulas and conversions

TopicReference relation
Three-phase powerP ≈ √3 × V × I × PF × efficiency
EnergykWh = kW × hours
Thermal load1 kW ≈ 3,412 BTU/h
Ton of refrigeration1 TR ≈ 3.517 kW thermal
PUEtotal data center energy ÷ IT energy
Simple availabilityMTBF ÷ (MTBF + MTTR)
Compound growthfuture capacity = current × (1 + rate)^years

The relations above are simplifications for estimation. Executive design requires curves, correction factors, environmental conditions, losses, harmonics, coincidence, and normative criteria.

Conceptual sizing example: IT load of 600 kW, growth to 900 kW in five years, and modular expansion of 150 kW. The strategy can deploy 4 active modules + 1 redundant in the initial phase, reserving space and primary infrastructure for future modules. The analysis must verify whether the system maintains N during maintenance and failure, and whether the auxiliary loads keep pace with each stage.

Always consult the current edition, the contracted scope, and local legal requirements. The references below guide the study but do not replace acquisition, full reading, and application by qualified professionals.

ReferenceApplication
ISO/IEC 22237Data centre facilities and infrastructures — concepts, building, power, environment, telecom, security, and operation
ISO/IEC 30134Key performance indicators for data centers, including PUE
ASHRAE TC 9.9Thermal Guidelines for Data Processing Environments
Uptime Institute Tier StandardTopology criteria and certification programs
ANSI/TIA-942Telecommunications infrastructure for data centers
NFPA 70 / 75 / 76 / 855Electrical, IT and telecom equipment protection, and energy storage systems, as applicable
ABNT NBR 5410 / 14039 / 5419LV and MV electrical installations and lightning protection
ABNT NBR 14565Structured cabling for commercial buildings and data centers
ABNT NBR 17240 and applicable fire standardsDetection and alarm, plus local Fire Department requirements

Brazilian legislation, ABNT standards, Fire Department instructions, and environmental requirements must be confirmed for the location and date of the project.

Conclusion

Designing a data center means managing consequences. The best solution is not the one that accumulates the most redundant equipment, but the one that connects criticality, risk, architecture, execution, testing, and operation in a coherent and demonstrable way.

  • Hold a requirements and BIA workshop.
  • Consolidate inventory, capacity, and expansion horizon.
  • Perform due diligence on the site and utilities.
  • Produce a conceptual design with alternatives and TCO.
  • Define commissioning criteria before contracting.
  • Plan handover and operation from the start.

EnQ Digital: intelligent infrastructure for companies that cannot stop — strategy, engineering, deployment, operation, and evolution oriented toward risk and performance.