C&I Engineer – Open
April 14, 2025EC&I Engineer – OPEN
May 12, 2025Our client is looking for a senior experienced professional tasked with vetting our next data-center sites for AI and GPU workloads. They must master several key areas to ensure the selected locations meet the unique demands of high-performance computing, scalability, and operational efficiency.
1. Power Infrastructure Assessment
- Expertise in evaluating electrical capacity, reliability, and redundancy (e.g., access to high-voltage grids, backup generators, and UPS systems).
- Understanding of power usage effectiveness (PUE) and ability to estimate energy costs for GPU-intensive AI workloads.
- Knowledge of renewable energy options or local utility incentives to optimize long-term sustainability and cost.
- Experience working with power companies and getting power agreements / commitments
2. Cooling and Thermal Management
- Deep understanding of cooling requirements for high-density GPU clusters, including liquid cooling, air conditioning, or hybrid systems.
- Ability to assess site-specific factors like ambient climate, humidity, and airflow to ensure efficient heat dissipation.
- Experience in calculating cooling costs and scalability for future expansion.
3. Network Connectivity and Latency
- Proficiency in evaluating fiber optic infrastructure, bandwidth availability, and proximity to internet exchanges for low-latency data transfer.
- Understanding of network redundancy and resilience to support continuous AI model training and inference.
- Ability to assess telecom provider options and negotiate service agreements.
4. Site Scalability and Space Planning
- Skill in analyzing physical space for current and future rack density, considering GPU server layouts and expansion potential.
- Knowledge of zoning laws, building codes, and land availability for constructing or retrofitting facilities.
- Experience in forecasting growth needs based on AI workload trends (e.g., larger models, more GPUs).
5. Risk Evaluation and Environmental Factors
- Ability to identify risks such as natural disasters (earthquakes, floods, hurricanes) that could disrupt operations.
- Expertise in assessing local environmental regulations, noise restrictions, or community impact concerns.
- Skill in incorporating contingency plans for power outages, network failures, or supply chain disruptions.
6. Cost Analysis and Budgeting
- Mastery of estimating total cost of ownership (TCO), including land acquisition, construction, energy, cooling, and maintenance.
- Ability to balance upfront capital expenditures (CapEx) with operational expenses (OpEx) for long-term viability.
- Experience in benchmarking costs against industry standards and alternative sites.
7. Stakeholder Collaboration and Communication
- High EQ, Strong interpersonal skills to work with engineers, architects, local governments, and utility providers during site evaluation.
- Ability to clearly articulate site pros and cons to decision-makers, aligning recommendations with business goals.
- Experience in managing negotiations for permits, tax incentives, or land deals.
8. Technical Standards and Compliance
- Knowledge of data-center standards (e.g., Uptime Institute Tiers, ASHRAE guidelines) and GPU-specific requirements (e.g., NVIDIA DGX compatibility).
- Familiarity with security protocols (physical and cyber) to protect sensitive AI data and intellectual property.
- A willingness to understand and problem solve issues with local regulations for energy usage, emissions, or data sovereignty.
9. Market and Location Intelligence
- Insight into regional advantages, such as proximity to talent pools, tech hubs, or research institutions for AI development.
- Awareness of the competitive landscape, including where other AI/GPU data centers are located and why.
- Ability to adapt site selection based on economic factors like tax breaks, labor costs, or energy price trends.
10. Attention to Detail and Problem-Solving
- Precision in reviewing site data (e.g., geotechnical surveys, utility reports) to avoid costly oversights.
- Creative problem-solving to address challenges like limited power capacity or suboptimal cooling through innovative solutions (e.g., modular designs, edge computing integration).