We are building our own GPU infrastructure to run AI models, support AI-driven software development, and power AI capabilities in our software products. We are looking for a hands-on engineer to build this infrastructure in a colocation data centre and take responsibility for its ongoing operation.
You will turn workload requirements into a practical infrastructure setup-from selecting servers and coordinating installation to configuring Linux, deploying model-serving environments, and keeping systems secure and reliable.
Your Responsibilities
Plan and Build the Infrastructure
• Translate AI workload requirements into GPU, compute, storage, and networking configurations.
• Evaluate hardware compatibility, supplier proposals, and costs.
• Coordinate procurement, delivery, installation, and commissioning.
• Confirm rack space, power, cooling, connectivity, and access requirements with the colocation provider.
• Install and configure servers, networking equipment, cabling, and remote management.
• Maintain infrastructure documentation and asset inventories.
Configure GPU and AI Systems
• Configure and maintain Linux servers, GPU drivers, CUDA environments, and container runtimes.
• Deploy and maintain model-serving infrastructure in collaboration with the AI and engineering teams.
• Diagnose GPU, memory, storage, and network performance issues.
• Benchmark workloads and improve resource utilisation.
• Automate deployment and configuration to make environments reproducible.
Manage Networking and Security
• Configure network segmentation, firewalls, VPNs, and secure administrative access.
• Implement access controls, credential management, system hardening, and patching.
• Separate workloads and environments according to product and security requirements.
• Troubleshoot connectivity issues with infrastructure providers.
Own Reliability and Operations
• Establish monitoring and alerting for hardware health, GPU utilisation, capacity, and service availability.
• Implement backup and recovery procedures and test restoration.
• Maintain operational runbooks and troubleshoot infrastructure incidents.
• Coordinate hardware replacements, maintenance windows, and remote-hands support.
• Identify single points of failure and propose practical improvements.
• Establish clear support coverage and escalation procedures with the team.
Manage Capacity and Costs
• Track utilisation, operating costs, and capacity constraints.
• Recommend upgrades based on measured demand and performance.
• Assess trade-offs between owned hardware, rented GPU capacity, and cloud services.
• Plan expansion while avoiding unnecessary complexity and overprovisioning.
What You Bring
• Hands-on experience deploying and operating physical servers in a data centre or colocation environment.
• Strong Linux administration and troubleshooting skills.
• Practical experience with GPU servers, NVIDIA drivers, CUDA, and AI workloads.
• Solid networking knowledge, including switching, routing, VLANs, firewalls, and VPNs.
• Experience with containers, monitoring, backups, and infrastructure automation.
• An understanding of rack power, cooling, hardware compatibility, and remote server management.
• A disciplined approach to security, documentation, maintenance, and recovery.
• The ability to work independently and coordinate with vendors and data centre providers.
• Professional English for technical documentation and collaboration.
• Willingness to perform on-site installation and maintenance when required.
Useful Additional Experience
• Operating AI inference and model-serving systems.
• Managing multi-GPU workloads and high-speed networking.
• Using infrastructure-as-code and configuration-management tools.
• Supporting environments with workload isolation and availability requirements.
How You Will Work
• You will own the implementation and reliable operation of our AI infrastructure, working closely with the AI and
software engineering teams to align infrastructure with application and model requirements.
• This is a hands-on role covering both the initial build and ongoing operations, with direct involvement in model deployment, performance optimisation, capacity planning, and infrastructure decisions.
What Success Looks Like
• Infrastructure meets agreed workload, security, and operational requirements.
• AI workloads run reliably with measurable performance and utilisation.
• Monitoring, tested recovery procedures, and operational documentation are in place.
• Infrastructure costs are transparent and expansion decisions are based on evidence.
• Incidents are resolved effectively, with clear coordination across suppliers and service providers.