We are recruiting for a AI/HPC Consultant for a leading IT service provider based in London.
The AI/HPC Consultant will support the definition, design and validation of advanced AI infrastructure solutions based on NVIDIA reference architectures. The role is heavily focused on high-performance networking and physical connectivity design across GPU-accelerated platforms, with regular involvement in bill-of-materials reviews, design assurance and customer-facing technical guidance.
Typical activities
- Validate NVIDIA reference architecture choices against customer scale, rack layout, cabling distances, data centre constraints and growth expectations.
- Review proposed switch, optic and cable selections to confirm correct speed, form factor, reach, fibre type, connector type and supported use case.
- Compare design options such as InfiniBand versus Spectrum-X Ethernet Back End, MMF versus SMF, DAC versus optical and centralised versus distributed switch placement.
- Produce clear diagrams, port maps, cable matrices, BoM annotations and design assumptions to support sales, delivery and customer review.
- Engage with NVIDIA and OEM partners to validate emerging design patterns, component availability and product-specific constraints.
- Support handover into delivery, including explanation of design decisions, known constraints, dependencies, risks and operational considerations.
- Desirable certifications and background
- NVIDIA certifications or demonstrable equivalent experience in AI infrastructure, networking or data centre technologies.
- Cisco CCNP/CCIE, NVIDIA Networking, HPE, Aruba, Juniper, Arista or other relevant data centre networking certifications.
- Experience in professional services, consultancy, pre-sales architecture or customer-facing technical design roles.
- Experience with high-performance storage platforms used in AI/HPC environments.
- Familiarity with automation, monitoring, observability or infrastructure-as-code tooling used to deploy and operate data centre fabrics.
Essential experience and skills
- Strong networking background with demonstrable experience in data centre, HPC, AI infrastructure or large-scale low-latency fabric design.
- Excellent knowledge of high-performance AI/HPC networking technologies, including InfiniBand, RoCEv2, Spectrum-X Ethernet, congestion control concepts and lossless Ethernet design principles.
- Strong understanding of VXLAN EVPN, BGP, leaf-spine fabric design, routing design, resiliency models and operational troubleshooting in modern data centre networks.
- Good understanding of physical infrastructure design, including switch placement, row-level design, structured cabling, patching strategy, fibre polarity, cable routing and data centre implementation constraints.
- Practical knowledge of optical and copper connectivity choices, including MMF, SMF, DAC, ACC/AEC, breakout cables, MPO/MTP, LC, QSFP, QSFP112, OSFP and QSFP-DD form factors.
- Ability to interpret and validate NVIDIA reference architectures and translate them into customer-specific designs and commercially accurate Bills of Material.
- Strong written communication skills, including the ability to produce clear HLDs, LLDs, design notes, BoM justifications and customer-facing technical explanations.
- Confident stakeholder engagement skills, including working with customers, partners, vendors and internal delivery teams.
- Analytical mindset with the ability to spot design risks, interoperability issues, unnecessary cost, unavailable components or avoidable cabling complexity.
Technical capabilities
- NVIDIA AI/HPC platforms
- GB300, B300, DGX, HGX, NVL72, GPU node connectivity, ConnectX, BlueField/SuperNIC concepts and NVIDIA reference architecture interpretations
- InfiniBand Back End fabrics
- NDR/XDR/HDR concepts, rail-optimised design, spine/leaf sizing, congestion and resiliency considerations, and Quantum-class switching such as QM3400.
- Spectrum-X Ethernet
- Spectrum-X Front End and Back End use cases, RoCEv2, multi-plane designs, Spectrum switching such as SN5610, Ethernet optics and cabling patterns.
- Data centre networking
- Leaf-spine design, VXLAN EVPN, BGP, resilient routing, underlay/overlay separation, management, storage and in-band/out-of-band network segmentation.
- Optics and cabling
- MMF versus SMF, DR4/FR4/LR4, SR4, DAC/ACC/AEC, breakout/splitter assemblies, MPO-12/APC, LC duplex, reach limits, form factors and compatibility risks.
- BoM optimisation
- Ability to identify practical cost, availability and implementation improvements without compromising performance, supportability or reference architecture alignment.
Essential
- Physical infrastructure
- Rack/row placement, switch placement, cable length planning, patching, cooling impact, power density and layout considerations for high-density AI clusters.
Strongly preferred
- Liquid cooling
- Awareness of direct liquid cooling, in-row versus in-rack CDU options, operational constraints and how cooling decisions affect rack layout and networking design.
Desirable
- Storage and management integration
- Understanding of high-speed storage connectivity, management fabrics, OOB networks and how they integrate with AI infrastructure designs.
This is an umbrella contract, the role is Inside IR35