NVIDIA Logo

NVIDIA

AI Test Architect

Reposted 3 Days Ago
Be an Early Applicant
Remote
6 Locations
Expert/Leader
Remote
6 Locations
Expert/Leader
The AI Test Architect will profile, benchmark, and optimize deep learning models and training pipelines while focusing on high-performance networking for large-scale supercomputing solutions.
The summary above was generated by AI

We are looking for an AI Test Architect joining E2E Verification group to profile Innovative large scale Distributed training on NVIDIA AI End-to-End solutions in a large scale supercomputing clusters.

Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs. You will work with the latest Accelerated Computing and Deep Learning software and hardware platforms, with researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, Switch, HCA, CPU and GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.

What you’ll be doing:

  • Profiling, benchmarking, and analyzing deep learning models to identify areas for optimization and improvement in terms of performance, efficiency, and accuracy, with a strong emphasis on networking aspects.

  • Collaborating closely with data scientists, researchers, development, automation teams to design and implement scalable training pipelines and frameworks that demonstrate large scale high -performance networking capabilities.

  • Staying up-to-date with the latest advancements in deep learning algorithms, architectures, NVIDIA GPU technologies, and high-performance networking solutions.

  • Optimizing deep learning models for performance, memory usage, and power efficiency while maximizing high-performance networking features on NVIDIA supercomputers.

  • Providing insights and recommendations based on the analysis of large-scale training results, specifically focusing on networking bottlenecks and optimizations, to improve model outcomes and achieve business objectives.

  • Collaborating with hardware engineers to guide the development and integration of efficient networking solutions for deep learning, including exploring network architecture optimizations and bringing to bear technologies such as RDMA or InfiniBand.

What we need to see:

  • B.Sc in Computer Science, Software Engineering, or equivalent experience.

  • Strong understanding and practical experience with machine learning algorithms and techniques, with a specialization in deep learning and expertise in high-performance networking.

  • 8+ years of overall experience, with CUDA programming for deep learning frameworks like TensorFlow, PyTorch, combined with expertise in networking libraries and protocols.

  • Ability to profile and optimize deep learning workflows, focusing on networking-related bottlenecks and optimizations, to improve overall performance and efficiency.

  • Exceptional analytical and problem-solving skill, with a keen attention to detail, particularly in identifying and resolving networking performance issues.

  • Excellent communication and collaboration skills, enabling effective teamwork and cooperation.

  • Familiarity with supercomputers, parallel computing, distributed systems, and high- performance networking technologies like RDMA or InfiniBand.

Ways to stand out from the crowd:

  • Demonstrated experience in successfully profiling and optimizing large-scale deep learning training on NVIDIA supercomputers, with a significant focus on high-performance networking enhancements.

  • Experience with distributed deep learning, distributed training frameworks, or large-scale data pipelines enhanced by high-performance networking solutions.

  • Expertise in optimizing networking parameters, such as bandwidth, latency, or congestion control, for deep learning workloads.

  • Familiarity with NVIDIA's networking technologies, such as Mellanox InfiniBand, and their integration with deep learning workflows.

  • Strong understanding of high-performance networking protocols and standards and their application to deep learning.

Top Skills

Cuda
Infiniband
PyTorch
Rdma
TensorFlow

Similar Jobs

17 Hours Ago
Remote
18 Locations
Mid level
Mid level
Blockchain • Software • Cryptocurrency • NFT • Web3 • App development
The Blockchain Engineer will design, develop, and deploy smart contracts, ensure multi-chain interoperability, and build secure staking systems while collaborating with various teams.
Top Skills: AnchorEthers.JsFoundryHardhatNode.jsPythonRustSolidityTruffleTypescriptWeb3.Js
Yesterday
Remote or Hybrid
Italy
Mid level
Mid level
Artificial Intelligence • Hardware • Information Technology • Security • Software • Cybersecurity • Big Data Analytics
The Key Account Manager will manage existing and new customers, drive sales strategies, maintain relationships, and ensure effective communication to achieve revenue targets in Italy's telecommunications sector.
Top Skills: ExcelOutlookPowerPointWord
Yesterday
Remote or Hybrid
Rome, ITA
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
The Cloud Security Professional will enhance service security, support sales with customer inquiries, manage security incidents, and drive security initiatives and presentations.
Top Skills: AIAWSAzureCloud SecurityGCPItilItsmLlm

What you need to know about the Edinburgh Tech Scene

From traditional pubs and centuries-old universities to sleek shopping malls and glass-paneled office buildings, Edinburgh's architecture reflects its unique blend of history and modernity. But the fusion of past and future isn't just visible in its buildings; it's also shaping the city's economy. Named the United Kingdom's leading technology ecosystem outside of London, Edinburgh plays host to major global companies like Apple and Adobe, as well as a growing number of innovative startups in fields like cybersecurity, finance and healthcare.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account