inhousefyi
← Back to listings

LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect LimitedSeattle, Washington, United States · Posted 4 months ago
Full-timeEst. 141,000 USD
Apply now

Description

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.

Responsibilities:

  • Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
  • Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
  • Automate checkpointing and failure recovery during month-long training runs.

Required Skills:

  • Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
  • Experience managing SLURM or Kubernetes-based GPU clusters.
  • Strong systems engineering background (C++, CUDA, Python).

Similar jobs

Hyphen Connect LimitedOregon, United States

Est. 141,000 USD

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…

Full-time
Hyphen Connect LimitedSan Francisco, California, United States

Est. 140,000 USD

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…

Full-time
Hyphen Connect LimitedBoston, Massachusetts, United States

Est. 140,000 USD

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…

Full-time

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…

Full-time
Hyphen Connect LimitedSingapore, Singapore

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…

Full-time

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…

Full-time
Hyphen Connect LimitedHong Kong, Hong Kong

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…

Full-time
Scale AILondon, United Kingdom

Est. 120,000 GBP

As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs. Our platform powers cutting-edge research and production systems, supportin…

Full-time
Hyphen Connect LimitedSeattle, Washington, United States

Est. 141,000 USD

We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…

Full-time
TenstorrentBoston, Massachusetts, United States

Est. 250,000 USD

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify inn…

Full-time
Hyphen Connect LimitedOregon, United States

Est. 140,000 USD

We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…

Full-time
Scale AISan Francisco, California, United States

Est. 202,500 USD

As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs. Our platform powers cutting-edge research and production systems, supportin…

Full-time
Lightning AIRemote

Est. 237,500 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-timeRemote
Scale AISan Francisco, California, United States

Est. 297,900 USD

AI is becoming vitally important in every function of our society. At Scale, our mission is to accelerate the development of AI applications. For 9 years, Scale has been the leading AI data foundry, helping fuel the most…

Full-time
Scale AISan Francisco, California, United States

Est. 326,700 USD

Scale's LLM post-training platform team builds our internal distributed framework for large language model training. The platform powers MLEs, researchers, data scientists, and operators for fast and automatic training a…

Full-time
FigureSan Jose, California, United States

Est. 300,000 USD

Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of…

Full-time
Lightning AIPuyallup, Washington, United States

Est. 85,000 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-time
Hyphen Connect LimitedSan Francisco, California, United States

Est. 141,000 USD

We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…

Full-time
Lightning AILondon, England, United Kingdom

Est. 200,000 GBP

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-time

We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…

Full-time
Hyphen Connect LimitedBoston, Massachusetts, United States

Est. 141,000 USD

We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…

Full-time

We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…

Full-time
Lightning AINew York, New York, United States

Est. 127,500 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-time
HarkSan Jose, California, United States

Est. 315,000 USD

About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…

Full-time
Scale AILondon, United Kingdom

Est. 120,000 GBP

Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and…

Full-time
Lightning AILondon, England, United Kingdom

Est. 200,000 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-time
Lightning AIRemote

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-timeRemote
NebiusRemote

Est. 120,000 EUR

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-timeRemote
Lightning AINew York, New York, United States

Est. 215,000 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-time
NebiusRemote

Est. 120,000 EUR

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-timeRemote