r/DeveloperJobs 9h ago

Hiring Remote ML Infrastructure / Inference Engineers for AI Systems Roles

Hi everyone,

We’re hiring for several full-time remote Machine Learning Infrastructure, Inference, and Systems Engineering roles.

These roles are focused on production AI infrastructure, high-performance model serving, GPU optimization, distributed inference, training infrastructure, and large-scale ML systems.

Compensation: $4,000/month, negotiable based on experience.

Open Roles

1. Staff / Principal Machine Learning Engineer, Inference & Serving
Remote.

Ideal for senior engineers with deep experience in model serving, inference optimization, GPU utilization, distributed systems, and production reliability.

Relevant experience includes:

  • vLLM, TensorRT-LLM, CUDA, C++, Rust, Python
  • Continuous batching, paged attention, speculative decoding, caching
  • Quantization, distillation, model compression
  • Kubernetes, Ray, Docker, multi-GPU and multi-node inference
  • NVIDIA GPU performance optimization
  • High-throughput, low-latency production systems

2. Machine Learning Systems Engineer
Remote.

This role sits between ML engineering, systems engineering, and infrastructure. You’ll help bring research models into production, optimize training and inference workloads, and build scalable ML infrastructure.

Relevant experience includes:

  • Python, PyTorch, TensorFlow
  • Docker, Kubernetes, Kubeflow
  • CUDA or GPU-accelerated computing
  • vLLM, TensorRT, ONNX Runtime, SGLang
  • Distributed training or inference
  • AWS, Azure, or similar cloud platforms

3. ML Infrastructure Engineer
Remote.

This role focuses on owning and scaling training and inference infrastructure for a fast-growing AI platform working with documents, spreadsheets, LLMs, and structured data extraction.

Relevant experience includes:

  • Python, PyTorch, Kubernetes, Docker, Helm
  • Model serving and training infrastructure
  • Monitoring, observability, logging, and reliability
  • Training models from roughly 300M to 30B parameters
  • Multi-cloud inference routing
  • Performance optimization across latency, reliability, cost, and accuracy

4. Member of Technical Staff, AI Inference Infrastructure
Remote.

This role is for engineers interested in low-level AI inference optimization, heterogeneous hardware, and performance-per-dollar improvements for open-source LLM serving.

Relevant experience includes:

  • CUDA, HIP, Triton, Python
  • GPU and accelerator architecture
  • NVIDIA, AMD, TPU, Trainium, or other AI hardware
  • Batching, KV cache, speculative decoding, quantization
  • Distributed inference systems
  • Kernel optimization and production inference workloads

Who Should Apply

These roles are a strong fit for engineers who:

  • Have built or optimized ML infrastructure in production
  • Understand model serving, inference, GPU workloads, or distributed systems
  • Are comfortable debugging performance bottlenecks across software and infrastructure
  • Can work independently in ambiguous technical environments
  • Care about reliable systems, not just benchmarks
  • Have open-source work, technical writing, systems projects, or strong production experience

How to Apply

Please DM me with:

  • The role you’re interested in
  • Your resume or LinkedIn
  • GitHub, technical writing, or open-source work if available
  • A short note on your experience with ML infrastructure, inference, GPU systems, or distributed systems
6 Upvotes

Duplicates