Principal Machine Learning Engineer

  • -
  • Full-Time
  • On-Site

Job Description:

Our client is an AI-native product company built to replace how billions of people manage their digital lives — starting with email, notes, and task tools that were never designed to be AI-native. It's backed by multi-million-dollar investment and built remote-first from day one, with a clear target: cut the time it takes users to get everyday things done by roughly 90%. That means solving the problems most AI products avoid — long-running workflows, persistent context, and reliable behavior under real-world, non-deterministic conditions. The team is small and high-talent-density by design, not by necessity, and moves at a pace that matches the scale of what it's building.


About the Role


You'll be responsible for turning research direction into working, production-grade ML systems. This role owns the execution layer of the company's intelligence — training pipelines, inference systems, evaluation tooling, and deployment.


What You'll Work On


  • Build and own end-to-end ML pipelines spanning data, training, evaluation, inference, and deployment
  • Fine-tune and adapt models using state-of-the-art methods — LoRA, QLoRA, SFT, DPO, distillation
  • Architect and operate scalable inference systems, balancing latency, cost, and reliability
  • Design and maintain data systems for high-quality synthetic and real-world training data
  • Implement evaluation pipelines covering performance, robustness, safety, and bias, in partnership with research leadership
  • Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies
  • Collaborate closely with application engineering to integrate ML systems cleanly into backend, mobile, and desktop products
  • Make pragmatic trade-offs and ship improvements quickly, learning from real usage
  • Work under real production constraints: latency, cost, reliability, and safety


Requirements


  • Strong background in deep learning and transformer-based architectures
  • Hands-on experience training, fine-tuning, or deploying large-scale ML models in production
  • Proficiency with at least one modern ML framework (PyTorch, JAX), and ability to pick up others quickly
  • Experience with distributed training and inference frameworks (DeepSpeed, FSDP, Megatron, ZeRO, Ray)
  • Strong software engineering fundamentals — you write robust, maintainable, production-grade systems
  • Experience with GPU optimization, including memory efficiency, quantization, and mixed precision
  • Comfort owning ambiguous, zero-to-one ML systems end-to-end
  • A bias toward shipping, learning fast, and improving systems through iteration


Location


Remote, open to candidates in Poland, Portugal, Spain, Sweden, Switzerland, Germany, Ireland, Indonesia, Seoul, or Singapore.


Nice to Have


  • Experience with LLM inference frameworks such as vLLM, TensorRT-LLM, or FasterTransformer
  • Contributions to open-source ML or systems libraries
  • Background in scientific computing, compilers, or GPU kernels
  • Experience with RLHF pipelines (PPO, DPO, ORPO)
  • Experience training or deploying multimodal or diffusion models
  • Experience with large-scale data processing (Apache Arrow, Spark, Ray)


What to Expect


The best products in the world are built by small, world-class teams. Decisions are made collectively, at rapid speed — balancing high-quality shipping with fast learning. You'll be expected to bring structure, exercise judgment, and execute independently.


Compensation and benefits are competitive and vary by location; the package includes base salary and equity, discussed openly with you as part of the process.


If there's a fit, expect 3–4 interviews total, followed by a prompt, transparent decision. To be considered, please make sure to answer the screening questions in the application form — we're not able to review submissions that skip them.