Technical Lead, Machine Learning
Job Description:
Our client is an AI-native product company built to replace how billions of people manage their digital lives — starting with email, notes, and task tools that were never designed to be AI-native. It's backed by multi-million-dollar investment and built remote-first from day one, with a clear target: cut the time it takes users to get everyday things done by roughly 90%. That means solving the problems most AI products avoid — long-running workflows, persistent context, and reliable behavior under real-world, non-deterministic conditions. The team is small and high-talent-density by design, not by necessity, and moves at a pace that matches the scale of what it's building.
About the Role
As Technical Lead, Machine Learning, you own the execution layer of the company's intelligence — translating research direction into reliable, scalable, production-grade ML systems. This role sits at the intersection of research, infrastructure, and product: you're responsible for making models trainable, deployable, observable, and performant under real-world constraints.
What You'll Do
- Own end-to-end ML system execution: data pipelines, training workflows, evaluation systems, inference architecture, deployment
- Fine-tune and adapt models using state-of-the-art methods — LoRA, QLoRA, SFT, DPO, distillation
- Architect and operate scalable inference systems, balancing latency, cost, and reliability
- Design and maintain data systems for high-quality synthetic and real-world training data
- Implement evaluation pipelines covering performance, robustness, safety, and bias, in partnership with research leadership
- Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies
- Collaborate closely with application engineering to integrate ML systems cleanly into backend, mobile, and desktop products
- Make pragmatic trade-offs and ship improvements quickly, learning from real usage
- Work under real production constraints: latency, cost, reliability, and safety
Requirements
- You've built or shipped real ML systems used by people, not just demos
- You're comfortable working with large models and understanding their failure modes
- You write strong, production-grade code and care about system correctness
- You're self-directed, pragmatic, and take full ownership of outcomes
- You communicate clearly and collaborate well in small, high-trust teams
Location
Remote, open to candidates in Poland, Portugal, Spain, Sweden, Switzerland, Germany, Ireland, Indonesia, Seoul, or Singapore.
Tech Stack
Python · PyTorch / JAX · GPU-based training and inference systems
What Success Looks Like
- Research and models reliably translate into production-ready solutions with clear performance and quality targets
- ML pipelines, training loops, and inference systems are stable, efficient, and maintainable
- Production issues are detected, debugged, and resolved quickly, minimizing user impact
- Team members are supported, aligned, and able to deliver high-impact ML work with minimal friction
- Iterations on models and systems are measurable, safe, and improve user experience over time
What to Expect
The best products in the world are built by small, world-class teams. Decisions are made collectively, at rapid speed — balancing high-quality shipping with fast learning. You'll be expected to bring structure, exercise judgment, and execute independently.
Compensation and benefits are competitive and vary by location; the package includes base salary and equity, discussed openly with you as part of the process.
If there's a fit, expect 3–4 interviews total, followed by a prompt, transparent decision. To be considered, please make sure to answer the screening questions in the application form — we're not able to review submissions that skip them.