Senior Machine Learning Engineer, LLM Inference Optimization
Nebius Group
ph3About Nebius: /h3 /brpNebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. /p /brpBuilt by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. /p /brpListed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with RD hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI RD. /p /brh3The role /h3 /brpNebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, inference engines, serving architecture, and benchmarking, with a focus on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability. /p /brpThis is a hands-on role in which you will work on complex optimization projects, diagnose difficult serving problems, and deliver measurable improvements in production. Working closely with kernel and platform engineers, you will evaluate serving configurations, resolve performance and quality regressions, and optimize inference for real-world workloads, supported by reproducible benchmarks and safe production rollouts. /p /brh3Your responsibilities: /h3 /brul /brliOwn optimization work for specific model families, customer endpoints, or serving backends. /li /brliRun engine comparisons and recommend practical serving configurations for specific workloads. /li /brliDebug model quality or performance regressions during production rollouts. /li /brliOptimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. /li /brliDeploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. /li /brliBuild and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. /li /brliImplement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. /li /brliBuild reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. /li /brliPartner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. /li /brliWrite clear design docs, performance reports, rollout plans, and customer-facing technical explanations. /li /br /ul /brh3Must-haves: /h3 /brul /brliStrong Python and PyTorch engineering skills. /li /brliHands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems. /li /brliPractical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. /li /brliStrong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. /li /brliAbility to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. /li /brliStrong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. /li /br /ul /brh3Nice-to-have: /h3 /brul /brliExperience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques. /li /brliExperience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. /li /brliExperience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. /li /brliCUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. /li /brliOpen-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects. /li /br /ul /brh3Benefits Perks: /h3 /brul /brliCompetitive compensation /li /brliCareer growth and learning opportunities /li /brliFlexibility and ownership /li /brliCollaborative and innovative culture /li /brliOpportunity to work on impactful AI projects /li /brliInternational environment and talented teams /li /br /ul /brh3What's it like to work at Nebius: /h3 /brpFast moving- Bold thinking- Constant growth- Meaningful impact- Trust and real ownership- Opportunity to shape the future of AI /p /brh3Equal Opportunity Statement: /h3 /brpNebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. /p /brpApplicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. /p /brpIf you need accommodations during the application process, please let us know. /p /p #J-18808-Ljbffr
- pNebius Group in Zürich seeks a Senior Machine Learning Engineer to own model and endpoint optimization from artifacts through production deployment. You will work on model internals, inference engines, serving architecture, and benchmarking to improve latency, throughput...Senior
- ...Founded by former Google search engineers with PhDs in AI from ETH... ...RoleDeepJudge is seeking a Machine Learning Engineer to design, build, evaluate... ...ML: Train, evaluate, and optimize a wide range of machine... ...learning systems. Drive efficient inference and deployment while...Empfohlen
- ...first humanoid, AEON, was launched in June 2025 and is already in pilots with five customers. /p pWe are hiring a bSenior Machine Learning Engineer /b to join our Foundation Model team, focused on building, evaluating, and deploying state‑of‑the‑art foundation models for...Senior
- ...uncertainty is the norm and failure is not an option. As a Machine Learning Engineer, you will push the limits of computer vision by designing,... ...where it matters most. What You´ll Do Design, train, and optimize deep learning models for computer vision tasks in highly dynamic...Empfohlen
- ...Collaborate closely with Data Scientists, Data Engineers, Software Engineers, Product Managers,... ..., and governance throughout the machine learning lifecycle.Monitor and optimise model performance... ...such as PyTorch, TensorFlow, or Scikit-learn.Solid understanding of MLOps practices,...Empfohlen
- ...already in pilots with five customers. We are looking for a Machine Learning Engineer to develop and improve learning-based behaviors for our... ...proactive working style, with a strong drive to experiment, learn, and solve problem Nice to have Experience working with real...Flexible Arbeitszeit
- ...As an ML Research Engineer in the... ...them into highly optimized, task-specific components... ...tackling the hardware inference constraints of... ...Mentorship: Depending on seniority, support the team... ...). Deep Learning Computer Vision:... ...downstream users and senior stakeholders....Senior
- Bist Du ein begeisterter Software Engineer, der sein Fachwissen in einem spannenden und modernen Arbeitsumfeld einbringen möchte? SUMEX bietet Dir eine hervorragende Karrierechance, umgeben von erfahrenen, motivierten und dynamischen Teammitgliedern. Bei uns kannst Du...Im FrühlingHomeofficeFlexible Arbeitszeit
- ...execution are expected. About the Role As a Machine Learning Engineer on our Foundational team in Paris, you... ...multi-modal foundational models that learn robust representations of the... ...Minimum of 5-6 years of experience for senior levels. Experience training and scaling...Senior
- ...construction industry, turning heavy construction machines into autonomous robots. Gravis began... ...spin-out, and our unique combination of learning-based automation and augmented remote... ...industry. The Gravis RACK is a machine-agnostic retrofit kit that adds autonomy...SeniorVollzeitRemote job
- ...algorithms. /lili2 years of experience with machine learning algorithms and tools (e.g., TensorFlow... ...ulpAbout the job /ppGoogle's software engineers develop the next-generation... ...Feedback (RLHF), and iterative process optimization (IPO) establishing data flywheel, etc....
- Machine Learning Engineer, 80-100% Wir sind ein agiles Software-Entwicklungsunternehmen... ...und Orchestrierung von LLM-basierten Agenten in Produktionsumgebungen... ..., inklusive APIs, Inference‑Services, Workflow‑... ...in ML-Frameworks wie scikit‑learn, PyTorch oder TensorFlow vorweisen...HomeofficeFlexible Arbeitszeit
- ...generalisation.You benchmark learned approaches against a tuned classical... ....You work with FPGA and DSP engineers to translate successful... ...mathematics, applied physics, machine learning or a related... ...experience with PyTorch, scikit-learn or equivalent frameworksExperience...
- Summary Apple's Security Engineering Architecture organization is responsible for the security... ...to do so. We are seeking an Applied Machine Learning Engineer who will help us invent and... ...career where you feel like you belong. Learn about accessibility in Apple’s workplace...
- ..., turning heavy construction machines into autonomous robots. Gravis... ...our unique combination of learning-based automation and augmented... ...excavation that generalize across machine models and soil conditions... ...and motion planning engineers Build tools for analysing and...SeniorRemote job
- pNVIDIA Switzerland AG in Zürich is seeking a Senior Software Developer to join our AI networking acceleration team, developing a highly optimized inference framework and contributing to open-source work with hardware offloads and RDMA networking. /ppYou will work on Linux...Senior
- ...software design and architecture. /li li5 years of experience with Machine Learning (ML) models, ML infrastructure, Natural Language Processing... ...qualifications /h3 ul liMaster’s degree or PhD in Engineering, Computer Science, or a related technical field. /li li3 years...
- ...world's first general-purpose learning agent. Central to this... ...our prototypes. As a Software Engineer, you will be working with the... ...developed by our exceptional team of Machine Learning and Neuroscience... .../h3ulliDesign, develop, and optimize software tools to improve the...Senior
- ph3Senior Software Engineer, RTL Optimization, DeepMind /h3 h3Qualifications /h3 ul liBachelor’s degree... ...the world's first general-purpose learning agent. Central to this mission is the... ...developed by our exceptional team of Machine Learning and Neuroscience research scientists...Senior
- pGoogle DeepMind is seeking a software engineer to help design, develop, and optimize software tools for RTL generation and hardware compilation within our AI/EDA pipeline. You will collaborate with machine learning researchers, hardware architects, and software engineers...Senior
- ...equivalent practical experience. /liliExperience in computer vision, machine learning, and deep learning. /liliExperience with graphics and... ...humans. /liliCollaborate cross-functionally with other engineers, researchers, and technical artists, and product teams. /liliDesign...
- DeepJudge is seeking a Machine Learning Engineer to design, build, and refine AI systems powering our core products. You will own the ML lifecycle from research to production, including data analysis, modeling, implementation, and evaluation for customers.You will work...
- ...leading AI technology company in Zurich is seeking experienced engineers to innovate AI compiler technology for accelerated computing. You will lead compiler design and optimization efforts to enhance deep learning frameworks like PyTorch, collaborating closely with experts...Senior
- ...Description We are seeking an exceptional Staff Machine Learning Engineer to lead the development of audio and... ...and visual representations, mentor senior and junior engineers, and shape the... ...where you feel like you belong. Learn about accessibility in Apple's workplace...Senior
- ...back to production Optimize models for deployment... ...drone hardware, cloud inference, or both Spend time... ...practical depth in deep learning for visual perception:... ...alongside experienced engineers and researchers... ...and great coffee. Learn more about who we are,...SeniorVollzeit
- 3ap sucht einen Machine Learning Engineer, der an der cloudbasierten AI-in-a-Box-Plattform arbeitet und ML‑Algorithmen, Drittanbieter‑Modelle sowie API‑basierte Dienste integriert. Sie betreuen den kompletten ML‑Projektlebenszyklus von Forschung über Prototyping bis zur...Senior
- A Swiss company with nearly four decades of engineering pedigree behind the software that fights financial crime inside 1,500+ banks across... ...engineering that gets them there.Build and optimise real-time inference pipelines where latency and throughput are the constraint, not...
- pGoogle DeepMind is seeking a Senior Software Engineer specializing in RTL optimization to contribute to the software tools that accelerate hardware compilation and RTL generation. The role involves collaborating with ML researchers and hardware architects to optimize pipelines...Senior
- pSkydio Inc. in Zürich is hiring an Autonomy Engineer specializing in deep learning model acceleration for Computer Vision. You will work on high-performance solutions that improve ML inference and contribute to the development of groundbreaking autonomous flight technology...
- ph3Autonomy Engineer - Deep Learning Model Acceleration /h3pSkydio... ...deep learning inference for CV workloads that... ...bottlenecks and acceleration/optimization opportunities and... ...advanced Machine Learning knowledge to... ...employment eligibility. To learn more about E‑Verify,...
Wollen Sie mehr Stellenangebote erhalten?
Abonnieren Sie und erhalten Sie ähnliche Stellenangebote wie Senior Machine Learning Engineer, LLM Inference Optimization. Seien Sie der Erste, der sich bewirbt!
- analyste financier senior Zürich ZH
- juriste senior Zürich ZH
- senior compliance officer Zürich ZH
- senior graphic designer Zürich ZH
- senior operator Zürich ZH
- senior IT security consultant Zürich ZH
- senior berater private banking Zürich ZH
- senior catastrophe modeling analyst Zürich ZH
- pro seniore Zürich ZH
- senior research associate Zürich ZH
