About the Role
We're hiring a Software Engineer to join a fast ‑ moving , high ‑ ownership team building the next ‑ generation large language model (LLM) serving system. This role builds and optimizes the end-to-end path to get Microsoft AI (MAI )'s models running reliably in production through Microsoft Azure cloud platform. We work at the intersection of distributed model serving systems , inference engine optimization and integration, benchmarking, and evaluation to make MAI models shippable, measurable, and evaluable.
Job Overview
Model Serving & Integration
• Partner with engine teams to integrate, optimize, and deploy models to production serving environments with high reliability.
• Help implement and evaluate production serving features (e.g., content provenance/watermarking) while minimizing latency impact.
• Profile and optimize for production SLOs (TTFT, TPOT, throughput, QPS) on modern GPU hardware and distributed serving system topology.
• Work across platform, API, and research teams to streamline the path from model release to production availability.
Model Fine-tuning & Adaptation Support
• Support serving of adapted models (including multi-LoRA scenarios) with careful attention to memory and caching behavior.
• Help build tooling to package and validate fine-tuned models for production deployment.
• Improve the reliability and repeatability of deployment workflows.
Benchmarking, Evaluation & Observability
• Design and run benchmarks using real usage patterns to measure serving performance under realistic conditions.
• Make production endpoints evaluable, identify bottlenecks in evaluation workflows, and improve their speed and reliability.
• Build and improve dashboards and telemetry to track utilization, latency, cache efficiency, and SLA compliance.
• Translate measurement insights into concrete optimizations with quantifiable impact.
Required Qualifications
• Bachelor's Degree in Computer Science or related technical field AND 5+ years of technical engineering experience with coding in languages including Python, C++, or similar OR equivalent experience.
• Experience with LLM inference/serving systems (vLLM, SGLang, TensorRT-LLM, or custom inference engines).
• Strong systems/debugging skills and comfort analyzing performance on GPU clusters.
• Experience with distributed serving/inference concepts (TP/PP/DP, KV cache, speculative decoding).
Preferred Qualifications
• Master's Degree AND 5+ years of relevant experience OR equivalent experience.
• Experience deploying models to production serving platforms (Azure ML/Foundry, Kubernetes, containerized deployments).
• Familiarity with multi-adapter serving, disaggregated prefill/decode, or large-scale KV cache management.
• Experience with observability tools (Grafana, Prometheus or equivalent, Kusto) and analyzing production telemetry.
• Experience building evaluation/benchmarking harnesses for LLMs.
• Knowledge of GPU topology, memory management, and modern accelerator hardware.
Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.
Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Interested in this role?
$120k – $235k
Listed range
Market data coming soon for this role.
See your match score
Sign in to compare your skills and get AI-powered application help.
Get started freeAlready have an account? Sign inNo specific skills listed for this role. Check the job description for requirements.
Job DNA
groundedA deterministic fingerprint of this role, read from its description. No AI.
This role has closed — here are open roles you might like.
Compute Orchestration & Scheduling
Microsoft AI · London
Health Strategy & Commercialization
Microsoft AI · New York
Senior Network Production Engineer, AI Supercomputing
Microsoft AI · London
Product Manager, AI Safety
Microsoft AI · Mountain View
Software Engineer, Model API Infra
Microsoft AI · Mountain View
AI Reliability Engineer, HPC
Microsoft AI · London
Principal, Financial Planning and Analysis
Coupang · Mountain View
Principal, Security Engineer
Coupang · Mountain View
Principal, Machine Learning Engineer
Coupang · Mountain View
Staff Machine Learning Engineer, Personalization
Coupang · Mountain View
Staff Computer Vision Engineer - Search AI Product Engineering
Coupang · Mountain View
Staff ML Engineer – Search & Discovery Relevance
Coupang · Mountain View