About Microsoft AI
Microsoft AI is building AI systems and products that empower people’s lives. Our work is driven by a community of brilliant, interdisciplinary minds working across frontier model development, product engineering, and responsible AI. Within Microsoft AI, the Safety team develops the training methods, evaluations, runtime safeguards, monitoring, and infrastructure needed to make advanced AI systems safer, more reliable, and more useful. Our work spans text, multimodal, and agentic systems and is developed in close partnership with other Research teams, Production Inference, Security, Responsible AI, Microsoft product organizations, and external partners and customers.
About the Role
We are looking for a Member of Technical Staff to build, deploy, and operate the production systems that keep advanced AI products safe and reliable at global scale. You will work on safety-critical components in the inference path, including orchestration, model- and rules-based safeguards, configuration, telemetry, fail-safe behavior, and rollout mechanisms. You will also build production monitoring for large-scale training and evaluation runs, ensuring that pipelines, compute, and safety signals remain healthy and reliable.
This role is a strong fit for a software engineer with experience in high-scale cloud and distributed systems and balancing product safety and quality with latency, availability, and cost.
Responsibilities
• Design, build, deploy, and operate safety-critical services in the production inference path for text, multimodal, and agentic AI systems.
• Integrate model-based classifiers, policy engines, and other guardrails with model APIs and serving platforms.
• Build containerized services on Kubernetes, with automated CI/CD, staged regional rollout, rollback, and failover.
• Define and meet service-level objectives for availability, latency, throughput, correctness, and cost, including capacity planning and autoscaling.
• Build monitoring and debugging capabilities that make safety decisions and service health observable, and detect failures, stalls, and regressions in large-scale training and evaluation runs.
• Validate launches across development, staging, and production, and contribute to on-call, incident response, and root-cause analysis.
• Partner with Production Inference, Security, Privacy, Responsible AI, and product teams to resolve dependencies and meet launch requirements.
• Contribute reusable libraries, testing standards, deployment runbooks, and operational practices for safety-critical infrastructure.
Required Qualifications
• Bachelor’s degree in Computer Science, Engineering, a related technical field, or equivalent practical experience.
• Strong software engineering skills in one or more production languages such as C++, C#, Java, Go, Rust, or Python.
• Experience designing, building, deploying, and operating distributed services or other large-scale production systems.
• Experience with containerized workloads, Kubernetes, and automated CI/CD pipelines.
• Knowledge of service reliability fundamentals, including observability, capacity planning, failure isolation, graceful degradation, and incident response.
• Ability to reason carefully about correctness, security, privacy, and failure modes in systems that affect end users.
• Ability to collaborate across engineering, machine learning, product, and policy teams and communicate system tradeoffs clearly.
Preferred Qualifications
• Experience with online inference platforms, model serving, API gateways, policy enforcement systems, or other latency-sensitive infrastructure.
• Experience deploying or operating machine learning models, training pipelines, or evaluation systems in production, including versioning, monitoring, failure detection, and rollback or recovery.
• Experience with Azure technologies.
• Experience with observability tooling and capacity planning.
• Familiarity with or interest in AI safety, trust and safety, abuse prevention, content moderation, security engineering, privacy, or responsible AI.
Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Interested in this role?
$120k – $235k
Listed range
Market data coming soon for this role.
See your match score
Sign in to compare your skills and get AI-powered application help.
Get started freeAlready have an account? Sign inNo specific skills listed for this role. Check the job description for requirements.
Job DNA
groundedA deterministic fingerprint of this role, read from its description. No AI.
This role has closed — here are open roles you might like.
Compute Orchestration & Scheduling
Microsoft AI · London
Health Strategy & Commercialization
Microsoft AI · New York
Senior Network Production Engineer, AI Supercomputing
Microsoft AI · London
Product Manager, AI Safety
Microsoft AI · Mountain View
Software Engineer, Model Serving System
Microsoft AI · Mountain View
Software Engineer, Model API Infra
Microsoft AI · Mountain View
Principal, Financial Planning and Analysis
Coupang · Mountain View
Principal, Security Engineer
Coupang · Mountain View
Principal, Machine Learning Engineer
Coupang · Mountain View
Staff Machine Learning Engineer, Personalization
Coupang · Mountain View
Staff Computer Vision Engineer - Search AI Product Engineering
Coupang · Mountain View
Staff ML Engineer – Search & Discovery Relevance
Coupang · Mountain View