Site Reliability Engineer III -(AIML SRE)
- Define and refine Service Level Objectives (SLOs) for large language model serving and training systems, using metrics like accuracy, fairness, latency, drift targets, TTFT, and TPOT, while balancing reliability and development velocity.
- Design, implement, and continuously improve monitoring systems to track availability, latency, drift, and other key metrics for robust observability and rapid issue detection.
- Collaborate in the design and deployment of high-availability language model serving infrastructure that supports high-traffic internal workloads across multiple regions and cloud providers.
- Champion site reliability engineering practices, providing technical leadership and fostering a culture of reliability, resilience, and continuous improvement across teams.
- Develop and manage automated failover and recovery systems for model serving deployments, ensuring seamless operation and rapid recovery from failures.
- Create and lead AI-specific incident response playbooks for issues like model drift or bias spikes, including automated rollbacks, circuit breakers, and systematic post-incident improvements.
- Build and maintain cost optimization systems for large-scale AI infrastructure, leveraging load balancing, caching, optimized GPU scheduling, and AI Gateways to ensure efficient, secure, and scalable operations.
- Formal training or certification on AI reliability concepts and 3+ years applied experience.
- Demonstrate a strong sense of curiosity and a passion for continuous learning, especially in the rapidly evolving field of AI reliability.
- Show proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices.
- Possess deep knowledge and experience in observability, including white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, and Splunk.
- Be proficient with continuous integration and delivery tools like Jenkins, GitLab, or Terraform, as well as container and orchestration technologies such as ECS, Kubernetes, and Docker.
- Have experience troubleshooting common networking technologies and issues, and understand the unique challenges of operating AI infrastructure, including model serving, batch inference, and training pipelines.
- Communicate effectively and bridge the gap between ML engineers and infrastructure teams, with proven experience implementing and maintaining SLO/SLA frameworks for business-critical services, and working with both traditional and AI-specific metrics.
- Experience with AI-specific observability tools and platforms, such as OpenTelemetry and OpenInference.
- Familiarity with AI incident response strategies, including automated rollbacks and AI circuit breakers.
- Knowledge of AI-centric SLOs/SLAs, including metrics like accuracy, fairness, drift targets, TTFT (Time To First Token), and TPOT (Time Per Output Token).
- Expertise in engineering for scale and security, including load balancing, caching, optimized GPU scheduling, and AI Gateways.
Experience with continuous evaluation processes, including pre-deployment, pre-release, and post-deployment monitoring for drift and degradation. - Understand ML model deployment strategies and their reliability implications
- Have contributed to open-source infrastructure or ML tooling
- Have experience with chaos engineering and systematic resilience testing
JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans Base Pay/Salary
Jersey City,NJ $133,000.00 - $185,000.00 / year
Recommended Jobs
Confirmation Call Center Manager
Job Description Job Description Confirmation Call Center Manager Creating a fresh solution to bath remodeling, Bath Planet offers a stylish, cost-effective, low-maintenance bath improvement to…
Configuration Manager
Job Description Job Description NDI Engineering Company is seeking a full-time Configuration Management Analyst Support person to join our team in Philadelphia and supporting the US Navy Engineer…
Export Specialist
Job Description Job Description Role Description This is a full-time on-site role located in Moonachie, NJ for an Export Specialist. The Export Specialist will be responsible for handling expo…
Mechanical Engineer
Job Description Job Description Description: About the Company: At Lincoln Electric Products Co. Inc., we specialize in the design, manufacture, and distribution of custom equipment tailore…
Director, Statistics (Office-based)
Company Description AbbVie's mission is to discover and deliver innovative medicines and solutions that solve serious health issues today and address the medical challenges of tomorrow. We striv…
Membership Specialist
Job Description Job Description The Membership Specialist (MS) will represent UFC GYM by providing a welcoming, informative, and entertaining experience for all members and guests during their vi…
Insurance Customer Service Rep
Job Description Job Description McDyer Insurance Agency LLC has proudly served New Jersey communities since 2001. We are the largest Allstate agency in the state, protecting more than 7,000 hous…
Business Development Director - Market Access Technology & Solutions
Position Summary The Business Development Director will play a pivotal role in driving strategic growth and revenue generation for IQVIA's Market Access Technology and Services (MATS) practice. Th…
Studio Team Member
Job Description Job Description Company Description Alpha Fit Club is a premier group fitness training program for all fitness levels. This circuit-style concept features a blend of total body…
IT Support Technician
Job Description Job Description TITLE: IT Support Technician DURATION: Three (3+) months with possibility to be extended LOCATION: 210 Carnegie Center, Suite 103, Princeton, NJ 0854…