Design, develop, and operate scalable, resilient systems using Python and SRE practices. Implement monitoring, alerting, automated recovery, SLIs, SLOs, chaos engineering, and performance testing. Troubleshoot production services, conduct root cause analysis, improve system architecture and coding quality, and collaborate with product and engineering teams on reliability, security, and operational excellence.
As a Lead Software Engineer – Software Reliability at JPMorgan Chase as a part of our product team, you will design, build, and operate scalable, resilient systems using Python and modern reliability practices. You will apply Site Reliability Engineering (SRE) principles to improve availability, performance, and operational excellence, and you will help establish engineering standards that increase system reliability and security. You will collaborate with engineering and product partners to troubleshoot, optimize, and maintain production services while fostering a collaborative and inclusive team culture.
Job Responsibilities
- Design and develop scalable and resilient systems using Python to support continuous improvement and apply Site Reliability Engineering (SRE) concepts to enhance system reliability and performance
- Execute software solutions, including design, development, and technical troubleshooting and create secure, high-quality production code and maintain algorithms that run synchronously with appropriate systems
- Produce or contribute to architecture and design artifacts, ensuring design constraints are met
- Gather, analyze, and synthesize data to develop visualizations and reporting for software and system improvement
- Identify hidden problems and patterns in data to drive improvements in coding hygiene and system architecture
- Implement reliability engineering practices such as monitoring, alerting, and automated recovery
- Define and measure Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to track system health
- Conduct chaos engineering experiments to test system resiliency and identify weaknesses
- Perform performance testing using tools such as JMeter to ensure scalability and stability
- Collaborate with product teams to enhance system reliability, scalability, and performance and contribute to software engineering communities of practice and events exploring new and emerging technologies
- Foster a team culture of diversity, opportunity, inclusion, and respect and participate in post-incident reviews and drive root cause analysis for system failures
Required qualifications, capabilities, and skills
- Hands-on practical experience in system design, application development, testing, and operational stability and proficient in coding in Python
- Experience developing, debugging, and maintaining code in a large corporate environment with modern programming languages and database querying languages
- Knowledge of the Software Development Life Cycle, AWS cloud exposure, troubleshooting abilities, resiliency, and automation focus
- Understanding of agile methodologies such as CI/CD, application resiliency, and security
- Knowledge of software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning, mobile, etc.)
- Experience with AI and full understanding of the SDLC process, and MongoDB
- Familiarity with reliability engineering concepts, including monitoring, alerting, and automated recovery and ability to implement and maintain system health checks and performance metrics
- Commitment to writing maintainable, testable, and high-quality code
- Understanding of SRE principles, including SLIs, SLOs, and error budgets and experience with incident response and root cause analysis
- Experience with chaos engineering practices to test system resiliency and proficiency in performance testing tools such as JMeter
Preferred qualifications, capabilities, and skills
- Familiarity with modern front-end technologies
- Exposure to cloud technologies
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.
Similar Jobs
Financial Services
Lead design, operation, and reliability of large language model serving infrastructure. Build backend services, deploy and lifecycle-manage LLMs on cloud and GPU clusters, implement observability, tune performance/cost, run incident response/on-call, and drive AI-assisted engineering and safe-responsible AI practices.
Top Skills:
Amazon BedrockAmazon Sagemaker EndpointsAmazon Sagemaker JumpstartCrewaiGpuKserveKubernetesLangchainLanggraphLlm-DNvidia Triton Inference ServerOpentelemetryPythonRayRay ServeVllm
Financial Services
Leads cybersecurity architecture for cloud-based applications and enterprise systems. Defines target architecture, identifies and mitigates risks, automates recurring remediation, evaluates emerging technologies and vendors, and drives AI-assisted security validation across SDLC toolchains. Collaborates with technical teams and senior business stakeholders while ensuring security, resiliency, auditability, and data-sensitivity requirements are met.
Top Skills:
Agile MethodologiesArtificial IntelligenceCloud-Native TechnologiesContinuous DeliveryContinuous IntegrationCybersecurity ControlsMachine LearningPublic CloudSoftware Development Lifecycle
Fintech • Information Technology • Financial Services
Leads middle office service implementations for investment management clients, translating business requirements into scalable operational workflows. Partners with technology, production, and operational teams to improve processes, controls, and implementation efficiency. Performs data analysis, identifies and mitigates risks, communicates recommendations, and serves as a regional escalation point. The role also builds team capability to execute implementations independently and at scale.
Top Skills:
Aladdin
What you need to know about the Edinburgh Tech Scene
From traditional pubs and centuries-old universities to sleek shopping malls and glass-paneled office buildings, Edinburgh's architecture reflects its unique blend of history and modernity. But the fusion of past and future isn't just visible in its buildings; it's also shaping the city's economy. Named the United Kingdom's leading technology ecosystem outside of London, Edinburgh plays host to major global companies like Apple and Adobe, as well as a growing number of innovative startups in fields like cybersecurity, finance and healthcare.

