Site Reliability Engineer
Job Summary
Job Summary:
The Site Reliability Engineer (SRE) is embedded directly with our product teams working closely with them to design code test run and evolve the systems that help people around the world make payments. Works closely with internal teams to drive adoption of modern reliability practices like SLOs error budget policies actionable alerts follow-the-sun on-call incident retrospectives chaos testing and end-to-end ownership. SREs are crucial contributors to nearly every ACI product team that has big traffic big data high availability and performance requirements.
Job Responsibilities:
Design develop deploy and motivate the creation of software and systems to increase product reliability and organizational efficiency.
Guide reliability practices through the entire software development lifecycle through activities like architecture reviews code reviews creating platforms and frameworks capacity planning and chaos testing.
Maintain service health by implementing and evolving monitoring alerting self-healing and follow-the-sun incident response.
Improve service reliability through blameless post-incident reviews and using code to prevent or respond to problem recurrence.
Function as a key technical and culture leader throughout your assigned line of business
Drive and evolve the overall resilience strategy of your given line of business leveraging industry and internal tools
Ensure that local and cross site redundancy mechanisms are meeting requirements work as designed and ever evolving
Set maintain and enforce standards across deployment practices operations etc. Engage in change review as a key member
Function as a key contributor to overall capacity peak season and business continuity methodologies and testing for your space
Interface directly with key clients as needed
Support and help standardize sales responses for your space by helping to craft the go forward offers with business and DevOps teams aligning costs SLAs and technology
Perform other duties as assigned
Understand and adhere to all corporate policies to include but not limited to the ACI Code of Business Conduct and Ethics.
Knowledge Skills and Experience required for the job:
BS degree in Computer Science related technical field or equivalent practical experience.
Experience writing code in Java Go Shell Python or a similar language.
Experience in data structures database systems algorithms and software design.
Passion for debugging optimizing code and automating routine tasks.
Practical knowledge with RDBs (such as PostgreSQL Oracle) NoSQL KV stores (such as Cassandra) and messaging systems (such as Kafka RabbitMQ and MQ)
Strong interest in SRE topics like SLOs resilience scaling performance and more
Strong experience in Production level mission critical environments. Azure Cloud experience running Production workloads preferred
Strong experience working in virtualized Linux operating environments
5 years of real-world experience
Preferred Knowledge Skills and Experience needed for the job:
Previous experience in an SRE or Production Engineering role
Direct experience with standard automation tools such as Ansible Jenkins etc.
BASE24-eps or BASE24 knowledge a plus GOAFT ACI Desktop UI NCPCOM/XPNET ICE-XS
Experience with globally distributed teams
Takes initiative to solve problems using a scientific approach
Skilled in providing substantial feedback on distributed system designs and recommendations on new technologies and processes
Collaboration skills
Work Environment:
Office work environment
Collaborative team
Prolonged periods of sitting at a desk and working on a computer
15% Travel may be domestic or international
Weekend and off-hours support may be required
Required Experience:
IC
About Company
With more than 3,800 employees worldwide and offices in principal cities around the globe, ACI has one of the most diverse and robust product portfolios in the industry, with application software spanning the entire electronic payments chain.