Site Reliability Engineer (SRE) I
Job Summary
Thomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build operate and improve reliable production services.
TheSite Reliability Engineerwill support the tools processes and operational practices that help teams detect investigate respond to and prevent production reliability issues. You will work with observability platforms operational documentation automation deployment information and AI-enabled tools to improve the quality and availability of context used during day-to-day operations and incidents.
This is a hands-on engineering role for someone who enjoys learning how complex systems work improving operational readiness and contributing practical solutions to production challenges. You will work closely with experienced SREs Product Engineering teams platform teams and operations partners to maintain reliable services and reduce operational toil.
Rather than expecting you to know every architecture on day one this role will help you build familiarity across products and platforms through maintained documentation dashboards runbooks telemetry deployment data and operational tooling. When you identify gaps in that context you will help improve the systems and processes that keep it current.
You will also use AI-enabled engineering and investigation tools responsibly to accelerate analysis documentation and operational workflows. You will apply technical judgment validate outputs and escalate when additional expertise or review is needed.
Support and maintain SRE operational tooling including dashboards alerts runbooks service documentation telemetry baselines deployment visibility and dependency information.
Use observability toolsincluding logs metrics traces dashboards and alertsto investigate service-health issues identify trends and support incident response.
Participate in incident response by gathering relevant context reviewing recent changes following established runbooks documenting findings and helping coordinate technical follow-up actions.
Execute approved runbooks and mitigation procedures within established escalation change-management and decision-making processes.
Clearly document facts observations hypotheses actions and open questions during incidents handoffs and operational reviews.
Help improve the accuracy completeness and freshness of operational context used by engineering and operations teams during incidents and routine production support.
Contribute to automation and integration work that keeps operational information current such as CI/CD notifications deployment telemetry change-correlation data service ownership records and monitoring configuration.
Review and validate operational artifacts including runbooks diagrams dashboards alerts and AI-generated documentation with guidance from senior engineers and service owners.
Assist with root-cause analysis post-incident reviews and follow-up work by identifying gaps in monitoring documentation automation instrumentation or operational processes.
Treat missing runbooks outdated documentation incomplete telemetry and unclear service ownership as improvement opportunities; partner with the appropriate teams to help resolve those gaps.
Contribute to service-health and error-reduction initiatives using available SLO error budget incident alerting and operational data.
Partner with Product Engineering and platform teams to identify reliability and observability improvements including monitoring gaps alert quality deployment visibility capacity concerns and failure-mode coverage.
Contribute directly to code scripts infrastructure configuration dashboards alerts automation and documentation that improve service reliability and reduce manual operational work.
Use AI-enabled coding and investigation tools to accelerate log review documentation updates runbook drafting incident summarization and hypothesis generation while validating results before relying on them.
Provide actionable feedback when AI-enabled operational tools produce incomplete inaccurate or insufficiently supported outputs.
Participate in design reviews sprint planning and operational-readiness discussions helping ensure reliability and observability considerations are addressed before production deployment.
Support blameless post-incident reviews focused on learning systemic improvement and preventing recurring issues.
3 years of experience in Site Reliability Engineering DevOps cloud infrastructure platform engineering systems engineering production operations software engineering or a related technical field.
Working knowledge of at least two of the following areas: cloud infrastructure distributed systems observability and monitoring networking databases CI/CD containers or infrastructure automation.
Experience working with production telemetry including logs metrics traces dashboards monitoring platforms or alerting systems.
Experience troubleshooting production issues participating in incident response or supporting business-critical applications and services.
Experience writing maintaining or improving operational documentation runbooks knowledge articles or support procedures.
Experience with scripting or programming in one or more languages such as Python Bash PowerShell JavaScript Java Go or a comparable language.
Willingness and ability to contribute to code infrastructure configuration dashboards monitoring rules alerts automation or documentation.
Ability to communicate clearly about technical findings operational risks and next steps with engineers and operational stakeholders.
Ability to work effectively in a collaborative environment ask for help when needed and learn from more experienced engineers.
Familiarity with AI-enabled coding documentation investigation or operational-analysis tools along with an understanding that outputs must be reviewed and validated.
Experience with cloud platforms Kubernetes containers CI/CD tooling infrastructure-as-code or configuration-management tools.
Experience with observability platforms such as Datadog Dynatrace New Relic Splunk Grafana Prometheus Elastic or similar technologies.
Experience defining or working with Service Level Objectives Service Level Indicators error budgets service-health metrics or incident-management processes.
Experience supporting 24/7 production environments or participating in an on-call rotation.
Experience improving dashboards alerts runbooks deployment visibility change correlation service documentation or operational workflows.
Familiarity with AI agent workflows AI-assisted root-cause analysis or AI-enabled incident-management tools.
Experience contributing to automation that reduces repetitive operational work and improves response consistency.
Experience participating in blameless post-incident reviews and helping drive corrective actions to completion.
Relevant certifications in cloud infrastructure Kubernetes DevOps SRE observability or incident management are beneficial but not required.
#LI-LP2
Whats in it For You
- Hybrid Work Model: Weve adopted a flexible hybrid working environment for our office-based roles while delivering a seamless experience that is digitally and physically connected.
- Flexibility & Work-Life Balance: Flex My Way is a set of supportive workplace policies designed to help manage personal and professional responsibilities whether caring for family giving back to the community or finding time to refresh and reset. This builds upon our flexible work arrangements including work from anywhere for up to 8 weeks per year empowering employees to achieve a better work-life balance.
- Career Development and Growth: By fostering a culture of continuous learning and skill development we prepare our talent to tackle tomorrows challenges and deliver real-world solutions. Our Grow My Way programming and skills-first approach ensures you have the tools and knowledge to grow lead and thrive in an AI-enabled future.
- Industry Competitive Benefits: We offer comprehensive benefit plans to include flexible vacation two company-wide Mental Health Days off access to the Headspace app retirement savings tuition reimbursement employee incentive programs and resources for mental physical and financial wellbeing.
- Culture: Globally recognized award-winning reputation for inclusion and belonging flexibility work-life balance and more. We live by our values: Obsess over our Customers Compete to Win Challenge (Y)our Thinking Act Fast / Learn Fast and Stronger Together.
- Social Impact: Make an impact in your community with our Social Impact Institute. We offer employees two paid volunteer days off annually and opportunities to get involved with pro-bono consulting projects and Environmental Social and Governance (ESG) initiatives.
- Making a Real-World Impact:We are one of the few companies globally that helps its customers pursue justice truth and transparency. Together with the professionals and institutions we serve we help uphold the rule of law turn the wheels of commerce catch bad actors report the facts and provide trusted unbiased information to people all over the world.
In the United States Thomson Reuters offers a comprehensive benefits package to our employees. Our benefit package includes market competitive health dental vision disability and life insurance programs as well as a competitive 401k plan with company addition Thomson Reuters offers market leading work life benefits with competitive vacation sick and safe paid time off paid holidays (including two company mental health days off) parental leave sabbatical leave. These benefits meet or exceeds the requirements of paid time off in accordance with any applicable state or municipal laws. Finally Thomson Reuters offers the following additional benefits: optional hospital accident and sickness insurance paid 100% by the employee; optional life and AD&D insurance paid 100% by the employee; Flexible Spending and Health Savings Accounts; fitness reimbursement; access to Employee Assistance Program; Group Legal Identity Theft Protection benefit paid 100% by employee; access to 529 Plan; commuter benefits; Adoption & Surrogacy Assistance; Tuition Reimbursement; and access to Employee Stock Purchase Plan.Thomson Reuters complies with local laws that require upfront disclosure of the expected pay range for a position. The base compensation range varies across locations. For any eligible US locations unless otherwise noted the base compensation range for this role is $70800 USD - $131400 USD. Base pay is positioned within the range based on several factors including an individuals knowledge skills and experience with consideration given to internal equity. Base pay is one part of a comprehensive Total Reward program which also includes flexible and supportive benefits and other wellbeing programs. This role may also be eligible for an Annual Bonus based on a combination of enterprise and individual performance.
About Us
Thomson Reuters informs the way forward by bringing together the trusted content and technology that people and organizations need to make the right decisions. We serve professionals across legal tax accounting compliance government and media. Our products combine highly specialized software and insights to empower professionals with the data intelligence and solutions needed to make informed decisions and to help institutions in their pursuit of justice truth and transparency. Reuters part of Thomson Reuters is a world leading provider of trusted journalism and news.
We are powered by the talents of 26000 employees across more than 70 countries where everyone has a chance to contribute and grow professionally in flexible work environments. At a time when objectivity accuracy fairness and transparency are under attack we consider it our duty to pursue them. Sound exciting Join us and help shape the industries that move society forward.
As a global business we rely on the unique backgrounds perspectives and experiences of all employees to deliver on our business goals. To ensure we can do that we seek talented qualified employees in all our operations around the world regardless of race color sex/gender including pregnancy gender identity and expression national origin religion sexual orientation disability age marital status citizen status veteran status or any other protected classification under applicable law. Thomson Reuters is proud to be an Equal Employment Opportunity Employer providing a drug-free workplace.
Thomson Reuters makes reasonable accommodations for applicants with disabilities including veterans with disabilities and for sincerely held religious beliefs in accordance with applicable law. If you reside in the United States and require an accommodation in the recruiting process you may contact our Human Resources Department at. Disability accommodations in the recruiting process may include things like a sign language interpreter making interview rooms accessible providing assistive technology or other relevant accommodations. Please note this email is not intended for general recruitment questions and we will promptly respond to inquiries regarding accommodations. More information on requesting an accommodation here.
Learn more on how to protect yourself from fraudulent job postings here.
Required Experience:
IC
About Company
Core Print Solutions: Full-service book printer for book publishers. Complete printing services including binding, finishing, warehousing & order fulfillment. Trusted US book printer.