Enter a job title or keyword

Site Reliability Engineer II AI & Infrastructure (fmd)

Focused


Job Location:

Berlin - Germany

Monthly Salary: Not provided by the employer
Posted: 30 August 2026 (3 days ago)
Application Deadline: 27 November 2026
Vacancies: 1 Vacancy

Job Summary

Focused Energy is a pioneering international deep-tech company with locations in Germany and the US dedicated to commercializing laser-driven nuclear fusion. Our mission is to deliver clean safe and virtually limitless energy to the world. As we rapidly scale we seek talented individuals who thrive on bringing clarity and structure to fast-growing environments.

About the Role

Focused Energy is looking for a Site Reliability Engineer II to help build reliable secure observable and easy-to-operate systems for its internal applications. Youll own day-to-day reliability deployments incident response monitoring backup recovery and rollback processes across a broad technical environment spanning cloud services databases identity networking secrets and CI/CD.

This role is ideal for an engineer who enjoys solving complex operational problems and turning recurring incidents into lasting improvements. Alongside your core infrastructure responsibilities youll provide overflow support for AI-enablement workflows and tools when the dedicated AI Tech Enabler needs additional support.

What Youll Do
  • Own the reliability and day-to-day operation of multiple internal applications deployment platforms and supporting services.

  • Lead the investigation and resolution of incidents and deployment failures across cloud services databases identity networking secrets and CI/CD systems.

  • Drive automation improvements across build release deployment monitoring alerting backup recovery rollback and runbook processes.

  • Implement safe well-tested code configuration infrastructure and database fixes to restore or improve service reliability.

  • Develop proposals for cloud hosting database and runtime migrations as well as horizontal scaling approaches supported by appropriate testing and rollback plans.

  • Build and maintain clear operational documentation including runbooks recovery procedures incident fixes and system knowledge that can be reused by other engineers.

  • Partner with software engineering IT security identity and external platform-support stakeholders to plan and execute infrastructure changes safely.

  • Provide overflow diagnostic and troubleshooting support for AI workflows and tools including Langdock without losing focus on core reliability priorities.

Who You Are
  • You take ownership of well-scoped technical work from investigation through testing deployment and follow-up.

  • You are calm and methodical when responding to live incidents and can diagnose problems across multiple interconnected systems.

  • You focus on root-cause resolution rather than repeatedly applying temporary fixes.

  • You communicate incidents trade-offs risks and recovery plans clearly to both technical and non-technical stakeholders.

  • You are proactive about identifying reliability risks automation opportunities recurring issues and operational improvements.

  • You make independent decisions on tactical fixes and safe configuration changes while seeking appropriate sign-off for architecture migration and scaling decisions.

  • You document changes and operational knowledge so that other engineers can support systems effectively.

  • You ask for help early when risks or dependencies are unclear and keep the Head of IT informed about significant issues and progress.

Desirable Skills & Knowledge
  • Strong experience with Linux networking structured troubleshooting and cloud hosting concepts.

  • Experience operating internal applications and deployment platforms in Azure or a comparable cloud environment.

  • Practical knowledge of relational databases particularly PostgreSQL or an equivalent platform including safe migration and recovery practices.

  • Experience with infrastructure as code and CI/CD tools such as OpenTofu Terraform GitLab CI GitHub Actions or Azure DevOps.

  • Knowledge of monitoring alerting backup recovery rollback automation incident response and service reliability practices.

  • Experience with horizontal scaling stateless system design migrations and reversible change management.

  • Scripting and automation skills together with experience reducing repetitive operational work.

  • Familiarity with identity and access management least-privilege access secrets management FEs internal AI/SaaS stack and tools such as Langdock.


Focused Energy is an equal opportunity employer committed to creating an inclusive environment. Qualified applicants will receive consideration for employment without regard to race color religion sex sexual orientation gender perception or identity national origin age marital status protected veteran status or disability status.

Pursuant to the San Francisco Fair Chance Ordinance Focused Energy will consider for employment qualified applicants with arrest and conviction records.

Compensation offered will be determined by factors such as location level job-related knowledge skills and experience. Certain roles may be eligible for incentive compensation equity benefits.


Required Experience:

IC


About Company

Company Logo

Focused Energy is the leading Laser Fusion company building on NIF’s foundation & pursuing the most promising path to clean, limitless fusion power

View Profile View Profile