Senior AI Engineer (fmd) Remote in Germany
Job Summary
Most AI teams talk about evaluation. At Synera you will own it: end-to-end in production for multiple real agentic products used by companies like BMW Airbus and NASA.
WHAT YOU WILLDO
Youll join the Agentic Ants team as their second native AI engineer. Youll build and extend agentic systems alongside the rest of the team including new tools sub-agents and prompt iterations. But most importantly you will contribute to an evaluation framework that gates every change: Golden datasets LLM-as-judge pipelines regression suites production-trace mining. Your job is to prove that the agents the team is shipping can be trusted.
Youll not only improve our existing software but also shape the next AI-native generation of Synera. Partnering with Danny and the wider R&D department youll help the teams go deeper on customer insights data and come up with new ways to automate engineering work.
Monday: Join our company-wide all-hands meeting to stay in the loop then review Langfuse traces from the weekend and flag any new failure modes worth triaging.
Tuesday: Work on the golden dataset for the supervisor routing surface curating examples versioning the set and writing evaluators with Ahmed.
Wednesday: Join a cross-team sync with QA and product to align on new eval coverage for an upcoming agent feature then push a CI integration so eval regressions block the next PR.
Thursday: Show off your latest work to other departments in the sprint review use your lunch break to go for a quick run then kick off sprint planning with the Agentic Ants.
Friday: Pair with the AI and software engineers on extending our agentic system then write the eval that gates the change before it ships.
The evaluation framework is live and in use across multiple agent CI gates are blocking on regressions and the team can actually trust the results.
Youve shipped multiple meaningful changes to the agent graphs (new tools sub-agents or routing improvements) that are measurably better in production.
The agents improve themselves: Production traces feed continuously back into datasets and the improvement loop is running without manual intervention.
You co-own at least one agent surface with the team and have become the go-to person for evaluation methodology at Synera. The combination has measurably improved product quality.
Here is our team at the summer event - join us at the next one!
We believe in transparent conversations about compensation from the start. For this role our planned salary ranges are:
Experienced Level: 77000 - 97000 EUR annually
We determine the level we hire you for based on your experience the scope of responsibilities youll take on and the impact you can drive. While we typically hire within these bands were open to some flexibility for candidates who bring exceptional value to the role.
Check out this page to learn how we approach salary career growth and creating an environment where everyone can shine.
Flexible working: you decide when & where to work (as long as you have a residency in Germany).
Flexible public holidays: swap days off according to your values and beliefs!
Home office setup support access to our office in Bremen.
Personal development budget of 2000 to attend conferences and trainings or buy interesting books to improve in an area of your choice.
We dont count your vacation days as we trust all our team members to decide whats best for them and the company.
Prefer two wheels over four Weve got you covered with JobRad.
To support your personal and professional well-being we offer company fitness with Wellpass and mental health platform nilo.
Regular team events virtual coffee breaks and spontaneous afterworks. We also get together as a whole company for 2-3 day off-sites twice a year!
Even if you dont meet these criteria perfectly but believe you have lots to bring to the role we encourage you to apply. We know its tough but please keep in mind that you dont have to match all the listed requirements exactly to be considered for this role.
You resonate with Syneras Core Values - theyre central to how we work and well explore them together in your first interview.
Youve designed and shipped agent graphs in production with LangGraph or equivalent including supervisor / sub-agent patterns tool design prompt iteration.
Youve built or meaningfully contributed to an LLM evaluation pipeline: LLM-as-judge design calibration against human labels dataset versioning and the statistical reasoning behind it (CIs sample sizes false positives).
You write production-grade Python (using FastAPI and PostgreSQL).
Youve worked with at least two of the following AI platforms: Anthropic OpenAI Azure Froundry Bedrock or Vertex AI.
Youre familiar with Langfuse LangSmith or similar tracing tools.
You handle reliability when it shows up in your work implementing retries error handling or graceful degradation.
You communicate clearly and push back when you disagree.
P.S. Synera is a place where everyone can grow. So however you identify and whatever background you bring with you please apply if this is a role that would make you excited to come to work every day and be prepared to share with us how your perspective will bring something unique and valuable to our Agentic Ants team.
Your application has been successfully submitted!
Required Experience:
Senior IC