Incident Manager
Job Summary
You run Sev1s at KIBO. When one starts you are paged. You triage it and if it clears the bar you open the bridge and take command of it.
From that point the incident is yours. You decide who is on the call you hand out the work and track it and you release people by name when their part is done. You decide whether the client hears from us and when the client lane opens. You call the incident fixed close it and write the document that follows an RCA the pod delivers to the client or a triage document explaining why it was not a Sev1.
You do not debug. You run the incident. The engineers own the fix; you own the timeline the decisions and the record. The role rotates on a short cadence so the person holding it stays sharp and you will be part of a small team that covers it 24/7.
KIBO is a composable digital commerce platform for B2C D2C and B2B organizations who want to simplify the complexity in their businesses and deliver modern customer experiences. KIBO is the only modular modern commerce platform that supports experiences spanning B2B and B2C Commerce Order Management and Subscriptions. Companies like Ace Hardware Zwilling Jelly Belly Nivel and Honey Birdette trust KIBO to bring simplicity and sophistication to commerce operations and deliver experiences that drive value.
KIBOs cutting-edge solution is MACH Alliance Certified and has been recognized by Forrester Gartner IDC Internet Retailer and TrustRadius. KIBO has been named a leader in The Forrester Wave: Order Management Systems Q1 2025 and in the IDC MarketScape report Worldwide Enterprise Headless Digital Commerce Applications 2024 Vendor Assessment.
By joining KIBO you will be part of a team of Kibonauts all over the world in a remote-friendly environment. Whether your job is to build sell or support KIBOs commerce solutions we tackle challenges together with the approach of trust growth mindset and customer obsession.
Carry the on-call page and claim the incident by name inside the acknowledgement window
Triage alerts and client-raised tickets against our Sev1 criteria and make the call quickly with DevOps and Support helping you
Open the bridge and command it one driver one conversation status reported to you rather than shouted into the channel
Page the on-call engineers for each service involved and pull in the pod leader or PS architect when the cause looks like it sits in a clients custom implementation
Own the timeline: assign the work track it decide who is needed and let people go as soon as their part is done
Decide whether the client needs to hear from us and start the client lane in time for first contact inside the 30-minute commitment
Read core-function signals against baseline rather than against a fixed count: error rate above baseline share of expected volume affected and whether it holds across the window plus the rule that a collapse in good orders is an outage even when the errors look normal
Judge the ambiguous cases a security advisory with no confirmed impact an unverified third-party cause stuck orders that break no error threshold and decide what the incident needs before the metrics settle
Keep the record in shared channels so anyone joining catches up without asking and close the incident with T0 resolution time and services affected
Write the RCA for the client what broke what it hit what we did why it happened and what we are changing or the triage document when it was not a Sev1 and add the entry to the First Call playbook
Surface repeat causes across incidents to problem management and keep the triage runbooks current as the platform changes
8 years in incident management production support or technical operations
Experience as Incident Commander or the equivalent on major incidents with several teams on the call
The temperament to run a bridge: keep it calm keep it to one conversation and say no to a Sev1 that is not one
Comfortable reading logs and following a live technical debug closely enough to ask the next question and disciplined enough not to start debugging yourself
Sound judgment on incomplete information under a clock and the willingness to state what you do not know along with when you will know it
Hands-on familiarity with observability and paging tools (Splunk Grafana Prometheus PagerDuty Opsgenie or equivalent)
Working knowledge of ITIL-aligned incident and problem management practices
Familiarity with cloud platforms (AWS GCP or Azure) containerized environments (Kubernetes/Docker) and how APIs databases and messaging fit together
Clear writing under pressure client updates that say only what needs saying and RCAs an enterprise client will accept
Availability for a 24/7 on-call rotation
Commerce retail or payments platform experience checkout payments order routing inventory and availability fulfillment
Exposure to enterprise-approved AI/GenAI tooling for triage log analysis or RCA drafting with attention to data sensitivity and validating output
Mentoring junior incident responders or on-call engineers
Flexible schedule and time away programs
Paid company holidays and global volunteer day
Generous health wellness and benefit programs including 401(k) match and pet insurance
Opportunity for impact rapid career growth and intellectual stimulation
Passionate high-achieving teammates excited to help you succeed and learn
Company events and other activities (Holiday parties Meet-ups Volunteering)
At KIBO we celebrate and support all differences. KIBO is proud to be an equal opportunity workplace. We are committed to equal employment opportunity regardless of race color ancestry religion sex national origin sexual orientation age citizenship marital disability and veteran status.
Required Experience:
Manager
About Company
About This Role You run Sev1s at KIBO. When one starts, you are paged. You triage it, and if it clears the bar you open the bridge and take command of it. From that point the incident is yours. You decide who is on the call, you hand out the work and track