GPU Infrastructure NOC Engineer
Pittsburgh, PA - USA
Job Summary
Pay:$75000.00 - $140000.00 per year
Why This Is a Great Opportunity
- Monitor and support cutting-edge GPU and AI infrastructure powering demanding enterprise workloads.
- Go beyond traditional NOC monitoring by building automation improving tooling and making operations smarter.
- Get hands-on exposure to GPU HPC networking data center infrastructure and AI-enabled operations.
- Play a direct role in protecting customer uptime and meeting demanding SLA commitments.
- Help create and continuously improve runbooks SOPs monitoring processes and incident-response workflows.
- Work alongside technical teams OEMs data center operators and network providers to solve real infrastructure problems.
- Join a fast-growing environment where your ideas can directly improve how the NOC operates.
- Receive bonus and equity opportunities in addition to competitive base compensation.
Location:Remote nationwide. This is a fully remote role supporting a 24/7 global operations environment through rotating shifts including nights and weekends.
Note:Must have 2 years of NOC network operations or infrastructure monitoring experience hands-on GPU or HPC infrastructure experience and Python or Bash scripting experience. Candidates must be willing to work rotating 24/7 shifts including nights and weekends. These are hard requirements.
About Us
We are building next-generation AI infrastructure designed to give enterprise customers fast flexible access to high-performance GPU compute. Our operations team plays a critical role in keeping these environments reliable responsive and continuously improving. Confidential Employer.
Job Description
- Monitor live GPU cluster health power cooling networking and infrastructure status across production deployments.
- Triage troubleshoot and resolve incidents while maintaining SLA requirements.
- Identify infrastructure issues early and take proactive action before they become customer-impacting incidents.
- Escalate appropriate issues to OEMs data center operators network providers or other Tier 3 partners.
- Build and improve internal monitoring tools scripts and automation to reduce repetitive manual work.
- Use Python Bash or similar scripting tools to automate monitoring triage reporting and operational workflows.
- Explore and implement AI-enabled workflows that improve NOC speed accuracy and efficiency.
- Create maintain and continuously improve technical runbooks and standard operating procedures.
- Participate in incident reviews and turn recurring problems into permanent tooling process or automation improvements.
- Track SLA and incident metrics and identify opportunities to improve reliability and response times.
- Communicate clearly and proactively with customers and internal stakeholders during incidents.
- Coordinate with data center operators OEMs network providers and other third parties to resolve customer-impacting issues.
- Support a 24/7 rotating operations schedule including nights weekends and other assigned shifts.
Qualifications
- 2 years of experience in a NOC network operations infrastructure monitoring or related technical operations environment.
- Direct experience supporting or monitoring GPU HPC AI infrastructure or comparable high-performance computing environments.
- Hands-on Python or Bash scripting experience.
- Experience with infrastructure monitoring and alerting tools such as Datadog Grafana PagerDuty or similar platforms.
- Strong troubleshooting and incident-response skills.
- Experience working with network compute storage or data center infrastructure.
- Ability to understand technical issues quickly and communicate effectively during incidents.
- Demonstrated interest in automation scripting tool-building and continuous operational improvement.
- Must be comfortable working rotating 24/7 shifts including nights and weekends.
- Ability to work independently in a remote environment while collaborating effectively with global technical teams.
- Additional languages beyond English are a plus.
Why You Will Love Working Here
- Work directly with advanced GPU and HPC infrastructure rather than generic IT systems.
- Build technical depth across AI compute networking data centers monitoring and automation.
- Have a voice in how the NOC operates and help improve processes rather than simply following them.
- Work with modern monitoring automation and AI-enabled operations tools.
- Gain exposure to complex enterprise infrastructure and high-availability environments.
- Remote nationwide flexibility with multiple shift options.
- Medical dental and vision insurance.
- 401(k).
- Paid maternity and paternity leave.
- Bonus and equity opportunities.
JPC-1916
Benefits:
- Dental insurance
- Paid time off
- Retirement plan
- Vision insurance
Requirements: Must-have: remote nationwide. 2 yrs NOC network ops or infrastructure monitoring exp Must accept 24/7 rotating shifts including nights/weekends (Y/N hard deal break) must have GPU/HPC exp. Must have Python/Bash exp
Required Years as Associate: 2
Additional context: Relocation- no. Packages - no
Submission Email / Name: submissions are via Slack c/o Dil-Nash-Nina
Quick Recruiter Reference
Monitor and support live GPU/HPC clusters in a 24/7 NOC while handling incidents SLA performance automation and runbook improvements. Must have 2 years NOC/network/infrastructure monitoring GPU/HPC experience and Python or Bash. Must accept rotating nights/weekends.
Recruiter Submission: To submit cancel - GPU Infrastructure NOC Engineer - Axe Compute - JPC-1916 - source
New Job Order Alert
Client job title: GPU Infrastructure NOC Engineer
Location: Nationwide
On-site hybrid remote: Remote
Experience: 2
Good fit job titles/keywords for candidates: NOC Engineer NOC Analyst Network Operations Engineer Infrastructure Operations Engineer GPU NOC HPC Operations GPU Infrastructure Engineer AI Infrastructure Engineer Systems Operations SRE Data Center Operations Python Bash
# of hires needed: 1