Princ Engr-Systems Engrg
Temple Terrace, FL - USA
Job Summary
You want more out of a career. A place to share your ideas freely even if theyre daring or different. Where the true you can learn grow and thrive. At Verizon we power and empower how people live work and play by connecting them to what brings them joy. We do what we love driving innovation creativity and impact in the world. Our V Team is a community of people who anticipate lead and believe that listening is where learning crisis and in celebration we come together lifting our communities and building trust in how we show up everywhere & always. Want in Join the #VTeamLife.
We are seeking a highly skilled AI / GPU Engineer with expertise in infrastructure and infrastructure solutions for AWS GCP and on-premise environments. This role will focus on designing deploying and optimizing GPU-accelerated infrastructure to support AI/ML workloads. The ideal candidate will have deep experience with cloud-based AI infrastructure GPU provisioning cluster management and performance optimization for large-scale computing environments.
Key Responsibilities
- Bare-Metal Architecture & Leadership: Drive technical strategy and infrastructure delivery for on-prem NVIDIA DGX clusters managing hardware virtualization Multi-Instance GPU (MIG) slicing and GPU pooling to maximize compute utilization. Provide input and team collaboration regarding the design of bare-metal compute topologies GPU partitioning and hardware lifecycle management. Ability to Automate bare-metal provisioning OS base image customization driver updates and multi-cluster configurations.
- Orchestration & Batch Scheduling: Lead the deployment of enterprise Kubernetes and OpenShift AI environments. Optimize workload placement using Kueue KEDA and custom CRDs for gang scheduling multi-tenant GPU sharing and event-driven autoscaling. Set platform standards for bare-metal Kubernetes multi-tenancy access controls and operator development.
- Advanced Inference Engine Deployment: Build enterprise-grade LLM inference architectures utilizing engines like vLLM LiteLLM and llm-d. Implement prefill/decode disaggregation and prefix-cache-aware routing. Architect disaggregated LLM serving pipelines distributed KV-cache offloading and batch job queuing.
- High-Performance Storage & Interconnects: Architect ultra-low-latency network topologies using InfiniBand and RDMA (RoCE) alongside high-throughput parallel storage systems (VAST Weka Ceph) to eliminate data bottlenecks during large-scale model checkpoints and training runs. Design lossless fabric topologies and high-IOPS storage for multi-node training checkpoints and dataset streaming.
- Hybrid Cloud Readiness: Design infrastructure patterns via Infrastructure as Code (Terraform Ansible) that allow seamless bursting or workload migration to Google Cloud (GKE) and AWS (EKS) when necessary.
- Deep Telemetry & Automation: Standardize hardware-level monitoring and KPI collection using NVIDIA DCGM Prometheus Mimir and Grafana. Build automated alerting and dynamic auto-remediation for failing nodes or GPU memory leaks. Establish SLA metrics tracking cluster health down to the GPU-core level and driving automated metric-based alerts.
- Framework Profiling: Partner with AI research teams to profile CUDA operations and optimize deep learning frameworks (PyTorch TensorFlow JAX) for maximum hardware-level FLOPS and throughput.
What were looking for...
- Bachelors degree or four or more years of work experience.
- Six or more years of relevant experience required demonstrated through one or a combination of work and/or military experience or specialized training
- Knowledge of NVIDIA AI stack including CUDA cuDNN TensorRT and NVIDIA MIG.
- Experience with hybrid cloud solutions (on-prem AWS Google Cloud).
- Experience with DevOps & CI/CD for AI workloads including GitOps.
- Google Cloud Professional ML Engineer or Cloud Architect certification
- Hands-on experience with Google Cloud AI/ML offerings and on-prem GPU server deployment.
- Experience with multi-node GPU training using NCCL MPI or Horovod.
- Familiarity with alternative cloud providers (AWS OCI Azure) for AI workloads.
- Contributions to open-source AI/GPU infrastructure projects.
- Knowledge of AI security best practices (model encryption secure AI pipelines).
If Verizon and this role sound like a fit for you we encourage you to apply even if you dont meet every even better qualification listed above.
Verizon is an equal opportunity employer. We evaluate qualified applicants without regard to veteran status disability or other legally protected characteristics.
Our benefits are designed to help you move forward in your career and in areas of your life outside of Verizon. From health and wellness benefit options including: medical dental vision short and long term disability basic life insurance supplemental life insurance AD&D insurance identity theft protection pet insurance and group home & auto insurance. We also offer a matched 401(k) savings plan up to 8 company paid holidays per year and up to 6 personal days per year paid parental leave adoption assistance and tuition assistance plus other incentives weve got you covered with our award-winning total rewards package. Depending on the role employees have the opportunity to receive compensation in the form of premium pay such as overtime shift differential holiday pay allowances etc. Newly hired employees receive up to 15 days of vacation per year which grows with additional service. For part-timers your coverage will vary as you may be eligible for some of these benefits depending on your individual circumstances.
The salary will vary depending on your location and confirmed job-related skills and experience. This is an incentive based position with the potential to earn more. For part-time roles your compensation will be adjusted to reflect your hours.The annual salary range for the location(s) listed on this job requisition based on a full-time schedule is: $120500.00 - $231000.00.About Company
Shop Verizon smartphone deals and wireless plans on the largest 4G LTE network. First to 5G. Get Fios for the fastest internet, TV and phone service.