Member of Technical Staff | Inference Platform
Department:
Job Summary
At Avra every technical IC is a Member of Technical Staff (MTS). The title doesnt put anyone in a silo: you own systems and outcomes not steps in a function and you keep building depth in your area. Seniority shows up in your scope level and compensation not in titles.
In this role youll join the Platform team to own where our models execute. Customers consume our models through large batches of millions of records and through real-time APIs and they make business decisions on every response. Youll run governed model releases reliably and efficiently in our cloud and on customer-hosted Kubernetes and make inference fast predictable and cheap enough to serve both enterprise and mid-market customers.
Evolve Sophos our online and batch inference runtime built on Kubernetes.
Run large batch inference on ephemeral jobs with multi-dimensional admission control (CPU memory GPU) through Kueue.
Build and extend the Sophos controller and its Kubernetes custom resources.
Optimize each models inference engine and feature processing using vectorized columnar operations.
Serve graphs and data efficiently from Lance-based storage.
Own execution of training post-training and fine-tuning jobs in our cloud and in customer dataplanes.
Drive autoscaling GPU serving performance and cost optimization with telemetry for every model we run.
Solve open problems such as deterministic job sizing checkpointing and recovery for batch runs per-customer encryption and isolation resilience to difficult input files and automatic profiling when a new model is accepted.
99.9% serving availability.
p95/p99 latency for online inference and throughput for batch.
Cost per prediction and per training job.
GPU utilization: paid capacity versus capacity actually used.
Training and batch jobs that finish on time and succeed without manual retries.
Experience running model serving or large-scale batch compute on Kubernetes.
Experience building Kubernetes controllers or operators.
Skill at profiling and optimizing data-heavy Python pipelines.
A clear sense of cost: you treat compute efficiency as a product feature.
Production-quality code and reviews and a willingness to operate what you build.
Ray Ray Serve or KubeRay in production.
Kueue or other batch scheduling and admission-control systems.
GPU serving and performance optimization.
Arrow Parquet Lance or other columnar formats.
Shipping software to customer-hosted Kubernetes.
GCP/AWS and GKE/EKS and financial services or regulated environments.
Required Experience:
Staff IC
About Company
Our foundation model helps our clients bring the right SME to the top of the funnel, hyper-personalize offers, and reduce default. Request a demo.