AI Engineer, Model Training, Inference & Infra
Bellevue, WA - USA
Job Summary
Full-time On-site San Jose CA Austin TX or Taiwan
Agentrys is building the next generation of design automation for the semiconductor industry.
Our mission is to enable every engineering organization to build its own self-improving agentic design workforce. Agentrys Studio combines AI agents engineering knowledge agent-native tools advanced models and continuous learning to automate complex chip-design workflows.
Our team brings deep experience in artificial intelligence electronic design automation semiconductor design GPU-accelerated computing and production software systems. We work closely with leading semiconductor companies to turn advanced research into technology that improves engineering productivity design quality and time to market.
We are looking for an exceptional AI Engineer to own the model training inference and infrastructure that power Agentrys agentic design workforce.
You will drive the full model lifecycle: data pipelines pretraining and post-training reinforcement learning evaluation and high-performance serving. You will build and operate the GPU training and serving stack on the compute substrate our infrastructure team provides that keeps large-scale training reliable and low-latency inference efficient at production scale. You will work on evaluation systems model training and the self-improving self-evolving learning loops that let our models and agents get better over time from real execution feedback.
Your work directly determines how capable fast and cost-effective our agents are. This role is ideal for someone who combines strong research ability with exceptional systems and performance engineering skills and who wants the models they train and serve deployed in real semiconductor design environmentsnot left in notebooks or benchmarks.
Train post-train and fine-tune large language models for agentic engineering workflows including supervised fine-tuning RLHF/RLAIF reinforcement learning and distillation.
Build scalable data pipelines for pretraining post-training and evaluation including sparse private and domain-specific engineering data.
Design and operate distributed training on multi-node GPU clusters using data tensor pipeline and sequence parallelism (for example FSDP DeepSpeed or Megatron-style approaches).
Build high-throughput low-latency inference systems with continuous batching KV-cache management paged attention quantization and speculative decoding.
Write and optimize custom GPU kernels (CUDA Triton) and profile end-to-end performance across CPUs and GPUs.
Build core model infrastructure: cluster orchestration ML job scheduling checkpointing fault tolerance reproducibility model and environment management observability and cost and utilization tracking.
Build automated evaluation systems benchmarks and reward models that measure agent capability reliability and regression across complex engineering tasks including problems where design data is private or customer-specific.
Design self-improving and self-evolving algorithms and learning loops where models and agents learn from execution feedback outcomes and new data to improve continuously over time.
Integrate models with agent runtimes tool use retrieval and the production serving stack.
Improve reliability throughput and cost efficiency across the training and inference platform.
Translate promising research ideas into reliable scalable product capabilities.
Collaborate with research product platform and solutions teams across San Jose Austin and Taiwan.
Contribute to patents publications technical presentations and the broader development of Agentic Design Automation.
PhD or masters degree in Computer Science Electrical Engineering Computer Engineering or a related field or equivalent practical experience.
Strong programming skills in Python and proficiency in at least one systems language such as C or Rust.
Deep experience with machine learning frameworks such as PyTorch or JAX.
Hands-on experience with one or more of the following:
Large-scale or distributed model training
High-performance model inference and serving
GPU programming and performance optimization
ML infrastructure and platform engineering
Automated evaluation reward modeling or self-improving and continuous-learning systems
Ability to take a model from data and problem formulation through training evaluation and production deployment.
Strong analytical software engineering performance-optimization and debugging skills.
High ownership intellectual curiosity and willingness to work across research and product boundaries.
Clear written and verbal communication skills.
Pretraining or post-training large language models at scale.
Distributed training frameworks such as Megatron-LM DeepSpeed FSDP or Ray.
Production inference engines such as vLLM TensorRT-LLM SGLang or TGI.
Custom kernel development with CUDA Triton or CUTLASS.
Inference optimization techniques such as quantization (FP8 GPTQ AWQ) speculative decoding or KV-cache optimization.
Reinforcement learning RLHF or reward-model training for LLMs.
Automated evaluation benchmarking or LLM-as-judge systems for agents.
Self-improving self-evolving or continuous-learning systems including learning from execution feedback automated curricula or synthetic data generation.
GPU cluster infrastructure with Kubernetes Slurm or Ray and high-performance networking such as NCCL or InfiniBand.
Data pipelines and MLOps for training and continuous learning.
Experience deploying AI systems in enterprise or security-sensitive environments.
A strong record of implementation through research systems open-source projects production software or technical competitions.
At Agentrys you will have the opportunity to:
Help define a new category of semiconductor design technology.
Build the training and inference stack that powers autonomous engineering agents.
Develop GPU-accelerated systems that make large-scale training and low-latency serving practical and cost-effective.
Build the evaluation and self-improvement loops that let agents learn and get better from real engineering work.
Build AI systems that perform complex consequential engineering worknot just generate recommendations.
Work with real semiconductor workflows tools and private engineering knowledge.
See your models deployed directly with leading chip-design organizations.
Work in a small highly technical team where individual contributions can shape the product and company.
Collaborate with colleagues across San Jose Austin and Taiwan.
Change how chips are designed rather than focus on only one design or one point tool.
Agentrys is an equal opportunity employer. We welcome candidates from diverse backgrounds who are excited to combine ambitious research with meaningful engineering impact.