Enter a job title or keyword

Principal Edge AI Software Architect

NXP Semiconductors


Job Location:

Beijing - China

Monthly Salary: Not provided by the employer
Posted: 6 July 2026 (30+ days ago)
Application Deadline: 12 October 2026
Vacancies: 1 Vacancy

Job Summary

We are seeking an experiencedEdge AI Software Architectto lead the design and implementation of advanced machine learning solutions for edge devices and embedded systems. This role focuses on deploying and optimizing large language models (LLMs) and other AI models on resource-constrained hardware.

Responsibilities:

Architecture & Design

  • Design and architect scalable Edge AI inference engines for microcontrollers edge devices and embedded systems
  • Define technical roadmaps for deploying LLMs and foundation models on edge hardware
  • Lead the architecture of model compression quantization and optimization pipelines for resource-constrained devices

LLM & Large Model Optimization

  • Optimize and deploy Large Language Models (LLMs) on edge devices using techniques such as quantization (INT8 INT4) pruning and knowledge distillation
  • Implement model compression techniques to reduce model size while maintaining accuracy for edge deployment
  • Design and optimize inference pipelines for transformer-based models and other foundation models on low-power devices
  • Develop custom kernels and operators optimized for edge AI accelerators

Model Development & Deployment

  • Train fine-tune and optimize machine learning models using TensorFlow PyTorch and ONNX for edge deployment
  • Implement model conversion workflows (TensorFlow Lite ONNX Runtime TensorRT OpenVINO) for various edge platforms
  • Design and implement efficient model serving architectures for edge devices with latency and power constraints

Performance Optimization

  • Optimize ML algorithms and inference engines to meet strict performance power and memory constraints
  • Profile and optimize model performance on various edge AI accelerators (NPU DSP GPU)
  • Achieve low-latency high-throughput inference while minimizing power consumption

Requirements:

Education

  • Masters or Ph.D. degree in Computer Science Electrical Engineering Machine Learning or related field
  • Bachelors degree with 8 years of relevant experience may be considered

Technical Skills

  • Deep expertise in LLM optimization and deployment: quantization pruning distillation LoRA QLoRA
  • Strong proficiency in ML frameworks: TensorFlow PyTorch ONNX TensorFlow Lite PyTorch Mobile
  • Expert-level programming skillsin C/C and Python
  • Extensive experiencein embedded software development and real-time systems
  • Proven track recordof deploying ML models (especially LLMs) to production edge devices
  • Strong understanding of computer architecture memory hierarchies and hardware acceleration

Professional Experience

  • 5 years of experience in embedded ML or Edge AI development
  • Demonstrated experience optimizing and deploying large models (>1B parameters) on edge devices
  • Proven ability to architect and deliver complex ML systems from concept to production
  • Experience with model compression achieving >10x size reduction with minimal accuracy loss

Soft Skills

  • Excellent ability to read understand and implement research papers in English
  • Strong problem-solving skills and architectural thinking
  • Outstanding communication skills for technical documentation and cross-team collaboration in a global working environment
  • Experience with multimodal models (vision-language audio-text) on edge devices
  • Contributions to ML optimization frameworks or edge inference engines
  • Understanding of security considerations for edge AI deployment


More information about NXP in Greater China...

#LI-d6f4

Required Experience:

Staff IC


About Company

Company Logo

NXP is a global semiconductor company creating solutions that enable secure connections for a smarter world.

View Profile View Profile