Enter a job title or keyword

AI Engineer (Voice & Speech)


Job Location:

Hanoi - Vietnam

Salary: Not provided by the employer
Experience Required: 2-3years
Posted: 15 August 2026 (18 days ago)
Application Deadline: 12 November 2026
Vacancies: 1 Vacancy

Job Summary

ACG3756JOB
Our client is a leading software company in Vietnam who is looking for a qualified candidate to join their firm.

JOB OVERVIEW
  • The AI Engineer will work at the intersection of AI Research Product Engineering and Production focusing on the development of advanced Speech AI and Voice AI capabilities.
  • The role involves researching and developing speech technologies building end-to-end AI pipelines deploying models into production environments and continuously optimizing AI systems for accuracy scalability latency and real-world performance.
KEY RESPONSIBILITIES

Speech AI Research & Development

  • Explore and prototype new approaches to solve complex speech and audio AI problems.
  • Develop fine-tune and adapt speech models for practical real-world applications.
  • Evaluate different modeling approaches and conduct systematic experiments to improve model performance.
  • Research and reproduce state-of-the-art Speech AI methods and translate promising research findings into production-ready solutions.
  • Design rigorous experiments benchmarks and error analyses to understand model limitations and identify opportunities for improvement.
  • Explore alternative modeling approaches when existing solutions are insufficient.

Speech AI System Development

  • Develop and optimize intelligent voice technologies including Automatic Speech Recognition Text-to-Speech audio processing speech enhancement speech understanding and voice interaction systems.
  • Improve speech recognition accuracy robustness and performance across different environments and use cases.
  • Improve speech synthesis quality naturalness and reliability.
  • Design AI solutions that can operate effectively under real-world audio conditions.

End-to-End Audio AI Pipelines

  • Design develop and optimize production-ready audio AI pipelines that connect AI models with real-world applications.
  • Build audio preprocessing and feature extraction pipelines.
  • Develop noise suppression and speech enhancement capabilities.
  • Build and optimize real-time audio streaming and processing systems.
  • Develop model serving and inference pipelines.
  • Integrate AI models into cloud mobile and other deployment environments.
  • Design and implement end-to-end AI workflows from data preparation and model development through deployment and monitoring.

AI-Powered Product Development

  • Collaborate closely with Product Mobile Backend and other Engineering teams to transform AI capabilities into practical product experiences.
  • Contribute to voice recording and transcription capabilities.
  • Develop AI-powered voice assistants and voice-driven workflows.
  • Build natural conversational experiences and speech-enabled product features.
  • Integrate AI capabilities into mobile applications and business workflows.

AI Model Production & Scalability

  • Take AI models from research prototypes through deployment into reliable production systems.
  • Optimize inference latency and computational resource consumption.
  • Optimize AI models for different deployment environments and hardware constraints.
  • Build scalable stable and reliable AI services.
  • Monitor model performance in production environments and use real-world feedback to continuously improve models and systems.
  • Identify performance degradation model limitations and opportunities for optimization after deployment.

Data & Experimentation

  • Build and improve high-quality datasets for model training and evaluation.
  • Contribute to data collection synthetic data generation data augmentation and quality validation processes.
  • Develop appropriate evaluation datasets and benchmarks for Speech AI applications.
  • Continuously improve models and datasets based on experimental results and production feedback.

Requirements
  • At least 3 years of experience developing AI/ML systems or AI-powered products.
  • Strong Python programming skills.
  • Hands-on experience with deep learning frameworks such as PyTorch TensorFlow or equivalent technologies.
  • Proven experience deploying AI models into production environments.
  • Solid understanding of speech processing audio machine learning and deep learning techniques.
  • Hands-on experience with at least one area of Speech AI including Automatic Speech Recognition Text-to-Speech Audio AI or Voice AI applications.
  • Ability to design and implement end-to-end AI and machine learning workflows.
  • Experience conducting model experimentation evaluation benchmarking and performance optimization.
  • Ability to translate research concepts into practical and scalable engineering solutions.
  • Strong analytical and problem-solving skills.
  • Product-oriented mindset with the ability to balance model performance engineering constraints and user experience.
  • Ability to collaborate effectively with cross-functional engineering and product teams.
Nice to have
  • Experience working with modern Speech AI foundation models frameworks or technologies such as Whisper wav2vec 2.0 NVIDIA NeMo XTTS VITS SpeechBrain Kaldi or equivalent solutions.
  • Experience deploying AI models on mobile or edge devices.
  • Experience with real-time audio streaming and processing systems.
  • Knowledge of speaker recognition and speaker diarization.
  • Experience with voice cloning technologies.
  • Strong knowledge of digital audio and signal processing.
  • Experience integrating Large Language Models into AI applications.
  • Experience developing conversational AI or intelligent voice interaction systems.


Benefits
  • Competitive compensation and performance-based incentives based on capabilities and contribution.
  • Full statutory benefits in accordance with applicable labor regulations and company policies.
  • Opportunities to work on challenging AI problems and apply state-of-the-art research to real-world applications.
  • Ownership of the complete AI development lifecycle from research and experimentation to model development deployment monitoring and continuous optimization.
  • Opportunities to collaborate with experienced AI Engineers Software Engineers and Product professionals.
  • Strong exposure to production-scale AI engineering and the development of practical Speech AI and Voice AI products.
  • Continuous learning and professional development opportunities in artificial intelligence machine learning speech technologies and production AI systems.
  • Opportunities to create measurable product impact by developing AI capabilities that directly improve user experiences.

Contact: Giang Van or Thao Phan or Thuong Le

Due to the immense number of applications only shortlisted candidates will be contacted.