Staff Backend Engineer K8 (Envoy)
Job Summary
CompanyIntroduction:
We exist to wow our customers. We know were doing the right thing when we hear our customers say How did I ever live without Coupang Born out of an obsession to make shopping eating and living easier than ever were collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.
We are proud to have the best of both worlds a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang you will see yourself your colleagues your team and the company grow every day.
Our mission to build the future of commerce is real. We push the boundaries of whats possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on high-tech and hyper-connected world.
RoleOverview:
As a Staff Backend Engineer you will work closely with platform and product leaders to design and deliver solutions for complex infrastructure problems. You will drive the development of highly scalable reliable and efficient platform services while providing technical direction across teams working with Java AWS Kafka Kubernetes Kubeflow Argo CD and gRPC.
WhatYouWillDo:
- ArchitectandbuildCoupangsnext-generationAIInferenceGatewayplatformthatservesmission-criticalmachinelearningandgenerativeAIworkloadsatscale.
- Designanddevelophigh-performancerequestroutingmodelendpointabstractiontrafficshapingloadbalancingfailovercachingandpolicyenforcementmechanismsforinferenceservices.
- DrivethetechnicalvisionandroadmapforscalablesecureandreliableAIinferenceinfrastructureacrosscloudandon-premenvironments.
- DevelopcriticalinfrastructurecomponentsinGoJavaorPythonwithastrongfocusonperformanceresiliencyandoperationalexcellence.
- Designmulti-tenantplatformcapabilitiesincludingauthenticationauthorizationquotamanagementcostattributionratelimitingandgovernancecontrols.
- PartnercloselywithMLPlatformModelServingDataandProductEngineeringteamstoenableseamlessdeploymentandoperationofAIworkloads.
- Leadarchitectureanddesignreviewsraiseengineeringstandardsandmentorseniorengineersacrossmultipleteams.
- Optimizesystemperformancelatencythroughputandinfrastructureefficiencyforlarge-scaleinferenceworkloads.
- DefineobservabilitystandardsthroughmetricstracingloggingandSLO-basedoperationsforbusiness-criticalAIservices.
- Investigatecomplexproductionissuesdriveroot-causeanalysisandimplementlong-termarchitecturalsolutions.
- CollaboratewithengineeringleadersacrossCoupangtoestablishcommonplatformstandardsandunlockAIinnovationacrosstheorganization.
BasicQualifications:
- 8yearsofprofessionalsoftwaredevelopmentexperience.
- 5yearsofexperiencedesigningandoperatinglarge-scaledistributedsystemsinproduction.
- Stronghands-onprogrammingexpertiseinoneormoreofGoJavaorPython.
- Proventrackrecordofbuildinghighlyavailablemission-criticalplatformorinfrastructureservices.
- ExperiencedesigningAPIplatformsservicegatewaysservicemeshorlarge-scalenetworkinginfrastructure.
- Deepunderstandingofmicroservicesarchitecturedistributedsystemsdesignandcloud-nativetechnologies.
- ExperiencewithKubernetesandcontainerizedworkloadsinproductionenvironments.
- ExperienceoperatingservicesonAWSAzureorGCP.
- Strongunderstandingofobservabilityreliabilityengineeringcapacityplanningandproductionoperations.
PreferredQualifications:
- ExperiencebuildingAI/MLinferenceplatformsLLMgatewaysmodelservinginfrastructureorGPU-acceleratedworkloads.
- DeepexpertiseinKubernetesecosystemtechnologiessuchasGatewayAPIIngressControllersServiceMesh(IstioLinkerdEnvoy)andplatformnetworking.
- ExperiencewithinferenceservingframeworkssuchasvLLMTritonInferenceServerTensorRT-LLMRayServeKServeSGLangorsimilartechnologies.
- Strongunderstandingofhigh-performancenetworkinggRPCHTTP/2streamingprotocolsAPIgatewaysandserviceproxyarchitectures.
- Experiencedesigninglarge-scaletrafficmanagementsystemsincludingroutingretriescircuitbreakingratelimitingandrequestprioritization.
- ExperienceoptimizinglatencythroughputandresourceutilizationforCPUandGPUworkloads.
- FamiliaritywithGenAIandLLMecosystemsincludingmodeldeploymentpromptroutingRAGsystemsmodelobservabilityandAIgovernance.
- ExperiencewithdistributeddatasystemssuchasKafkaCassandraRedisMongoDBorsimilartechnologies.
- Strongunderstandingofconcurrencysynchronizationasynchronousprogrammingandnon-blockingI/O.
Typeofworkmodel:
Hybrid/Onsite/Remoteworking
- OurHybridworkmodel:.
Detailstoconsider:
Those eligible for employment protection (recipients of veterans benefits the disabled etc.) may receive preferential treatment for employment in accordance with applicable laws.
PrivacyNotice
- :// Experience:
Staff IC
About Company
Join us to innovate. Rocket your career. Collaborate with teams across the globe. Find your role and learn more about our culture.