Job Description
One of our largest Telecom customers is seeking a skilled software engineer to join their Data Platform team within the Chief Data Office. This Software Engineer will play a key role in building, operating, and scaling a production-grade AI inferencing platform that supports high-throughput, low-latency large language model (LLM) workloads. The position combines software engineering, platform engineering, and reliability engineering responsibilities, with a focus on developing Python-based services, automating model deployments, and optimizing model serving performance using vLLM, Kubernetes, KServe, and Knative. The engineer will be responsible for containerization with Docker and Podman, maintaining and tuning PostgreSQL data models, troubleshooting production issues, and driving platform reliability through observability and automation. Working closely with ML engineers and platform architects, this individual will help scale mission-critical AI services, support new model rollouts, and continuously improve the infrastructure that powers enterprise AI applications.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Required Skills & Experience
-Strong proficiency in Python, with experience building and maintaining production services.
-Hands-on experience with Docker and Podman for containerization and image management.
-Practical experience with Postgres, including query optimization and schema design.
•-Solid working knowledge of Kubernetes, including deployments, services, autoscaling, and
troubleshooting workloads in production.
-Direct experience with vLLM for LLM inference serving.
-Experience deploying models with KServe and experience with Knative for serverless workloads or event-driven scaling.
-Comfortable working in a Linux environment and using standard CLI tooling.
-Cloud platform experience, such as Azure/AKS, AWS/EKS, or GCP/GKE.
-Strong debugging skills across the stack — from application code to infrastructure
Nice to Have Skills & Experience
-Experience with GPU-aware scheduling and resource management in Kubernetes.
-Familiarity with LLM-specific optimizations: speculative decoding, quantization (FP8/INT8), KV cache
management, and continuous batching.
-Experience with API gateways or LLM routing layers, such as LiteLLM or similar.
-Background in SRE practices: SLOs/SLAs, incident response, and on-call tooling.
-Experience with Helm, ArgoCD, or other GitOps deployment tooling.
-Familiarity with observability stacks, including Prometheus, Grafana, and OpenTelemetry.
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.