Job Description
Role Overview
We are looking for a Software Architect (12+ years experience) to lead the application/framework layer and deployment stack for the Next Generation Accelerator AI platform. This role owns how models run on Next Generation Accelerator—from vLLM / PyTorch / TensotFlow/XLA to production deployment—ensuring correctness, performance, and scalability.
Key Responsibilities
- Architect integration of vLLM, PyTorch, and TensorFlow, JAX/XLA into Next Generation Accelerator stack
- Define framework → compiler → runtime APIs and contracts
- Own LLM execution behavior (batching, KV cache, streaming inference)
- Design and implement end-to-end deployment workflows (packaging, versioning, reproducibility)
- Drive performance optimization across model → framework → runtime
- Work cross-functionally with compiler, runtime, and low-level SW teams
- Support customer workloads, model onboarding, and debugging
Impact
Own customer-visible AI execution and deployment on Next Generation Accelerator, closing the gap between models and system performance, and enabling enterprise-grade AI solutions