Key responsibilities
- Team leadership and org build
- Hire, mentor, and develop a high-performing team; set the technical bar, operating rhythms, and code/research review practices
- Organize sub-teams (e.g., Core Modeling, AI Platform/Infra, Integrations) with clear ownership, SLOs, andon-call
- Manage roadmap, capacity planning, and delivery across parallel initiatives
- Architecture and platform
- Own the LLM gateway: unified APIs and proxy layers for multi-provider routing (OpenAI, Gemini, Bedrock), with rate limits, fallbacks, and cost tracking
- Build high-performance RAG pipelines (ingestion, embeddings, vector stores, caching) with robust observability and safety guardrails
- Partner with Java/NestJSteams to define clean async contracts, schemas, and eventing patterns; drive low-latency, scalable inference
- Model lifecycle and operations
- Lead end-to-end model and prompt lifecycle: data curation, training/fine-tuning, evaluation, deployment, rollback
- Establish LLMOps/MLOps: model/prompt registries, CI/CD, canary/A/B tests, offline/online evals, drift and cost monitoring
- Optimizeinference throughput and cost (autoscaling, batching, quantization/distillation, caching)
- Strategy and collaboration
- Translate company goals into an AI/ML roadmap with measurable outcomes; balance exploration with reliability and cost
- Own build-vs-buy/vendor strategy for models, infrastructure, and data services; manage budgets and SLAs
- Governance and security
- Implement data privacy, security, and compliance practices (RBAC, secrets, auditability); track prompt/model lineage and reproducibility
- Define incident response, runbooks, and postmortems for AI features
