Skip to content
All jobs

Sr. Engineer, AI Platform

  • Abbott
  • United States - Wisconsin - Madison, United States of America
  • Full time

Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Position Overview The Sr AI Platform Infrastructure Engineer is a pivotal position for driving the strategic vision and technical execution for the Cancer Diagnostics AI/ML platform. The role is accountable for the architecture, reliability, scalability, security and cost of every layer beneath the platform's applications, from the cloud, network and compute foundation to the serving, orchestration and streaming systems that workloads run on. This engineering leader partners closely with the AI Platform Software and Safety leads to own and develop the AI/ML platform. This role architects, builds and operates AI infrastructure as software, supporting large language models, predictive models, agentic systems and traditional software applications across the MLOps lifecycle. This is a hands-on engineering role that requires a technical leader to demonstrate ownership, accountability and leadership. This role is based in Madison, WI. Essential Duties Include, but are not limited to, the following: Design and own the platform's cloud and Kubernetes foundation, including account and network topology, VPC and subnet design, private connectivity, ingress and egress control, DNS and certificates, and the provisioning, upgrade, migration, multi-tenant isolation and security of production clusters. Lead the platform's expansion across clusters and regions, covering topology, identity federation, state and data placement, traffic management and failure domains, and evaluate multi-cloud where business need warrants it. Own infrastructure-as-code and GitOps for the platform, including module design, state management, environment promotion, drift control, and the testing and rollback of infrastructure changes; build and operate the CI/CD and deployment tooling that platform software releases run on in partnership with the AI Platform Software lead. Architect, build and operate production model serving and GPU infrastructure for large language and predictive models, including scheduling, multi-tenancy, autoscaling, low-latency serving, inference optimization (quantization, batching, graph compilation), and capacity planning and cost attribution for shared compute. Partner with the AI Platform Software and Safety leads on the model lifecycle by providing the infrastructure it runs on, including reproducible data, feature and training pipelines with dataset-to-model lineage, promotion environments, the traffic-shifting and rollback mechanisms behind canary, shadow and A/B rollout, and pipeline enforcement of the evaluation and promotion gates. Architect, build and operate workflow orchestration and event-streaming infrastructure, including workflow and state-machine design, event-driven triggers and scheduling, retries and idempotency, topic and partition design, schema management, delivery semantics and cross-region replication. Own the observability stack and service level objectives for infrastructure services, with proactive detection across the platform; lead disaster-recovery strategy as a business tradeoff between cost and recovery time and recovery point objectives; and serve as the escalation point for complex infrastructure incidents, leading the post-incident analysis that drives systemic reliability improvements. Own infrastructure security and governance, including least-privilege identity, secrets management, network policy, workload isolation and auditability, and partner with the safety lead to build the governance controls into the paved road as automated checks. Design, develop, test, review and deploy production software as a hands-on contributor, mentor engineers, and represent the infrastructure pillar in design reviews and long-term technology planning. Uphold company mission and values through accountability, innovation, integrity, quality and teamwork. Maintain regular and reliable attendance. Ability to act with an inclusion mindset and model these behaviors for the organization. Minimum Qualifications Bachelor’s degree in Computer Science, Artificial Intelligence, Data Science, or a related field. Demonstrated expertise in designing and deploying cutting-edge AI solutions that drive measurable business impact. Demonstrated understanding of ethical considerations in AI systems. 3+ years of project leadership experience including Agile project management, Scaled Agile Frameworks (SAFE). Advanced ability working with machine learning frameworks, programming languages like Python, and cloud platforms. Strong analytical and problem-solving skills with expertise in AI/ML techniques. Preferred Qualifications Production experience with Python, Go and/or Rust. Experience leading platform transformations and/or building large-scale AI infrastructure. Multi-cloud experience, or experience migrating significant workloads between cloud providers. Kubernetes-native serverless and event-driven autoscaling (for example Knative or KEDA), service mesh, API gateway and zero-trust network patterns. Cost ownership of a GPU fleet, or comparable FinOps accountability for shared compute. Experience operating inside enterprise guardrails, where network, identity and change management are owned by other teams, and knowledge of compliance frameworks such as HIPAA, FDA and GxP. Experience guiding cross-functional engineering groups or mentoring senior-level engineers. The base pay for this position is $99,300.00 – $198,700.00 In specific locations, the pay range may vary from the range posted. JOB FAMILY: Product Development DIVISION: ONCO Cancer Diagnostics LOCATION: United States > Madison : 5505 Endeavor Ln ADDITIONAL LOCATIONS: WORK SHIFT: Standard TRAVEL: No MEDICAL SURVEILLANCE: No SIGNIFICANT WORK ACTIVITIES: Continuous sitting for prolonged periods (more than 2 consecutive hours in an 8 hour day), Keyboard use (greater or equal to 50% of the workday) Abbott is an Equal Opportunity Employer of Minorities/Women/Individuals with Disabilities/Protected Veterans. EEO is the Law link - English: EEO is the Law link - Espanol: