Start Your Search Here

Job Search

Accenture

Assago / Global

AI Infrastructure Architect

Job Description

Overview

In this role you design and optimize AI/ML infrastructure across cloud and on‑prem environments. You will lead architecture workstreams, implement scalable compute and data pipelines, and drive production deployment of AI systems. You partner with cross‑functional teams to ensure secure, efficient, and compliant operations while advancing InfraOps and MLOps practices. Your work supports high‑impact, real‑world AI applications and requires hands‑on coding, deployment, and mentoring. This is a ownership‑driven opportunity to shape the AI infrastructure roadmap at scale.

Responsabilità Write, review, and debug code, scripts, and infrastructure‑as‑code for AI infrastructure and tooling; set quality standards

Architect, configure, and provision cloud and on‑prem compute resources (GPU clusters, distributed training) for performance and utilization

Design and maintain deployment automation and CI/CD pipelines for reliable releases of AI systems

Deploy AI systems, models, and data pipelines into production and codify best practices

Lead container orchestration and model serving using Docker, Kubernetes, and deployment frameworks

Architect and optimize the computational stack for performance, power, cost, and scalability

Evaluate and select tools and platforms to shape the infrastructure roadmap

Integrate AI models into enterprise systems ensuring interoperability, security, and compliance

Own AI monitoring and infra health across InfraOps and MLOps, drive remediation

Troubleshoot complex issues across hardware, networking, software, and models and perform root‑cause analysis

Mentor junior engineers and lead code reviews; define architecture standards, processes, and cost‑efficient practices

Requisiti fondamentali Practical experience in coding, building, monitoring, and troubleshooting AI/ML applications and deploying/running them on premises or in public cloud

Strong understanding of AI/ML concepts

Strong understanding of computing infrastructure; preferred AI infra knowledge

Proficiency in Python, Java, or C++

Experience with data pipeline/workflow tools (e.g., Apache Airflow, Kubeflow)

Strong problem‑solving skills and fast pace

Excellent communication and collaboration

Proven experience in AI/ML infra engineering or related roles on a hyperscaler platform for deploying large scale solutions

strong communication

collaboration

mentoring

Python

Java

C++

Applica Now

Similar Opportunities

View all jobs