AI Platform Engineer (Cloud)
Salary
Not disclosed
Job Type
full time
Posted
about 2 hours ago
Closing date
12 Oct 2026
Job Description
Absa Bank is seeking a skilled AI Platform Engineer with a cloud focus to join their Chief Data Analytics and Applied AI Office. This is a critical role aimed at bolstering the bank's enterprise-wide artificial intelligence capabilities, ensuring they are robust, secure, and cost-effective across various business units and geographies. The successful candidate will play a key part in the infrastructure that underpins advanced AI use cases within Corporate and Investment Banking, Personal and Private Banking, Business Banking, and Absa’s regional operations.
The position requires a deep understanding of cloud technologies and infrastructure-as-code principles, with a specific emphasis on deploying and managing AI workloads. You will be instrumental in building and maintaining the multi-cloud environment, leveraging platforms like AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, and Hugging Face. Experience with GPU-based infrastructure is also essential, reflecting the demanding nature of modern AI development.
About the Role
This role is central to the functioning and evolution of Absa’s enterprise AI platform. You will be responsible for the entire lifecycle of AI infrastructure, from initial design and deployment through to ongoing operation and continuous improvement. This involves ensuring the platform is not only technically sound but also meets stringent security, risk, and responsible AI requirements set by the bank. You will collaborate closely with a variety of stakeholders, including senior engineers, architects, security experts, FinOps specialists, and AI Solution Engineers, to deliver high-quality platform services.
Your responsibilities will span several key areas:
- **AI Platform Engineering and Architecture:** You will support the design, deployment, and operational management of the multi-cloud AI infrastructure, ensuring it scales effectively across different cloud environments and supports a wide range of AI applications.
- **AI FinOps and Cost Optimisation:** A significant aspect of the role involves monitoring AI infrastructure usage, optimising costs related to token consumption, Databricks usage, and GPU resources, and contributing to cost allocation and reporting frameworks.
- **Platform Observability and Reliability:** Implementing and maintaining robust monitoring, alerting, and dashboarding solutions will be crucial to guarantee the availability, performance, and resilience of production AI services.
- **AI Security and Zero-Trust:** You will be involved in embedding security controls within the AI platform, specifically for APIs, model endpoints, data pipelines, and agentic AI services, adhering to Absa’s security policies and regulatory mandates.
- **Agentic AI Infrastructure:** Supporting the deployment and operation of infrastructure that enables AI agents, including tool-calling services, autonomous workflows, and orchestration frameworks, will be a key responsibility.
- **Agile Engineering and Collaboration:** Working within an agile framework, you will contribute to platform enhancements while fostering strong collaboration across diverse teams and with third-party technology providers.
Key Responsibilities
Your day-to-day activities will involve a hands-on approach to building and managing complex cloud infrastructure. You will be tasked with supporting the design, deployment, configuration, and operation of Absa’s multi-cloud AI stack. This includes developing and maintaining reusable platform components, such as AI Gateway configurations, model serving environments, vector databases, and API integrations.
A core element of the role is the development and maintenance of infrastructure-as-code using tools like Terraform or Pulumi, ensuring repeatable and auditable deployments across various cloud environments. You will also configure and support agentic AI infrastructure, integrating it with enterprise systems. A focus on cloud-agnostic model serving patterns will be important to enhance workload portability.
Furthermore, you will contribute to the evaluation and implementation of new platform technologies and engineering patterns. This involves participating in architectural reviews, technical design sessions, and peer reviews, as well as creating and maintaining essential documentation, such as architectural diagrams and operational procedures. You will take ownership of the quality and operational readiness of assigned platform components and proactively escalate any significant risks.
In terms of FinOps, you will monitor and analyse platform consumption, assist with cost attribution for various AI services, and develop cost reporting dashboards. Identifying and implementing cost optimisation strategies, such as workload scheduling, right-sizing infrastructure, and utilising spot instances, will be a key contribution.
What They're Looking For
The ideal candidate will possess practical experience in cloud platform engineering, with a strong grasp of infrastructure-as-code principles and the intricacies of deploying AI workloads. Proficiency in relevant cloud technologies such as AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and Kubernetes is essential. Experience with GPU-based infrastructure is also a prerequisite.
Key technical competencies include:
- **Cloud Platform Engineering:** Proven ability to manage and scale cloud infrastructure.
- **Infrastructure-as-Code:** Hands-on experience with tools like Terraform, Pulumi, or AWS CDK.
- **AI Workload Deployment:** Familiarity with deploying and managing AI models and applications in cloud environments.
- **Platform Observability:** Experience with monitoring, logging, and alerting tools for cloud-native applications.
- **Cloud Cost Optimisation:** Understanding of cloud financial management principles and techniques.
- **Security Controls:** Knowledge of implementing security best practices in cloud and AI environments.
- **Agentic AI Infrastructure:** Familiarity with infrastructure supporting AI agents and autonomous systems.
Beyond technical skills, the role demands strong critical thinking, design thinking, and problem-solving abilities within an agile engineering context. You should be comfortable working collaboratively with diverse teams and stakeholders, and possess the ability to take accountability for assigned platform components while contributing to broader platform goals.
For individuals in South Africa, a role like this typically falls within the specialist to senior engineer salary bands, often ranging from R700,000 to over R1,200,000 per annum, depending on experience and specific responsibilities. Major financial institutions and large technology companies are the most common employers for such specialised cloud and AI engineering talent. Career progression often leads to senior engineering, architecture, or team lead positions within these organisations.
Requirements
Qualifications and Experience
Education and Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Data Science, Mathematics, Statistics, Engineering, or a related quantitative discipline is essential.
- A postgraduate qualification is advantageous.
- Relevant practical experience may be considered where supported by a strong record of cloud and platform engineering delivery.
Advantageous Certifications
One or more of the following certifications would be advantageous:
Cloud
- AWS Certified Solutions Architect
- AWS Certified Machine Learning Engineer
- Microsoft Certified: Azure AI Engineer Associate
- Microsoft Certified: Azure Solutions Architect Expert
- Databricks Certified Data Engineer or Machine Learning certification
Infrastructure-as-Code
- HashiCorp Certified: Terraform Associate
- Equivalent Terraform, Pulumi, or cloud infrastructure certification
FinOps
- FinOps Certified Practitioner
- Equivalent cloud cost management or financial operations certification
Security
- Certified Cloud Security Professional
- AWS Certified Security
- Microsoft Security, Compliance, and Identity certification
- Equivalent cloud or cybersecurity certification
Work Experience
- Approximately 4 to 6 years of relevant experience in cloud engineering, platform engineering, DevOps, MLOps, infrastructure engineering, or AI platform engineering.
- At least 2 years of practical experience supporting cloud-based data, machine learning, generative AI, or AI platform workloads in a production environment.
- Production experience with at least two of the following:
- AWS Bedrock or Amazon SageMaker
- Databricks
- Microsoft Azure AI Foundry or Azure Machine Learning
- Hugging Face
- Kubernetes-based model serving
- Practical infrastructure-as-code experience using Terraform, Pulumi, AWS CDK, or an equivalent technology.
- Experience building or supporting CI/CD pipelines for cloud infrastructure, platform components, data services, or machine learning workloads.
- Experience with Docker, Kubernetes, Helm, APIs, identity integration, and cloud-native platform services.
- Experience implementing monitoring, dashboards, alerts, and operational support processes for production platforms.
- Working knowledge of cloud cost management, cost allocation, capacity monitoring, and infrastructure optimisation.
- Experience applying cloud security controls, identity and access management, secrets management, and secure API integration.
- Experience working within enterprise risk, architecture, security, and change management processes.
- Experience in financial services, telecommunications, healthcare, insurance, or another regulated industry is advantageous.
Knowledge and Skills
- Multi-Cloud AI Platform Engineering - Practical knowledge of designing, deploying, and supporting AI services across AWS, Microsoft Azure, Databricks, Hugging Face, or Kubernetes-based environments.
- Agentic AI Infrastructure - Working knowledge of agent orchestration frameworks, tool-calling API patterns, agent memory, state management, tracing, and multi-agent workflows.
- AI FinOps and Cost Management - Knowledge of cloud consumption models, token-based pricing, Databricks DBUs, provisioned throughput, GPU utilisation, chargeback and showback reporting, and spend anomaly detection.
- AI Security and Zero Trust - Working knowledge of OAuth 2.0, OIDC, JWT, RBAC, ABAC, API security, managed identities, secrets management, prompt injection controls, data loss prevention, and secure agent tool access.
- Infrastructure-as-Code - Strong practical experience with Terraform, Pulumi, AWS CDK, or equivalent infrastructure automation technologies.
- Containerisation and Orchestration - Experience with Docker, Kubernetes, Helm, container registries, workload scheduling, resource allocation, and production container operations.
- Platform Observability - Experience with Prometheus, Grafana, Datadog, OpenTelemetry, cloud-native monitoring tools, or Databricks Lakehouse Monitoring.
- Cloud-Agnostic Model Serving - Working knowledge of containerised model deployment and serving technologies such as ONNX, BentoML, Triton Inference Server, Kubernetes, or equivalent frameworks.
- MLOps Tooling - Working knowledge of MLflow, Kubeflow, Airflow, model registries, feature stores, automated testing, and CI/CD for machine learning workloads.
- GPU Infrastructure - Understanding of GPU workload deployment, capacity management, right-sizing, spot instance strategies, and cost optimisation for model training and inference.
- Enterprise Risk and Governance - Working knowledge of information security, technology risk, architecture governance, responsible AI, privacy, data residency, and change management requirements within a regulated environment.
- Agile Delivery - Experience working in agile engineering teams using sprint planning, backlog management, iterative delivery, peer review, testing, and continuous improvement practices.
Education
Bachelor's Degree: Information TechnologyAbsa Bank Limited is an equal opportunity, affirmative action employer. In compliance with the Employment Equity Act 55 of 1998, preference will be given to suitable candidates from designated groups whose appointments will contribute towards achievement of equitable demographic representation of our workforce profile and add to the diversity of the Bank.
Absa Bank Limited reserves the right not to make an appointment to the post as advertised
About the employer
Absa
Absa is a hiring organisation operating in Sandton within the banking sector. They are currently recruiting for the AI Platform Engineer (Cloud) role advertised on this page. Visit the official application link for more about the company, its culture and the team you would be joining.
Interested in this role at Absa?
JobVault never charges job seekers to apply.
More Banking jobs
See all →- View →
Senior AI Platform Engineer (Cloud)
Absa · Sandton
- View →
Specialist: Operations (Absa Rewards/Loyalty Programmes)
Absa · Johannesburg
- View →
Executive: Corporate Real Estate Strategy and Portfolio Management
Absa · Johannesburg
- View →
Senior Manager: Risk Assurance, Governance and Reporting
Absa · Johannesburg
- View →
Financial Manager
FNB · Randburg, South Africa
- View →
Business Specialist
FNB · Johannesburg, South Africa
Get ready for your application
Free career guides written for South African job seekers.
