Job Description:
At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work is core to how we drive Responsible Growth. This includes our commitment to being an inclusive workplace, attracting and developing exceptional talent, supporting our teammates’ physical, emotional, and financial wellness, recognizing and rewarding performance, and how we make an impact in the communities we serve.
Bank of America is committed to an in-office culture with specific requirements for office-based attendance and which allows for an appropriate level of flexibility for our teammates and businesses based on role-specific considerations.
At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!
Job Description:
This job is responsible for tool and service designs within a technical domain that enable business strategies in accordance with architectural governance, standards and policies. Key responsibilities include creating infrastructure tools and their integration as a service, facilitating deployment of technical solutions by developing templates, playbooks and automation used during implementation. Job expectations include looking for opportunities to improve efficiency when implementing and maintaining tools/services and embracing a culture of innovation and continuous improvement.
We are seeking an experienced AI/ML Infrastructure and Platform Engineer to design, build, validate, and deliver enterprise-scale AI infrastructure supporting model development, training, deployment, and inference workloads. This is an architecture and engineering role focused on secure, scalable, resilient, and production-ready AI platforms—not an operational support role.
Responsibilities:
The role requires hands-on expertise across AI-HPC platforms, GPU-enabled infrastructure, Slurm workload orchestration, OpenShift/Kubernetes, containerized AI runtimes, model-serving platforms, storage integration, observability, security, automation, and production-readiness validation.
The engineer will collaborate with AI/ML, data science, CTO, CIO, infrastructure, database, network, security, and operations teams to translate enterprise requirements into robust platform solutions.
Provides subject matter expertise and consulting services on a range of technologies and assists Technical Analysts and Infrastructure Engineers to ensure that technology solutions comply with enterprise system design and engineering standards
Assists with translating business requirements into technical definitions, reference models, blueprints, and playbooks for deployment in compliance with architecture standards and policies
Assists in the evaluation of reference models, blueprints and playbooks to ensure they are fit for purpose
Develops software solutions to address manual and repeatable work or inefficient processes
Conducts on-site evaluations of third-party products being considered for firm adoption
Promotes an inclusive and healthy working environment and helps to resolve organizational impediments/blockers
Contributes to the creation/selection of functional and non-functional product evaluation requirements within and across domains
Required Qualifications:
Design, engineer, build, and deliver AI infrastructure and platform solutions for model development, model deployment, and inference workloads.
Administer, expand, and validate AI-HPC clusters across development, disaster recovery, and production environments, including compute node onboarding, high-availability validation, environment readiness, and cluster lifecycle activities.
Install, configure, and manage Slurm-based workload orchestration, including node registration, partition management, GPU allocation, GRES configuration, job scheduling, and troubleshooting.
Manage NVIDIA GPU-enabled infrastructure, including CUDA, TensorRT, Triton Inference Server, GPU drivers, compatibility validation, performance tuning, and troubleshooting for AI/ML workloads.
Deploy and validate containerized AI platforms using Docker, Pyxis, Enroot, Podman, OpenShift, Kubernetes, operators, Helm charts, pods, services, and deployments.
Onboard, validate, and operationalize AI/ML models across GPU-enabled HPC and OpenShift environments, including model artifact integration, container runtime validation, dependency verification, endpoint testing, and production-readiness checks.
Integrate AI infrastructure with enterprise storage, data, and connectivity services, including NAS, qtrees, HPC storage, S3 object storage, model repositories, Redis, load balancers, FQDN/LTM configurations, and lifecycle workflows.
Implement observability, monitoring, logging, and alerting for AI/ML platforms using tools such as Prometheus, Grafana, Splunk, Elastic, or equivalent platforms; analyze GPU utilization, token throughput, latency, and resource consumption.
Apply secure platform engineering practices, including Kubernetes security, RBAC, network and pod security policies, authentication, authorization, JWT management, access controls, change management, auditability, and compliance controls.
Develop and maintain automation, coding standards, configuration standards, CI/CD and GitOps patterns, SOPs, technical documentation, operational runbooks, architecture strategy materials, and release-readiness validation procedures.
Collaborate across AI/ML, data science, CTO, CIO, infrastructure, database, network, security, and operations teams to translate business and technical requirements into scalable platform solutions.
Demonstrate strong analytical thinking, troubleshooting skills, documentation discipline, communication, ownership, attention to detail, and production-readiness mindset.
Desired Qualifications:
Experience supporting enterprise AI/ML platforms, model deployment pipelines, Model-as-a-Service capabilities, GPUaaS patterns, shared inference endpoints, vLLM, embeddings, reranking, and OpenAI-compatible inference endpoints.
Familiarity with AI/ML frameworks and development environments such as PyTorch, TensorFlow, Jupyter Notebook, Python, JSON, PyPI, UV, virtual environments, and dependency management.
Experience working in regulated enterprise environments with strong documentation, risk management, audit readiness, operational controls, and production change management standards.
Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical discipline, or equivalent practical experience.
7+ years of infrastructure engineering experience across AI/ML platforms, HPC systems, GPU-enabled environments, Kubernetes/OpenShift, or enterprise-scale Linux infrastructure.
Skills:
Shift:
1st shift (United States of America)
Hours Per Week:
40
Pay Transparency details
US - NJ - Jersey City - 101 Hudson St - 101 Hudson (NJ2101)Pay and benefits informationPay range$104,200.00 - $155,300.00 annualized salary, offers to be determined based on experience, education and skill set.Discretionary incentive eligibleThis role is eligible to participate in the annual discretionary plan. Employees are eligible for an annual discretionary award based on their overall individual performance results and behaviors, the performance and contributions of their line of business and/or group; and the overall success of the Company.BenefitsThis role is currently benefits eligible. We provide industry-leading benefits, access to paid time off, resources and support to our employees so they can make a genuine impact and contribute to the sustainable growth of our business and the communities we serve.