Skip to content
Silicon Roles

Senior AI Infrastructure Engineer - EDA Infrastructure

NvidiaUS, CA, RemoteEDA & CAD
Apply on Nvidia’s site →
First seen on this site 2026-09-01, open 25 days.
Posted 2026-09-21 · read from workday

AI Infrastructure Engineers at NVIDIA build the systems, tooling, and data infrastructure that enable operation of our GPU cloud services. We are enabling engineering teams to innovate while proactively identifying, tracking, and mitigating risks across the entire technical task. This role is ideal for engineers who thrive at the intersection of product, infrastructure, and software engineering and who want to build automated, intelligence-driven systems that protect NVIDIA’s most critical AI platforms. What you’ll be doing: Telemetry Build and operate scalable telemetry pipelines for metrics, logs, traces, and events across on-premise, CSP, and NCP clusters. Establish common instrumentation, collection, storage, and access patterns so teams can generate and consume telemetry consistently. Deliver dashboards, alerting, and analysis capabilities that improve service visibility, detection, and troubleshooting. Operational Excellence Standardize and automate incident, maintenance, service on-call, and support on-call workflows across HWInf. Integrate operational data and lifecycle signals to improve ownership, escalation, communication, and post-incident learning. Build reporting and AI-assisted tooling that reduces manual toil and improves operational responsiveness. Cataloging and Inventory Build and maintain physical hardware and software catalogs as trusted sources of truth for infrastructure inventory, service ownership, dependencies, and documentation.