Skip to content
Silicon Roles

Software Engineer, GPU Inference

CerebrasUS and Canada OfficesReliability & Quality
Apply on Cerebras’s site →
First seen on this site 2026-08-10.
Posted 2025-11-25 · read from ashby

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership https://openai.com/index/cerebras-partnership/ with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. ABOUT THE ROLE Cerebras is building a new generation of disaggregated AI inference systems https://www.cerebras.ai/press-release/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference that combine GPU-accelerated prefill with ultra-fast decode on the Cerebras Wafer-Scale Engine. We are hiring a Software Engineer to productionize and optimize our GPU serving stack, working across our custom inference APIs, the vLLM serving runtime, the AMD ROCm software stack, and rack-scale AMD GPU infrastructure, to make this new serving path reliable, numerically correct, observable, and exceptionally performant. You will write production code, establish operational practices for a new accelerator fleet, and drive improvements in time to first token, throughput, tail latency, and capacity efficiency.