Senior Embedded Performance Engineer
Apply on Fractile’s site →London or Bristol, 3 days in the office, 2 days WFH At Fractile, we’re building what we believe will be the world’s fastest AI inference chip from the ground up. We’re balanced across hardware and software engineering, and HW/SW co-design is real here. We move fast, and we help each other move fast. We care about each other, the software we ship, and the people who rely on it. On the device, close to the metal, we write the runtime software that orchestrates work across the chip and runs performance-critical ML kernels. This is where performance gets real and the wins compound. Your work directly influences trade-offs for the silicon, system deployment, and the compiler. You'll drive the first accelerator compute runs, evaluating performance on silicon, running early benchmarks, and feeding results back into the hardware and software roadmap. What you’ll do Write and optimise performance-critical ML kernels in C, with assembly where it matters (RISC-V and our own ISA) Build the low-level control paths that feed those kernels, including scheduling, synchronisation, and data movement Write targeted validation workloads and microbenchmarks to keep simulation and hardware behaviour aligned and performance measurable. Profile, benchmark, and track regressions so performance improvements are real and repeatable Work closely with simulation, hardware, ML, compiler, firmware, and runtime engineers in a tight loop, turning profiling data into architecture feedback and real performance wins.